Similar Jobs
See allForward Deployed Engineer - SRE
Andromeda Cluster
North America
GPU Clusters
InfiniBand
Kubernetes
Senior Solutions Engineer
TensorWave
US
Kubernetes
GPU Infrastructure
Linux
Senior AI Infrastructure & Platform Operations Engineer
Mirantis
Europe
Linux
Kubernetes
Networking
Senior Manager, AI Infrastructure Operations
Jobgether
United States
Linux
Kubernetes
Terraform
AI Infrastructure & Platform Operations Engineer (remote in the US)
Mirantis
US
Linux
Kubernetes
NVIDIA GPU
About the role:
- You will be the first line of defense to support customers building training and inference solutions.
- You will dive deep into complex technical challenges, providing swift solutions while serving as a product expert.
Responsibilities:
- Monitor GPU cluster health and proactively communicate hardware issues to customers.
- Operate and maintain production infrastructure including fleet rebalancing and workload management.
- Investigate and resolve storage and networking issues such as Weka filesystem degradation and InfiniBand failures.
Requirements:
- 3+ years of customer-facing technical role with at least 1 year supporting an AI service or mission-critical API.
- Experience as an SRE or DevOps engineer working with Kubernetes and high-performance computing.
- Strong understanding of GPU technologies, SLURM, and high-performance network fabrics.
Together AI
Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.