Similar Jobs
See allTechnical Support Engineer (GPU Clusters)
Together AI
US
Kubernetes
GPU
Ansible
Staff AI Scheduling & Orchestration Engineer
Bitdeer
US
Kubernetes
Distributed Systems
Terraform
Senior Software Engineer (Storage)
Mirantis
US
Go
Kubernetes
Infrastructure As Code
Senior Software Engineer - Cloud Platform
Shadeform
US
Kubernetes
Golang
Cloud Platforms
Sr Software Engineer - Cloud Platform—Kubernetes, Hyperscalers & Distributed Systems
ServiceNow
APAC
Go
Kubernetes
AWS
What You’ll Do:
- Design and build a managed Slurm service on Kubernetes
- Write clean, reliable, and maintainable Go code, developing scheduling and orchestration for GPU workloads
- Build observability and automated remediation for GPU, node, network, and control-plane failures
What We're Looking For:
- Hands-on experience with Slurm in production, including submitting and debugging workloads
- Strong proficiency in Go and building Kubernetes operators, controllers, CRDs
- Product mindset and ability to take end-to-end ownership of complex distributed-system challenges
Nice to Have:
- Experience operating large-scale HPC or GPU clusters for external customers
- Experience with distributed training and high-performance infrastructure
- Contributions to Slurm, Kubernetes, or other cloud-native HPC open-source projects
Benefits:
- Competitive compensation and flexible working hours
- Private medical insurance and extra paid vacation/sick leave
- Language courses and modern offices with snacks and activities
Gcore
Gcore provides infrastructure and software solutions for AI, cloud, network, and security, powering digital experiences worldwide. With over 550 professionals, they build and support the global digital ecosystem.