Similar Jobs
See allAI Infrastructure & Platform Operations Engineer (remote in the US)
Mirantis
US
Linux
Kubernetes
Networking
Infrastructure/GPU Cluster/Platform Operations Lead
ELEKS
Canada
Kubernetes
CUDA
Cloud Infrastructure
Staff Platform Engineer
Postscript
Global
Kubernetes
AWS
Terraform
Senior Manager, AI Infrastructure
Vultr
Kubernetes
Terraform
Ansible
Senior Solutions Engineer
TensorWave
US
Kubernetes
Linux
Python
About Andromeda:
- Provides scaled AI infrastructure for startups with 80+ customers and tens of thousands of GPUs.
- Founded by Nat Friedman and Daniel Gross to give startups access to hyperscaler-level compute.
The Role:
- Build and operate the control plane that runs the GPU fleet.
- Automate cluster deployment from bare metal to customer-ready.
- Manage machine lifecycle, Kubernetes, Postgres, and custom operators at scale.
Requirements:
- 2+ years of on-call experience for critical production services.
- Deep Kubernetes experience and strong Linux fundamentals.
- Experience operating databases, monitoring, CI/CD, and cloud infrastructure.
Andromeda
Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.