What you'll own:
- Own our platform: GCP, Kubernetes, Temporal, the GPU fleet behind cloud export, and the deploy and rollback machinery everything ships through.
- The AI enablement substrate: GPU capacity, training and inference pipelines, and the reliability and cost of the systems serving models in production.
- Security comes with the systems you run: identity and access, secrets management, least-privilege boundaries, and supply-chain integrity.
What you bring:
- 8+ years building and operating production distributed systems or equivalent server-side engineering with a heavy infrastructure focus.
- Experience running systems where failure was expensive, using SLOs and error budgets as operating tools.
- Proficiency with a major cloud provider and Kubernetes in production, with infrastructure-as-code as your default.
Required and preferred experience:
- Required: 8+ years experience, distributed systems, cloud provider, Kubernetes, infrastructure-as-code.
- Preferred: GPU and ML infrastructure, production security engineering, cloud cost modeling, CI/CD at monorepo scale.
- Small teams owning a large surface, whether at a Series B to D company or on an internal platform team.