Similar Jobs

See all

The Role:

  • Set the standard for building, shipping, and operating ML and AI systems at scale.
  • Tackle hard problems: model serving reliability, inference cost and latency, reproducible pipelines, and agentic workload operations.

What You'll Work On:

  • Build and operate model and inference serving infrastructure with a focus on latency, throughput, autoscaling, and reliability.
  • Own the ML deployment lifecycle including model registry, versioning, promotion workflows, and rollout strategies.
  • Operate agentic and LLM workloads in production, managing inference providers, gateways, quotas, and guardrails.

Must Have:

  • 5+ years in platform engineering, SRE, MLOps, or infrastructure with production systems.
  • Deep Terraform and Kubernetes expertise with hands-on GKE and multi-project GCP experience.
  • Strong automation-first mindset and experience with agentic coding tools.

ReadyOn

ReadyOn is an AI-native Labor Operating System that redefines how enterprises manage frontline labor by matching workers to shifts in real time. Headquartered in San Francisco with over 100 employees, it grew revenue 8x year over year in 2025.

Apply for This Position