Similar Jobs

See all

Production Infrastructure Ownership:

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
  • Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.

Scalability and Reliability:

  • Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
  • Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly.

Developer Experience and Tooling:

  • Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
  • Write clean, maintainable production code to automate systems, improve backend services, and create internal developer tooling.

AI/ML Platform Startup

A fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. The engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues.

Apply for This Position