Similar Jobs

See all

Responsibilities:

  • Design and evolve cloud infrastructure on GCP for scale and resilience.
  • Build internal tooling and automation that promote team autonomy and developer productivity.
  • Advance observability platform with metrics, logging, tracing, and alerting to reduce recovery time.

Requirements:

  • 3+ years of SRE or production SRE experience.
  • Proficiency with GCP including cost optimization and governance.
  • Hands-on experience with Kubernetes, Terraform, and observability stack.

Benefits:

  • Fully remote role in EU, UK, or North America.
  • Compensation discussed during interview.
  • Opportunity to work on cutting-edge AI/ML infrastructure.

Unnamed AI/ML Company

The company is a well-funded AI/ML company at the intersection of geospatial intelligence and climate technology, building products on scalable cloud infrastructure. The engineering team fosters a culture of reliability and continuous improvement, operating with a focus on SLOs, error budgets, and DORA metrics.

Apply for This Position