Source Job

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Kubernetes Python Go Distributed Systems Observability

20 jobs similar to Staff Site Reliability Engineer

Jobs ranked by similarity.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

Portugal

  • Work with teams to define SLIs and SLOs, and create systems for observability.
  • Analyze failure scenarios, create runbooks, and reduce work that does not add value.
  • Participate in incident management and facilitate on-call duty to ensure reliable production environments.

Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.

US

  • Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
  • Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
  • Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.

Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

$154,700–$208,000/yr
US Canada

  • Design and build automated reliability and self-healing systems to protect production at scale.
  • Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
  • Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.

Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.

$118,800–$237,600/yr
Europe

  • Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
  • Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
  • Manage distributed systems, observability, incident response, and automation with a security-first mindset.

Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.

$200,000–$215,000/yr
US

  • Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
  • Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
  • Automate toil and set SRE standards while mentoring engineers on observability tooling.

WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

UK

  • Lead Cloud Platform and SRE teams, driving infrastructure strategy and ownership including Kubernetes, GCP, and Terraform.
  • Champion SRE culture, define SLOs, SLAs, and enhance observability and incident management.
  • Own the developer-facing platform as a product, ensuring self-service infrastructure and security compliance (SOC2, ISO-27001).

Prolific builds human data infrastructure for AI development, connecting researchers with a global participant pool. They foster a culture of high performance, ownership, and cross-functional collaboration.

India

  • Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
  • Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
  • Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.

$150,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Design the human side of reliability: on-call rotations, incident roles, blameless postmortems, and change management.
  • Build and operate the core reliability toolkit: observability, CI/CD, infrastructure as code, and incident response.
  • Embed with product engineers to raise operational literacy and define SLOs grounded in patient needs.

Novellia is a health tech startup that provides patients with access to their health records and enables biopharma innovation through real-world data. They are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors, with a focus on building a diverse and inclusive workplace.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

$62,640–$104,760/yr
Europe 4w PTO

  • Design, build, and run distributed cloud architectures and large-scale production systems.
  • Ensure reliability, observability, performance, and cost efficiency of the platform.
  • Collaborate with product and backend teams to design system architecture and optimize resource use.

Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.

US

  • Design, scale, and maintain enterprise monitoring and alerting ecosystems across multi-cloud and native systems.
  • Bridge development and operations to ensure high availability, performance tuning, and deep visibility.
  • Automate infrastructure and build robust observability pipelines using cloud-native tools like Prometheus, Grafana, and GCP.

Ontrac Solutions is a leading technology consulting firm specializing in cutting-edge solutions that drive business transformation. Their team is committed to innovation, collaboration, and excellence, empowering clients to succeed in an evolving digital landscape.

Canada Europe

  • Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
  • Partnering with development teams to establish production readiness and operational readiness.
  • Building tooling to automate observability and operational workflows, eliminating manual toil.

Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.

$260,000–$280,000/yr
US

  • Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
  • Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
  • Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.

We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.

$75,600–$124,200/yr
Europe

  • Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
  • Drive automation and eliminate operational toil through self-service tooling and process improvements.
  • Lead incident response and mentor engineers to improve reliability practices.

Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.

UK

  • Help define and mature Engineering Operations by improving application health visibility, service reliability, and operational analytics.
  • Build and implement scalable processes for Incident, Problem, and Change Management that engineers actually want to use.
  • Connect engineering systems, data, and teams to reduce fragmentation and improve operational visibility across the organization.

Turnitin is a recognized innovator in global education, developing learning integrity solutions that help educators and institutions uphold academic integrity. With over 16,000 academic institutions using our services in more than 185 countries, we foster a remote-first culture and a diverse community of colleagues across 35+ countries.