Source Job

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Python Go Kubernetes AWS CI/CD

20 jobs similar to Senior Site Reliability Engineer

Jobs ranked by similarity.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

$160,000–$208,000/yr
US

  • Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
  • Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
  • Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.

Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.

$192,000–$192,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable and reliable systems.
  • Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
  • Develop and maintain automation tools for deployment, monitoring, and system health checks.

LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.

Portugal

  • Work with teams to define SLIs and SLOs, and create systems for observability.
  • Analyze failure scenarios, create runbooks, and reduce work that does not add value.
  • Participate in incident management and facilitate on-call duty to ensure reliable production environments.

Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

Spain

  • You will own and deliver quarterly goals for your team, leading engineers through ambiguity to solve open-ended problems.
  • You will proactively identify technical solutions and operational processes that strengthen incident readiness and response.
  • You will foster a culture of quality and ownership by setting or improving code review and design standards.

Affirm is reinventing credit to make it more honest and friendly, offering consumers the flexibility to buy now and pay later. The company has a strong engineering culture focused on reliability and ownership.

$150,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Design the human side of reliability: on-call rotations, incident roles, blameless postmortems, and change management.
  • Build and operate the core reliability toolkit: observability, CI/CD, infrastructure as code, and incident response.
  • Embed with product engineers to raise operational literacy and define SLOs grounded in patient needs.

Novellia is a health tech startup that provides patients with access to their health records and enables biopharma innovation through real-world data. They are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors, with a focus on building a diverse and inclusive workplace.

Canada Europe

  • Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
  • Partnering with development teams to establish production readiness and operational readiness.
  • Building tooling to automate observability and operational workflows, eliminating manual toil.

Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

UK

  • Collaborate with engineering teams to design scalable, secure systems.
  • Establish SLOs, manage incident response, and drive reliability improvements.
  • Leverage expertise in Go, Python, Kubernetes, and cloud platforms.

ClickHouse is a leading real-time analytics company recognized on the 2025 Forbes Cloud 100 list. With over 3,000 customers and rapid growth, the company offers a remote-friendly, globally distributed culture.

$166,500–$291,400/yr
North America And Canada

  • Design and build cloud-native engineering platforms for software validation, release validation, and production readiness.
  • Develop automation solutions that improve engineering productivity and reduce manual toil through shift-left practices.
  • Foster a culture of reliability, automation, and operational excellence while mentoring engineers.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help organizations work smarter. They serve 85% of the Fortune 500 and foster an AI-native culture where technology and talent are unstoppable.

$75,600–$124,200/yr
Europe

  • Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
  • Drive automation and eliminate operational toil through self-service tooling and process improvements.
  • Lead incident response and mentor engineers to improve reliability practices.

Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.

Ireland

  • Design, build, and deploy production systems with focus on scalability, reliability, and security.
  • Develop and maintain automation to streamline operations and eliminate toil.
  • Proactively monitor systems and implement automated incident response to minimize downtime.

Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.

Global

  • Support the deployment, operation, and maintenance of the Karuna service running on Kubernetes.
  • Monitor production environments to ensure high availability, reliability, and performance.
  • Investigate, troubleshoot, and resolve production incidents, performing root cause analysis.

Software Mind develops solutions that make an impact for companies around the globe. They build cross-functional engineering teams with a culture of openness, respect, grit, and enjoyment.

$160,000–$208,000/yr
US Unlimited PTO

  • Design, build, and optimize reliable infrastructure for healthcare technology.
  • Improve scalability, reliability, and performance across distributed systems.
  • Collaborate with engineers and data professionals to shape modern infrastructure practices.

This company provides innovative healthcare technology solutions. It fosters a remote-first culture with a focus on engineering excellence and collaboration.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

$190,000–$225,000/yr
US

  • Design, build, and operate core cloud infrastructure on AWS, including compute, networking, and container orchestration.
  • Own the CI/CD platform used across engineering teams, including build pipelines, environment promotion, and progressive rollout.
  • Build and maintain the observability stack across the organization, including logging, metrics, distributed tracing, and alerting.

RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company is an Equal Opportunity and Affirmative Action employer committed to diversity and collaboration.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.