Source Job

$192,000–$192,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable and reliable systems.
  • Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
  • Develop and maintain automation tools for deployment, monitoring, and system health checks.

Python Go AWS Kubernetes Terraform

20 jobs similar to Senior Site Reliability Engineer (SRE)

Jobs ranked by similarity.

Canada

  • Design, implement, and maintain highly available and scalable infrastructure solutions.
  • Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
  • Automate infrastructure deployment, configuration management, and operational workflows.

The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.

$160,000–$208,000/yr
US

  • Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
  • Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
  • Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.

Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.

US

  • Design and implement scalable cloud infrastructure to support growth.
  • Develop monitoring, alerting, and incident response for system reliability.
  • Automate deployment pipelines and ensure high availability and security.

Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

Ireland

  • Design, build, and deploy production systems with focus on scalability, reliability, and security.
  • Develop and maintain automation to streamline operations and eliminate toil.
  • Proactively monitor systems and implement automated incident response to minimize downtime.

Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.

UK

  • Collaborate with engineering teams to design scalable, secure systems.
  • Establish SLOs, manage incident response, and drive reliability improvements.
  • Leverage expertise in Go, Python, Kubernetes, and cloud platforms.

ClickHouse is a leading real-time analytics company recognized on the 2025 Forbes Cloud 100 list. With over 3,000 customers and rapid growth, the company offers a remote-friendly, globally distributed culture.

$160,000–$208,000/yr
US Unlimited PTO

  • Design, build, and optimize reliable infrastructure for healthcare technology.
  • Improve scalability, reliability, and performance across distributed systems.
  • Collaborate with engineers and data professionals to shape modern infrastructure practices.

This company provides innovative healthcare technology solutions. It fosters a remote-first culture with a focus on engineering excellence and collaboration.

US Unlimited PTO

  • Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
  • Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
  • Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.

NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

$75,600–$124,200/yr
Europe

  • Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
  • Drive automation and eliminate operational toil through self-service tooling and process improvements.
  • Lead incident response and mentor engineers to improve reliability practices.

Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.

Canada

  • Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
  • Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
  • Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.

Our partner is a technology company focused on high-availability platforms and mission-critical infrastructure. The team is collaborative and works with modern cloud technologies.

$120,000–$150,000/yr
US Unlimited PTO

  • Optimize new and existing systems by increasing reliability, performance, and scalability.
  • Automate routine operational tasks to reduce toil and improve efficiency.
  • Ensure infrastructure security compliance and implement least-privilege access controls.

Prove provides phone-centric identity tokenization and passive cryptographic authentication solutions to reduce friction and enhance security across digital channels. With over 1,000 enterprise customers processing 20 billion requests annually, they foster a fast-paced, collaborative culture focused on impact and tenacity.

$260,000–$280,000/yr
US

  • Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
  • Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
  • Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.

We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

Portugal

  • Work with teams to define SLIs and SLOs, and create systems for observability.
  • Analyze failure scenarios, create runbooks, and reduce work that does not add value.
  • Participate in incident management and facilitate on-call duty to ensure reliable production environments.

Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

US

  • Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
  • Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
  • Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.

Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

$150,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Design the human side of reliability: on-call rotations, incident roles, blameless postmortems, and change management.
  • Build and operate the core reliability toolkit: observability, CI/CD, infrastructure as code, and incident response.
  • Embed with product engineers to raise operational literacy and define SLOs grounded in patient needs.

Novellia is a health tech startup that provides patients with access to their health records and enables biopharma innovation through real-world data. They are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors, with a focus on building a diverse and inclusive workplace.