Source Job

$100,000–$150,000/yr
US

  • Design, build, and maintain reliable distributed systems to improve availability and performance.
  • Develop automation tools using Python, Go, or Java to reduce operational complexity.
  • Lead incident response and implement observability solutions with Prometheus and Grafana.

Python Kubernetes Linux Prometheus Grafana

20 jobs similar to Systems Reliability Engineer

Jobs ranked by similarity.

$140,400–$372,300/yr
US

  • Partner with engineering teams to improve reliability, scalability, and operational health of production systems.
  • Investigate and resolve complex production incidents, designing sustainable long-term solutions.
  • Design, build, and maintain automation tools and infrastructure to enhance developer productivity.

The company builds and maintains highly reliable, scalable production systems supporting millions of users worldwide. It fosters a remote-first culture that values innovation, collaboration, and engineering excellence.

$124,000–$170,500/yr
US

  • Perform operational deployments, implementations, and maintenance for production systems.
  • Implement and maintain monitoring, reporting, and alerting systems for Core Speech products.
  • Be part of an on-call rotation and work collaboratively to improve system performance and architecture.

Solventum is a new healthcare company with a long legacy of solving big challenges to improve lives and enable healthcare professionals to perform at their best. They are a large company that values empathy, insight, and clinical intelligence, collaborating with top minds in healthcare.

$160,000–$208,000/yr
US Unlimited PTO

  • Design, build, and optimize reliable infrastructure for healthcare technology.
  • Improve scalability, reliability, and performance across distributed systems.
  • Collaborate with engineers and data professionals to shape modern infrastructure practices.

This company provides innovative healthcare technology solutions. It fosters a remote-first culture with a focus on engineering excellence and collaboration.

United States

  • Ensure reliability, scalability, and security of mission-critical cloud infrastructure and CI/CD environments.
  • Develop and implement automation solutions to streamline operational tasks and improve deployment efficiency.
  • Monitor system health and performance, proactively resolving incidents and contributing to continuous improvement.

Jobgether uses AI-powered matching to connect candidates with hiring companies. They are a platform that processes applications and shares top-fitting candidates with employers, operating remotely.

$100,000–$150,000/yr
US

  • Design, build, and operate scalable infrastructure platforms for large-scale AI model training and inference.
  • Manage and optimize GPU clusters, distributed training environments, and scheduling systems for machine learning applications.
  • Develop software solutions and automation tools using Python and systems programming languages like Go or C++.

Our partner builds and operates foundational technology powering advanced AI training and inference workloads at scale. They offer a collaborative culture focused on innovation, engineering excellence, and continuous learning.

US Unlimited PTO

  • Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
  • Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
  • Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.

NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.

$62,640–$104,760/yr
Europe 4w PTO

  • Design, build, and run distributed cloud architectures and large-scale production systems.
  • Ensure reliability, observability, performance, and cost efficiency of the platform.
  • Collaborate with product and backend teams to design system architecture and optimize resource use.

Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.

Canada

  • Design, implement, and maintain highly available and scalable infrastructure solutions.
  • Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
  • Automate infrastructure deployment, configuration management, and operational workflows.

The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.

US

  • Design, develop, and enhance backend systems for enterprise AI governance.
  • Build scalable services and contribute to architectural decisions.
  • Collaborate with product and engineering teams to deliver impactful features.

The company builds a governance platform that helps organizations manage and scale responsible AI initiatives. Its size is not specified, but it fosters a collaborative, innovative culture focused on learning, continuous improvement, and diversity.

UK

  • Design, build, and operate reliable infrastructure supporting AI-powered products.
  • Own and improve Kubernetes environments and cloud infrastructure.
  • Enhance production reliability through observability, automation, and incident response.

The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.

$141,000–$230,000/yr
US

  • Design, build, operate, and maintain large-scale distributed systems for product metrics.
  • Develop backend services using Golang within a cloud-native, Kubernetes-based environment.
  • Ensure high performance, reliability, and cost efficiency of critical data processing systems.

The company builds a large-scale distributed data platform for real-time analytics and customer-facing insights. It is a remote-first organization with a strong engineering culture focused on ownership, learning, and technical excellence.

US

  • Own the architecture, deployment, and operation of a mission-critical platform for secure, scalable delivery in regulated environments.
  • Design Infrastructure-as-Code foundations and manage containerized infrastructure using Docker, Kubernetes, and K3s.
  • Partner with cross-functional teams to ensure reliable, secure, and automated infrastructure solutions.

The company provides mission-critical platform solutions for secure, scalable delivery in regulated environments. It is a growing technology firm with a collaborative culture focused on automation and security.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

US

  • Lead and mentor a team of software engineers while setting technical architecture and engineering standards.
  • Design and deliver reference applications, cloud-native solutions, and reusable implementation patterns.
  • Collaborate with cross-functional teams to solve complex engineering challenges and improve developer practices.

US Security & IT delivers secure, scalable IT solutions focused on cloud-native and containerized systems. They offer a fully remote, collaborative environment with an emphasis on innovation and professional growth.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

$127,008–$152,410/yr
UK Sweden Spain Germany Ireland 6w PTO

  • Partner with product engineering squads to own production reliability for high-SLA customer environments, designing automation and defining per-tenant SLOs.
  • Serve as a primary escalation point for incidents, leading response, post-incident reviews, and reducing SLO burn to prevent repeats.
  • Influence feature design for scalability and operability, improve alert quality, and eliminate toil through automation.

Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform for organizations to see, understand, and act on their data. With over 35 million users, 7,000+ customers, and 1,600+ team members across 40+ countries, we foster a remote, collaborative culture rooted in open-source values.

$160,000–$208,000/yr
US

  • Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
  • Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
  • Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.

Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.

$20,258–$25,660/yr
India

  • Manage and maintain cloud infrastructure environments across development, staging, and production to ensure high availability and operational stability.
  • Administer Kubernetes clusters and containerized workloads, including deployment, scaling, and troubleshooting.
  • Develop and maintain automation tools and infrastructure-as-code frameworks to improve efficiency and scalability.

The company is a partner organization focused on building scalable, cloud-native infrastructure. The team is growing and operates in a remote-first, collaborative environment with a focus on continuous learning.

US

  • Design, deploy, and maintain cloud infrastructure across AWS, Azure, and/or Google Cloud environments.
  • Develop and manage CI/CD pipelines to automate application build, testing, and deployment processes.
  • Implement Infrastructure as Code solutions using tools such as Terraform, CloudFormation, or similar technologies.