Source Job

US Unlimited PTO

  • Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
  • Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
  • Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.

Kubernetes Docker Datadog Python CI/CD

20 jobs similar to Site Reliability Engineering Manager

Jobs ranked by similarity.

$260,000–$280,000/yr
US

  • Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
  • Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
  • Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.

We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.

US

  • Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
  • Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
  • Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.

Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.

Canada

  • Design, implement, and maintain highly available and scalable infrastructure solutions.
  • Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
  • Automate infrastructure deployment, configuration management, and operational workflows.

The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

Global Unlimited PTO

  • Lead a high-impact infrastructure team, evolving internal platforms and CI/CD systems to support large-scale engineering operations.
  • Drive automation initiatives and AI-driven practices to reduce operational complexity and improve developer experience.
  • Define and execute strategies for scalable infrastructure, cloud environments, and platform engineering.

The partner company is a technology organization focused on building infrastructure platforms that enable engineering teams to deliver software faster. It is a remote-first company with a collaborative culture and a focus on innovation and scalability.

Latin America

  • Define and implement SLOs, SLIs, and Error Budgets to ensure production system reliability.
  • Lead incident command during major outages and drive blameless postmortems.
  • Develop observability strategies, including monitoring, logging, tracing, and alerting.

Oowlish is a rapidly expanding software development company in Latin America. It is certified as a Great Place to Work and offers a nurturing environment with professional development opportunities.

Portugal

  • Work with teams to define SLIs and SLOs, and create systems for observability.
  • Analyze failure scenarios, create runbooks, and reduce work that does not add value.
  • Participate in incident management and facilitate on-call duty to ensure reliable production environments.

Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.

$124,000–$170,500/yr
US

  • Perform operational deployments, implementations, and maintenance for production systems.
  • Implement and maintain monitoring, reporting, and alerting systems for Core Speech products.
  • Be part of an on-call rotation and work collaboratively to improve system performance and architecture.

Solventum is a new healthcare company with a long legacy of solving big challenges to improve lives and enable healthcare professionals to perform at their best. They are a large company that values empathy, insight, and clinical intelligence, collaborating with top minds in healthcare.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

$160,000–$208,000/yr
US

  • Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
  • Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
  • Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.

Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.

US Unlimited PTO

  • Own the US-only production environment end-to-end, including infrastructure deployment, maintenance, scaling, and reliability.
  • Lead and grow the US-based DevOps team, design scalable AWS infrastructure, and build CI/CD pipelines for safe, fast shipping.
  • Partner with engineering on application error investigations, improve monitoring and alerting, and coordinate with the Tel Aviv team on shared platform standards.

Zafran de-risks 90% of critical vulnerabilities overnight across hybrid environments using existing security tools. Backed by Sequoia Capital and Cyberstarts, it is one of the fastest-growing companies in cybersecurity, scaling to meet demand from advanced organizations.

United States

  • Ensure reliability, scalability, and security of mission-critical cloud infrastructure and CI/CD environments.
  • Develop and implement automation solutions to streamline operational tasks and improve deployment efficiency.
  • Monitor system health and performance, proactively resolving incidents and contributing to continuous improvement.

Jobgether uses AI-powered matching to connect candidates with hiring companies. They are a platform that processes applications and shares top-fitting candidates with employers, operating remotely.

Global

  • Lead the engineering team responsible for cloud infrastructure powering a global SaaS platform.
  • Guide design and evolution of Kubernetes-based, multi-cloud environments ensuring reliability and scalability.
  • Collaborate with cross-functional teams to enable safe software delivery and improve automation.

The company operates a mission-critical SaaS platform. It is a fully remote global organization with a focus on engineering excellence and responsible AI adoption.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

Global

  • Lead a high-leverage remote team of four infrastructure engineers, driving the evolution toward a scalable zero-toil platform.
  • Guide the team through an AI-driven engineering approach to reduce manual work and achieve zero-touch, scalable infrastructure.
  • Prepare and execute the strategy for CI/CD and artifact distribution systems to scale during a quality surge without increasing engineering toil.

Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex business processes. Trusted by over 700 organizations worldwide, including 9 of top 10 US banks, Camunda is a fully remote and global company with 150+ engineers across 20+ teams, and is transforming into an AI-first organization.

$160,000–$205,000/yr
US

  • Lead, mentor, and develop a platform engineering team to build scalable infrastructure and developer enablement solutions.
  • Combine engineering leadership, DevOps expertise, and strategic thinking to improve reliability and automation.
  • Partner with product and engineering leaders to define technical roadmaps and align platform capabilities with organizational goals.

The company is a mission-driven organization focused on innovation and sustainability. It fosters a collaborative culture of trust, accountability, learning, and professional growth.

US

  • Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
  • Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
  • Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.

Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.

Europe Middle East Asia North America

  • Build and operate frameworks to ensure reliable, sustainable solution delivery across Mistral-hosted and customer-hosted environments.
  • Operate Tier-1 customer environments, ensure SLO compliance, manage on-call and incident response.
  • Productize deployment, security, and scaling of Applied AI solutions with automation and security guardrails.

Mistral provides full-stack AI solutions from frontier models to developer tools, applications, and compute, partnering with enterprises across high-stakes industries. It is a dynamic, collaborative team with a diverse workforce distributed globally, known for being creative, low-ego, and team-spirited.

$127,008–$152,410/yr
UK Sweden Spain Germany Ireland 6w PTO

  • Partner with product engineering squads to own production reliability for high-SLA customer environments, designing automation and defining per-tenant SLOs.
  • Serve as a primary escalation point for incidents, leading response, post-incident reviews, and reducing SLO burn to prevent repeats.
  • Influence feature design for scalability and operability, improve alert quality, and eliminate toil through automation.

Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform for organizations to see, understand, and act on their data. With over 35 million users, 7,000+ customers, and 1,600+ team members across 40+ countries, we foster a remote, collaborative culture rooted in open-source values.

UK

  • Design, build, and operate reliable infrastructure supporting AI-powered products.
  • Own and improve Kubernetes environments and cloud infrastructure.
  • Enhance production reliability through observability, automation, and incident response.

The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.