Source Job

US

  • Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
  • Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
  • Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.

Python Go Kubernetes Linux

20 jobs similar to Senior Site Reliability Engineer

Jobs ranked by similarity.

$165,000–$216,000/yr
US Unlimited PTO

  • Develop internal tools and automate infrastructure using AWS, Kubernetes, and programming languages.
  • Research and design solutions to increase website robustness, availability, and cost efficiency.
  • Collaborate on documentation, code reviews, and rollout of new processes.

Angi powers the future of the home services industry, connecting homeowners with skilled pros. With 9 brands in 8 countries and employees worldwide, Angi has helped homeowners with over 300 million home projects.

Poland

  • Design, write and deliver software to implement and support large web-scale, highly-performant, highly-available infrastructure on GCP/AWS.
  • Monitor infrastructure, respond to incidents, correct and improve systems to prevent incidents, and plan capacity.
  • Tune large-scale clusters for optimal performance and efficiency and support system deployments and product releases.

OpenX develops digital advertising marketplaces and technologies to optimize ad delivery for publishers and advertisers. The company operates a large-scale cloud infrastructure in Poland and values teamwork, customer centricity, and continuous learning.

Canada

  • Design, implement, and maintain highly available and scalable infrastructure solutions.
  • Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
  • Automate infrastructure deployment, configuration management, and operational workflows.

The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

US Unlimited PTO

  • Lead a global SRE team of ~10 engineers, owning day-to-day operations and long-term technical direction.
  • Drive strategic partnerships with product engineering to shift from reactive support to proactive reliability ownership.
  • Scale multi-tenant infrastructure, manage cloud costs, and champion developer self-service.

Counterpart Health is transforming healthcare by providing an AI-enabled primary care tool that supports physicians in early diagnosis and management of chronic conditions. As a subsidiary of Clover Health, it has a remote-first culture that emphasizes collaboration and innovation.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

US Unlimited PTO

  • Architect and improve cloud foundations on Google Cloud Platform to support scalable, secure, and well-governed workloads.
  • Design and build platform capabilities across GCP, Kubernetes, CI/CD, GitOps, and developer tooling.
  • Mentor engineers and raise the technical bar through code review, architecture guidance, and direct implementation.

Wpromote is a digital marketing agency focused on performance marketing and technology. The company fosters a diverse, inclusive culture with a remote-friendly environment and office hubs in Los Angeles, Chicago, and New York.

US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

US

  • Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
  • Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
  • Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.

Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.

$150,000–$175,000/yr
US

  • Design and manage GCP project structure, networking, and core infrastructure.
  • Own Kubernetes cluster infrastructure and define infrastructure-as-code standards.
  • Improve system architecture for scalability, resilience, and performance.

UJET provides an AI-powered contact center platform that redefines customer experience with cloud-native architecture and mobile-first approach. The company is a growing tech firm with a collaborative culture focused on innovation and security.

Global

  • Manage Kubernetes clusters and maintain infrastructure in the cloud.
  • Administer Linux servers and implement configuration management using Puppet or Ansible.
  • Troubleshoot and ensure observability of systems with CI/CD integration.

Xsolla is a global commerce company providing tools and services to help developers solve challenges in the video game industry. They employ over 1,500 developers and cultivate a supportive, collaborative culture focused on creativity and professional growth.

$120,000–$150,000/yr
US Unlimited PTO

  • Optimize new and existing systems by increasing reliability, performance, and scalability.
  • Automate routine operational tasks to reduce toil and improve efficiency.
  • Ensure infrastructure security compliance and implement least-privilege access controls.

Prove provides phone-centric identity tokenization and passive cryptographic authentication solutions to reduce friction and enhance security across digital channels. With over 1,000 enterprise customers processing 20 billion requests annually, they foster a fast-paced, collaborative culture focused on impact and tenacity.

Brazil

  • Act as a technical reference for SRE, DevOps, and cloud infrastructure, analyzing cloud environments in GCP/AWS for improvements and cost optimization.
  • Implement FinOps strategies, manage CI/CD pipelines, and maintain infrastructure as code using Terraform and Kubernetes.
  • Provide consultative support and communicate technical recommendations to engineering and business stakeholders.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The company uses technology to streamline the application process and promote fair evaluation.

Unlimited PTO

  • Own and evolve cloud infrastructure and CI/CD pipelines, driving projects from inception to deployment.
  • Gain deep knowledge of the backend stack (Java Spring Boot) and optimize system reliability and scalability.
  • Actively adopt AI tools to enhance productivity, automate infrastructure provisioning, and ensure security.

Archy is a Series B vertical SaaS company that provides AI-powered software to revolutionize dental practice management. The company has a growing, collaborative team with a remote-friendly culture.

$118,000–$151,000/yr
US 4w PTO

  • Build and improve platform services, including CI/CD pipelines and cloud infrastructure.
  • Collaborate with senior engineers to design scalable solutions and enhance developer experience.
  • Participate in incident response and retrospectives to drive continuous improvement.

Octopus Energy is a tech-powered energy company focused on renewable energy and customer experience. The company culture emphasizes ownership, collaboration, and making a tangible impact across teams.

India

  • Lead infrastructure strategy and cloud architecture for scalability and reliability.
  • Define enterprise CI/CD strategy and champion Infrastructure as Code (IaC) practices.
  • Mentor engineering teams and drive SRE principles for operational excellence.

Anaplan optimizes business decision-making through its AI-infused scenario planning and analysis platform. With over 2,400 global customers including Fortune 50 companies, it fosters a culture of innovation, diversity, and winning.

$170,000–$210,000/yr
US

  • Design and deliver solutions for cloud-hosted production infrastructure and automate CI/CD pipelines.
  • Build and support resilient, observable, and cost-optimized cloud infrastructure and security technologies.
  • Communicate effectively, share knowledge, participate in planning, and perform technology evaluation.

Ping Identity provides an intelligent cloud identity platform that enables secure and seamless digital experiences. The company serves more than half of the Fortune 100, has offices globally, and fosters a culture of respect and individuality.

US

  • Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
  • Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
  • Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.

Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.

Germany Unlimited PTO

  • Lead the technical direction of infrastructure security initiatives across cloud-native platforms and distributed systems.
  • Design and establish secure architectural patterns, reference implementations, and automation frameworks.
  • Conduct comprehensive security assessments, threat modeling, and risk analysis to drive effective remediation.

The partner company focuses on building secure, scalable cloud infrastructure and modern software platforms. They foster a collaborative engineering culture with a focus on security and innovation.

Canada Unlimited PTO

  • Develop and maintain a new CI/CD pipeline for multiple products and services using Golang and Kubernetes.
  • Collaborate with cross-functional teams to integrate systems and automate tasks in a cloud-native environment.
  • Participate in on-call rotation and provide product support to internal and external stakeholders.

Acquia is an open-source digital experience company that helps brands create customer moments that matter. Headquartered in Boston, they have been named one of North America's fastest growing software companies and a Best Place to Work.