Source Job

$109,800–$252,500/yr
US Unlimited PTO 16w maternity 8w paternity

  • Own the full platform stack for Veeam Data Cloud in Government and Sovereign Cloud environments, including incident response, reliability, and observability.
  • Design and implement high-availability, fault-tolerant infrastructure on Azure (including Azure Government) with SLIs, SLOs, and error budgets.
  • Drive reliability improvements through automation, chaos engineering, and cross-team collaboration, with a focus on compliance and security.

Azure Kubernetes Terraform CI/CD Observability

20 jobs similar to Site Reliability Engineer

Jobs ranked by similarity.

US

  • You will ensure the reliability and high availability of Tenable's cloud products in cloud environments.
  • You will respond to support escalations and troubleshoot complex technical problems.
  • You will develop software, tools, and scripts to automate deployment and monitoring of production systems.

We are the Exposure Management company, trusted by over 40,000 organizations to understand and reduce cyber risk. Our global team supports 65% of the Fortune 500 and 50% of the Global 2000, with a culture of belonging, respect, and excellence.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

US

  • Design, implement, and evolve Azure-based infrastructure using Infrastructure as Code with Terraform, Bicep, or ARM templates.
  • Build and maintain scalable CI/CD pipelines using Azure DevOps Pipelines for application, infrastructure, and platform workloads.
  • Drive reliability engineering practice, including incident responses, root cause analysis, and automated remediation.

Sonatype is the software supply chain security company, providing end-to-end security solutions including proactive open source protection and SBOM management. Trusted by over 2,000 organizations, including 70% of the Fortune 100, we are pioneers in open-source security and DevSecOps with a culture that values diversity and innovation.

$180,000–$195,000/yr
US

  • Architect, build, and evolve secure, scalable cloud infrastructure on AWS and AWS GovCloud to power the Hypori SaaS platform.
  • Independently own ambiguous, high-impact infrastructure problems, guide technical direction, and act as a senior escalation point during production incidents.
  • Drive strategy and execution of Infrastructure as Code, observability, CI/CD, and operational frameworks while mentoring engineers and raising the technical bar across the organization.

Hypori Inc. is a high-growth cybersecurity SaaS company transforming secure mobility through a virtual workspace platform that enables users to access enterprise apps and data from any mobile device with zero data on the endpoint and total personal privacy. Backed by $55M in funding from investors including UBS, AE Industrial Partners, Hale Capital Partners, and GreatPoint Ventures, the company is expanding into new commercial and regulated markets.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

India

  • Design, implement, and manage scalable, highly available systems on Azure Cloud, optimizing Kubernetes and containerized workloads.
  • Build and maintain robust CI/CD pipelines with GitHub Actions and implement infrastructure as code using Helm Charts.
  • Monitor system performance, troubleshoot issues, ensure uptime, and perform root cause analysis to improve reliability.

Resilinc is pioneering intelligent, autonomous systems that redefine supply chain risk management using agentic AI, trusted by top companies in life sciences, aerospace, high tech, and automotive. We are a fully remote, mission-led team with a collaborative culture focused on high-impact work.

US

  • Lead a Dedicated Tenant Site Reliability Engineering organization, driving complex initiatives and operational excellence across multiple teams.
  • Oversee delivery and operation of PingOne Advanced Identity Cloud and Advanced Services, improving consistency and reliability.
  • Partner with SRE, Security, and Development teams to manage dependencies and evolve software delivery strategies.

Ping Identity provides an intelligent cloud identity platform that secures and streamlines digital experiences. Headquartered in Denver, Colorado, the company serves more than half of the Fortune 100 and fosters a culture that champions individuality and digital freedom.

$120,000–$150,000/yr
US Unlimited PTO

  • Optimize new and existing systems by increasing reliability, performance, and scalability.
  • Automate routine operational tasks to reduce toil and improve efficiency.
  • Ensure infrastructure security compliance and implement least-privilege access controls.

Prove provides phone-centric identity tokenization and passive cryptographic authentication solutions to reduce friction and enhance security across digital channels. With over 1,000 enterprise customers processing 20 billion requests annually, they foster a fast-paced, collaborative culture focused on impact and tenacity.

Global

  • Ensure availability, performance, scalability, and resilience of production services in AWS.
  • Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
  • Collaborate with development, architecture, security, and product teams to promote reliability best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.

$118,800–$237,600/yr
Europe

  • Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
  • Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
  • Manage distributed systems, observability, incident response, and automation with a security-first mindset.

Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

$166,500–$291,400/yr
North America And Canada

  • Design and build cloud-native engineering platforms for software validation, release validation, and production readiness.
  • Develop automation solutions that improve engineering productivity and reduce manual toil through shift-left practices.
  • Foster a culture of reliability, automation, and operational excellence while mentoring engineers.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help organizations work smarter. They serve 85% of the Fortune 500 and foster an AI-native culture where technology and talent are unstoppable.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.

India

  • Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
  • Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
  • Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

$100,000–$122,222/yr
Canada 8w PTO

  • End-to-end ownership of internal orchestration platform built on event-driven architecture with Redpanda, including code, architecture, and roadmap.
  • Own infrastructure-as-code using Terraform Cloud, manage Kubernetes workloads with Helm, and provide self-service tooling for engineering teams.
  • Set SLOs, handle production on-call, lead incident response, author design docs, and operate AI-natively using tools like Cursor and Notion AI.

Velora unifies Aplos, Raisely, and Keela into one company with a shared mission to help nonprofit organizations thrive by offering fundraising, donor management, financial tracking, and communications tools. We are a financially solid company with a combined team dedicated to making nonprofit work easier, more impactful, and more sustainable.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

Portugal

  • Work with teams to define SLIs and SLOs, and create systems for observability.
  • Analyze failure scenarios, create runbooks, and reduce work that does not add value.
  • Participate in incident management and facilitate on-call duty to ensure reliable production environments.

Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.