Source Job

$135,000–$170,000/yr
US Unlimited PTO

  • Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
  • Build and own the observability layer for the fleet, surfacing data in shared dashboards and alerting.
  • Design and build automated recovery and self-healing for production systems.

Kubernetes Azure Observability Automation Incident Response

13 jobs similar to Senior Site Reliability Engineer (US)

Jobs ranked by similarity.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

Canada USA Unlimited PTO

  • Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
  • Build and maintain automation for lifecycle management of Kubernetes clusters.
  • Identify and root-fix reliability bottlenecks before they become incidents.

Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.

US

  • Design, build, and operate shared platform foundations including GCP, Kubernetes, networking, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems running at high request volume.
  • Raise the reliability bar through dashboards, alerting, on-call readiness, and automation.

Sanity.io builds an AI-powered content operating system that helps teams model, create, and automate content. The company has 200+ employees and a positive, flexible, trust-based culture that supports growth and work-life balance.

US

  • Design, build, and operate high-scale observability pipelines for logs, metrics, traces, and exceptions.
  • Lead cross-functional initiatives to resolve scaling bottlenecks and evolve production infrastructure safely.
  • Partner with engineering teams to improve observability tools and provide technical leadership across teams.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It emphasizes learning, professional growth, and an inclusive work environment.

$191,000–$226,000/yr
US Unlimited PTO

  • Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
  • Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
  • Build and maintain observability systems and automate operational toil using AI tools.

Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

Europe

  • Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
  • Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
  • Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.

Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.

Global

  • Contribute to platform and harness engineering, including CI/CD and developer tooling.
  • Build systems to reduce toil and maintain production infrastructure under conversational AI traffic.
  • Participate in on-call rotation and incident management to ensure platform uptime.

Replicant builds an AI-powered customer service platform that helps contact centers resolve requests and improve agent performance. The company is distributed, with a focus on ownership and collaboration, and serves Fortune 500 companies.

EMEA

  • Lead and develop a global team of SRE leaders, managers, and engineers, driving reliability strategy and operating model.
  • Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, and service health.
  • Drive cloud modernization initiatives, advancing containerization and Kubernetes-based operating models.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent are unstoppable together.

Global

  • Build distributed C#/.NET microservices on Azure and Kubernetes with event-driven architecture.
  • Own services end-to-end, from design to production, with observability and on-call feedback.
  • Embrace AI-native engineering, using tools like Claude for design, code review, and testing.

M-KOPA is a fintech company providing flexible financing to under-banked customers in Africa for smartphones, e-motorbikes, and clean energy. With 10M+ customers served, it's a mission-driven, remote-first company and an FT-recognized fastest-growing firm.

EU

  • Own and operate production infrastructure across Kubernetes, Linux, networking, and virtualization.
  • Lead incident response and implement observability to improve availability and performance.
  • Define SLOs and automate infrastructure with Ansible, Bash, Python, and GitOps.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective, data-driven processes. They foster a collaborative, international, and fully remote work environment, emphasizing autonomy and ownership for their small to mid-sized team.

$114,800–$150,000/yr
US 4w PTO

  • Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution alongside reliability work.
  • Partner with Software Engineering to investigate complex production issues, identify root causes and reliability risks, and drive permanent solutions.
  • Build observability across products, services, and critical customer workflows using meaningful metrics, logs, traces, dashboards, and actionable alerts.

Bloomerang provides a powerful giving platform and support for nonprofits to raise more, recruit more, and retain more. The company fosters a mission-driven culture built on core values of Simplify, Care and Act, and is home to innovative and skilled individuals.

US

  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, networking, storage, and workload schedulers.
  • Translate requirements from diverse customers into clear product direction and partner with engineering to define requirements.
  • Manage the observability backlog using feedback from production deployments and design partners to refine priorities.

Mirantis is a Kubernetes-native AI infrastructure company that enables organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI workloads. It is a distributed team committed to openness and technical excellence.