Source Job

$217,000–$303,900/yr
US 17w maternity 17w paternity

  • Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing.
  • Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
  • Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events.

Go Kubernetes Distributed Systems Observability Incident Management

20 jobs similar to Staff Site Reliability Engineer, Ads

Jobs ranked by similarity.

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

$217,000–$303,900/yr
US Unlimited PTO

  • Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute upstream changes to open source projects and share on-call responsibilities.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

$0–$150,000/yr
US EU UK

  • Help design, build, and operate the Kubernetes platform used across PulsePoint.
  • Own reliability, observability, and incident response across platform services.
  • Build infrastructure automation and GitOps workflows to reduce operational toil.

PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health signals using real-world data. With over 300 employees, the company is a post-acquisition profitable leader in the US healthcare ad market, known for a flat hierarchy and high engineering bar.

$164,200–$229,900/yr
United States

  • Design, write, and deliver software (primarily in Go and Python) to improve availability, scalability, latency, and efficiency of Reddit's products.
  • Dive deep into the codebase of Go services and the Python monolith legacy stack to implement complex system-level improvements.
  • Collaborate with cross-functional teams and share on-call responsibilities to ensure reliability and scalability of Tier-0 services.

Reddit is a community of communities, built on shared interests, passion, and trust. With over 100,000 active communities and approximately 130 million daily active users, it is one of the internet's largest sources of information.

$217,000–$303,000/yr
US

  • Lead the team that architects, builds, and operates the Notifications Platform at Reddit.
  • Set technical direction for shared platform capabilities spanning notification APIs, eventing, and delivery pipelines.
  • Collaborate with product and engineering teams to define platform contracts and improve developer experience.

Reddit is a community of communities where users submit, vote, and comment on topics they care about. With 100,000+ active communities and approximately 130 million daily active unique visitors, it is one of the internet's largest sources of information.

United States

  • Shape the future of Reddit by adapting platforms to evolving privacy, security, and regulatory landscapes.
  • Partner with product, design, and engineering teams to build trusted, compliant experiences for millions of users.
  • Own product areas end-to-end from technical design to launch, influencing technical and product strategy.

Reddit is a community of communities built on shared interests and authentic conversations. With over 100,000 active communities and approximately 130 million daily active users, Reddit is one of the largest sources of information on the internet.

$192,000–$192,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable and reliable systems.
  • Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
  • Develop and maintain automation tools for deployment, monitoring, and system health checks.

LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

US

  • Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
  • Drive platform security initiatives with preventative controls and resilient architecture.
  • Lead incident response and recovery, including root-cause analysis and preventative measures.

This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.

Spain

  • You will own and deliver quarterly goals for your team, leading engineers through ambiguity to solve open-ended problems.
  • You will proactively identify technical solutions and operational processes that strengthen incident readiness and response.
  • You will foster a culture of quality and ownership by setting or improving code review and design standards.

Affirm is reinventing credit to make it more honest and friendly, offering consumers the flexibility to buy now and pay later. The company has a strong engineering culture focused on reliability and ownership.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

$240,000–$240,000/yr
US

  • Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
  • Drive SLOs, observability, alerting, and on-call processes across teams.
  • Build the platform engineering function from the ground up and influence cross-cutting architecture.

First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.

US

  • Apply SRE principles to improve reliability, scalability, and performance of production systems.
  • Design and implement automation to reduce operational toil and improve engineering efficiency.
  • Lead incident response and develop sustainable solutions for complex production issues.

The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

$235,000–$275,000/yr

  • Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
  • Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.

Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

$241,000–$270,000/yr
US Unlimited PTO

  • Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
  • Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
  • Build observability platforms to detect issues proactively and mentor engineers on reliability standards.

Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.

US

  • Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
  • Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
  • Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.

WellSaid Labs is the leading AI voiceover studio for enterprise and professional use, providing ultra-realistic voices that the world’s biggest brands trust. We are a fully distributed team across the U.S. with a focus on responsible AI and an inclusive culture.

Spain

  • Define and implement reliability strategy including SLOs, SLIs, error budgets, and incident practices.
  • Manage cloud infrastructure on AWS using Infrastructure as Code and ensure Kubernetes scalability.
  • Lead incident response and establish chaos engineering practices to strengthen platform resilience.

This partner company builds a globally scaled, AI-native platform with a focus on reliability and event-driven systems. They offer a collaborative international culture with significant technical ownership and continuous improvement.

$74,000–$111,000/yr
Canada Unlimited PTO

  • Define and implement observability strategies, standards, and governance across applications and platforms.
  • Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
  • Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.

Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.