Source Job

US

  • Provide technical leadership for reliability across a large-scale advertising technology ecosystem
  • Lead reliability initiatives across ad serving, auctions, targeting, reporting, and billing systems
  • Mentor engineers and influence technical decisions to improve system resilience and developer productivity

Go Kubernetes Distributed Systems Observability Incident Management

20 jobs similar to Staff Site Reliability Engineer, Ads

Jobs ranked by similarity.

$217,000–$303,900/yr
US 17w maternity 17w paternity

  • Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing.
  • Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
  • Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet's largest sources of information.

$0–$150,000/yr
US EU UK

  • Help design, build, and operate the Kubernetes platform used across PulsePoint.
  • Own reliability, observability, and incident response across platform services.
  • Build infrastructure automation and GitOps workflows to reduce operational toil.

PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health signals using real-world data. With over 300 employees, the company is a post-acquisition profitable leader in the US healthcare ad market, known for a flat hierarchy and high engineering bar.

$145,000–$177,000/yr
US

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, and networking.
  • Design and implement solutions that improve availability, scalability, performance, and resilience of the platform.

Everbridge empowers enterprises and government organizations to anticipate, mitigate, respond to, and recover from critical events. The company focuses on building resilient systems and fostering a culture of ownership, continuous improvement, and operational excellence.

  • Drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering.
  • Lead adoption of service criticality models to set reliability expectations based on business impact.
  • Partner with engineering teams to improve system design, reduce incident risk, and strengthen operational readiness.

LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. We foster a culture of trust, care, inclusion, and fun, investing in employee growth to transform the way the world works.

$217,000–$303,900/yr
US Unlimited PTO

  • Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute upstream changes to open source projects and share on-call responsibilities.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.

$240,000–$240,000/yr
US

  • Lead centralization of DevOps, SRE, database reliability, incident management, and developer experience practices.
  • Drive SLOs, observability, alerting, and on-call processes across teams.
  • Build the platform engineering function from the ground up and influence cross-cutting architecture.

First Due provides fire and EMS agencies with transformative, end-to-end software solutions to improve safety and effectiveness. The company offers a fully remote workplace with a comprehensive benefits package and opportunities for advancement.

$217,000–$303,000/yr
US

  • Lead the team that architects, builds, and operates the Notifications Platform at Reddit.
  • Set technical direction for shared platform capabilities spanning notification APIs, eventing, and delivery pipelines.
  • Collaborate with product and engineering teams to define platform contracts and improve developer experience.

Reddit is a community of communities where users submit, vote, and comment on topics they care about. With 100,000+ active communities and approximately 130 million daily active unique visitors, it is one of the internet's largest sources of information.

$151,000–$206,000/yr
US Canada Unlimited PTO

  • Build large-scale real-time services and applications leveraging massive datasets.
  • Develop and maintain data pipelines, messaging systems, databases, and cloud services.
  • Work with Machine Learning Engineers and Security Researchers on security solutions.

Censys provides real-time Internet intelligence and threat insights to global governments and Fortune 500 companies. It is a growing company with a focus on comprehensive internet mapping and security solutions.

$200,000–$350,000/yr
US

  • Define architecture for complex, high-impact systems.
  • Lead company-critical engineering initiatives.
  • Mentor senior and staff-level engineers.

They build sophisticated infrastructure and software systems for critical business operations. They are a rapidly scaling company with a collaborative, fast-moving culture where ownership is encouraged and decisions are made quickly.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

US

  • Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
  • Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments.

Tech Holding is a full-service consulting firm that delivers predictable outcomes and high-quality solutions to clients. The company was founded by experienced industry professionals who have held senior positions at startups to Fortune 50 firms, fostering a culture of deep expertise, integrity, transparency, and dependability.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

$74,000–$111,000/yr
Canada Unlimited PTO

  • Define and implement observability strategies, standards, and governance across applications and platforms.
  • Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
  • Establish and drive SRE best practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.

Valtech is the experience innovation company that helps brands unlock new value in an increasingly digital world by blending crafts, categories, and cultures. They have a workplace culture that fosters creativity, diversity, and autonomy, with a borderless global framework enabling seamless collaboration.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

$163,000–$263,670/yr
US

  • Design, build, and operate the distributed systems that deliver feature-flag configuration to LaunchDarkly SDKs.
  • Own and improve the reliability, latency, and scalability of our streaming and polling infrastructure.
  • Debug and resolve complex production issues while sharing the team's on-call rotation.

LaunchDarkly builds a feature management platform that enables safe and gradual software releases. The company is growing and fosters a humble, open, collaborative culture.

US

  • Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
  • Drive platform security initiatives with preventative controls and resilient architecture.
  • Lead incident response and recovery, including root-cause analysis and preventative measures.

This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.

$241,000–$270,000/yr
US Unlimited PTO

  • Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
  • Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
  • Build observability platforms to detect issues proactively and mentor engineers on reliability standards.

Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.

US Unlimited PTO

  • Lead architecture and technical direction for AI-powered Loyalty product within Zeta Marketing Platform.
  • Own roadmap delivery and drive engineering best practices across multiple teams.
  • Coach and develop engineers while ensuring product reliability and scalability.

Zeta Global is an AI-powered marketing cloud that leverages consumer signals to help marketers acquire, grow, and retain customers. Founded in 2007 and headquartered in New York City, the company operates globally and fosters a culture of trust, belonging, and inclusion.

$200,000–$240,000/yr
US Unlimited PTO 16w maternity 16w paternity

  • You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.

Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.

$160,000–$200,000/yr
US Unlimited PTO

  • Deliver production-ready Temporal implementations by co-building workflows with customer and partner engineering teams.
  • Identify and remove early activation blockers and establish operational standards.
  • Define observability, reliability, and deployment strategies for production systems.

Temporal is an open source programming model that simplifies code, makes applications more reliable, and helps developers focus on delivering features faster. The company is mission-driven, building a team that values curiosity, drive, collaboration, genuineness, and humility.