Similar Jobs
See allStaff Site Reliability Engineer, Ads
US
Go
Kubernetes
Distributed Systems
Site Reliability Engineer
PulsePoint
US
Kubernetes
Terraform
Prometheus
Senior Site Reliability Specialist II
Everbridge
US
Cloud Infrastructure
Kubernetes
Observability
Principal Staff Software Engineer, Systems Infrastructure
Distributed Systems
Site Reliability Engineering
Observability
Staff Software Engineer, Observability
US
Distributed Systems
Kubernetes
Prometheus
Accountabilities:
- Lead reliability initiatives across multiple advertising technology domains, including ad serving, auctions, targeting, reporting, measurement, and billing.
- Partner with engineering leadership to define and execute a roadmap focused on reliability, scalability, operational excellence, and developer productivity.
- Design and build platforms, tooling, automation, and engineering solutions that improve system resilience and developer efficiency at scale.
Requirements:
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related discipline, with experience operating large-scale distributed systems.
- Strong experience evolving and supporting high-traffic, user-facing production environments.
- Deep expertise in distributed systems, scale engineering, cloud-native architectures, and highly available system design.
Benefits:
- Comprehensive health benefits, including medical, dental, and vision coverage.
- 401(k) program with employer matching.
- Equity compensation in the form of restricted stock units.
Undisclosed Company
The company is a partner organization operating a large-scale advertising technology ecosystem. Its size and culture are not detailed, but the role emphasizes reliability and operational excellence in a high-traffic environment.