Similar Jobs
See allLead Site Reliability Engineer (Performance & Scalability)
Tech Holding
US
Site Reliability Engineering
Distributed Systems
Cloud Infrastructure
Site Reliability Engineer
Veeam
US
Azure
Kubernetes
Terraform
Principal Staff Software Engineer, Systems Infrastructure
Distributed Systems
Site Reliability Engineering
Observability
Senior Site Reliability Specialist II
Everbridge
US
Cloud Infrastructure
Kubernetes
Observability
Sr Mgr, Reliability Engineering
ServiceNow
EMEA
Site Reliability Engineering
Kubernetes
AWS
About the Role:
- This is a senior engineering opportunity focused on building and maintaining reliable infrastructure for government-focused technology environments.
- You will apply deep expertise in site reliability engineering to improve availability, resilience, observability, and operational performance.
- As a Staff-level engineer, you will contribute beyond individual systems by shaping architecture, engineering standards, and operational strategy.
Key Responsibilities:
- Lead the design, implementation, and continuous improvement of highly reliable, scalable, and resilient systems.
- Establish and influence SRE practices, standards, and architectural approaches across engineering teams.
- Identify and address systemic reliability, availability, scalability, performance, and operational risks.
What We’re Looking For:
- Extensive professional experience in SRE, infrastructure engineering, platform engineering, or DevOps.
- Staff-level technical leadership with ability to influence architecture and technical strategy across multiple teams.
- Strong understanding of distributed systems, cloud infrastructure, production operations, and automation.
Not specified
This company is a research and development organization focused on building and maintaining reliable infrastructure for government-focused technology environments. It operates with a collaborative, remote-first culture and values technical innovation, scalability, and operational excellence.