Similar Jobs
See allStaff Site Reliability Engineer
Filevine
Kubernetes
Python
Go
Senior Director, Site Reliability Engineering
Ping Identity
US
Site Reliability Engineering
Kubernetes
Infrastructure As Code
Staff Reliability Engineer
ServiceNow
North America And Canada
Kubernetes
Python
Go
Software Engineer III, Site Reliability
MyFitnessPal
US
Go
Python
AWS
Principal Site Reliability Engineer
Experian
Brazil
Kubernetes
AWS
Terraform
Key Responsibilities:
- Define and execute the technical strategy for observability, alerting, platform infrastructure, and operational excellence.
- Lead the design and evolution of scalable, secure, reliable, and cost-efficient cloud-native platforms.
- Drive adoption of AI and machine learning for observability, anomaly detection, and automated remediation.
What You'll Need:
- 12+ years of experience in software engineering, infrastructure engineering, or SRE.
- Deep expertise in Kubernetes, observability platforms like Datadog, and programming in Python, Go, or Bash.
- Proven experience in automation, Infrastructure-as-Code, and cloud-native architectures.
Why Join Us:
- Competitive compensation and comprehensive medical, dental, and vision insurance.
- Maternity and paternity leave programs, and disability coverage.
- Collaborative culture with focus on professional development and mentorship.
Unknown
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.