Similar Jobs
See allSenior Software Engineer II, DevEx, OPX
Samsara
US
Go
Python
Datadog
Software Engineer III, Site Reliability
MyFitnessPal
US
Go
Python
AWS
Sr/Staff Site Reliability Engineer, Consumer Apps
Attain
US
AWS
GCP
Kubernetes
Principal Site Reliability Engineer
Experian
Brazil
Kubernetes
AWS
Terraform
Sr. Site Reliability Engineer - SRE
Redzone
Europe
Kubernetes
AWS
Terraform
Role Overview:
- Focus on modernizing reliability engineering through automation, observability, and AI-powered operations.
- Build resilient cloud platforms, improve system performance, and reduce operational complexity.
- Work at the intersection of SRE practices, AIOps, automation, and AI to improve incident response.
Accountabilities:
- Provide end-to-end reliability ownership for production systems, including on-call support and incident response.
- Build and maintain observability solutions across metrics, logs, and traces with intelligent monitoring.
- Design AIOps workflows for event ingestion, correlation, and alert noise reduction.
Requirements:
- 6+ years in SRE, AIOps, DevOps, or production engineering within large-scale cloud environments.
- Expert-level experience with ELK/OpenSearch and observability practices.
- Hands-on experience with AIOps solutions and Infrastructure as Code tools like Terraform.
Jobgether
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.