Similar Jobs
See allPrincipal Site Reliability Engineer
Experian
Brazil
Kubernetes
AWS
Terraform
Sr. Site Reliability Engineer - SRE
Redzone
Europe
Kubernetes
AWS
Terraform
Site Reliability Engineering Manager
NationsBenefits
US
Kubernetes
Docker
Datadog
Director of Platform Engineering
First Due
US
CI/CD
SRE
Observability
Site Reliability Engineer 3
Jobgether
India
SRE
AIOps
AWS
Team Leadership:
- Lead, mentor, and grow a team of SRE/DevOps engineers.
- Work closely with engineering leadership to assess team needs and develop talent.
- Guide and support the team’s day-to-day work, balancing feature delivery, reliability work, security, and cost efficiency.
Incident Management:
- Oversee the incident management process end to end.
- Manage on-call rotations, escalation paths, incident command, postmortems, and follow-through on RCA and preventative actions.
- Ensure effective handoffs, coverage, and communication across IST and US time zones.
Reliability Practices:
- Define and drive SRE principles including SLIs, SLOs, error budgets, capacity planning, and observability standards.
- Champion a culture of reliability and operational excellence.
- Partner with engineering leadership to identify and implement process improvements across operational priorities.
Eltropy
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.