Similar Jobs
See allSenior Site Reliability Engineer - AI Experience Framework
Software Mind
Global
Kubernetes
Node.js
Prometheus
Site Reliability Engineer (SRE)
Software Mind
Global
Kubernetes
Splunk
CI/CD
Site Reliability Engineer II
NationsBenefits
US
Kubernetes
Docker
Datadog
Principal Site Reliability Engineer
Experian
Brazil
Kubernetes
AWS
Terraform
Staff Site Reliability Engineer
Unknown
US
Kubernetes
Python
Go
Project:
- Build and maintain the platform powering ServiceNow's AI-first user interfaces.
- Own production reliability for an SSR runtime and Java platform layer.
Position:
- Support deployment and operation of production services on Kubernetes.
- Monitor health, investigate incidents, and improve reliability.
- Collaborate with engineering teams on troubleshooting and CI/CD.
Expectations:
- 5+ years in SRE, DevOps, or Platform Engineering with strong Kubernetes experience.
- Production incident response, observability with Splunk, Prometheus, Grafana.
- Strong Linux and networking fundamentals, Node.js/JVM troubleshooting.
Additional Skills:
- Web Components/Lit experience for first-level debugging of UI issues.
- Canary rollout, distributed tracing, and event-driven autoscaling experience.
Software Mind
Software Mind develops innovative solutions for global companies, partnering with tech giants and unicorns on transformative projects. They foster cross-functional engineering teams with a culture of openness, respect, and passion, combining employment with enjoyment.