Similar Jobs

See all

Responsibilities:

  • Be on an on-call rotation responding to production availability incidents.
  • Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes.
  • Design, build, and maintain core infrastructure that scales to hundreds of thousands of concurrent users.

Qualifications:

  • Think cloud-first, security-first, and about systems edge cases.
  • Know your way around Linux and Windows and config-management systems.
  • Have strong programming skills in Python, Java, Golang, or Node.js.

Projects you could work on:

  • Coding infrastructure automation with Ansible and Terraform.
  • Improving Prometheus monitoring or building new metrics.
  • Planning and executing migration from AWS VMs to Kubernetes on EKS.

Our Client

Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.

Apply for This Position