Source Job

$175,000–$185,000/yr
US

  • Consolidate Terraform and establish conventions for state management, modules, and CI checks.
  • Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
  • Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.

Terraform GCP Kubernetes Python Datadog

20 jobs similar to Site Reliability Engineer III

Jobs ranked by similarity.

Brazil 4w PTO

  • Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
  • Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
  • Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.

Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.

Canada

  • Design and evolve highly available cloud architecture on Google Cloud using Terraform and GitOps.
  • Build and maintain secure CI/CD pipelines for Infrastructure-as-Code and develop self-service developer platforms.
  • Strengthen platform observability and apply SRE principles to improve reliability and operational maturity.

They are a financial technology company that provides production-critical infrastructure. They have a globally distributed team and a culture of autonomy and async-first collaboration.

France

  • Evolve an Internal Developer Platform enabling development teams to deploy and operate applications securely with high self-service.
  • Play a crucial role in infrastructure architecture in a multi-cloud environment with a focus on GCP.
  • Build and maintain the platform used by over 50 internal clients, then support their concrete use by teams.

Lifen believes medical data can transform healthcare by reducing administrative burden, improving care coordination, and accelerating scientific discovery. Since 2015, the company has connected 800 hospitals and 150,000 healthcare professionals, with over 150 employees working remotely and from offices to unlock the potential of health data.

$191,000–$226,000/yr
US Unlimited PTO

  • Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
  • Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
  • Build and maintain observability systems and automate operational toil using AI tools.

Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.

India

  • Build and maintain scalable cloud infrastructure for high availability.
  • Enhance observability and monitoring frameworks for accurate alerts.
  • Support on-call rotations and incident response with post-mortems.

GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.

US Unlimited PTO

  • Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers.
  • Provide technical leadership and mentorship across the Systems Engineering team.

Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.

Brazil

  • Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
  • Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
  • Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

$185,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Coordinate with technical and non-technical staff across departments, including workflow automation that bridges infrastructure and business processes.
  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in GCP.
  • Maintain incident response process and tooling, and build automation that reduces toil and enables self-healing infrastructure.

Branch empowers workers with financial freedom by helping companies accelerate payments and providing accessible, free financial services. It is a remote-first, award-winning FinTech with employees across the U.S., fostering a culture of transparency, accountability, and trust.

UK

  • Lead Cloud Platform and SRE teams to scale securely and efficiently.
  • Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
  • Champion SRE culture with SLOs, error budgets, and observability.

Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.

$180,000–$220,000/yr
US

  • Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
  • Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
  • Implement observability solutions and enforce security best practices to ensure uptime and system performance.

Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.

$75,450–$169,700/yr
Global Unlimited PTO

  • Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
  • Balance 60% hands-on technical work with 40% people leadership and career growth.
  • Drive a maturing reliability practice including SLOs, incident response, and on-call.

Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.

UK Ireland Estonia Netherlands Sweden Israel Eastern Europe Portugal Unlimited PTO

  • Develop and evolve foundational software and services enabling product and development teams.\n- Architect, design, and implement Infrastructure as Code using Terraform.\n- Deploy, manage, and optimize Kubernetes clusters on GCP (GKE) and AWS (EKS).

DoiT is a global technology company that helps cloud-driven organizations leverage cloud for business growth and innovation through data, technology, and human expertise. They work with over 4,000 customers worldwide and foster a remote-first, entrepreneurial culture.

$138,000–$138,000/yr
US Unlimited PTO

  • Architect, automate, and manage cloud infrastructure using Terraform.
  • Oversee system reliability, performance, and monitoring with self-hosted Grafana, Sentry, and Laravel Nightwatch.
  • Drive security and compliance efforts, including SOC2 and PCI DSS audits, and manage AI/ML infrastructure.

Clair provides a digital banking platform that gives workers instant access to their earned wages, embedded within the apps they already use. As a fintech startup, Clair is mission-driven, with a culture that values innovation and employee well-being.

$150,000–$165,000/yr
US Unlimited PTO

  • Design and manage high-availability platforms using Kubernetes, Terraform, and Ansible with native-AI capabilities.
  • Develop and operate the observability stack: Grafana, Mimir, Loki, Tempo, and Prometheus on Kubernetes via GitLab CI/CD.
  • Build automation scripts in Python, maintain GitOps pipelines, and mentor mid-level engineers.

Flexential builds and operates critical IT platforms including observability, DevOps, and ITSM technologies. The company fosters a collaborative engineering culture and values diversity.

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.

US

  • Own core platform infrastructure including Terraform migration, CI/CD pipelines, and environment provisioning as part of a growing team.
  • Build and maintain tooling and abstractions that let product engineers ship and run code without solving infrastructure problems from scratch.
  • Advance observability foundation, standardize monitoring and alerting, and support GCP infrastructure scaling.

Astra builds mission-critical infrastructure for moving money at scale, processing billions in annual transaction volume with 99.9%+ uptime. We are a remote-first company hiring within the U.S., with a small team focused on thoughtful collaboration and clarity.

$143,200–$243,400/yr
North America

  • Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
  • Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
  • Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.

ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

Canada USA Unlimited PTO

  • Own the observability, logging and alerting for Kubernetes clusters and critical workloads.
  • Build and maintain automation for lifecycle management of Kubernetes clusters.
  • Identify and root-fix reliability bottlenecks before they become incidents.

Wrapbook is an AI platform for production finance, built for feature films and TV, trusted by Netflix and Paramount. Backed by top investors, our team of over 350 employees uses AI to transform how finance teams work.