Source Job

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Kubernetes Terraform Go Python AWS

20 jobs similar to Senior SRE Engineer

Jobs ranked by similarity.

US

  • Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
  • Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
  • Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.

Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

India

  • Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
  • Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
  • Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.

Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.

$75,600–$124,200/yr
Europe

  • Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
  • Drive automation and eliminate operational toil through self-service tooling and process improvements.
  • Lead incident response and mentor engineers to improve reliability practices.

Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.

Argentina

  • Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
  • Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
  • Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.

Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.

$180,000–$250,000/yr
US Unlimited PTO

  • Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
  • Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
  • Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.

We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

$140,000–$165,000/yr
Global Unlimited PTO

  • Design, build, and optimize cloud infrastructure (AWS/Kubernetes/EKS) and CI/CD pipelines across multiple teams.
  • Troubleshoot and resolve production incidents of varying scope, ensuring reliability and performance.
  • Drive infrastructure projects end-to-end, mentor engineers, and establish standards that improve developer productivity.

Pacvue is a leading Commerce Media OS powering over $12B in advertising spend across 100+ global retail media networks. It enables over 70,000 brands and agencies with an inclusive global community that fosters innovation and career growth.

Latin America

  • Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.

US Unlimited PTO

  • Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare.
  • Deliver SRE and DevOps initiatives to improve reliability, scalability, observability, and deployment safety.
  • Build reusable infrastructure modules, automation, and self-service workflows to reduce manual work and improve developer experience.

YipitData is the leading market research and analytics firm for the disruptive economy, recently raising up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points daily and have been recognized as one of Inc’s Best Workplaces, cultivating a people-centric culture focused on mastery, ownership, and transparency.

Global

  • Be on an on-call rotation responding to production incidents and support service engineers.
  • Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
  • Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.

Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.

$241,000–$270,000/yr
US Unlimited PTO

  • Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
  • Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
  • Build observability platforms to detect issues proactively and mentor engineers on reliability standards.

Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.

US

  • Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
  • Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
  • Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.

$120,000–$155,000/yr
Global

  • Own infrastructure as code across development, staging, and production environments
  • Build, maintain, and improve CI/CD pipelines for reliable and efficient deployments
  • Manage cloud infrastructure, establish scalable engineering practices, and lead incident response

CelebriOS is a software company building B2B SaaS products that help businesses make better decisions and streamline operations. The company has a remote-first working environment and a benefits package designed to support their team.

India

  • Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
  • Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
  • Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.