Source Job

Global

  • Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
  • Standardize and instrument pre-production deployment workflows for trustworthy signal.
  • Drive disaster-recovery readiness and reduce mean-time-to-detect and recover for deploy-related incidents.

SRE AWS Kubernetes Observability Scripting

20 jobs similar to Release Engineer (SRE)

Jobs ranked by similarity.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

Canada Europe

  • Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
  • Partnering with development teams to establish production readiness and operational readiness.
  • Building tooling to automate observability and operational workflows, eliminating manual toil.

Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.

Global

  • Ensure availability, performance, scalability, and resilience of production services in AWS.
  • Automate infrastructure provisioning and management using Infrastructure as Code (IaC).
  • Collaborate with development, architecture, security, and product teams to promote reliability best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide, operating in markets such as financial services, healthcare, automotive, and insurance. The company has over 25,200 employees across 32 countries and is recognized as a Top 25 global workplace by Fortune.

$62,640–$104,760/yr
Europe 4w PTO

  • Design, build, and run distributed cloud architectures and large-scale production systems.
  • Ensure reliability, observability, performance, and cost efficiency of the platform.
  • Collaborate with product and backend teams to design system architecture and optimize resource use.

Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.

$127,008–$152,410/yr
UK Sweden Spain Germany Ireland 6w PTO

  • Partner with product engineering squads to own production reliability for high-SLA customer environments, designing automation and defining per-tenant SLOs.
  • Serve as a primary escalation point for incidents, leading response, post-incident reviews, and reducing SLO burn to prevent repeats.
  • Influence feature design for scalability and operability, improve alert quality, and eliminate toil through automation.

Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform for organizations to see, understand, and act on their data. With over 35 million users, 7,000+ customers, and 1,600+ team members across 40+ countries, we foster a remote, collaborative culture rooted in open-source values.

US

  • Design and implement monitoring and alerting systems using tools like Prometheus, Grafana, and DataDog to ensure high availability and reliability.
  • Optimize performance and reliability of healthcare payment applications, lead incident response, and develop SLOs/SLIs.
  • Automate CI/CD pipelines, infrastructure provisioning with Terraform, and manage cloud infrastructure on AWS with Kubernetes.

LMI is a digital solutions provider accelerating government impact with innovation and speed, bringing commercial-grade platforms and mission-ready AI to federal agencies. Headquartered in Tysons, Virginia, LMI serves the defense, space, healthcare, and energy sectors, focusing on agility and collaboration to drive impactful results.

Global

  • Embed with product and platform teams from early stages to ensure reliability is designed in from the start.
  • Define production-readiness standards and measurable SLIs/SLOs to guide operational excellence.
  • Build tooling and infrastructure across AWS, GCP, and Azure using Terraform, and share on-call rotation.

We build WebContainers and Bolt.new, an AI-powered app builder that lets you create, edit, and deploy full-stack apps instantly in your browser. We are a fully remote, globally distributed team of passionate engineers serving over 1 million developers monthly.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.

Latin America

  • Define and implement SLOs, SLIs, and Error Budgets to ensure production system reliability.
  • Lead incident command during major outages and drive blameless postmortems.
  • Develop observability strategies, including monitoring, logging, tracing, and alerting.

Oowlish is a rapidly expanding software development company in Latin America. It is certified as a Great Place to Work and offers a nurturing environment with professional development opportunities.

US

  • Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
  • Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
  • Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.

Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.

SRE Engineer

IPSY
Mexico Colombia

  • Build and maintain observability across the platform in Datadog, including dashboards, monitors, APM, and log pipelines.
  • Participate in on-call rotation and incident response, driving blameless post-incident reviews and automating toil.
  • Leverage AI tools to accelerate debugging, generate runbooks, and build automation for operational efficiency.

IPSY is a beauty subscription platform that connects brands and consumers through curated beauty products. It is a remote-first company with a focus on community and engagement.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

$154,700–$208,000/yr
US Canada

  • Design and build automated reliability and self-healing systems to protect production at scale.
  • Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
  • Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.

Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.

Global 4w PTO

  • Own and improve service reliability for the product team: design for HA/performance/scale, define SLIs/SLOs
  • Align with org standards, implement DevOps-driven updates: support processes, templates, services, breaking changes, security fixes
  • Build and evolve GitLab CI/CD for build, test, security scans, and progressive delivery; speed up and harden pipelines

Plata Card is a fintech company focused on cards and accounts services. They foster a high-tech environment with a supportive team and innovative spirit.

$164,000–$218,000/yr
US Unlimited PTO

  • Lead design and evolution of secure cloud infrastructure and deployment systems for critical decentralized applications.
  • Drive improvements across CI/CD pipelines, deployment workflows, and engineering productivity practices.
  • Collaborate with developers, security specialists, product leaders, and infrastructure teams in a remote-first environment.

Our partner is building and scaling secure, high-performance infrastructure powering one of the most widely used decentralized technology platforms in the world. They operate as a fully remote, globally distributed team with a focus on DevOps, security, and blockchain technology.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

$150,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Design the human side of reliability: on-call rotations, incident roles, blameless postmortems, and change management.
  • Build and operate the core reliability toolkit: observability, CI/CD, infrastructure as code, and incident response.
  • Embed with product engineers to raise operational literacy and define SLOs grounded in patient needs.

Novellia is a health tech startup that provides patients with access to their health records and enables biopharma innovation through real-world data. They are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors, with a focus on building a diverse and inclusive workplace.

UK

  • Collaborate with engineering teams to design scalable, secure systems.
  • Establish SLOs, manage incident response, and drive reliability improvements.
  • Leverage expertise in Go, Python, Kubernetes, and cloud platforms.

ClickHouse is a leading real-time analytics company recognized on the 2025 Forbes Cloud 100 list. With over 3,000 customers and rapid growth, the company offers a remote-friendly, globally distributed culture.

$260,000–$280,000/yr
US

  • Establish an SRE function, building a team of 4-6 and professionalizing incident management with clear processes and tooling.
  • Drive engineering excellence through design reviews, code reviews, and blameless retrospectives to foster a quality culture.
  • Lead people and projects, balancing incident response with a roadmap of observability and reliability engineering initiatives.

We are a fast-growing, remote cybersecurity company dedicated to enabling organizations to proactively find and fix exploitable attack vectors. Our team is a fusion of former special operations cyber operators and engineers, committed to a culture of respect, collaboration, and ownership.