Source Job

Americas

  • You'll operate production day-to-day, including oncall, incident response, and postmortems.
  • You'll own reliability practice by defining SLIs/SLOs and error budgets.
  • You'll ship infrastructure through code in a GitOps workflow for cloud and Kubernetes.

PostgreSQL Kubernetes Go Python GitOps

20 jobs similar to Senior Site Reliability Engineer

Jobs ranked by similarity.

US 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
  • Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.

Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.

US Unlimited PTO

  • Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
  • Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
  • Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.

$120,000–$165,000/yr
US Unlimited PTO

  • Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
  • Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
  • Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.

MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.

$75,600–$124,200/yr
Europe

  • Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
  • Drive automation and eliminate operational toil through self-service tooling and process improvements.
  • Lead incident response and mentor engineers to improve reliability practices.

Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.

UK

  • Collaborate with engineering teams to design scalable, secure systems.
  • Establish SLOs, manage incident response, and drive reliability improvements.
  • Leverage expertise in Go, Python, Kubernetes, and cloud platforms.

ClickHouse is a leading real-time analytics company recognized on the 2025 Forbes Cloud 100 list. With over 3,000 customers and rapid growth, the company offers a remote-friendly, globally distributed culture.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

Portugal

  • Work with teams to define SLIs and SLOs, and create systems for observability.
  • Analyze failure scenarios, create runbooks, and reduce work that does not add value.
  • Participate in incident management and facilitate on-call duty to ensure reliable production environments.

Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.

Global

  • Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
  • Standardize and instrument pre-production deployment workflows for trustworthy signal.
  • Drive disaster-recovery readiness and reduce mean-time-to-detect and recover for deploy-related incidents.

Supabase is the Postgres development platform providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. They are a globally distributed team of ~400 members across 60+ countries, with over $1B raised and 540,000+ community members.

Canada

  • Design, implement, and maintain highly available and scalable infrastructure solutions.
  • Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
  • Automate infrastructure deployment, configuration management, and operational workflows.

The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.

$150,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Design the human side of reliability: on-call rotations, incident roles, blameless postmortems, and change management.
  • Build and operate the core reliability toolkit: observability, CI/CD, infrastructure as code, and incident response.
  • Embed with product engineers to raise operational literacy and define SLOs grounded in patient needs.

Novellia is a health tech startup that provides patients with access to their health records and enables biopharma innovation through real-world data. They are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors, with a focus on building a diverse and inclusive workplace.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

$140,000–$170,000/yr
US

  • Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
  • Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
  • Serve as the on-account technical presence, partnering with customer infrastructure and security teams.

We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.

6w PTO 18w maternity 4w paternity

  • Become owner/main contributor to PostgreSQL stack
  • Maintain and upgrade DB infrastructure using GCP, AWS, and IaC
  • Manage product infrastructure with Kubernetes and ensure 99.9% uptime SLA

Wrike is a powerful work management platform. With over 1,000 employees across 10 global hubs, it fosters a culture of innovation and collaboration.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.

$62,640–$104,760/yr
Europe 4w PTO

  • Design, build, and run distributed cloud architectures and large-scale production systems.
  • Ensure reliability, observability, performance, and cost efficiency of the platform.
  • Collaborate with product and backend teams to design system architecture and optimize resource use.

Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.

North America Latin America

  • Design, build, and evolve the core data platform infrastructure including distributed query engines, orchestration, and warehousing.
  • Own the lakehouse infrastructure as code, managing deployments through Terraform and Ansible on Kubernetes.
  • Build and maintain low-latency streaming and CDC ingestion pipelines, as well as batch ingestion paths landing in Iceberg.

Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, and more. The company has a global team of 400+ members, is backed by $400M in funding, and fosters a culture of curiosity, empathy, and accountability.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

$192,000–$192,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable and reliable systems.
  • Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
  • Develop and maintain automation tools for deployment, monitoring, and system health checks.

LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.