Source Job

$140,000–$170,000/yr
US

  • Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
  • Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
  • Serve as the on-account technical presence, partnering with customer infrastructure and security teams.

Kubernetes Terraform Python Observability Infrastructure As Code

20 jobs similar to Site Reliability Engineer

Jobs ranked by similarity.

$109,800–$252,500/yr
US Unlimited PTO 16w maternity 8w paternity

  • Own the full platform stack for Veeam Data Cloud in Government and Sovereign Cloud environments, including incident response, reliability, and observability.
  • Design and implement high-availability, fault-tolerant infrastructure on Azure (including Azure Government) with SLIs, SLOs, and error budgets.
  • Drive reliability improvements through automation, chaos engineering, and cross-team collaboration, with a focus on compliance and security.

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide.

$185,000–$280,000/yr
US 4w PTO

  • Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
  • Scale single-tenant deployments and build observability, incident response, and compliance practices.
  • Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.

Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.

$152,000–$195,000/yr
US Unlimited PTO

  • Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
  • Build and operate AI tooling infrastructure, including MCP servers and secure AI access.
  • Optimize CI/CD pipelines, implement progressive delivery, and advance Infrastructure as Code.

SecurityScorecard is the global leader in cybersecurity ratings, rating over 12 million companies across 64 countries. Headquartered in New York, it is recognized as a best workplace and funded by top investors.

Europe 6w PTO

  • Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
  • Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
  • Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.

Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.

US

  • Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
  • Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
  • Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.

The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.

US

  • Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
  • Use AI agents as force multipliers to automate manual processes and improve developer experience.
  • Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.

Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.

US

  • You will ensure the reliability and high availability of Tenable's cloud products in cloud environments.
  • You will respond to support escalations and troubleshoot complex technical problems.
  • You will develop software, tools, and scripts to automate deployment and monitoring of production systems.

We are the Exposure Management company, trusted by over 40,000 organizations to understand and reduce cyber risk. Our global team supports 65% of the Fortune 500 and 50% of the Global 2000, with a culture of belonging, respect, and excellence.

Europe Middle East Asia North America

  • Build and operate frameworks to ensure reliable, sustainable solution delivery across Mistral-hosted and customer-hosted environments.
  • Operate Tier-1 customer environments, ensure SLO compliance, manage on-call and incident response.
  • Productize deployment, security, and scaling of Applied AI solutions with automation and security guardrails.

Mistral provides full-stack AI solutions from frontier models to developer tools, applications, and compute, partnering with enterprises across high-stakes industries. It is a dynamic, collaborative team with a diverse workforce distributed globally, known for being creative, low-ego, and team-spirited.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

UK

  • Lead Cloud Platform and SRE teams, driving infrastructure strategy and ownership including Kubernetes, GCP, and Terraform.
  • Champion SRE culture, define SLOs, SLAs, and enhance observability and incident management.
  • Own the developer-facing platform as a product, ensuring self-service infrastructure and security compliance (SOC2, ISO-27001).

Prolific builds human data infrastructure for AI development, connecting researchers with a global participant pool. They foster a culture of high performance, ownership, and cross-functional collaboration.

$118,800–$237,600/yr
Europe

  • Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
  • Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
  • Manage distributed systems, observability, incident response, and automation with a security-first mindset.

Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.

Brazil

  • Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
  • Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
  • Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.

Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.

$120,000–$150,000/yr
US Unlimited PTO

  • Optimize new and existing systems by increasing reliability, performance, and scalability.
  • Automate routine operational tasks to reduce toil and improve efficiency.
  • Ensure infrastructure security compliance and implement least-privilege access controls.

Prove provides phone-centric identity tokenization and passive cryptographic authentication solutions to reduce friction and enhance security across digital channels. With over 1,000 enterprise customers processing 20 billion requests annually, they foster a fast-paced, collaborative culture focused on impact and tenacity.

$166,500–$291,400/yr
North America And Canada

  • Design and build cloud-native engineering platforms for software validation, release validation, and production readiness.
  • Develop automation solutions that improve engineering productivity and reduce manual toil through shift-left practices.
  • Foster a culture of reliability, automation, and operational excellence while mentoring engineers.

ServiceNow is the AI control tower for business reinvention, bringing together AI, data, and workflows to help organizations work smarter. They serve 85% of the Fortune 500 and foster an AI-native culture where technology and talent are unstoppable.

Global

  • Own and optimize CI/CD pipelines, Kubernetes deployment, and infrastructure for model serving and inference.
  • Build telemetry, observability, and alerting to catch real problems and reduce noise.
  • Eliminate toil through thoughtful automation and improve developer and agent productivity.

Obvious is building an AI-native workspace that serves as an operating system for work, putting co-intelligence at the center. They are a small, talent-dense team with founders and leaders from top tech companies.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

Costa Rica

  • Ensure reliability, performance, and scalability of Backcountry's multi-cloud platform.
  • Drive incident resolution, postmortems, and automation to reduce operational toil.
  • Leverage AI-assisted engineering tools and collaborate with teams to build and maintain observability and SLI/SLO instrumentation.

Backcountry is an online retailer of outdoor gear and apparel, rooted in adventure and the outdoor lifestyle. The company fosters a culture of recognition, wellbeing, and connection, with a lean, fast-paced engineering team.

UK

  • Design, build, and operate reliable infrastructure supporting AI-powered products.
  • Own and improve Kubernetes environments and cloud infrastructure.
  • Enhance production reliability through observability, automation, and incident response.

The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.

$100,000–$122,222/yr
Canada 8w PTO

  • End-to-end ownership of internal orchestration platform built on event-driven architecture with Redpanda, including code, architecture, and roadmap.
  • Own infrastructure-as-code using Terraform Cloud, manage Kubernetes workloads with Helm, and provide self-service tooling for engineering teams.
  • Set SLOs, handle production on-call, lead incident response, author design docs, and operate AI-natively using tools like Cursor and Notion AI.

Velora unifies Aplos, Raisely, and Keela into one company with a shared mission to help nonprofit organizations thrive by offering fundraising, donor management, financial tracking, and communications tools. We are a financially solid company with a combined team dedicated to making nonprofit work easier, more impactful, and more sustainable.