Source Job

$171,000–$192,000/yr
US

  • Design, implement, and maintain scalable and reliable systems for space domain awareness.
  • Automate deployment, monitoring, and incident response to ensure high service reliability.
  • Collaborate with development teams to streamline deployment and enhance product reliability.

Python Go AWS Kubernetes Terraform

20 jobs similar to Senior Site Reliability Engineer (SRE)

Jobs ranked by similarity.

India

  • Design, deploy, and maintain the reliability, availability, and performance of critical systems and APIs across AWS and GCP.
  • Build observability frameworks, define SLIs/SLOs, and implement monitoring using Datadog and Kubernetes.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews to drive systemic improvements.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce by consolidating identity, device, and access management. The company is remote-first with teams in over 15 countries and values building connections, thinking big, and continuous improvement.

India

  • Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
  • Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
  • Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.

US

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil and improve platform reliability.
  • Build observability and monitoring solutions with Grafana and AWS CloudWatch, and implement Infrastructure as Code using Terraform.
  • Develop CI/CD pipelines, define SRE standards and metrics, and participate in incident management and on-call rotation.

Peraton is a next-generation national security company that delivers mission-critical solutions and transformative IT services to government agencies and the U.S. armed forces. The company operates across land, sea, space, air, and cyberspace, with employees solving the most daunting challenges facing customers worldwide.

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.

$150,000–$200,000/yr
Global Unlimited PTO

  • Design, implement, and manage cloud infrastructure and deployment pipelines.
  • Build automation tools for build, test, deployment, and infrastructure provisioning.
  • Mentor engineers and drive DevOps best practices for security, scalability, and reliability.

Accrete creates advanced AI solutions and autonomous agents that turn complex data into actionable insights for businesses and government organizations. The team is collaborative and innovative, focused on pushing the boundaries of AI technology.

$165,000–$165,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
  • Build and manage CI/CD pipelines, automate operational tasks, and improve deployment processes.
  • Monitor production systems, participate in incident response, and champion DevOps best practices.

First Due provides transformative end-to-end software solutions for fire and EMS agencies, helping them run safer, smarter, and more effective operations. The company offers a fully remote workplace, comprehensive benefits, and opportunities for advancement, with a culture focused on respect, inclusivity, and equal opportunity.

Unlimited PTO

  • Architect and scale AWS and Kubernetes infrastructure.
  • Build and maintain Terraform modules and drive capacity planning.
  • Own observability, CI/CD, and networking security.

Scribe is a Specialized Intelligence platform that automatically captures how work happens and turns it into a living asset for people and AI agents. Based in San Francisco, we're a LinkedIn Top Startup valued at over $1B with 7 million users across 600,000 businesses.

$124,750–$178,215/yr
US

  • Design, build, and maintain automation pipelines for workload packaging, validation, testing, deployment, and monitoring across AWS environments.
  • Collaborate with data engineers to operationalize workloads within a data lakehouse ecosystem and develop reusable infrastructure-as-code constructs using Terraform, AWS CDK, or CloudFormation.
  • Ensure pipelines meet enterprise security, compliance, and scalability standards while mentoring junior engineers and contributing to DevSecOps practices.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The company uses AI to review applications and ensure fair, objective candidate evaluation, and operates in a distributed, fully remote environment.

US 4w PTO 16w maternity 4w paternity

  • Proactively identify, triage, and resolve performance issues across our Ruby on Rails stack and infrastructure.
  • Enhance system observability through monitoring metrics, SLOs, and SLIs across Ruby, Rails, and database systems.
  • Build and maintain AI agents and automations that reduce operational toil across incident response and routine maintenance.

Fleetio is a modern software platform that helps thousands of organizations worldwide manage their fleet operations. We raised $450M in Series D funding in 2025 and maintain a remote-friendly, engineering-driven culture.

Europe

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
  • Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.

ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.

Brazil

  • Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
  • Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
  • Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.

UK 5w PTO

  • Design, implement, and maintain cloud infrastructure on AWS.
  • Develop and maintain CI/CD pipelines and deployment architectures.
  • Automate infrastructure provisioning and ensure security best practices.

Nearform is an independent team of data & AI experts, engineers, and designers who build intelligent digital solutions and capability at pace. Our team of 500 experts in 20+ countries is trusted by leading enterprises including Lululemon, Puma, Sun Life, Starbucks, and Walmart.

$175,000–$185,000/yr
US

  • Consolidate Terraform and establish conventions for state management, modules, and CI checks.
  • Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
  • Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.

Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.

Brazil

  • Own the day-to-day operation of a monitoring platform, including dashboards and alerts.
  • Proactively analyze logs, traces, and metrics to detect failures and drive resolution.
  • Define and measure SLIs, improve reliability, and automate operational tasks.

CI&T helps large enterprises transform AI potential into real business impact with AI deployment and tech-integrated solutions. With 30 years of experience and 8,000 employees across 25 countries, we collaborate to build solutions with real impact.

$168,675–$229,900/yr
US

  • Establish and employ continuous integration and delivery (CI/CD) patterns for successful software solutions.
  • Design secure, operationally sound solutions across AWS, Azure, OpenShift, and IBM Cloud.
  • Implement observability stacks, manage Kubernetes clusters, and automate infrastructure with Terraform.

Conga unifies commercial operations by aligning pricing, quoting, contracting, rebates, and communications so companies run as connected, smarter businesses. With more than 10,000 customers worldwide, including over 50% of the Fortune 100, Conga fosters a collaborative culture where every voice is heard.

UK 4w PTO

  • Build and operate reliable, scalable cloud infrastructure on AWS and Kubernetes.
  • Own production infrastructure, containerized applications, deployment workflows, and monitoring.
  • Collaborate with development teams to streamline CI/CD and drive high availability.

Our partner is a fast-growing AdTech and e-commerce platform. They offer a flexible, remote-first culture that values ownership, proactive problem-solving, and continuous improvement.

$170,000–$235,000/yr
US

  • Design and implement backend services for licensing, entitlements, feature access, and usage limits across NodeZero's product and APIs.
  • Build and evolve provisioning, admin experience, MSP/MSSP capabilities, and audit logging for a multi-tenant SaaS platform.
  • Operate production services with monitoring, incident response, and a high bar for design quality and test coverage.

Horizon3 is a fast-growing, remote cybersecurity company that helps organizations proactively find, fix, and verify exploitable attack vectors through its NodeZero autonomous pentesting platform. The team is a fusion of former special operations cyber operators and startup engineers, fostering a culture of respect, collaboration, ownership, and results.

US

  • Own ScaleOps' infrastructure end-to-end, including self-hosted product, SaaS platform, and AI infrastructure.
  • Manage cloud infrastructure across AWS, GCP, and Azure, covering networking, security, SSO, and compute.
  • Collaborate with customers and internal teams to ensure reliable feature delivery and eliminate operational toil.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps and platform engineers from manual resource management. We are the category leader backed by over $210M in funding, trusted by leading enterprises including Adobe, Coinbase, and Fortune 100 companies.

Europe Unlimited PTO

  • Design and build scalable, reliable cloud infrastructure on GCP and AWS.
  • Manage Kubernetes environments and infrastructure as code with Terraform.
  • Drive CI/CD automation, platform reliability, and developer self-service.

The company builds and operates scalable cloud infrastructure and internal developer platforms. It is a globally distributed, fully remote engineering team with a collaborative and inclusive culture.

Global 4w PTO

  • Champion SRE culture and best practices to improve production reliability and system resilience.
  • Communicate with stakeholders at all stages and bring fresh ideas to the table.
  • Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.

Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.