Source Job

US 4w PTO 16w maternity 4w paternity

  • Proactively identify, triage, and resolve performance issues across our Ruby on Rails stack and infrastructure.
  • Enhance system observability through monitoring metrics, SLOs, and SLIs across Ruby, Rails, and database systems.
  • Build and maintain AI agents and automations that reduce operational toil across incident response and routine maintenance.

Ruby On Rails AWS Kubernetes Datadog Terraform

20 jobs similar to Senior Site Reliability Engineer

Jobs ranked by similarity.

India

  • Design, deploy, and maintain the reliability, availability, and performance of critical systems and APIs across AWS and GCP.
  • Build observability frameworks, define SLIs/SLOs, and implement monitoring using Datadog and Kubernetes.
  • Participate in on-call rotations, incident response, and blameless post-incident reviews to drive systemic improvements.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce by consolidating identity, device, and access management. The company is remote-first with teams in over 15 countries and values building connections, thinking big, and continuous improvement.

US Unlimited PTO

  • Design, build, and maintain core services and infrastructure that power large-scale logistics solutions.
  • Identify and address performance bottlenecks, ensuring reliability and scalability across the platform.
  • Collaborate closely with Product, Data Science, and other engineering teams to deliver impactful features.

Roadie is a logistics and delivery platform that helps businesses tackle delivery complexities with same-day and last-mile solutions. With over 310,000 independent drivers nationwide and a remote-first culture, Roadie values innovation, collaboration, and flexibility.

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.

India

  • Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
  • Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
  • Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.

JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.

US 12w maternity 12w paternity

  • Drive the technical vision and architectural evolution of the Core Platform, including foundational components like the notification system, onboarding, and APIs.
  • Design systems that empower all engineering teams to build faster and more securely, acting as a technical thought leader for scaling challenges.
  • Mentor engineers, partner with management, and eliminate cross-team dependencies to ensure smooth delivery of complex features.

Huntress is a cybersecurity company founded in 2015 by former NSA operators, providing enterprise-grade security to businesses of all sizes. They are a remote-first team of over 500 employees, securing more than 5 million endpoints and 11 million identities worldwide.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

Brazil

  • Own the day-to-day operation of a monitoring platform, including dashboards and alerts.
  • Proactively analyze logs, traces, and metrics to detect failures and drive resolution.
  • Define and measure SLIs, improve reliability, and automate operational tasks.

CI&T helps large enterprises transform AI potential into real business impact with AI deployment and tech-integrated solutions. With 30 years of experience and 8,000 employees across 25 countries, we collaborate to build solutions with real impact.

Global 4w PTO

  • Champion SRE culture and best practices to improve production reliability and system resilience.
  • Communicate with stakeholders at all stages and bring fresh ideas to the table.
  • Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.

Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.

UK 4w PTO

  • Build and operate reliable, scalable cloud infrastructure on AWS and Kubernetes.
  • Own production infrastructure, containerized applications, deployment workflows, and monitoring.
  • Collaborate with development teams to streamline CI/CD and drive high availability.

Our partner is a fast-growing AdTech and e-commerce platform. They offer a flexible, remote-first culture that values ownership, proactive problem-solving, and continuous improvement.

US

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil and improve platform reliability.
  • Build observability and monitoring solutions with Grafana and AWS CloudWatch, and implement Infrastructure as Code using Terraform.
  • Develop CI/CD pipelines, define SRE standards and metrics, and participate in incident management and on-call rotation.

Peraton is a next-generation national security company that delivers mission-critical solutions and transformative IT services to government agencies and the U.S. armed forces. The company operates across land, sea, space, air, and cyberspace, with employees solving the most daunting challenges facing customers worldwide.

$150,000–$175,000/yr
United States

  • Design, build, and maintain automation and tooling to reduce operational toil.
  • Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
  • Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.

Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.

Brazil

  • Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
  • Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
  • Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.

$75,450–$169,700/yr
Global Unlimited PTO

  • Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
  • Balance 60% hands-on technical work with 40% people leadership and career growth.
  • Drive a maturing reliability practice including SLOs, incident response, and on-call.

Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.

Europe

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
  • Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.

ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.

US 12w maternity 12w paternity

  • Own and drive projects end-to-end, from design and coding to testing and rollout.
  • Work across the full stack to deliver secure, scalable solutions on the Core Platform.
  • Mentor teammates and maintain high standards for code quality through reviews and pairing.

Huntress is a remote-first cybersecurity company founded in 2015 by former NSA cyber operators, making enterprise-grade security accessible to businesses of all sizes. The company secures over 5 million endpoints and 15 million identities worldwide, with a culture focused on impactful work and customer commitment.

  • Own the technical direction and architecture of the releases platform, including the canonical release model.
  • Lead key technical decisions around security, reliability, scalability, and release governance.
  • Build and scale tools that support large-scale engineering teams and critical software delivery workflows.

This company builds a foundational release control platform used by leading engineering teams to ship software safely and efficiently. It fosters a collaborative, inclusive, and remote-first culture, with significant autonomy and opportunities for technical leadership.

$163,000–$185,000/yr
US

  • Own the technical roadmap for Platform Engineering, architecting scalable backend services and infrastructure.
  • Build and operate Kubernetes infrastructure, CI/CD pipelines, and developer tooling to boost engineering productivity.
  • Design and maintain enterprise integrations, billing platforms, and AI infrastructure including LLM orchestration and observability.

EasyLlama is an AI-powered Human Risk Management platform replacing outdated compliance training with bite-sized learning. Trusted by 6,000+ organizations, it's bootstrapped, profitable, and Inc. 5000-recognized with a fast-paced, ownership-driven culture.

$170,000–$235,000/yr
US

  • Design and implement backend services for licensing, entitlements, feature access, and usage limits across NodeZero's product and APIs.
  • Build and evolve provisioning, admin experience, MSP/MSSP capabilities, and audit logging for a multi-tenant SaaS platform.
  • Operate production services with monitoring, incident response, and a high bar for design quality and test coverage.

Horizon3 is a fast-growing, remote cybersecurity company that helps organizations proactively find, fix, and verify exploitable attack vectors through its NodeZero autonomous pentesting platform. The team is a fusion of former special operations cyber operators and startup engineers, fostering a culture of respect, collaboration, ownership, and results.

$165,000–$165,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
  • Build and manage CI/CD pipelines, automate operational tasks, and improve deployment processes.
  • Monitor production systems, participate in incident response, and champion DevOps best practices.

First Due provides transformative end-to-end software solutions for fire and EMS agencies, helping them run safer, smarter, and more effective operations. The company offers a fully remote workplace, comprehensive benefits, and opportunities for advancement, with a culture focused on respect, inclusivity, and equal opportunity.

$175,000–$185,000/yr
US

  • Consolidate Terraform and establish conventions for state management, modules, and CI checks.
  • Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
  • Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.

Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.