Source Job

Europe

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
  • Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.

Go Python Kubernetes Terraform AWS

20 jobs similar to Site Reliability Engineer

Jobs ranked by similarity.

Global Unlimited PTO

  • Architect and build robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on public cloud and automate resource management.
  • Improve reliability, security, and cost efficiency of cloud services.

ClickHouse builds a real-time analytics database platform and manages ClickHouse Cloud data plane end-to-end with compute, networking, and security. As a rapidly scaling global startup, the company operates across 25+ countries and fosters a flexible, remote-friendly culture with equity and healthcare benefits.

$175,000–$185,000/yr
US

  • Consolidate Terraform and establish conventions for state management, modules, and CI checks.
  • Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
  • Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.

Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.

$191,000–$226,000/yr
US Unlimited PTO

  • Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
  • Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
  • Build and maintain observability systems and automate operational toil using AI tools.

Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.

Global Unlimited PTO

  • Enhance and scale a high-performance data onboarding platform to handle real-time data at petabyte scale.
  • Own and evolve infrastructure across multiple cloud providers, including AWS, GCP, and Azure.
  • Drive architectural decisions, improve deployment pipelines, and participate in on-call rotations.

ClickHouse is the company behind the open-source ClickHouse database, providing a high-performance database platform. It is a globally distributed start-up operating in over 25 countries, with a culture that values autonomy, innovation, and impact.

India

  • Build and maintain scalable cloud infrastructure for high availability.
  • Enhance observability and monitoring frameworks for accurate alerts.
  • Support on-call rotations and incident response with post-mortems.

GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.

US

  • Design, build, and operate shared platform foundations including GCP, Kubernetes, networking, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems running at high request volume.
  • Raise the reliability bar through dashboards, alerting, on-call readiness, and automation.

Sanity.io builds an AI-powered content operating system that helps teams model, create, and automate content. The company has 200+ employees and a positive, flexible, trust-based culture that supports growth and work-life balance.

Global Unlimited PTO

  • Build a cutting-edge cloud-native database platform on top of the public cloud.
  • Work on our in-house Kubernetes operator to support seamless infrastructure management.
  • Architect and build a robust, scalable, and highly available distributed infrastructure.

ClickHouse builds a cloud-native database platform and is transforming its open-source columnar database into a serverless solution. The company is a rapidly scaling global startup with a remote-friendly culture operating in over 25 countries.

$165,000–$165,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
  • Build and manage CI/CD pipelines, automate operational tasks, and improve deployment processes.
  • Monitor production systems, participate in incident response, and champion DevOps best practices.

First Due provides transformative end-to-end software solutions for fire and EMS agencies, helping them run safer, smarter, and more effective operations. The company offers a fully remote workplace, comprehensive benefits, and opportunities for advancement, with a culture focused on respect, inclusivity, and equal opportunity.

Brazil Unlimited PTO

  • Build and maintain the company's internal platform, driving operational excellence.
  • Collaborate with engineering squads to ensure applications are safe and reliable.
  • Take ownership of software infrastructure projects and provide off-hours support.

Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.

Global 4w PTO

  • Champion SRE culture and best practices to improve production reliability and system resilience.
  • Communicate with stakeholders at all stages and bring fresh ideas to the table.
  • Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.

Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.

US Unlimited PTO

  • Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers.
  • Provide technical leadership and mentorship across the Systems Engineering team.

Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.

$150,000–$185,000/yr
US Unlimited PTO

  • You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
  • You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
  • You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.

Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.

$75,450–$169,700/yr
Global Unlimited PTO

  • Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
  • Balance 60% hands-on technical work with 40% people leadership and career growth.
  • Drive a maturing reliability practice including SLOs, incident response, and on-call.

Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.

EU

  • Own and operate production infrastructure across Kubernetes, Linux, networking, and virtualization.
  • Lead incident response and implement observability to improve availability and performance.
  • Define SLOs and automate infrastructure with Ansible, Bash, Python, and GitOps.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective, data-driven processes. They foster a collaborative, international, and fully remote work environment, emphasizing autonomy and ownership for their small to mid-sized team.

Europe Unlimited PTO

  • Design and build scalable, reliable cloud infrastructure on GCP and AWS.
  • Manage Kubernetes environments and infrastructure as code with Terraform.
  • Drive CI/CD automation, platform reliability, and developer self-service.

The company builds and operates scalable cloud infrastructure and internal developer platforms. It is a globally distributed, fully remote engineering team with a collaborative and inclusive culture.

US Unlimited PTO

  • Define and drive the product roadmap for ClickHouse Cloud, focusing on cloud-native database capabilities.
  • Collaborate with engineering to prioritize and deliver improvements to the core database engine and storage layer.
  • Engage directly with customers and cross-functional teams, defining KPIs and guiding product decisions.

ClickHouse develops an open-source columnar database and offers ClickHouse Cloud, a cloud-native database service. They are a rapidly scaling startup with a globally distributed, remote-friendly team operating in over 25 countries.

UK 4w PTO

  • Build and operate reliable, scalable cloud infrastructure on AWS and Kubernetes.
  • Own production infrastructure, containerized applications, deployment workflows, and monitoring.
  • Collaborate with development teams to streamline CI/CD and drive high availability.

Our partner is a fast-growing AdTech and e-commerce platform. They offer a flexible, remote-first culture that values ownership, proactive problem-solving, and continuous improvement.

$170,000–$235,000/yr
US

  • Design and implement backend services for licensing, entitlements, feature access, and usage limits across NodeZero's product and APIs.
  • Build and evolve provisioning, admin experience, MSP/MSSP capabilities, and audit logging for a multi-tenant SaaS platform.
  • Operate production services with monitoring, incident response, and a high bar for design quality and test coverage.

Horizon3 is a fast-growing, remote cybersecurity company that helps organizations proactively find, fix, and verify exploitable attack vectors through its NodeZero autonomous pentesting platform. The team is a fusion of former special operations cyber operators and startup engineers, fostering a culture of respect, collaboration, ownership, and results.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

$200,700–$250,900/yr
US Canada

  • Embed with product teams to improve operational maturity through on-call, monitoring, and alerting practices.
  • Run game day exercises and implement reliability techniques in Haskell & TypeScript code.
  • Champion reliability practices through design reviews and advocate for SLOs tied to customer outcomes.

Mercury is a fintech company that provides banking services for startups. They are building a modern banking platform and value reliability and innovation.