Source Job

$190,000–$230,000/yr
US

  • Design, implement, and manage scalable cloud infrastructure using Kubernetes and Pub/Sub.
  • Refactor systems for scalability, including transforming stateful components to stateless ones.
  • Drive platform reliability initiatives like alerting, health checking, and incident management.

Python Ruby Kubernetes GCP Terraform

20 jobs similar to Staff SWE, Distributed Systems

Jobs ranked by similarity.

$151,000–$206,000/yr
US Canada Unlimited PTO

  • Build large-scale real-time services and applications leveraging massive datasets.
  • Develop and maintain data pipelines, messaging systems, databases, and cloud services.
  • Work with Machine Learning Engineers and Security Researchers on security solutions.

Censys provides real-time Internet intelligence and threat insights to global governments and Fortune 500 companies. It is a growing company with a focus on comprehensive internet mapping and security solutions.

$126,290–$190,000/yr
United States 18w maternity 12w paternity

  • Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
  • Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
  • Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.

Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.

Global

  • Own and scale cloud infrastructure including compute, networking, storage, and data systems.
  • Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
  • Establish reliability through service-level objectives, observability, and incident response processes.

A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.

US

  • Design and implement production-quality infrastructure and internal platforms with end-to-end ownership.
  • Build and expand automated testing infrastructure for software deployed across diverse edge hardware.
  • Improve CI/CD pipelines, build and release systems, and overall platform reliability.

The company is a technology organization that builds infrastructure for reliable software delivery across complex edge environments. It is a small, fully remote engineering team with a culture of ownership, technical judgment, and fast execution.

UK

  • Architect and build a robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on top of the public cloud and automate cloud resource management.
  • Work closely with core database development and security teams to produce the SaaS offering.

ClickHouse is a real-time analytics and data warehousing company recognized on the Forbes Cloud 100 list. With over 4,000 customers and rapid growth, the company is a leader in its space.

Europe

  • Manage and troubleshoot complex distributed large-scale software systems
  • Build scalable, secure and reliable container-based infrastructure
  • Automate software delivery processes with CI/CD pipelines

Coinspaid Dev is the engineering brand behind the technology, infrastructure, and R&D expertise built within Coinspaid, focusing on advancing blockchain infrastructure engineering. With over 120 engineers and more than 11 years of industry experience, they bring together teams building distributed systems and blockchain infrastructure across 20+ blockchain networks.

$186,700–$255,000/yr
US Canada 18w maternity 12w paternity

  • Create and test reliable cloud infrastructure services supporting Webflow's product range.
  • Lead initiatives to reduce triage load, increase reliability, and handle growing customer scale.
  • Collaborate with product engineering teams to deliver new solutions and improve existing services.

Webflow is an agentic web marketing platform that helps modern marketing teams build, manage, and optimize high-performing web experiences. The company values grit, speed, and craft, fostering a culture of ownership and continuous improvement.

$0–$150,000/yr
US EU UK

  • Help design, build, and operate the Kubernetes platform used across PulsePoint.
  • Own reliability, observability, and incident response across platform services.
  • Build infrastructure automation and GitOps workflows to reduce operational toil.

PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health signals using real-world data. With over 300 employees, the company is a post-acquisition profitable leader in the US healthcare ad market, known for a flat hierarchy and high engineering bar.

$160,000–$208,000/yr
US 16w maternity 10w paternity

  • Lead CI/CD pipeline optimization and container orchestration to scale infrastructure across teams.
  • Architect secure infrastructure with secrets management, IAM, and vulnerability scanning as default.
  • Drive platform reliability through SLOs, observability, and incident response.

Evolve aims to make vacation rental easy for everyone. The team is high-performing, customer-obsessed, and runs on curiosity, communication, and accountability.

$128,000–$176,000/yr
US

  • Design and implement infrastructure using Terraform, Python, and Kubernetes on AWS.
  • Collaborate with engineering and data science teams to improve cloud infrastructure.
  • Automate CI/CD pipelines and enforce security governance and compliance.

Lyra Health is a mental health care provider serving 20 million people through employer and health plan partnerships. The company has delivered 15 million sessions and published 35 peer-reviewed studies, with a culture focused on clinical effectiveness.

Brazil 4w PTO

  • Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
  • Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
  • Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.

Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.

US

  • Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
  • Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
  • Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.

They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.

$180,000–$250,000/yr
US Unlimited PTO

  • Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
  • Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
  • Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.

We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.

US

  • Apply SRE principles to improve reliability, scalability, and performance of production systems.
  • Design and implement automation to reduce operational toil and improve engineering efficiency.
  • Lead incident response and develop sustainable solutions for complex production issues.

The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.

$97,976–$119,748/yr
Europe

  • Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
  • Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
  • Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.

Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.

Canada

  • Design and evolve highly available cloud architecture on Google Cloud using Terraform and GitOps.
  • Build and maintain secure CI/CD pipelines for Infrastructure-as-Code and develop self-service developer platforms.
  • Strengthen platform observability and apply SRE principles to improve reliability and operational maturity.

They are a financial technology company that provides production-critical infrastructure. They have a globally distributed team and a culture of autonomy and async-first collaboration.

Europe UK North America

  • Design and evolve cloud infrastructure on GCP for scale and resilience.
  • Build internal tooling and automation that promote team autonomy and developer productivity.
  • Advance observability platform with metrics, logging, tracing, and alerting to reduce recovery time.

The company is a well-funded AI/ML company at the intersection of geospatial intelligence and climate technology, building products on scalable cloud infrastructure. The engineering team fosters a culture of reliability and continuous improvement, operating with a focus on SLOs, error budgets, and DORA metrics.

$241,000–$270,000/yr
US Unlimited PTO

  • Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
  • Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
  • Build observability platforms to detect issues proactively and mentor engineers on reliability standards.

Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.

US Canada 16w maternity 16w paternity

  • Develop and operate cloud-based services that automate critical spacecraft operations and reduce complexity.
  • Build, maintain, and enhance automation, microservices, and cloud infrastructure for satellite operations.
  • Collaborate across teams to translate operational needs into effective technical solutions.

Our partner builds and operates satellite constellations and develops software for spacecraft operations. They have a mission-focused, collaborative culture with a strong emphasis on engineering excellence and team ownership.

US Unlimited PTO

  • Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers.
  • Provide technical leadership and mentorship across the Systems Engineering team.

Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.