Source Job

$190,000–$230,000/yr
US

  • Design and implement scalable cloud infrastructure using Kubernetes, Pub/Sub, and distributed systems technologies.
  • Collaborate with our AI team to optimize data pipelines and integrate AI to remove performance bottlenecks.
  • Drive platform reliability initiatives including alerting, health checking, and incident management.

Python Ruby GCP Kubernetes Terraform

20 jobs similar to Staff Software Engineer, Distributed Systems

Jobs ranked by similarity.

$220,000–$292,000/yr
US Unlimited PTO

  • Own the platform including GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
  • Contribute to AI enablement substrate: GPU capacity, training/inference pipelines, and cost optimization.
  • Strengthen team practices through tooling, standards, tests, observability, and release processes.

Descript is building a simple, intuitive, fully-powered editing tool for video and audio — an editing tool built for the age of AI. They are a team of 150 backed by top investors like OpenAI and Andreessen Horowitz, with a culture that values collaboration and serendipitous discovery.

$185,000–$200,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Coordinate with technical and non-technical staff across departments, including workflow automation that bridges infrastructure and business processes.
  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in GCP.
  • Maintain incident response process and tooling, and build automation that reduces toil and enables self-healing infrastructure.

Branch empowers workers with financial freedom by helping companies accelerate payments and providing accessible, free financial services. It is a remote-first, award-winning FinTech with employees across the U.S., fostering a culture of transparency, accountability, and trust.

$139,200–$235,200/yr
Canada United States Unlimited PTO

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
  • Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. Trusted by more than 50 million registered users and over 50% of the Fortune 100, GitLab fosters a high-performance culture driven by shared values and continuous knowledge exchange.

$175,000–$185,000/yr
US

  • Consolidate Terraform and establish conventions for state management, modules, and CI checks.
  • Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
  • Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.

Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.

Brazil 4w PTO

  • Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
  • Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
  • Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.

Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.

$135,000–$150,000/yr
US

  • Operate, scale, and troubleshoot Bitsight's SaaS cloud infrastructure with focus on reliability, efficiency, and security.
  • Tackle complex system-level designs and proactively anticipate performance and scalability issues.
  • Pioneer self-optimizing infrastructure systems using AI, ensuring manual and staging validation before production deployment.

Bitsight is a cyber risk management leader transforming how companies manage exposure, performance, and risk. Over 3,500 customers and 600 teammates work across Boston, Raleigh, New York, Lisbon, Singapore, and remote locations.

Global Unlimited PTO

  • Architect and build robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on public cloud and automate resource management.
  • Improve reliability, security, and cost efficiency of cloud services.

ClickHouse builds a real-time analytics database platform and manages ClickHouse Cloud data plane end-to-end with compute, networking, and security. As a rapidly scaling global startup, the company operates across 25+ countries and fosters a flexible, remote-friendly culture with equity and healthcare benefits.

$77,000–$106,700/yr
France

  • Design and develop scalable search and indexing systems for an AI search engine.
  • Ensure operational excellence by participating in on-call rotation and maintaining system quality.
  • Collaborate with a global remote team to solve distributed system challenges.

Algolia is a pioneer and market leader in AI Search, empowering over 18,000 businesses to deliver blazing-fast search experiences. With $150 million in Series D funding and a valuation of $2.25 billion, the company fosters a high-trust, flexible culture and values diversity and collaboration.

India

  • Design, build, and ship production services, APIs, and user-facing interfaces.
  • Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
  • Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.

Motive empowers people who run physical operations with tools to make their work safer, more productive, and more profitable. Serving nearly 100,000 customers across industries, the company values a diverse and inclusive workplace.

$59,400–$65,880/yr
Europe

  • Lead the design, implementation, and ongoing improvement of reliable, scalable, and secure production platforms and services.
  • Work closely with cross-functional teams to build and maintain resilient infrastructure and deployment patterns.
  • Provide technical leadership and mentorship, promoting strong engineering standards and operational best practices.

Cision is a global leader in PR, marketing and social media management technology and intelligence, helping brands connect with customers and stakeholders. They have offices in 24 countries, a network of over 1.1 billion influencers, and a culture that champions diversity, equity, and inclusion.

France

  • Evolve an Internal Developer Platform enabling development teams to deploy and operate applications securely with high self-service.
  • Play a crucial role in infrastructure architecture in a multi-cloud environment with a focus on GCP.
  • Build and maintain the platform used by over 50 internal clients, then support their concrete use by teams.

Lifen believes medical data can transform healthcare by reducing administrative burden, improving care coordination, and accelerating scientific discovery. Since 2015, the company has connected 800 hospitals and 150,000 healthcare professionals, with over 150 employees working remotely and from offices to unlock the potential of health data.

$126,000–$174,000/yr
US

  • Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems.
  • Architect and own scalable Kubernetes platforms, infrastructure as code, and DevSecOps implementation.
  • Drive platform reliability, performance SLAs, cost optimization, and lead complex migrations and AI/ML platform infrastructure.

Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, the company has delivery centers across Canada, the US, Eastern Europe, and Latin America, with teams averaging over 15 years of experience.

$180,000–$220,000/yr
US

  • Design, implement, and maintain reliable, scalable, and secure infrastructure to support applications and automation systems.
  • Automate infrastructure provisioning, configuration management, and deployment pipelines using tools like Terraform and ArgoCD.
  • Implement observability solutions and enforce security best practices to ensure uptime and system performance.

Bright Machines is a next-generation, AI-enabled manufacturer focused on data center infrastructure production, using proprietary AI-based robotics and software to assemble hardware products for hyperscalers and OEMs. The company is headquartered in San Francisco, California, with an integration center in Guadalajara, Mexico, and has been recognized by Forbes' AI 50 and other leading organizations.

Latin America

  • Build and operate model and inference serving infrastructure, managing latency, throughput, autoscaling, and reliability for real-time and batch inference.
  • Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies, and safe rollback.
  • Operate agentic and LLM workloads in production, managing inference providers, gateways, quotas, guardrails, and graceful degradation under load.

ReadyOn is an AI-native Labor Operating System that redefines how enterprises manage frontline labor by matching workers to shifts in real time. Headquartered in San Francisco with over 100 employees, it grew revenue 8x year over year in 2025.

UK

  • Lead Cloud Platform and SRE teams to scale securely and efficiently.
  • Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
  • Champion SRE culture with SLOs, error budgets, and observability.

Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.

US Unlimited PTO

  • Build and operate the Kubernetes platform supporting AI test and evaluation frameworks.
  • Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines.
  • Own platform reliability, observability, capacity planning, and operational readiness.

OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by the creator of NumPy and SciPy, the company is built by people with deep roots across the open-source ecosystem and maintains a remote-first culture.

Global

  • Contribute to platform and harness engineering, including CI/CD and developer tooling.
  • Build systems to reduce toil and maintain production infrastructure under conversational AI traffic.
  • Participate in on-call rotation and incident management to ensure platform uptime.

Replicant builds an AI-powered customer service platform that helps contact centers resolve requests and improve agent performance. The company is distributed, with a focus on ownership and collaboration, and serves Fortune 500 companies.

Europe

  • Lead end-to-end technical engagements: Partner directly with engineering teams to diagnose, unblock, and resolve complex infrastructure challenges.
  • Execute critical migrations: Develop reference implementations, tooling, and guidance to transition teams off deprecated systems seamlessly.
  • Accelerate platform adoption: Act as primary technical contact for new teams onboarding to Planet's core infrastructure.

Planet designs, builds, and operates the largest constellation of imaging satellites in history, delivering unprecedented dataset via a cloud-based platform for commercial, environmental, and humanitarian sectors. A global company with offices in the US, Europe, and Slovenia, Planet values a people-centric culture and community.

$145,000–$260,000/yr
US Canada Unlimited PTO

  • Design, build, and optimize multi-region, high-availability AWS infrastructure.
  • Drive resiliency and automation using GitOps, modern CI/CD, and Infrastructure as Code.
  • Build end-to-end telemetry and own incident management to harden reliability.

VGS is the world's leader in payment tokenization, trusted by the most innovative AI and Fortune 500 companies. They are a remote-first company with a culture of ownership, collaboration, and continuous learning.

  • Own the reliability, performance, and scalability of Runlayer's infrastructure across AWS and GCP.
  • Manage Kubernetes clusters, database reliability, and CI/CD pipelines for rapid deployments.
  • Lead incident response and partner with product engineers to design resilient systems for enterprise customers.

Runlayer builds a unified platform for MCPs, Skills, and AI Agents, providing enterprises with security, governance, and observability to deploy AI safely and at scale. Founded by engineers who built AI Actions for OpenAI and Zapier Agents, the team has raised $42M from Felicis and Khosla Ventures, serving companies like Gusto, Instacart, and Opendoor.