Source Job

$180,000–$240,000/yr
US

  • Own large slices of the system end to end, from approach to operation.
  • Turn Beads into a platform and take Gas City to the cloud.
  • Define SLOs, observability, backups, and security baseline for enterprise readiness.

Go Distributed Systems Kubernetes Reliability Observability

20 jobs similar to Staff Platform Engineer

Jobs ranked by similarity.

$180,000–$240,000/yr
US

  • Own large slices of the system end-to-end, from approach to shipping and operations.
  • Maintain the Beads Team Server and Hosted Gas City, focusing on state, concurrency, and reliability.
  • Be a maintainer of the open-source repos, representing community needs and compatibility with our roadmap.

Gas City builds the open-source stack that teams use to run coding agents at scale, including the work graph Beads and the agent orchestration tool Gas City. We are a small, flat organization of founding engineers who work 1:1 with the CTO and run our own product on itself, using agents to write and review code.

US Unlimited PTO

  • Design, build, and operate distributed systems that ingest, process, and store telemetry at very high scale.
  • Own the reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage systems.
  • Participate in the on-call rotation, help resolve production incidents, and drive root-cause fixes through to completion.

ClickHouse provides a real-time analytics database for data warehousing, observability, and AI workloads. Recognized on the 2025 Forbes Cloud 100 list, it has over 4,000 customers and significant year-over-year growth.

$139,200–$235,200/yr
Canada United States Unlimited PTO

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
  • Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. Trusted by more than 50 million registered users and over 50% of the Fortune 100, GitLab fosters a high-performance culture driven by shared values and continuous knowledge exchange.

US 6w PTO

  • Take an active role in influencing our roadmap and your own career objectives.
  • Design, build, operate, and maintain critical systems, owning reliability, performance, and availability.
  • Collaborate with your team to deliver new features and iterate based on results.

Grafana Labs is the company behind the open-source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries and 35 million users, the company thrives on a transparent, collaborative, and open-source culture.

$217,000–$303,900/yr
US 17w maternity 17w paternity

  • Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing.
  • Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
  • Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet's largest sources of information.

$176,000–$231,000/yr
US Unlimited PTO

  • Lead the design and delivery of business-critical systems powering customer acquisition, onboarding, billing, and revenue operations.
  • Drive architectural excellence for high-throughput, fault-tolerant metering and billing pipelines.
  • Mentor engineers and foster a culture of ownership, learning, and pragmatic excellence.

Temporal is an open source programming model that simplifies code, making applications more reliable and developers more productive. It is a growing, values-driven company with a collaborative culture focused on quality and pragmatism.

$140,000–$220,000/yr
North America LATAM Europe

  • Own and scale the cloud infrastructure behind our open-source platform: compute, networking, and the data layer.
  • Lead BYOC: turn customer-cloud deployments into a real product, with provisioning, upgrades, and observability that scale past bespoke work per deal.
  • Make reliability a product feature: meaningful SLOs, and an incident process people trust.

Nango is a developer infrastructure company that provides API access for agents and apps, enabling AI applications to connect to the real world through integrations. With over 400 paying customers and a team of 14 from top tech companies like AWS, GitHub, and Okta, they are a YC-backed, multi-million ARR company that values ownership and autonomy.

US

  • Provide technical leadership for reliability across a large-scale advertising technology ecosystem
  • Lead reliability initiatives across ad serving, auctions, targeting, reporting, and billing systems
  • Mentor engineers and influence technical decisions to improve system resilience and developer productivity

The company is a partner organization operating a large-scale advertising technology ecosystem. Its size and culture are not detailed, but the role emphasizes reliability and operational excellence in a high-traffic environment.

Global

  • Design and build agent runtime infrastructure with Firecracker, Rust, and Go
  • Define and enforce security boundaries for running untrusted AI agents
  • Architect global scale distributed systems for scheduling and orchestration

We build agent sandboxes—runtime infrastructure that is fast, durable, and secure by default for AI systems. We are a small, globally distributed team based in San Francisco backed by forward-thinking investors.

US

  • Co-own the MaaS system design as the deepest backend voice, driving architectural decisions and setting engineering standards.
  • Build a globally distributed, multi-tenant token service with high throughput, reliability, and observability.
  • Own end-to-end implementation, from code to migrations, ensuring zero-downtime upgrades and strict SLOs.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure, providing comprehensive Bitcoin mining solutions and AI computational infrastructure. Headquartered in Singapore, it has deployed data centers across the United States, Norway, Bhutan, and Ethiopia, and is committed to equal employment opportunities.

Poland UK Unlimited PTO

  • Own moderately sized to complex technical initiatives from problem definition through implementation, rollout, and operational follow-through.
  • Act as the directly responsible individual for projects by aligning stakeholders, communicating progress, and keeping execution moving.
  • Lead technical design for distributed storage, Git repository management, and performance using data to guide decisions.

GitLab is an intelligent orchestration platform for DevSecOps that enables organizations to increase developer productivity and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 as customers, GitLab fosters a high-performance culture driven by values and continuous knowledge exchange.

$150,000–$190,000/yr
US Canada Europe

  • Build and enhance SpiceDB, a distributed permissions database, and contribute to the open-source ecosystem.
  • Drive best practices in software development, testing, and CI/CD to ensure robust and scalable platform.
  • Collaborate with a high-performing engineering team to address complex challenges in distributed systems and authorization.

AuthZed creates and maintains SpiceDB, an authorization infrastructure used by companies globally to simplify permission management. As a Series A company, they have a fully remote, hardworking team with a software-driven culture across the US, Canada, and Europe.

Europe 6w PTO

  • Lead multi-quarter technical initiatives on Tempo's architecture, including trace aggregation APIs, autoscaling, and query engine improvements.
  • Drive operational excellence by owning SLOs, reducing toil, and ensuring Tempo operates reliably at scale across growing cell counts.
  • Design APIs for humans and agents, partner with product teams, and mentor engineers to raise the bar across the organization.

Grafana Labs is the company behind the open observability cloud, Grafana Cloud, a fully managed observability platform built for scale. With over 1,600 team members across 40+ countries, we are a 100% remote company backed by leading investors, fostering a global collaborative culture and a passion for meaningful work.

$151,000–$206,000/yr
US Canada Unlimited PTO

  • Build large-scale real-time services and applications leveraging massive datasets.
  • Develop and maintain data pipelines, messaging systems, databases, and cloud services.
  • Work with Machine Learning Engineers and Security Researchers on security solutions.

Censys provides real-time Internet intelligence and threat insights to global governments and Fortune 500 companies. It is a growing company with a focus on comprehensive internet mapping and security solutions.

$174,986–$209,983/yr
US Canada 6w PTO

  • Design and build core backend services for context ingestion, indexing, retrieval APIs, and agent-facing integrations.
  • Create a scalable multi-tenant SaaS foundation with tenant isolation, usage tracking, and reliable service boundaries.
  • Work across product and infrastructure to make practical tradeoffs between fast experimentation and long-term reliability.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we value an open-source legacy, global collaboration, and transparency.

Canada 6w PTO

  • Drive technical strategy and roadmap for adaptive telemetry databases.
  • Lead end-to-end delivery of large, cross-functional projects.
  • Own architecture, reliability, performance, and cost for critical systems.

Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we foster a collaborative, open-source culture.

UK

  • Architect and build a robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on top of the public cloud and automate cloud resource management.
  • Work closely with core database development and security teams to produce the SaaS offering.

ClickHouse is a real-time analytics and data warehousing company recognized on the Forbes Cloud 100 list. With over 4,000 customers and rapid growth, the company is a leader in its space.

$77,000–$106,700/yr
France

  • Design and develop scalable search and indexing systems for an AI search engine.
  • Ensure operational excellence by participating in on-call rotation and maintaining system quality.
  • Collaborate with a global remote team to solve distributed system challenges.

Algolia is a pioneer and market leader in AI Search, empowering over 18,000 businesses to deliver blazing-fast search experiences. With $150 million in Series D funding and a valuation of $2.25 billion, the company fosters a high-trust, flexible culture and values diversity and collaboration.

US Unlimited PTO

  • Build and deliver high-quality solutions that power Honeycomb's query and data storage infrastructure.
  • Scope and deliver projects independently, breaking down complex storage problems into achievable steps.
  • Support our services in production, participating in on-call rotations and reducing toil.

Honeycomb is a service defining observability and raising expectations for developer tools, working with companies like HelloFresh and Slack. We've scaled past 200 people, closed Series D funding, and were named to Forbes' America's Best Startups in 2022 and 2023.

Global

  • Design, build, and operate Go services powering core APIs and data access paths across HTTP and gRPC.
  • Own the reliability, performance, and scalability of services, including handling growing traffic and data volumes.
  • Contribute to code reviews, technical design discussions, and mentor other engineers to raise the technical bar.

Brightfield provides an AI-powered workforce analytics platform that helps large companies manage their extended workforce. We are a fully remote team of data-driven innovators, trusted by the Global 2000 since 2006.