Source Job

$139,200–$235,200/yr
Canada United States Unlimited PTO

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
  • Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.

Rust Kubernetes Distributed Systems AWS GCP

20 jobs similar to Senior Platform Engineer, GitLab Orbit

Jobs ranked by similarity.

Global

  • Design and build agent runtime infrastructure with Firecracker, Rust, and Go
  • Define and enforce security boundaries for running untrusted AI agents
  • Architect global scale distributed systems for scheduling and orchestration

We build agent sandboxes—runtime infrastructure that is fast, durable, and secure by default for AI systems. We are a small, globally distributed team based in San Francisco backed by forward-thinking investors.

US

  • Implement and manage the infrastructure stack to enable the engineering team to ship quickly and effectively.
  • Proactively identify and eliminate bottlenecks in the devops process to ensure optimal developer velocity.
  • Maintain Tempo chain reliability, validator infrastructure, and explorer reliability.

Tempo is a layer-1 blockchain purpose-built for stablecoins and real-world payments, born from Stripe and Paradigm. They are a team of crypto-optimists building infrastructure for onchain payments.

UK

  • Architect and build a robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on top of the public cloud and automate cloud resource management.
  • Work closely with core database development and security teams to produce the SaaS offering.

ClickHouse is a real-time analytics and data warehousing company recognized on the Forbes Cloud 100 list. With over 4,000 customers and rapid growth, the company is a leader in its space.

UK

  • Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
  • Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.

GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.

Canada

  • Build and operate backend platform services for APIs, Kafka-based event routing, workflow execution, and secrets management.
  • Design event-driven workflows and durable executions using technologies such as Kafka and Temporal.
  • Own production reliability, observability, and security across multi-tenant, multi-cloud environments.

This partner company builds core platform services for workflow execution, identity, event routing, secrets management, and audit infrastructure. It is a high-growth, engineering-led organization focused on innovation, accountability, and continuous improvement.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

$175,000–$250,000/yr
US 2w maternity 2w paternity

  • Design, develop, and maintain internal software, services, and automation using Go and Rust.
  • Build and operate Kubernetes-based infrastructure and improve developer workflows and CI/CD.
  • Collaborate across teams to solve ambiguous technical challenges and improve system reliability.

Our partner is a high-growth technology organization building internal platforms to enable efficient engineering. They operate with small, autonomous teams in a high-trust, collaborative remote environment with a focus on technical excellence.

Global

  • Evolving Supabase Edge Runtime, an open-source Rust-based host that runs Deno isolate and enforces per-request memory and CPU limits.
  • Implementing monitoring, alerting, and OpenTelemetry tracing to drive latency and reliability improvements.
  • Expanding functions for more use cases like AI inference, MCP servers, and improving developer experience with Supabase CLI.

Supabase is the Postgres development platform, providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. With a globally distributed team of ~400 members across 60+ countries, they are open-source-first and move fast, building in public.

Global

  • Design and improve the core orchestration engine for managing Supabase Branches lifecycle.
  • Provision ephemeral sandboxed execution environments on EKS/ECS for untrusted build workloads.
  • Implement job scheduling, pipeline optimization, and observability for thousands of concurrent builds.

Supabase is the Postgres development platform, built by developers for developers. They are a globally distributed team of ~400 members across 60+ countries, operating fully remote with an open-source-first culture.

$152,800–$259,200/yr
Americas Unlimited PTO

  • Lead cross-cutting initiatives and drive modular backend architecture to improve scalability and developer velocity.
  • Set standards for APIs, testing, security, and operability while mentoring engineers and shaping technical direction.
  • Own production systems, participate in on-call, and lead AI-assisted engineering workflows with quality and safety guardrails.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity and improve operational efficiency. With over 50 million users and trust from more than 50% of the Fortune 100, GitLab fosters a high-performance culture driven by values, continuous knowledge exchange, and AI integration.

Global

  • Own and scale cloud infrastructure including compute, networking, storage, and data systems.
  • Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
  • Establish reliability through service-level objectives, observability, and incident response processes.

A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.

US

  • Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
  • Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
  • Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.

WellSaid Labs is the leading AI voiceover studio for enterprise and professional use, providing ultra-realistic voices that the world’s biggest brands trust. We are a fully distributed team across the U.S. with a focus on responsible AI and an inclusive culture.

US Unlimited PTO

  • Own the infrastructure layer for AI workloads including inference serving, Kubernetes, and agent-sandboxing platforms.
  • Manage the serving tier for open-weight models, Kubernetes operators, and stateful data planes.
  • Oversee the sandbox runtime, control-plane services, and observability tooling.

AZX accelerates positive impact in critical industries through AI transformation, specializing in physics-informed ML and enterprise AI solutions for climate and sustainability. Founded in 2024, the company is a profitable public benefit corporation with a growing team working with category leaders in real estate, energy, logistics, and utilities.

$220,000–$292,000/yr
US Unlimited PTO

  • Own the platform including GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
  • Contribute to AI enablement substrate: GPU capacity, training/inference pipelines, and cost optimization.
  • Strengthen team practices through tooling, standards, tests, observability, and release processes.

Descript is building a simple, intuitive, fully-powered editing tool for video and audio — an editing tool built for the age of AI. They are a team of 150 backed by top investors like OpenAI and Andreessen Horowitz, with a culture that values collaboration and serendipitous discovery.

US

  • You'll contribute to infrastructure scaling to infinitely many apps, improving performance and reliability across backend services.
  • You'll support observability efforts, help implement SLOs, and build foundational services for next-generation cloud infrastructure.
  • You'll participate in triage and on-call processes to diagnose issues and implement changes to prevent recurrence.

Bubble is an AI visual development platform that empowers anyone to create software without code, from first-time entrepreneurs to enterprise teams. With over 6 million users in more than 100 countries and a mission to break down barriers to entrepreneurship, the company fosters a collaborative and inclusive culture focused on empowering builders worldwide.

Poland UK Unlimited PTO

  • Own moderately sized to complex technical initiatives from problem definition through implementation, rollout, and operational follow-through.
  • Act as the directly responsible individual for projects by aligning stakeholders, communicating progress, and keeping execution moving.
  • Lead technical design for distributed storage, Git repository management, and performance using data to guide decisions.

GitLab is an intelligent orchestration platform for DevSecOps that enables organizations to increase developer productivity and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 as customers, GitLab fosters a high-performance culture driven by values and continuous knowledge exchange.

Europe 6w PTO

  • Design, build, and operate infrastructure for real-time systems handling millions of concurrent connections and billions of monthly API requests.
  • Drive Kubernetes end to end: cluster architecture, workload design, and migration of existing services from AWS to GCP.
  • Own cloud cost and efficiency optimization, measuring impact against real spend and utilization data.

Stream powers real-time chat, video, activity feeds, and AI moderation for billions of end-users across thousands of apps. We are a Series B company with around 145 employees from over 35 countries, offering a fast-paced startup culture with real ownership.

$160,000–$200,000/yr
United States Canada Europe Unlimited PTO 12w maternity 12w paternity

  • Write reliable, secure, and scalable code for Lithic's product platform with minimal tech debt.
  • Improve system reliability and participate in team on-call rotation.
  • Own projects from planning to launch, including API key management, audit logging, and search.

Lithic is a modern card issuing and processing platform for financial companies to build the future of payments. They are a team of 170+ across 26 states and 7 countries, with a remote-first culture.

$140,000–$220,000/yr
North America LATAM Europe

  • Own and scale the cloud infrastructure behind our open-source platform: compute, networking, and the data layer.
  • Lead BYOC: turn customer-cloud deployments into a real product, with provisioning, upgrades, and observability that scale past bespoke work per deal.
  • Make reliability a product feature: meaningful SLOs, and an incident process people trust.

Nango is a developer infrastructure company that provides API access for agents and apps, enabling AI applications to connect to the real world through integrations. With over 400 paying customers and a team of 14 from top tech companies like AWS, GitHub, and Okta, they are a YC-backed, multi-million ARR company that values ownership and autonomy.

$180,000–$240,000/yr
US

  • Own large slices of the system end to end, from approach to operation.
  • Turn Beads into a platform and take Gas City to the cloud.
  • Define SLOs, observability, backups, and security baseline for enterprise readiness.

Gas City builds the open-source stack teams use to run coding agents at scale, including the Beads work graph and Gas City agent orchestration. It's a small, flat organization moving toward revenue with a focus on reliability and agent-driven development.