Source Job

North America

  • Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
  • Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
  • Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.

Kubernetes Linux Postgres CI/CD Cloud Infrastructure

20 jobs similar to Member of the Technical Staff, Platform

Jobs ranked by similarity.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

Canada

  • Lead the design and operation of GPU infrastructure for AI workloads.
  • Manage Kubernetes-based environments and optimize for AI training and inference.
  • Define operational standards, implement monitoring, and collaborate with AI engineering teams.

ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.

$190,000–$230,000/yr
Global Unlimited PTO

  • Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
  • Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
  • Lead incident response and encode every fix as a new runbook and regression test.

Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.

$150,000–$160,000/yr

  • Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
  • Oversee cluster reliability, GPU fleet operations, and incident response.
  • Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.

Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

Global Unlimited PTO

  • Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
  • Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
  • Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.

Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.

Europe 6w PTO

  • Run and evolve the Kubernetes landscape (Amazon EKS, on-prem via Rancher) for all deployments.
  • Automate deployment and scaling, and build observability to spot problems before users notice.
  • Design abstractions that let backend and data engineers ship without opening tickets.

Yazio is a nutrition app company that helps millions of users in over 150 countries lead healthier lives through diet tracking. The platform engineering team is small and senior, with a remote-first culture that values efficiency and work-life balance.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.

Global

  • Design and build infrastructure primitives that define how CI/CD, build systems, and developer environments scale across the engineering org.
  • Build and operate the Kubernetes-based control plane behind CI/CD, including GitHub Actions runners, GitOps workflows, and ephemeral environments.
  • Develop core infrastructure components like Kubernetes Operators and scaling automation that product teams use directly, reducing bespoke per-team tooling.

Chainlink is the industry-standard oracle platform that brings capital markets onchain and powers the majority of decentralized finance. The company has enabled tens of trillions in transaction value and is adopted by major financial institutions and top protocols.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

Global

  • Mature deployment artifacts for enterprise customers, including hardened images, Terraform modules, and Helm charts.
  • Own reference architectures and improve the upgrade experience for seamless, predictable operations.
  • Mentor other engineers and contribute to architecture decisions across engineering teams.

Coder is an AI software development company that empowers teams to build software faster and more securely through autonomous coding agents. They are a growing company that values innovation, collaboration, and high standards for code quality and security.

US

  • Build and maintain the Shadeform GPU platform and automated AI infrastructure services.
  • Work on novel solutions to GPU market challenges including provisioning, orchestration, and virtualization.
  • Own the systems that turn fragmented GPU capacity into a reliable, production-ready platform.

Shadeform provides a unified platform for deploying and managing GPU infrastructure across cloud providers, neoclouds, and data centers. They are a remote-first startup focused on innovation and working with the latest AI technologies.

Global

  • Lead the architecture and implementation of managed Kubernetes infrastructure across AWS, Azure, and GCP.
  • Own the systems that provision and manage cloud accounts and subscriptions across providers.
  • Design and implement the networking layer routing traffic into customer environments.

Ditto builds the world's leading edge sync platform, enabling applications to share data peer-to-peer with or without internet connectivity. With over $145 million in funding and trusted by major organizations, Ditto is a globally distributed, fast-growing startup committed to diversity and inclusion.

North America Unlimited PTO

  • Deploy and scale MCP-based AI agents on Kubernetes for enterprise customers across the US-East and EMEA regions.
  • Lead complex technical engagements, build reusable deployment patterns, and mentor engineers on the team.
  • Shape product roadmap by feeding back field insights from regulated industries and defining regional engagement standards.

Stacklok builds the control plane for enterprise AI agents, enabling organizations to run, govern, and secure them on Kubernetes and private cloud. Founded by two Kubernetes creators, the company is already adopted by leading tech and regulated industries, fostering a collaborative, AI-maximalist culture with deep open-source roots.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.

$160,000–$180,000/yr
US

  • Own the infrastructure and platform powering the marketplace, focusing on reliability, observability, security, and automation.
  • Manage production AWS and EKS clusters, infrastructure as code with Terraform and GitOps, and CI/CD pipelines via GitHub Actions.
  • Build automation and internal tooling in Python, Bash, Go, and Node.js/TypeScript, and operate PostgreSQL, MongoDB, and Temporal.

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. The company is hyper-growth, profitable, and expanding quickly, backed by top marketplace investors.

Unlimited PTO

  • Deploy and scale our AI-agent infrastructure.
  • Own observability end to end and define how we measure reliability.
  • Own CI/CD pipelines and make shipping fast and safe.

Footprint builds Percy, an AI agent that runs financial crime investigations end to end. The company is backed by QED, Index, and other investors, and its small, senior team ships fast and grew revenue 5x in the past year.

Global Unlimited PTO

  • Build enterprise-scale infrastructure using infrastructure-as-code and Kubernetes-native systems.
  • Sustain platform health and performance by owning critical systems in production.
  • Enable teams and customers to move faster with abstractions and tooling for AI/ML workloads.

Cake makes cutting-edge AI accessible to enterprise teams by removing infrastructure barriers, enabling 10x faster and cheaper AI/ML platform deployment. Backed by top investors, they have a small senior team focused on ownership and operational excellence.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms, owning the storage layer where Kubernetes meets bare metal.
  • Tune NFS data paths for high-throughput, low-latency GPU/AI workloads, integrating NFS-based storage into clusters via CSI, storage classes, and persistent volumes.
  • Automate storage provisioning with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines, and build monitoring and observability for storage performance and health.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The company combines open source innovation with deep Kubernetes expertise, empowering platform engineering teams across on-premises, cloud, edge, and sovereign data centers.

EMEA

  • Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
  • Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
  • Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.

They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.