Source Job

US

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.

Kubernetes Distributed Systems Terraform

20 jobs similar to Staff AI Scheduling & Orchestration Engineer

Jobs ranked by similarity.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

US

  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
  • Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
  • Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.

Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.

Global

  • Design and build a managed Slurm service on Kubernetes
  • Write clean, reliable, and maintainable Go code, developing scheduling and orchestration for GPU workloads
  • Build observability and automated remediation for GPU, node, network, and control-plane failures

Gcore provides infrastructure and software solutions for AI, cloud, network, and security, powering digital experiences worldwide. With over 550 professionals, they build and support the global digital ecosystem.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

India

  • Design end-to-end AI infrastructure solutions for scalable, high-performance AI and HPC environments.
  • Partner with Sales to qualify opportunities, conduct technical discovery, and serve as trusted advisor throughout the sales lifecycle.
  • Collaborate closely with Facilities, Delivery, OEM partners, and customer teams to ensure AI infrastructure aligns with data center capabilities.

Submer enables organizations scaling AI to overcome the limits of traditional datacenters in power, compute density, and efficiency. It is a fast-growing, international scale-up with a friendly, diverse, and hybrid-friendly work environment.

US

  • Design and maintain CI/CD and MLOps pipelines for AI and software applications, ensuring seamless deployment and automation.
  • Build and scale cloud-native infrastructure using Kubernetes, Docker, and GPU clusters to support high-performance AI workloads.
  • Champion Infrastructure as Code and observability practices to ensure high availability, security, and compliance across multi-cloud environments.

Bitdeer is a world-leading technology company providing AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company has a global presence with data centers in multiple countries and a culture focused on innovation and reliability.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

US

  • Lead end-to-end architecture design of AI Data Center Networks, high-performance DCI, and global backbone networks for large-scale GPU clusters.
  • Collaborate with NVIDIA, vendors, and partners to translate business requirements into top-level network designs including InfiniBand/RoCE, Spine-Leaf, and overlay integration.
  • Own congestion control tuning, produce architecture documentation, and identify risks to drive network evolution.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company operates data centers across multiple countries and is building AI computational infrastructure.

North America

  • Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
  • Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
  • Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.

Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

North America Unlimited PTO

  • Deploy and scale MCP-based AI agents on Kubernetes for enterprise customers across the US-East and EMEA regions.
  • Lead complex technical engagements, build reusable deployment patterns, and mentor engineers on the team.
  • Shape product roadmap by feeding back field insights from regulated industries and defining regional engagement standards.

Stacklok builds the control plane for enterprise AI agents, enabling organizations to run, govern, and secure them on Kubernetes and private cloud. Founded by two Kubernetes creators, the company is already adopted by leading tech and regulated industries, fostering a collaborative, AI-maximalist culture with deep open-source roots.

  • Lead and build the SE / Solutions Architect team, defining the pre-sales operating model as the org scales.
  • Own the technical win in large, complex deals, architecting solutions across compute, networking, storage, and orchestration.
  • Be the technical voice of the customer internally, feeding structured product requirements back to product and platform teams.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The company is committed to open standards and freedom from lock-in, empowering platform engineering teams to deliver composable, production-ready developer platforms across any environment.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms.
  • Own the storage layer where Kubernetes meets bare metal, tuning NFS data paths for high-throughput workloads.
  • Automate storage provisioning with infrastructure-as-code and GitOps, ensuring observability and reliability.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams across hybrid, edge, and sovereign environments, fostering a culture of open-source innovation and collaboration among passionate, talented colleagues.

Global

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
  • Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
  • Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms, owning the storage layer where Kubernetes meets bare metal.
  • Tune NFS data paths for high-throughput, low-latency GPU/AI workloads, integrating NFS-based storage into clusters via CSI, storage classes, and persistent volumes.
  • Automate storage provisioning with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines, and build monitoring and observability for storage performance and health.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The company combines open source innovation with deep Kubernetes expertise, empowering platform engineering teams across on-premises, cloud, edge, and sovereign data centers.

$100,000–$150,000/yr
US

  • Design, build, and operate scalable infrastructure platforms for large-scale AI model training and inference.
  • Manage and optimize GPU clusters, distributed training environments, and scheduling systems for machine learning applications.
  • Develop software solutions and automation tools using Python and systems programming languages like Go or C++.

Our partner builds and operates foundational technology powering advanced AI training and inference workloads at scale. They offer a collaborative culture focused on innovation, engineering excellence, and continuous learning.

US

  • Build and maintain the Shadeform GPU platform and automated AI infrastructure services.
  • Work on novel solutions to GPU market challenges including provisioning, orchestration, and virtualization.
  • Own the systems that turn fragmented GPU capacity into a reliable, production-ready platform.

Shadeform provides a unified platform for deploying and managing GPU infrastructure across cloud providers, neoclouds, and data centers. They are a remote-first startup focused on innovation and working with the latest AI technologies.

  • Own end-to-end infrastructure deployment programs for new capacity and site expansions, including hardware dependencies and commissioning gate frameworks.
  • Deliver crisp, data-driven executive updates and govern cross-organizational dependencies without escalation.
  • Coach junior TPMs and drive AI tool integration to improve program tracking and risk detection.

Evergrid builds frontier-level AI infrastructure for the fourth industrial revolution, designing, deploying, and operating large-scale AI systems. With over 1GW of capacity in active development and a 24/7 engineering team, they foster a culture of high conviction and high trust.

APAC

  • Design and deliver significant components and core subsystems of our Kubernetes platform, such as secrets management, workload identity, storage, or cluster networking, from design through production operation.
  • Contribute to the architecture of distributed workloads, working with dependent teams to get runtime and isolation models right, while spending most time hands-on in code.
  • Own operability of built systems including SLOs, failure modes, upgrades, migrations, and on-call, and mentor earlier-career engineers.

ServiceNow is the AI control tower for business reinvention, bringing together any AI, any data, and any workflow to help 85% of the Fortune 500 work smarter, faster, and better. The company fosters an AI-native culture where technology and talent are unstoppable together, with a focus on freeing people from busywork.