Source Job

$160,000–$230,000/yr
US

  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
  • Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
  • Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.

Kubernetes GPU Ansible InfiniBand

20 jobs similar to Technical Support Engineer (GPU Clusters)

Jobs ranked by similarity.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.

Canada

  • Lead the design and operation of GPU infrastructure for AI workloads.
  • Manage Kubernetes-based environments and optimize for AI training and inference.
  • Define operational standards, implement monitoring, and collaborate with AI engineering teams.

ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms.
  • Own the storage layer where Kubernetes meets bare metal, tuning NFS data paths for high-throughput workloads.
  • Automate storage provisioning with infrastructure-as-code and GitOps, ensuring observability and reliability.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams across hybrid, edge, and sovereign environments, fostering a culture of open-source innovation and collaboration among passionate, talented colleagues.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

Global

  • Define the networking strategy and roadmap for k0rdent AI, covering GPU cluster networking and multi-tenant cloud integration.
  • Partner with engineering and marketing to shape requirements, positioning, and competitive differentiation.
  • Represent Mirantis at events and with strategic accounts, driving product success in the AI cloud era.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With a world-class, distributed team, Mirantis empowers platform engineering teams and is committed to openness, collaboration, and continuous growth.

$125,000–$135,000/yr
Global

  • Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion.
  • Tune fabric performance parameters for distributed AI workloads such as NCCL and MPI.
  • Collaborate with GPU and networking teams to diagnose and resolve fabric-level issues and optimize performance.

Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is the largest privately-held cloud infrastructure company, offering a culture of innovation and growth.

$180,000–$200,000/yr
Global

  • Lead customers in designing and optimizing GPU-based solutions on Vultr's platform.
  • Collaborate with cross-functional teams to bring AI, ML, and GPU workloads into production.
  • Educate customers on the value of Vultr's cloud infrastructure and expand their possibilities.

Vultr provides high-performance cloud infrastructure solutions globally, making them easy to use, affordable, and locally accessible. It is a privately-held company with over a decade of self-funding, hundreds of thousands of customers across 185 countries, and a culture that emphasizes comprehensive benefits and employee growth.

US

  • Monitor, triage, and resolve customer-reported incidents within defined SLAs.
  • Serve as primary point of contact for technical issues related to installation, Helm configuration, and integrations.
  • Troubleshoot Kubernetes-related issues such as ingress, SSO, and other integrations.

ScaleOps redefines autonomous cloud and AI infrastructure, freeing DevOps engineers from manual resource management. The company is backed by over $210M in funding and trusted by leading enterprises including Adobe and Coinbase.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms, owning the storage layer where Kubernetes meets bare metal.
  • Tune NFS data paths for high-throughput, low-latency GPU/AI workloads, integrating NFS-based storage into clusters via CSI, storage classes, and persistent volumes.
  • Automate storage provisioning with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines, and build monitoring and observability for storage performance and health.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The company combines open source innovation with deep Kubernetes expertise, empowering platform engineering teams across on-premises, cloud, edge, and sovereign data centers.

Europe

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

Mirantis is a Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They serve many leading enterprises including Adobe, DocuSign, and PayPal, and are a leader in container management.

North America

  • Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
  • Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
  • Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.

Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.

US

  • Build and maintain the Shadeform GPU platform and automated AI infrastructure services.
  • Work on novel solutions to GPU market challenges including provisioning, orchestration, and virtualization.
  • Own the systems that turn fragmented GPU capacity into a reliable, production-ready platform.

Shadeform provides a unified platform for deploying and managing GPU infrastructure across cloud providers, neoclouds, and data centers. They are a remote-first startup focused on innovation and working with the latest AI technologies.

Global

  • Own the vision, roadmap, and priorities for k0rdent AI networking, spanning underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and network automation.
  • Translate requirements from GPU clouds, telcos, and enterprise platform teams into clear product direction, partnering with engineering to define requirements.
  • Track and shape response to emerging interconnect standards like Ultra Ethernet, UALink, and congestion control, and represent Mirantis with customers and partners.

Mirantis is the leading AI-infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They are a global, distributed team committed to openness and technical excellence, serving clients like Adobe, PayPal, and Volkswagen.

North America Unlimited PTO

  • Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against quality metrics.
  • Define the qualification bar: build the acceptance test suite, benchmark methodology, and quality thresholds to formalize tribal knowledge.
  • Guide providers through technical onboarding: work with their engineers to remediate gaps and bring clusters onto the network cleanly.

Andromeda Cluster provides early-stage startups access to scaled AI infrastructure that was once reserved for hyperscalers. It is a small, high-growth team at the center of the AI infrastructure boom, founded by Nat Friedman and Daniel Gross.