Source Job

Europe

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

InfiniBand TCP/IP Linux

14 jobs similar to Principal HPC Network Engineer

Jobs ranked by similarity.

$125,000–$135,000/yr
Global

  • Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion.
  • Tune fabric performance parameters for distributed AI workloads such as NCCL and MPI.
  • Collaborate with GPU and networking teams to diagnose and resolve fabric-level issues and optimize performance.

Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is the largest privately-held cloud infrastructure company, offering a culture of innovation and growth.

Europe

  • Design, implement, and operate high-performance GPU networking fabrics and classical datacenter networking components such as routing, security, and external connectivity.
  • Own the long-term technical direction and operational strategy for AI interconnect networks, architecting scalable RoCE and Ethernet fabrics for distributed training and inference.
  • Collaborate cross-functionally with infrastructure, platform, SRE, and operations teams to integrate networking into the overall platform architecture and drive operational excellence.

Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. It is a fast-growing scale-up with a global team across the USA, Australia, Central Europe, Malaysia, Singapore and Japan.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

US

  • Serve as a subject matter expert for NVIDIA Networking technologies including Ethernet and InfiniBand-based data center and AI fabric solutions.
  • Design and deploy complex solutions, create written deliverables, and conduct client workshops while communicating architecture strategy to senior management.
  • Provide technical leadership for data center modernization, AI infrastructure networking, and high-performance network design initiatives.

AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery to help enterprises deliver on digital transformation. The company prioritizes a culture of belonging where all perspectives are valued and is an equal opportunity employer committed to diversity.

Global

  • Own the vision, roadmap, and priorities for k0rdent AI networking, spanning underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and network automation.
  • Translate requirements from GPU clouds, telcos, and enterprise platform teams into clear product direction, partnering with engineering to define requirements.
  • Track and shape response to emerging interconnect standards like Ultra Ethernet, UALink, and congestion control, and represent Mirantis with customers and partners.

Mirantis is the leading AI-infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They are a global, distributed team committed to openness and technical excellence, serving clients like Adobe, PayPal, and Volkswagen.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

$225,000–$325,000/yr
Global Unlimited PTO

  • Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
  • Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
  • Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

Global

  • Define the networking strategy and roadmap for k0rdent AI, covering GPU cluster networking and multi-tenant cloud integration.
  • Partner with engineering and marketing to shape requirements, positioning, and competitive differentiation.
  • Represent Mirantis at events and with strategic accounts, driving product success in the AI cloud era.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With a world-class, distributed team, Mirantis empowers platform engineering teams and is committed to openness, collaboration, and continuous growth.

$85,000–$100,000/yr
Global

  • Work directly with customers to onboard BYO-BGP and troubleshoot network performance issues.
  • Research network events to identify technical debt and drive improvements across platforms.
  • Compose, review, and test procedure documentation for scheduled maintenance to improve customer experience.

Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. As the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, Vultr emphasizes employee care with comprehensive benefits and a culture of inclusion.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

  • Lead and mentor a high-performing field engineering team, defining deployment workflows and integration playbooks for repeatability and reliability.
  • Set technical strategy for field integrations, implementing scalable solutions and improving field engineering tooling with scripts and automation.
  • Partner across engineering, product, security, and mission operations to ensure secure, reliable deployments in customer-owned environments.

TurbineOne builds Mission-AI for the Frontlines, providing a Frontline Perception System that helps military and national security operators detect threats and accelerate decision-making at the tactical edge. The team is composed of experienced technologists, veterans, and operators committed to advancing national security through responsible innovation.

North America Unlimited PTO

  • Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against quality metrics.
  • Define the qualification bar: build the acceptance test suite, benchmark methodology, and quality thresholds to formalize tribal knowledge.
  • Guide providers through technical onboarding: work with their engineers to remediate gaps and bring clusters onto the network cleanly.

Andromeda Cluster provides early-stage startups access to scaled AI infrastructure that was once reserved for hyperscalers. It is a small, high-growth team at the center of the AI infrastructure boom, founded by Nat Friedman and Daniel Gross.