Source Job

North America Unlimited PTO

  • Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against quality metrics.
  • Define the qualification bar: build the acceptance test suite, benchmark methodology, and quality thresholds to formalize tribal knowledge.
  • Guide providers through technical onboarding: work with their engineers to remediate gaps and bring clusters onto the network cleanly.

HPC GPU Clusters InfiniBand Slurm Benchmarking

8 jobs similar to HPC Architect

Jobs ranked by similarity.

$125,000–$135,000/yr
Global

  • Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion.
  • Tune fabric performance parameters for distributed AI workloads such as NCCL and MPI.
  • Collaborate with GPU and networking teams to diagnose and resolve fabric-level issues and optimize performance.

Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is the largest privately-held cloud infrastructure company, offering a culture of innovation and growth.

Global Unlimited PTO

  • Own the technical evaluation end-to-end, from discovery to POC, ensuring evaluations are scoped and tied to customer ROI.
  • Take customers from signature to first successful production training run and serve as the technical owner post-launch.
  • Build the SA function by creating demo environments, benchmarking harnesses, and reference architectures.

Andromeda Cluster gives early-stage startups access to scaled AI infrastructure, partnering with leading AI labs and cloud providers. It is a high-growth company building an inclusive environment for all employees.

$225,000–$325,000/yr
Global Unlimited PTO

  • Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
  • Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
  • Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

$150,000–$160,000/yr

  • Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
  • Oversee cluster reliability, GPU fleet operations, and incident response.
  • Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.

Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

Europe

  • Design, implement, and operate high-performance GPU networking fabrics and classical datacenter networking components such as routing, security, and external connectivity.
  • Own the long-term technical direction and operational strategy for AI interconnect networks, architecting scalable RoCE and Ethernet fabrics for distributed training and inference.
  • Collaborate cross-functionally with infrastructure, platform, SRE, and operations teams to integrate networking into the overall platform architecture and drive operational excellence.

Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. It is a fast-growing scale-up with a global team across the USA, Australia, Central Europe, Malaysia, Singapore and Japan.