Similar Jobs

See all

The Role:

  • You own the technical relationship with compute providers, vetting their clusters and defining the quality bar.
  • You work closely with procurement and SREs to ensure only qualified clusters join the network.
  • You are the first HPC Architect, building the function from the ground up.

What You’ll Do:

  • Vet prospective providers and assess their architecture, hardware, and fabric against Andromeda's metrics.
  • Build the acceptance test suite and quality thresholds, formalizing what currently exists as tribal knowledge.
  • Guide providers through onboarding and remediation to bring clusters up to standard.

What We’re Looking For:

  • Deep HPC experience with GPU clusters at scale, strong fabric knowledge (InfiniBand and RoCE).
  • Experience with distributed orchestration (Slurm, Kubernetes) and data-center literacy.
  • The ability to write testable standards and deliver critical feedback to providers.

Andromeda Cluster

Andromeda Cluster provides early-stage startups access to scaled AI infrastructure that was once reserved for hyperscalers. It is a small, high-growth team at the center of the AI infrastructure boom, founded by Nat Friedman and Daniel Gross.

Apply for This Position