Similar Jobs

See all

About the Role:

  • This is a forward deployed SRE role, not generalist or support, embedding with teams running large-scale training and inference.
  • You'll own onboarding, diagnose failures, and improve performance across GPU clusters.

What You'll Do:

  • Serve as primary technical contact, working inside customer environments to fix issues like NCCL timeouts and degraded links.
  • Build automation for cluster provisioning, health checks, and preflight validation.
  • Lead incident response and ensure reliability outcomes for your accounts.

Why You'll Love It Here:

  • High-growth environment at the center of AI infrastructure boom, with early involvement.
  • Competitive compensation with meaningful equity, and comprehensive benefits including unlimited PTO.

Andromeda Cluster

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

Apply for This Position