Remote Information technology Jobs · InfiniBand

Job listings

  • Lead end-to-end architecture design of AI Data Center Networks, high-performance DCI, and global backbone networks for large-scale GPU clusters.
  • Collaborate with NVIDIA, vendors, and partners to translate business requirements into top-level network designs including InfiniBand/RoCE, Spine-Leaf, and overlay integration.
  • Own congestion control tuning, produce architecture documentation, and identify risks to drive network evolution.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company operates data centers across multiple countries and is building AI computational infrastructure.

  • Support the deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
  • Assist in troubleshooting network issues, including connectivity, latency, and performance degradation.
  • Collaborate with compute and storage teams to support HPC and AI workloads.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. They serve many of the world’s leading enterprises and are committed to open standards and freedom from lock-in.

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

Mirantis is a Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They serve many leading enterprises including Adobe, DocuSign, and PayPal, and are a leader in container management.

North America Unlimited PTO

  • Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against quality metrics.
  • Define the qualification bar: build the acceptance test suite, benchmark methodology, and quality thresholds to formalize tribal knowledge.
  • Guide providers through technical onboarding: work with their engineers to remediate gaps and bring clusters onto the network cleanly.

Andromeda Cluster provides early-stage startups access to scaled AI infrastructure that was once reserved for hyperscalers. It is a small, high-growth team at the center of the AI infrastructure boom, founded by Nat Friedman and Daniel Gross.