Lead end-to-end architecture design of AI Data Center Networks, high-performance DCI, and global backbone networks for large-scale GPU clusters.
Collaborate with NVIDIA, vendors, and partners to translate business requirements into top-level network designs including InfiniBand/RoCE, Spine-Leaf, and overlay integration.
Own congestion control tuning, produce architecture documentation, and identify risks to drive network evolution.
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company operates data centers across multiple countries and is building AI computational infrastructure.
Design, operate, and troubleshoot multi-tenant SDN, VPC/overlays, and east-west/north-south traffic.
Build and tune InfiniBand/RoCE fabrics for GPU training, including lossless config and congestion control.
Drive network-provisioning automation and define architecture standards with compute and control plane teams.
Bitdeer Technologies Group is a world-leading technology company for AI and Bitcoin mining infrastructure. It is a global team headquartered in Singapore with data centers across the US, Norway, Bhutan, and Ethiopia.
Deploy and commission GPU cloud infrastructure across regional and core datacenters, covering hardware, networking, storage, and platform software.
Act as a hands-on technical escalation point, troubleshooting complex issues across physical and software layers and driving incidents to resolution.
Establish deployment standards, validation procedures, documentation, and operational practices for a rapidly evolving AI infrastructure environment.
The company builds and operates cutting-edge GPU cloud and AI infrastructure at significant scale. It is a fast-growing international scale-up with a diverse and flexible working environment.
Support the deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
Assist in troubleshooting network issues, including connectivity, latency, and performance degradation.
Collaborate with compute and storage teams to support HPC and AI workloads.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. They serve many of the world’s leading enterprises and are committed to open standards and freedom from lock-in.
Design, build, validate, and operate public and private networks.
Use AI to automate network provisioning, configuration, and monitoring.
Manage BGP routing, IP address allocation, and private peering.
fal is a generative media ecosystem building infrastructure, tools, and model access for AI products. The company provides a unified platform for high-performance inference, orchestration, and observability, enabling teams to scale from idea to production.
Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
Mirantis is a Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They serve many leading enterprises including Adobe, DocuSign, and PayPal, and are a leader in container management.
Build and own the end-to-end zero-touch provisioning (ZTP) and automation platform for GPU network fabrics.
Lead a small team of engineers and SREs, ensuring clean integration with broader platform tooling.
Bring software engineering rigor to network automation including code review, testing, CI/CD, and release management.
TensorWave delivers seamless, secure, reliable, and resilient AI compute at scale by building a versatile cloud platform that eliminates infrastructure barriers. They empower builders to focus on innovation instead of fighting their stack, and are committed to creating an inclusive environment for all employees.
Own the architecture, design, deployment, and operational lifecycle of corporate LAN/WAN, wireless, OT core, and physical security networks across global data center campuses.
Lead network planning and execution from due diligence through commissioning, including design documentation, scope definition, standards development, and multi-site deployments.
Partner with engineering, construction, operations, vendors, and contractors to integrate network infrastructure with data center systems while ensuring compliance, quality, safety, and security.
Fleet Data Centers designs, builds and operates mega-scale data center campuses. The company is led by industry veterans and is headquartered in Denver, Colorado, with satellite offices in Seattle and Arlington.
Own the north-south fabric connecting AI cloud to the world, including EVPN-VxLAN fabrics, DCI, and WAN infrastructure.
Drive automation to turn network operations from tickets into policy using Ansible, Terraform, and Nautobot.
Ensure network reliability by integrating with AIOps substrate for predictive stability and automated remediation.
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure, providing comprehensive mining solutions and AI computational infrastructure. Headquartered in Singapore, it has deployed data centers across the United States, Norway, Bhutan, and Ethiopia.
Own the architecture and technical roadmap for Mount Thor’s production network.
Design and operate data center fabrics, site connectivity, and host networking.
Build software for network provisioning, configuration, and lifecycle management.
Mount Thor builds an infrastructure platform to make Apple hardware available at datacenter scale for AI workloads. The company operates as a startup providing elastic compute solutions.
Design scalable storage architectures for edge and core GPU deployments using platforms like StorPool, NVMe, VAST Data, and Weka.
Optimize storage for AI workloads including distributed training, fine-tuning, and inference with GPU Direct Storage and RDMA.
Act as primary storage design authority, influencing platform architecture and mentoring engineers across infrastructure domains.
Radian Arc builds scalable storage architectures powering AI and GPU infrastructure across edge and core environments. They are a growing company with a remote-friendly work model and an inclusive environment focused on next-generation AI infrastructure.
Define and drive innovative technical vision for intelligent networking and agentic AI platforms.
Design and build high-performance, production-ready services for real-time data and AI processing.
Mentor engineers, lead technical discussions, and uphold engineering excellence across the team.
We provide end-to-end, cloud-driven networking solutions trusted by over 50,000 customers globally. With double-digit growth and a culture of inclusion, we foster an innovative workplace where all employees thrive.
Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.
Design, deploy, and operate large-scale Linux infrastructure including bare metal servers, enterprise storage, and GPU clusters for AI/ML workloads.
Manage and optimize AI Factory environments with NVIDIA GPU technologies such as A100, H100, and H200, including provisioning, monitoring, and performance tuning.
Ensure high availability and reliability through expert-level Linux administration, storage management with Ceph and high-performance platforms, and networking in data centers.
The company is seeking a senior Linux infrastructure engineer with expertise in bare metal, storage, and AI Factory platforms. The size, employees, and culture are not specified.
Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. It is a privately-held company valued at $3.5 billion, trusted by hundreds of thousands of customers across 185 countries, and known for its self-funded growth and inclusive culture.
Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.
Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.