Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.
Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.
Design and operate high-performance GPU networking fabrics for AI workloads (RoCE/RDMA, Spectrum-X).
Own datacenter and edge networking architecture, security, and inter-datacenter backbone.
Lead automation, observability, and incident response while mentoring adjacent teams.
Radian Arc is a fast-growing scale-up at the intersection of AI, cloud infrastructure, GPU computing, and telecommunications. The company offers a lean, remote-first, and international culture with significant technical ownership and career growth opportunities.
Design end-to-end AI infrastructure solutions for scalable, high-performance AI and HPC environments.
Partner with Sales to qualify opportunities, conduct technical discovery, and serve as trusted advisor throughout the sales lifecycle.
Collaborate closely with Facilities, Delivery, OEM partners, and customer teams to ensure AI infrastructure aligns with data center capabilities.
Submer enables organizations scaling AI to overcome the limits of traditional datacenters in power, compute density, and efficiency. It is a fast-growing, international scale-up with a friendly, diverse, and hybrid-friendly work environment.
Design scalable storage architectures for edge and core GPU deployments using platforms like StorPool, NVMe, VAST Data, and Weka.
Optimize storage for AI workloads including distributed training, fine-tuning, and inference with GPU Direct Storage and RDMA.
Act as primary storage design authority, influencing platform architecture and mentoring engineers across infrastructure domains.
Radian Arc builds scalable storage architectures powering AI and GPU infrastructure across edge and core environments. They are a growing company with a remote-friendly work model and an inclusive environment focused on next-generation AI infrastructure.
Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.
Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.
Design, deploy, and operate large-scale Linux infrastructure including bare metal servers, enterprise storage, and GPU clusters for AI/ML workloads.
Manage and optimize AI Factory environments with NVIDIA GPU technologies such as A100, H100, and H200, including provisioning, monitoring, and performance tuning.
Ensure high availability and reliability through expert-level Linux administration, storage management with Ceph and high-performance platforms, and networking in data centers.
The company is seeking a senior Linux infrastructure engineer with expertise in bare metal, storage, and AI Factory platforms. The size, employees, and culture are not specified.
Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.
Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.
Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. It is a privately-held company valued at $3.5 billion, trusted by hundreds of thousands of customers across 185 countries, and known for its self-funded growth and inclusive culture.
Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.
Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
Mirantis is a Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They serve many leading enterprises including Adobe, DocuSign, and PayPal, and are a leader in container management.
Lead end-to-end architecture design of AI Data Center Networks, high-performance DCI, and global backbone networks for large-scale GPU clusters.
Collaborate with NVIDIA, vendors, and partners to translate business requirements into top-level network designs including InfiniBand/RoCE, Spine-Leaf, and overlay integration.
Own congestion control tuning, produce architecture documentation, and identify risks to drive network evolution.
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company operates data centers across multiple countries and is building AI computational infrastructure.
Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.
Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.
Support the deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
Assist in troubleshooting network issues, including connectivity, latency, and performance degradation.
Collaborate with compute and storage teams to support HPC and AI workloads.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. They serve many of the world’s leading enterprises and are committed to open standards and freedom from lock-in.
Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.
Own end-to-end infrastructure deployment programs for new capacity and site expansions, including hardware dependencies and commissioning gate frameworks.
Deliver crisp, data-driven executive updates and govern cross-organizational dependencies without escalation.
Coach junior TPMs and drive AI tool integration to improve program tracking and risk detection.
Evergrid builds frontier-level AI infrastructure for the fourth industrial revolution, designing, deploying, and operating large-scale AI systems. With over 1GW of capacity in active development and a 24/7 engineering team, they foster a culture of high conviction and high trust.
Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.
Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.
Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.
Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.
The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.