Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
Mirantis is a Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They serve many leading enterprises including Adobe, DocuSign, and PayPal, and are a leader in container management.
Design, implement, and operate high-performance GPU networking fabrics and classical datacenter networking components such as routing, security, and external connectivity.
Own the long-term technical direction and operational strategy for AI interconnect networks, architecting scalable RoCE and Ethernet fabrics for distributed training and inference.
Collaborate cross-functionally with infrastructure, platform, SRE, and operations teams to integrate networking into the overall platform architecture and drive operational excellence.
Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. It is a fast-growing scale-up with a global team across the USA, Australia, Central Europe, Malaysia, Singapore and Japan.
Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion.
Tune fabric performance parameters for distributed AI workloads such as NCCL and MPI.
Collaborate with GPU and networking teams to diagnose and resolve fabric-level issues and optimize performance.
Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is the largest privately-held cloud infrastructure company, offering a culture of innovation and growth.
Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.
Serve as a subject matter expert for NVIDIA Networking technologies including Ethernet and InfiniBand-based data center and AI fabric solutions.
Design and deploy complex solutions, create written deliverables, and conduct client workshops while communicating architecture strategy to senior management.
Provide technical leadership for data center modernization, AI infrastructure networking, and high-performance network design initiatives.
AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery to help enterprises deliver on digital transformation. The company prioritizes a culture of belonging where all perspectives are valued and is an equal opportunity employer committed to diversity.
Work directly with customers to onboard BYO-BGP and troubleshoot network performance issues.
Research network events to identify technical debt and drive improvements across platforms.
Compose, review, and test procedure documentation for scheduled maintenance to improve customer experience.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. As the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, Vultr emphasizes employee care with comprehensive benefits and a culture of inclusion.
Own the vision, roadmap, and priorities for k0rdent AI networking, spanning underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and network automation.
Translate requirements from GPU clouds, telcos, and enterprise platform teams into clear product direction, partnering with engineering to define requirements.
Track and shape response to emerging interconnect standards like Ultra Ethernet, UALink, and congestion control, and represent Mirantis with customers and partners.
Mirantis is the leading AI-infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They are a global, distributed team committed to openness and technical excellence, serving clients like Adobe, PayPal, and Volkswagen.
Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
Drive reliability, monitoring, automation, and incident response for AI infrastructure.
Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.
Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.
Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.
Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.
Define the networking strategy and roadmap for k0rdent AI, covering GPU cluster networking and multi-tenant cloud integration.
Partner with engineering and marketing to shape requirements, positioning, and competitive differentiation.
Represent Mirantis at events and with strategic accounts, driving product success in the AI cloud era.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With a world-class, distributed team, Mirantis empowers platform engineering teams and is committed to openness, collaboration, and continuous growth.
Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.
Lead and mentor a high-performing field engineering team, defining deployment workflows and integration playbooks for repeatability and reliability.
Set technical strategy for field integrations, implementing scalable solutions and improving field engineering tooling with scripts and automation.
Partner across engineering, product, security, and mission operations to ensure secure, reliable deployments in customer-owned environments.
TurbineOne builds Mission-AI for the Frontlines, providing a Frontline Perception System that helps military and national security operators detect threats and accelerate decision-making at the tactical edge. The team is composed of experienced technologists, veterans, and operators committed to advancing national security through responsible innovation.
Design and implement core networking services for Docker's Sandboxes platform across local and public cloud environments.
Build scalable networking infrastructure for microVM orchestration, workload scheduling, and lifecycle management.
Develop high-performance networking components and ensure reliability, observability, and performance across the infrastructure.
Docker is a developer tooling platform trusted by over 20 million monthly users and 20 billion container image pulls, enabling developers to build, share, and run applications. The company is a globally distributed, remote-first team with offices in Seattle and Paris, focused on innovation in AI and security.
Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.
Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.
Design, implement, and support enterprise LAN, WAN, WLAN, SD-WAN, and cloud networking solutions.
Provide Tier 2/3 support for complex network incidents and outages, ensuring network availability and performance.
Lead technical aspects of network projects, mentor junior engineers, and collaborate with cybersecurity teams.
Sutherland is a global provider of business process and technology management services, serving clients across various industries. With over 60,000 employees worldwide, we foster a culture of innovation, collaboration, and continuous improvement.
Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Own high-severity technical escalations from intake through resolution or engineering handoff.
Partner with Support and Product/Engineering to close the gap between customer problems and engineering fixes.
Monitor patterns across escalations to catch systemic issues and translate into product improvements.
Tailscale builds software that makes it easy to securely interconnect people and devices. Founded in 2019, the company is fully distributed and backed by Accel, CRV, and others.