Design, implement, and operate high-performance GPU networking fabrics and classical datacenter networking components such as routing, security, and external connectivity.
Own the long-term technical direction and operational strategy for AI interconnect networks, architecting scalable RoCE and Ethernet fabrics for distributed training and inference.
Collaborate cross-functionally with infrastructure, platform, SRE, and operations teams to integrate networking into the overall platform architecture and drive operational excellence.
Define and own the architectural roadmap for enterprise, data center, and cloud networks.
Co-author secure design and lifecycle management of high-performance data center networking.
Build an AI-agent based analysis framework for network changes and self-improvement.
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, delivering industry-leading training and inference speeds. The company partners with top model labs, global enterprises, and AI-native startups, including OpenAI, and has a non-corporate work culture that values continuous learning and growth.
Serve as a subject matter expert for NVIDIA Networking technologies including Ethernet and InfiniBand-based data center and AI fabric solutions.
Design and deploy complex solutions, create written deliverables, and conduct client workshops while communicating architecture strategy to senior management.
Provide technical leadership for data center modernization, AI infrastructure networking, and high-performance network design initiatives.
AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery to help enterprises deliver on digital transformation. The company prioritizes a culture of belonging where all perspectives are valued and is an equal opportunity employer committed to diversity.
Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion.
Tune fabric performance parameters for distributed AI workloads such as NCCL and MPI.
Collaborate with GPU and networking teams to diagnose and resolve fabric-level issues and optimize performance.
Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is the largest privately-held cloud infrastructure company, offering a culture of innovation and growth.
Own the vision, roadmap, and priorities for k0rdent AI networking, spanning underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and network automation.
Translate requirements from GPU clouds, telcos, and enterprise platform teams into clear product direction, partnering with engineering to define requirements.
Track and shape response to emerging interconnect standards like Ultra Ethernet, UALink, and congestion control, and represent Mirantis with customers and partners.
Mirantis is the leading AI-infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They are a global, distributed team committed to openness and technical excellence, serving clients like Adobe, PayPal, and Volkswagen.
Define the networking strategy and roadmap for k0rdent AI, covering GPU cluster networking and multi-tenant cloud integration.
Partner with engineering and marketing to shape requirements, positioning, and competitive differentiation.
Represent Mirantis at events and with strategic accounts, driving product success in the AI cloud era.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With a world-class, distributed team, Mirantis empowers platform engineering teams and is committed to openness, collaboration, and continuous growth.
Design, deploy, and maintain high-performance network infrastructures for HPC environments with a focus on InfiniBand fabrics.
Troubleshoot complex network issues across InfiniBand and Ethernet, manage Fortinet solutions, and perform performance tuning.
Collaborate with compute, storage, and platform teams to support HPC workloads and document network architecture and procedures.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises including Adobe, DocuSign, PayPal, and Volkswagen, fostering a collaborative and innovative work environment.
Plan, implement, and manage data center layouts and power density for maximum efficiency and reliability.
Oversee server capacity deployments and integrate infrastructure with automation to support global scalability.
Collaborate with cross-functional teams and provide expert technical guidance for complex data center infrastructure issues.
Zscaler accelerates digital transformation by providing a cloud-native Zero Trust Exchange platform that protects customers from cyberattacks and data loss. The company fosters a culture of innovation, execution, and collaboration, with a focus on impact and accountability.
Work directly with customers to onboard BYO-BGP and troubleshoot network performance issues.
Research network events to identify technical debt and drive improvements across platforms.
Compose, review, and test procedure documentation for scheduled maintenance to improve customer experience.
Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. As the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, Vultr emphasizes employee care with comprehensive benefits and a culture of inclusion.
Support deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
Troubleshoot network issues including connectivity, latency, and performance degradation.
Collaborate with compute and storage teams to support HPC and AI workloads.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Serving leading enterprises like Adobe, PayPal, and Volkswagen, the company fosters a culture of open standards and engineering excellence.
Design and operate firewall, segmentation, and zero-trust controls across data center, corporate, and cloud (AWS) networks.
Build and maintain network security infrastructure as code, including firewall rules, policy automation, and CI/CD-driven deployment.
Lead network lifecycle management with design review, configuration baselines, change automation, and ongoing rule hygiene.
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, using a novel wafer-scale architecture. The company serves top model labs, global enterprises, and AI-native startups, with a non-corporate culture that respects individual beliefs and fosters continuous learning.
Design and implement enterprise-scale Network-as-Code and GitOps solutions using Terraform, Ansible, and Python.
Build and optimize CI/CD pipelines for network infrastructure with automated testing and zero-downtime deployments.
Lead design and troubleshooting of large-scale network architectures including BGP, OSPF, and Kubernetes networking.
Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises. With nearly 1000 full-time professionals and a 99% project success rate, the company operates in over 25 countries with a culture of Relentless Performance.
Design and develop advanced computing architectures for AI accelerators.
Analyze and optimize computational workloads for low power and high performance.
Collaborate with cross-functional teams to translate system requirements into architectural solutions.
Axelera AI develops next-generation AI platforms to advance humanity. With 220+ employees and $370 million raised, they foster a collaborative, innovative culture.
Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.
Design, implement, and support enterprise network infrastructures including campus, data center, WAN, cloud, and security.
Serve as senior technical resource providing architecture leadership, implementation expertise, and customer consulting.
Lead complex infrastructure projects while mentoring engineers and promoting technical excellence.
Myriad360 challenges and enables employees to achieve great things through a culture of empathy, ownership, and continuous improvement. They are consistently listed among Inc & Crain's 'Best Places to Work' with an accessible executive team and a commitment to inclusion.