Source Job

Europe

  • Design, implement, and operate high-performance GPU networking fabrics and classical datacenter networking components such as routing, security, and external connectivity.
  • Own the long-term technical direction and operational strategy for AI interconnect networks, architecting scalable RoCE and Ethernet fabrics for distributed training and inference.
  • Collaborate cross-functionally with infrastructure, platform, SRE, and operations teams to integrate networking into the overall platform architecture and drive operational excellence.

BGP EVPN/VXLAN Python

14 jobs similar to Staff Network Engineer (AI Fabric, Datacenter and Edge Networking)

Jobs ranked by similarity.

US Canada

  • Define and own the architectural roadmap for enterprise, data center, and cloud networks.
  • Co-author secure design and lifecycle management of high-performance data center networking.
  • Build an AI-agent based analysis framework for network changes and self-improvement.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, delivering industry-leading training and inference speeds. The company partners with top model labs, global enterprises, and AI-native startups, including OpenAI, and has a non-corporate work culture that values continuous learning and growth.

US

  • Serve as a subject matter expert for NVIDIA Networking technologies including Ethernet and InfiniBand-based data center and AI fabric solutions.
  • Design and deploy complex solutions, create written deliverables, and conduct client workshops while communicating architecture strategy to senior management.
  • Provide technical leadership for data center modernization, AI infrastructure networking, and high-performance network design initiatives.

AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery to help enterprises deliver on digital transformation. The company prioritizes a culture of belonging where all perspectives are valued and is an equal opportunity employer committed to diversity.

$125,000–$135,000/yr
Global

  • Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion.
  • Tune fabric performance parameters for distributed AI workloads such as NCCL and MPI.
  • Collaborate with GPU and networking teams to diagnose and resolve fabric-level issues and optimize performance.

Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is the largest privately-held cloud infrastructure company, offering a culture of innovation and growth.

Global

  • Own the vision, roadmap, and priorities for k0rdent AI networking, spanning underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and network automation.
  • Translate requirements from GPU clouds, telcos, and enterprise platform teams into clear product direction, partnering with engineering to define requirements.
  • Track and shape response to emerging interconnect standards like Ultra Ethernet, UALink, and congestion control, and represent Mirantis with customers and partners.

Mirantis is the leading AI-infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They are a global, distributed team committed to openness and technical excellence, serving clients like Adobe, PayPal, and Volkswagen.

Global

  • Define the networking strategy and roadmap for k0rdent AI, covering GPU cluster networking and multi-tenant cloud integration.
  • Partner with engineering and marketing to shape requirements, positioning, and competitive differentiation.
  • Represent Mirantis at events and with strategic accounts, driving product success in the AI cloud era.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With a world-class, distributed team, Mirantis empowers platform engineering teams and is committed to openness, collaboration, and continuous growth.

EU

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments with a focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet, manage Fortinet solutions, and perform performance tuning.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and document network architecture and procedures.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises including Adobe, DocuSign, PayPal, and Volkswagen, fostering a collaborative and innovative work environment.

Netherlands

  • Plan, implement, and manage data center layouts and power density for maximum efficiency and reliability.
  • Oversee server capacity deployments and integrate infrastructure with automation to support global scalability.
  • Collaborate with cross-functional teams and provide expert technical guidance for complex data center infrastructure issues.

Zscaler accelerates digital transformation by providing a cloud-native Zero Trust Exchange platform that protects customers from cyberattacks and data loss. The company fosters a culture of innovation, execution, and collaboration, with a focus on impact and accountability.

$85,000–$100,000/yr
Global

  • Work directly with customers to onboard BYO-BGP and troubleshoot network performance issues.
  • Research network events to identify technical debt and drive improvements across platforms.
  • Compose, review, and test procedure documentation for scheduled maintenance to improve customer experience.

Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. As the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, Vultr emphasizes employee care with comprehensive benefits and a culture of inclusion.

Global

  • Support deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
  • Troubleshoot network issues including connectivity, latency, and performance degradation.
  • Collaborate with compute and storage teams to support HPC and AI workloads.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Serving leading enterprises like Adobe, PayPal, and Volkswagen, the company fosters a culture of open standards and engineering excellence.

North America

  • Design and operate firewall, segmentation, and zero-trust controls across data center, corporate, and cloud (AWS) networks.
  • Build and maintain network security infrastructure as code, including firewall rules, policy automation, and CI/CD-driven deployment.
  • Lead network lifecycle management with design review, configuration baselines, change automation, and ongoing rule hygiene.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, using a novel wafer-scale architecture. The company serves top model labs, global enterprises, and AI-native startups, with a non-corporate culture that respects individual beliefs and fosters continuous learning.

Global

  • Design and implement enterprise-scale Network-as-Code and GitOps solutions using Terraform, Ansible, and Python.
  • Build and optimize CI/CD pipelines for network infrastructure with automated testing and zero-downtime deployments.
  • Lead design and troubleshooting of large-scale network architectures including BGP, OSPF, and Kubernetes networking.

Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises. With nearly 1000 full-time professionals and a 99% project success rate, the company operates in over 25 countries with a culture of Relentless Performance.

Europe

  • Design and develop advanced computing architectures for AI accelerators.
  • Analyze and optimize computational workloads for low power and high performance.
  • Collaborate with cross-functional teams to translate system requirements into architectural solutions.

Axelera AI develops next-generation AI platforms to advance humanity. With 220+ employees and $370 million raised, they foster a collaborative, innovative culture.

$225,000–$325,000/yr
Global Unlimited PTO

  • Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
  • Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
  • Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.

$175,000–$195,000/yr
US Unlimited PTO

  • Design, implement, and support enterprise network infrastructures including campus, data center, WAN, cloud, and security.
  • Serve as senior technical resource providing architecture leadership, implementation expertise, and customer consulting.
  • Lead complex infrastructure projects while mentoring engineers and promoting technical excellence.

Myriad360 challenges and enables employees to achieve great things through a culture of empathy, ownership, and continuous improvement. They are consistently listed among Inc & Crain's 'Best Places to Work' with an accessible executive team and a commitment to inclusion.