Source Job

US

  • Design, operate, and troubleshoot multi-tenant SDN, VPC/overlays, and east-west/north-south traffic.
  • Build and tune InfiniBand/RoCE fabrics for GPU training, including lossless config and congestion control.
  • Drive network-provisioning automation and define architecture standards with compute and control plane teams.

InfiniBand Network Automation

14 jobs similar to Cloud Network / SDN Engineer

Jobs ranked by similarity.

US

  • Lead end-to-end architecture design of AI Data Center Networks, high-performance DCI, and global backbone networks for large-scale GPU clusters.
  • Collaborate with NVIDIA, vendors, and partners to translate business requirements into top-level network designs including InfiniBand/RoCE, Spine-Leaf, and overlay integration.
  • Own congestion control tuning, produce architecture documentation, and identify risks to drive network evolution.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company operates data centers across multiple countries and is building AI computational infrastructure.

US

  • Own the north-south fabric connecting AI cloud to the world, including EVPN-VxLAN fabrics, DCI, and WAN infrastructure.
  • Drive automation to turn network operations from tickets into policy using Ansible, Terraform, and Nautobot.
  • Ensure network reliability by integrating with AIOps substrate for predictive stability and automated remediation.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure, providing comprehensive mining solutions and AI computational infrastructure. Headquartered in Singapore, it has deployed data centers across the United States, Norway, Bhutan, and Ethiopia.

Europe

  • Design and operate high-performance GPU networking fabrics for AI workloads (RoCE/RDMA, Spectrum-X).
  • Own datacenter and edge networking architecture, security, and inter-datacenter backbone.
  • Lead automation, observability, and incident response while mentoring adjacent teams.

Radian Arc is a fast-growing scale-up at the intersection of AI, cloud infrastructure, GPU computing, and telecommunications. The company offers a lean, remote-first, and international culture with significant technical ownership and career growth opportunities.

$185,000–$225,000/yr
US Unlimited PTO

  • Own the architecture, design, deployment, and operational lifecycle of corporate LAN/WAN, wireless, OT core, and physical security networks across global data center campuses.
  • Lead network planning and execution from due diligence through commissioning, including design documentation, scope definition, standards development, and multi-site deployments.
  • Partner with engineering, construction, operations, vendors, and contractors to integrate network infrastructure with data center systems while ensuring compliance, quality, safety, and security.

Fleet Data Centers designs, builds and operates mega-scale data center campuses. The company is led by industry veterans and is headquartered in Denver, Colorado, with satellite offices in Seattle and Arlington.

$180,000–$250,000/yr
Global

  • Design, build, validate, and operate public and private networks.
  • Use AI to automate network provisioning, configuration, and monitoring.
  • Manage BGP routing, IP address allocation, and private peering.

fal is a generative media ecosystem building infrastructure, tools, and model access for AI products. The company provides a unified platform for high-performance inference, orchestration, and observability, enabling teams to scale from idea to production.

US Unlimited PTO

  • Build and own the end-to-end zero-touch provisioning (ZTP) and automation platform for GPU network fabrics.
  • Lead a small team of engineers and SREs, ensuring clean integration with broader platform tooling.
  • Bring software engineering rigor to network automation including code review, testing, CI/CD, and release management.

TensorWave delivers seamless, secure, reliable, and resilient AI compute at scale by building a versatile cloud platform that eliminates infrastructure barriers. They empower builders to focus on innovation instead of fighting their stack, and are committed to creating an inclusive environment for all employees.

Global

  • Support the deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
  • Assist in troubleshooting network issues, including connectivity, latency, and performance degradation.
  • Collaborate with compute and storage teams to support HPC and AI workloads.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. They serve many of the world’s leading enterprises and are committed to open standards and freedom from lock-in.

Europe

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

Mirantis is a Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They serve many leading enterprises including Adobe, DocuSign, and PayPal, and are a leader in container management.

  • Own the architecture and technical roadmap for Mount Thor’s production network.
  • Design and operate data center fabrics, site connectivity, and host networking.
  • Build software for network provisioning, configuration, and lifecycle management.

Mount Thor builds an infrastructure platform to make Apple hardware available at datacenter scale for AI workloads. The company operates as a startup providing elastic compute solutions.

Germany

  • Design and develop core networking components for a large-scale Virtual Private Cloud, including control-plane and data-plane systems.
  • Improve scalability, performance, and resilience of critical network services, implementing capabilities like IPv6, VPC peering, and load balancing.
  • Collaborate with engineering teams to deliver reliable, secure networking and participate in design reviews and coding interviews.

Our partner is a leading cloud infrastructure platform, enabling large-scale AI and cloud services. They foster a highly technical, international engineering environment with a trust-based culture and strong emphasis on ownership and continuous improvement.

India

  • Act as a subject matter expert for SDN technologies, leading design and deployment of data center network solutions.
  • Work directly with technical teams and senior stakeholders to translate business requirements into scalable network architecture.
  • Maintain strong focus on customer success, proactively identifying risks and opportunities for improvement.

The company is a technology services firm specializing in Software Defined Networking solutions for enterprise data centers. It fosters a collaborative, remote-friendly culture and values continuous learning and professional development.

$200,000–$230,000/yr
Global

  • Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
  • Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
  • Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.

Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. It is a privately-held company valued at $3.5 billion, trusted by hundreds of thousands of customers across 185 countries, and known for its self-funded growth and inclusive culture.

Bulgaria

  • Design, implement, and support enterprise LAN, WAN, WLAN, SD-WAN, and cloud networking solutions.
  • Provide Tier 2/3 support for complex network incidents and outages, ensuring network availability and performance.
  • Lead technical aspects of network projects, mentor junior engineers, and collaborate with cybersecurity teams.

Sutherland is a global provider of business process and technology management services, serving clients across various industries. With over 60,000 employees worldwide, we foster a culture of innovation, collaboration, and continuous improvement.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.