Source Job

Global

  • Design, deploy, and operate large-scale Linux infrastructure including bare metal servers, enterprise storage, and GPU clusters for AI/ML workloads.
  • Manage and optimize AI Factory environments with NVIDIA GPU technologies such as A100, H100, and H200, including provisioning, monitoring, and performance tuning.
  • Ensure high availability and reliability through expert-level Linux administration, storage management with Ceph and high-performance platforms, and networking in data centers.

Linux Ceph NVIDIA GPU Networking

15 jobs similar to Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

Jobs ranked by similarity.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

Global

  • Support the deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
  • Assist in troubleshooting network issues, including connectivity, latency, and performance degradation.
  • Collaborate with compute and storage teams to support HPC and AI workloads.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. They serve many of the world’s leading enterprises and are committed to open standards and freedom from lock-in.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms.
  • Own the storage layer where Kubernetes meets bare metal, tuning NFS data paths for high-throughput workloads.
  • Automate storage provisioning with infrastructure-as-code and GitOps, ensuring observability and reliability.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams across hybrid, edge, and sovereign environments, fostering a culture of open-source innovation and collaboration among passionate, talented colleagues.

US

  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
  • Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
  • Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.

Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

India

  • Design end-to-end AI infrastructure solutions for scalable, high-performance AI and HPC environments.
  • Partner with Sales to qualify opportunities, conduct technical discovery, and serve as trusted advisor throughout the sales lifecycle.
  • Collaborate closely with Facilities, Delivery, OEM partners, and customer teams to ensure AI infrastructure aligns with data center capabilities.

Submer enables organizations scaling AI to overcome the limits of traditional datacenters in power, compute density, and efficiency. It is a fast-growing, international scale-up with a friendly, diverse, and hybrid-friendly work environment.

Europe

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

Mirantis is a Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They serve many leading enterprises including Adobe, DocuSign, and PayPal, and are a leader in container management.

$200,000–$230,000/yr
Global

  • Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
  • Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
  • Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.

Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. It is a privately-held company valued at $3.5 billion, trusted by hundreds of thousands of customers across 185 countries, and known for its self-funded growth and inclusive culture.

Global

  • Design, develop, and maintain automated diagnostic, validation, and remediation frameworks for production GPU hardware using Python and infrastructure automation tools.
  • Engineer and support Python-based agents, APIs, and Ansible automation for hardware provisioning, telemetry, and health monitoring.
  • Analyze workload performance, thermals, and diagnostic output to identify hardware issues and improve validation methodologies.

Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators worldwide. We are the world's largest privately-held cloud infrastructure company, self-funded for over a decade, with 33 global data centers and hundreds of thousands of active customers.

US

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.

Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.

$80,000–$100,000/yr
US

  • Build, maintain, and operate Vultr's Ceph-based Cloud Storage environments with a focus on stability and scalability.
  • Develop automation frameworks and optimize metrics collection and alerting systems.
  • Collaborate with cross-functional teams for capacity expansion, maintenance, and design decisions.

Vultr makes high-performance cloud infrastructure easy, affordable, and locally accessible for enterprises and AI innovators globally. As the world’s largest privately-held cloud infrastructure company, Vultr has over 30 data centers, hundreds of thousands of customers, and a $3.5 billion valuation.

US

  • Lead end-to-end architecture design of AI Data Center Networks, high-performance DCI, and global backbone networks for large-scale GPU clusters.
  • Collaborate with NVIDIA, vendors, and partners to translate business requirements into top-level network designs including InfiniBand/RoCE, Spine-Leaf, and overlay integration.
  • Own congestion control tuning, produce architecture documentation, and identify risks to drive network evolution.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company operates data centers across multiple countries and is building AI computational infrastructure.

$180,000–$250,000/yr
Global

  • Design, build, validate, and operate public and private networks.
  • Use AI to automate network provisioning, configuration, and monitoring.
  • Manage BGP routing, IP address allocation, and private peering.

fal is a generative media ecosystem building infrastructure, tools, and model access for AI products. The company provides a unified platform for high-performance inference, orchestration, and observability, enabling teams to scale from idea to production.

  • Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
  • Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
  • Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.