Source Job

US

  • Deploy and integrate high-performance NFS-based storage into Kubernetes clusters via CSI for GPU-accelerated workloads.
  • Automate storage provisioning and monitoring using infrastructure-as-code tools like Terraform and GitOps pipelines.
  • Tune Linux and network settings to optimize throughput and latency for demanding AI and machine learning applications.

Kubernetes Linux Infrastructure As Code

17 jobs similar to Senior Software Systems Engineer (Storage)

Jobs ranked by similarity.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms.
  • Own the storage layer where Kubernetes meets bare metal, tuning NFS data paths for high-throughput workloads.
  • Automate storage provisioning with infrastructure-as-code and GitOps, ensuring observability and reliability.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams across hybrid, edge, and sovereign environments, fostering a culture of open-source innovation and collaboration among passionate, talented colleagues.

US

  • Design and build control-plane services and drivers for storage integration with Kubernetes-based AI workloads.
  • Write production-quality Go code with strong testing and operational rigor.
  • Deliver storage integration for k0s-based Kubernetes via Cluster API and K0rdent topologies.

Mirantis is a Kubernetes-native AI infrastructure company that builds scalable, secure infrastructure for AI and data-intensive applications. It is committed to open standards and freedom from lock-in, empowering platform engineering teams.

Global

  • Design, deploy, and operate large-scale Linux infrastructure including bare metal servers, enterprise storage, and GPU clusters for AI/ML workloads.
  • Manage and optimize AI Factory environments with NVIDIA GPU technologies such as A100, H100, and H200, including provisioning, monitoring, and performance tuning.
  • Ensure high availability and reliability through expert-level Linux administration, storage management with Ceph and high-performance platforms, and networking in data centers.

The company is seeking a senior Linux infrastructure engineer with expertise in bare metal, storage, and AI Factory platforms. The size, employees, and culture are not specified.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

$80,000–$100,000/yr
US

  • Build, maintain, and operate Vultr's Ceph-based Cloud Storage environments with a focus on stability and scalability.
  • Develop automation frameworks and optimize metrics collection and alerting systems.
  • Collaborate with cross-functional teams for capacity expansion, maintenance, and design decisions.

Vultr makes high-performance cloud infrastructure easy, affordable, and locally accessible for enterprises and AI innovators globally. As the world’s largest privately-held cloud infrastructure company, Vultr has over 30 data centers, hundreds of thousands of customers, and a $3.5 billion valuation.

  • Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
  • Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
  • Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.

US

  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
  • Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
  • Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.

Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

Global Unlimited PTO

  • Operate the Monad node fleet, including health, sync, upgrades, and incident response for validators, full nodes, and archive nodes.
  • Own infrastructure-as-code with Ansible, Terraform, and Kubernetes, and build observability with Prometheus, Grafana, and Loki.
  • Design and build AI agent tooling for automated operations, including runbooks-as-code and deterministic guardrails.

Category Labs designs and builds decentralized technology, including the Monad blockchain, a high-performance EVM-compatible Layer 1. The team raised $225M in series A funding and is a lean, collaborative group of engineers and researchers with a culture of low ego and high-quality output.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

$100,000–$120,000/yr
Europe 6w PTO

  • Manage physical "Metal" environments from bare metal to Kubernetes, including cluster networking and scheduling.
  • Maintain Crossplane compositions and Terraform modules for cloud service provider resources.
  • Work with application teams to understand needs and invest in right capabilities.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations. They are a 100% remote company with team members across 40+ countries, backed by leading investors, and known for an open-source legacy and global collaborative culture.

North America

  • Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
  • Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
  • Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.

Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.

US

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.

Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.

US

  • Design and maintain CI/CD and MLOps pipelines for AI and software applications, ensuring seamless deployment and automation.
  • Build and scale cloud-native infrastructure using Kubernetes, Docker, and GPU clusters to support high-performance AI workloads.
  • Champion Infrastructure as Code and observability practices to ensure high availability, security, and compliance across multi-cloud environments.

Bitdeer is a world-leading technology company providing AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company has a global presence with data centers in multiple countries and a culture focused on innovation and reliability.

Canada

  • Define and lead platform engineering strategy across complex, multi-environment cloud systems.
  • Architect scalable Kubernetes platforms, own IaC standards, and drive DevSecOps implementation.
  • Mentor engineers, partner with leadership on infrastructure direction, and lead complex migrations.

Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, with delivery centers in Canada, the US, Eastern Europe, and Latin America, we are a nimble team of senior engineers averaging 15+ years of experience.

UK

  • Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
  • Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
  • Drive AI-specific observability, FinOps, and security practices across the platform.

We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.