Source Job

Europe

  • Design scalable storage architectures for edge and core GPU deployments using platforms like StorPool, NVMe, VAST Data, and Weka.
  • Optimize storage for AI workloads including distributed training, fine-tuning, and inference with GPU Direct Storage and RDMA.
  • Act as primary storage design authority, influencing platform architecture and mentoring engineers across infrastructure domains.

AI Infrastructure Kubernetes Linux

20 jobs similar to Staff Storage Platform Engineer (AI Storage)

Jobs ranked by similarity.

Europe

  • Deploy and commission GPU cloud infrastructure across regional and core datacenters, covering hardware, networking, storage, and platform software.
  • Act as a hands-on technical escalation point, troubleshooting complex issues across physical and software layers and driving incidents to resolution.
  • Establish deployment standards, validation procedures, documentation, and operational practices for a rapidly evolving AI infrastructure environment.

The company builds and operates cutting-edge GPU cloud and AI infrastructure at significant scale. It is a fast-growing international scale-up with a diverse and flexible working environment.

India

  • Design end-to-end AI infrastructure solutions for scalable, high-performance AI and HPC environments.
  • Partner with Sales to qualify opportunities, conduct technical discovery, and serve as trusted advisor throughout the sales lifecycle.
  • Collaborate closely with Facilities, Delivery, OEM partners, and customer teams to ensure AI infrastructure aligns with data center capabilities.

Submer enables organizations scaling AI to overcome the limits of traditional datacenters in power, compute density, and efficiency. It is a fast-growing, international scale-up with a friendly, diverse, and hybrid-friendly work environment.

US

  • Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms.
  • Own the storage layer where Kubernetes meets bare metal, tuning NFS data paths for high-throughput workloads.
  • Automate storage provisioning with infrastructure-as-code and GitOps, ensuring observability and reliability.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams across hybrid, edge, and sovereign environments, fostering a culture of open-source innovation and collaboration among passionate, talented colleagues.

US

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.

Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.

US Unlimited PTO

  • Partner with the business-facing GSI leader to define and execute MinIO's global GSI strategy.
  • Architect and evangelize end-to-end AI Factory reference architectures spanning compute, GPU, Kubernetes, storage, and AI frameworks.
  • Influence C-level executives and lead technical workshops with strategic partners like NVIDIA and Databricks.

MinIO provides the data and memory foundation for enterprise AI, offering AIStor and MemKV to unify data across core, edge, and cloud. Trusted by 77% of the Fortune 100, the company is redefining how AI factories secure, persist, and unlock data value.

Europe

  • Lead and mentor a small team of engineers focused on AI platform and product engineering.
  • Drive architectural direction for AI product with emphasis on inference performance and cost efficiency.
  • Contribute hands-on to design and optimization of core AI systems including model serving and inference pipelines.

They operate a cloud and infrastructure platform. The team is small and high-leverage, focusing on engineering excellence and ownership.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

US

  • Deploy and integrate high-performance NFS-based storage into Kubernetes clusters via CSI for GPU-accelerated workloads.
  • Automate storage provisioning and monitoring using infrastructure-as-code tools like Terraform and GitOps pipelines.
  • Tune Linux and network settings to optimize throughput and latency for demanding AI and machine learning applications.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build scalable, secure infrastructure for AI and data-intensive workloads. With deep expertise in open source and Kubernetes orchestration, they enable platform engineering teams across on-premises, cloud, edge, and sovereign environments.

Europe

  • Design and operate high-performance GPU networking fabrics for AI workloads (RoCE/RDMA, Spectrum-X).
  • Own datacenter and edge networking architecture, security, and inter-datacenter backbone.
  • Lead automation, observability, and incident response while mentoring adjacent teams.

Radian Arc is a fast-growing scale-up at the intersection of AI, cloud infrastructure, GPU computing, and telecommunications. The company offers a lean, remote-first, and international culture with significant technical ownership and career growth opportunities.

Global

  • Design, deploy, and operate large-scale Linux infrastructure including bare metal servers, enterprise storage, and GPU clusters for AI/ML workloads.
  • Manage and optimize AI Factory environments with NVIDIA GPU technologies such as A100, H100, and H200, including provisioning, monitoring, and performance tuning.
  • Ensure high availability and reliability through expert-level Linux administration, storage management with Ceph and high-performance platforms, and networking in data centers.

The company is seeking a senior Linux infrastructure engineer with expertise in bare metal, storage, and AI Factory platforms. The size, employees, and culture are not specified.

Europe

  • Lead, mentor, and develop a small team of AI engineers, fostering a culture of ownership and engineering excellence.
  • Define the technical vision and architectural direction of the AI product, ensuring high standards for performance and scalability.
  • Contribute hands-on to the design and optimization of production AI systems, including model serving and inference pipelines.

The partner company is developing a cutting-edge AI platform that powers innovative products at scale. They foster a collaborative, remote-first environment and focus on engineering excellence.

India

  • Act as a trusted technical advisor to C-level executives, articulating the strategic value of high-performance object storage.
  • Design and validate complex, large-scale architectures for enterprise customers integrating MinIO into AI/ML pipelines.
  • Represent MinIO at industry conferences and write technical whitepapers on object storage, generative AI, and modern analytics.

MinIO is the data and memory foundation for enterprise AI, built for speed, scale, and economics. Trusted by 77% of the Fortune 100, MinIO is redefining how AI factories and intelligent applications secure and unlock data.

India

  • Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
  • Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
  • Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.

The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.

$92,000–$133,000/yr
Global

  • Collaborate with your team to identify and scope out well-defined tasks and deliver on time with quality.
  • Design systems for rapid growth, promote operational excellence through testing and monitoring, and reduce technical debt.
  • Uphold a culture of collaboration, transparency, creativity, inclusion, and data-driven decisions.

Mapbox is the leading real-time location platform for a new generation of location-aware businesses, equipping organizations with tools to power navigation of people, packages, and vehicles. With over 4 million registered developers, the company values flexibility, security, privacy, and building a diverse, inclusive team.

$80,000–$100,000/yr
US

  • Build, maintain, and operate Vultr's Ceph-based Cloud Storage environments with a focus on stability and scalability.
  • Develop automation frameworks and optimize metrics collection and alerting systems.
  • Collaborate with cross-functional teams for capacity expansion, maintenance, and design decisions.

Vultr makes high-performance cloud infrastructure easy, affordable, and locally accessible for enterprises and AI innovators globally. As the world’s largest privately-held cloud infrastructure company, Vultr has over 30 data centers, hundreds of thousands of customers, and a $3.5 billion valuation.

Europe

  • Own the AI architecture across Flip Fusion, Ask AI agent gateway, and Flip Flows.
  • Lead and grow a team of three AI engineers, staying hands-on with the hardest problems.
  • Shape a young AI engineering discipline with enormous design space and no inherited playbook.

Flip is the leading AI-powered employee experience platform for frontline workers in retail, manufacturing, and logistics. The company is a young tech company with a relaxed working atmosphere and highly motivated people.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

US

  • Design and build control-plane services and drivers for storage integration with Kubernetes-based AI workloads.
  • Write production-quality Go code with strong testing and operational rigor.
  • Deliver storage integration for k0s-based Kubernetes via Cluster API and K0rdent topologies.

Mirantis is a Kubernetes-native AI infrastructure company that builds scalable, secure infrastructure for AI and data-intensive applications. It is committed to open standards and freedom from lock-in, empowering platform engineering teams.

$200,000–$230,000/yr
Global

  • Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
  • Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
  • Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.

Vultr provides high-performance cloud infrastructure solutions, including Cloud Compute, GPU, Bare Metal, and Storage, with 33 global data centers. It is a privately-held company valued at $3.5 billion, trusted by hundreds of thousands of customers across 185 countries, and known for its self-funded growth and inclusive culture.

Europe 6w PTO

  • Build and deliver AI-powered features that help users manage large datasets and enhance analytics-focused AI agents.
  • Rapidly prototype, test, and iterate with real users, shipping LLM- or agent-driven workflows for data engineering.
  • Collaborate across teams to integrate AI components with internal tools, taking full ownership of scalable solutions.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations. It is a 100% remote company with team members across 40+ countries, backed by leading investors and known for its open-source legacy and collaborative culture.