Source Job

$200,000–$230,000/yr
Global

  • Lead product engagement with strategic AI infrastructure customers to define requirements and drive execution from discovery to production readiness.
  • Collaborate cross-functionally with engineering, infrastructure, and operations teams to deliver customer-ready solutions.
  • Translate complex customer needs into clear product priorities, technical specifications, and scalable AI infrastructure offerings.

GPU Compute Kubernetes Slurm AI/ML Infrastructure

20 jobs similar to Principal Product Manager

Jobs ranked by similarity.

Europe

  • Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.

Europe

  • Lead technical operations for large-scale AI infrastructure environments powered by NVIDIA GPUs and Kubernetes.
  • Act as a senior escalation point for critical incidents and drive root cause analysis and long-term corrective actions.
  • Mentor team members and shape operational standards, automation, and reliability practices for next-generation platform services.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, PayPal, Reliance Jio, Societe Generale, Splunk, and Volkswagen.

India

  • Design end-to-end AI infrastructure solutions for scalable, high-performance AI and HPC environments.
  • Partner with Sales to qualify opportunities, conduct technical discovery, and serve as trusted advisor throughout the sales lifecycle.
  • Collaborate closely with Facilities, Delivery, OEM partners, and customer teams to ensure AI infrastructure aligns with data center capabilities.

Submer enables organizations scaling AI to overcome the limits of traditional datacenters in power, compute density, and efficiency. It is a fast-growing, international scale-up with a friendly, diverse, and hybrid-friendly work environment.

  • Lead and build the SE / Solutions Architect team, defining the pre-sales operating model as the org scales.
  • Own the technical win in large, complex deals, architecting solutions across compute, networking, storage, and orchestration.
  • Be the technical voice of the customer internally, feeding structured product requirements back to product and platform teams.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The company is committed to open standards and freedom from lock-in, empowering platform engineering teams to deliver composable, production-ready developer platforms across any environment.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.

US

  • Engage directly with customers to resolve complex technical challenges involving Kubernetes GPU clusters.
  • Act as a customer-facing SRE to ensure Kubernetes clusters remain healthy and stable.
  • Become a product expert in GPU Cluster service, serving as the last line of technical defense before escalation.

Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems. The team has contributed to leading open-source research and aims to build the next generation AI infrastructure.

US

  • Build the technical product marketing function for the Provider business, creating collateral like white papers, reference architectures, and demo environments.
  • Directly support pipeline development by partnering with sales and solution architects through technical storytelling and proof-points.
  • Develop competitive intelligence and represent Mirantis as a credible technical voice on GPU infrastructure and sovereign AI.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It empowers platform engineering teams across any environment with a strong benefits plan and professional development.

Global

  • Support the deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructure.
  • Assist in troubleshooting network issues, including connectivity, latency, and performance degradation.
  • Collaborate with compute and storage teams to support HPC and AI workloads.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. They serve many of the world’s leading enterprises and are committed to open standards and freedom from lock-in.

US

  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Collaborate with GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.

Bitdeer provides comprehensive Bitcoin mining solutions and AI computational infrastructure. The company operates globally with data centers in multiple countries and focuses on AI and blockchain technology.

  • Own end-to-end infrastructure deployment programs for new capacity and site expansions, including hardware dependencies and commissioning gate frameworks.
  • Deliver crisp, data-driven executive updates and govern cross-organizational dependencies without escalation.
  • Coach junior TPMs and drive AI tool integration to improve program tracking and risk detection.

Evergrid builds frontier-level AI infrastructure for the fourth industrial revolution, designing, deploying, and operating large-scale AI systems. With over 1GW of capacity in active development and a 24/7 engineering team, they foster a culture of high conviction and high trust.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

US

  • Lead end-to-end architecture design of AI Data Center Networks, high-performance DCI, and global backbone networks for large-scale GPU clusters.
  • Collaborate with NVIDIA, vendors, and partners to translate business requirements into top-level network designs including InfiniBand/RoCE, Spine-Leaf, and overlay integration.
  • Own congestion control tuning, produce architecture documentation, and identify risks to drive network evolution.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Headquartered in Singapore, the company operates data centers across multiple countries and is building AI computational infrastructure.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

Global

  • Partner with Head of Product and engineering leadership to define technical direction and product strategy.
  • Conduct hands-on validation of alpha and beta builds to catch integration issues before customers.
  • Engage directly with platform engineers at AI Cloud operators and regulated enterprises to drive market positioning.

vCluster Labs is a venture-backed startup pioneering Kubernetes virtualization for the AI era. With over $30M in funding and a remote-first culture, the company is in hyper-growth phase, building the leading platform for operating GPU infrastructure.

US

  • Own a strategic domain of the platform and define the product roadmap.
  • Translate customer needs, technical requirements, and business goals into clear, well-scoped product features.
  • Lead customer discovery, define and track KPIs, and shape product strategy as the first PM.

ScaleOps provides an autonomous cloud and AI infrastructure management platform that reduces cloud costs by up to 80%. Backed by over $210M in funding, the company is trusted by Fortune 100 companies and leading enterprises.

Europe

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

Mirantis is a Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. They serve many leading enterprises including Adobe, DocuSign, and PayPal, and are a leader in container management.

Global

  • Own the technical path from customer interest to working deployment, integrating the platform into production AI environments.
  • Build and operate AI/MLOps pipelines, debug complex environments, and create prototypes and demos.
  • Translate customer needs into product improvements, partnering with Sales, Product, and Engineering.

Neuromorphic Labs is a Seed-stage AI startup building a trust layer for production AI, making security, governance, and control intrinsic to every model and deployment. Backed by top-tier VCs, the team is small and fast-paced, emphasizing ownership, high standards, and collaboration with exceptional builders.

  • Design and build LLM serving infrastructure on Kubernetes, including deployment, GPU scheduling, and model lifecycle management.
  • Package the platform for enterprise environments with Helm-based installs and support for restricted or offline networks.
  • Integrate the serving layer with API gateway, identity, and metering services, and build observability for GPU inference in production.

Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build and operate scalable infrastructure for AI and data-intensive applications. It is part of an IREN company and empowers platform engineering teams with open-source innovation and deep expertise in Kubernetes orchestration.

$200,000–$350,000/yr
US

  • Define architecture and technical strategy for highly scalable AI systems.
  • Lead complex model serving and inference initiatives while optimizing performance and cost.
  • Mentor senior engineers and influence cross-functional AI strategy and product direction.

The company builds critical infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. It is a fast-growing organization with a collaborative, fast-moving culture where engineers have meaningful influence on architecture and product direction.

US

  • Define and drive the technical architecture for AI, data, and platform initiatives supporting Anomali's product strategy.
  • Translate product strategy and customer outcomes into scalable architecture and implementation plans, balancing near-term delivery with long-term maintainability.
  • Partner closely with Product Management, Engineering, Data Science, and UX to ensure technical decisions align with the five-level maturity model.

Anomali delivers the first Intelligence-Native Agentic SOC Platform, unifying a security data lake, threat intelligence, and agentic AI into a single experience. The company is headquartered in Silicon Valley and fosters a culture of innovation and security.