Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Lead the design and operation of GPU infrastructure for AI workloads.
Manage Kubernetes-based environments and optimize for AI training and inference.
Define operational standards, implement monitoring, and collaborate with AI engineering teams.
ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.
Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
Drive reliability, monitoring, automation, and incident response for AI infrastructure.
Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.
Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.
Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
Oversee cluster reliability, GPU fleet operations, and incident response.
Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.
Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.
Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.
They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.
Define the networking strategy and roadmap for k0rdent AI, covering GPU cluster networking and multi-tenant cloud integration.
Partner with engineering and marketing to shape requirements, positioning, and competitive differentiation.
Represent Mirantis at events and with strategic accounts, driving product success in the AI cloud era.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. With a world-class, distributed team, Mirantis empowers platform engineering teams and is committed to openness, collaboration, and continuous growth.
Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.
TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.
Own the vision, roadmap, and priorities for k0rdent AI networking, spanning underlay fabric management, tenant connectivity, RDMA, DNS/IPAM, and network automation.
Translate requirements from GPU clouds, telcos, and enterprise platform teams into clear product direction, partnering with engineering to define requirements.
Track and shape response to emerging interconnect standards like Ultra Ethernet, UALink, and congestion control, and represent Mirantis with customers and partners.
Mirantis is the leading AI-infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI. They are a global, distributed team committed to openness and technical excellence, serving clients like Adobe, PayPal, and Volkswagen.
Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.
Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.
Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.
Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
Lead incident response and encode every fix as a new runbook and regression test.
Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.
Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms, owning the storage layer where Kubernetes meets bare metal.
Tune NFS data paths for high-throughput, low-latency GPU/AI workloads, integrating NFS-based storage into clusters via CSI, storage classes, and persistent volumes.
Automate storage provisioning with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines, and build monitoring and observability for storage performance and health.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The company combines open source innovation with deep Kubernetes expertise, empowering platform engineering teams across on-premises, cloud, edge, and sovereign data centers.
Diagnose and resolve complex production issues across Linux, Kubernetes, networking, storage, and GPU systems.
Act as a senior escalation point for critical incidents, collaborating with engineering teams on root cause analysis.
Develop tools and automation in Python, Bash, or Go to improve troubleshooting efficiency and observability.
The partner company provides advanced AI and cloud infrastructure solutions, supporting large-scale distributed computing and AI workloads. They operate in a fast-moving, collaborative environment with highly skilled engineering teams focused on cutting-edge technology and operational excellence.
Design and run Kubernetes environments optimized for AI inference, retrieval, and agent execution in secure settings.
Deploy and operate open-source model stacks, model gateways, vector databases, and supporting platform components.
Build reproducible platform automation using Infrastructure as Code and GitOps approaches for stable, auditable delivery.
Deutsche Telekom IT Solutions Slovakia provides innovative information and communication technology services. It has grown to become the second largest employer in eastern Slovakia with over 3900 employees, focusing on continuous transformation and improvement.
Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.
Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.
Design and run Kubernetes environments optimized for AI inference, retrieval, experimentation, and agent execution in secure or isolated settings.
Deploy and operate open-source or open-weight model stacks, model gateways, vector databases, and supporting platform components.
Build reproducible platform automation using Infrastructure as Code and GitOps approaches for stable, auditable delivery.
Deutsche Telekom IT Solutions is a subsidiary of Deutsche Telekom Group, providing IT and telecommunications services. It has over 5300 employees and has been recognized as Hungary's most attractive employer and most ethical multinational company.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Build and maintain the Shadeform GPU platform and automated AI infrastructure services.
Work on novel solutions to GPU market challenges including provisioning, orchestration, and virtualization.
Own the systems that turn fragmented GPU capacity into a reliable, production-ready platform.
Shadeform provides a unified platform for deploying and managing GPU infrastructure across cloud providers, neoclouds, and data centers. They are a remote-first startup focused on innovation and working with the latest AI technologies.
Develop core positioning, messaging, and value propositions for AI infrastructure and cloud-native platform portfolio.
Produce high-quality technical content including solution briefs, white papers, blog posts, and reference architectures.
Support product launches, sales enablement, competitive intelligence, and go-to-market strategy.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable infrastructure for AI and data-intensive applications. With deep expertise in open source and Kubernetes orchestration, Mirantis empowers platform engineering teams worldwide.