Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
Oversee cluster reliability, GPU fleet operations, and incident response.
Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.
Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.
Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.
TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.
Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.
Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
Drive reliability, monitoring, automation, and incident response for AI infrastructure.
Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.
Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Lead the design and operation of GPU infrastructure for AI workloads.
Manage Kubernetes-based environments and optimize for AI training and inference.
Define operational standards, implement monitoring, and collaborate with AI engineering teams.
ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.
Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.
Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.
Monitor, operate, and support production AI infrastructure platforms including NVIDIA GPU environments.
Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
Collaborate with engineering teams, vendors, and datacenter personnel to improve operational processes.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build scalable, secure, and sovereign infrastructure for modern AI and data-intensive applications. Serving enterprises like Adobe and PayPal, the company combines open source innovation with deep Kubernetes expertise to deliver composable developer platforms across any environment.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.
Validate and troubleshoot InfiniBand and RoCE fabrics during GPU cluster bring-up and expansion.
Tune fabric performance parameters for distributed AI workloads such as NCCL and MPI.
Collaborate with GPU and networking teams to diagnose and resolve fabric-level issues and optimize performance.
Vultr makes high-performance cloud infrastructure easy to use and affordable for enterprises and AI innovators worldwide. With 33 global data centers and hundreds of thousands of customers, it is the largest privately-held cloud infrastructure company, offering a culture of innovation and growth.
Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.
Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.
Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.
They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.
Participate in sales meetings, provide architectural recommendations, and build proof-of-concept solutions for onboarding high-spending customers.
Troubleshoot and resolve complex technical issues using code analysis, scripting, and log analysis.
Create and maintain technical documentation and deliver training sessions, webinars, and demos.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We are a small, remote-first team that takes ownership seriously, moves fast, and ships work relied on by more than a million developers daily.
Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms, owning the storage layer where Kubernetes meets bare metal.
Tune NFS data paths for high-throughput, low-latency GPU/AI workloads, integrating NFS-based storage into clusters via CSI, storage classes, and persistent volumes.
Automate storage provisioning with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines, and build monitoring and observability for storage performance and health.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The company combines open source innovation with deep Kubernetes expertise, empowering platform engineering teams across on-premises, cloud, edge, and sovereign data centers.
Build and operate the infrastructure behind AI-powered products, improving reliability, security, scalability, and cost efficiency.
Write code, automate infrastructure, investigate production issues, and design systems that reduce operational complexity.
Take ownership of unfamiliar systems, identify highest-leverage improvements, and balance immediate production needs with long-term platform investments.
Zencoder builds and orchestrates AI agents that ship real work across code, research, and operations. It is a growing platform where people and agents collaborate, with a high-caliber team and a culture that values individual contributors.
Own the technical evaluation end-to-end, from discovery to POC, ensuring evaluations are scoped and tied to customer ROI.
Take customers from signature to first successful production training run and serve as the technical owner post-launch.
Build the SA function by creating demo environments, benchmarking harnesses, and reference architectures.
Andromeda Cluster gives early-stage startups access to scaled AI infrastructure, partnering with leading AI labs and cloud providers. It is a high-growth company building an inclusive environment for all employees.
Productize deployment, security, and scaling of Applied AI solutions with automation and security guardrails.
Mistral provides full-stack AI solutions from frontier models to developer tools, applications, and compute, partnering with enterprises across high-stakes industries. It is a dynamic, collaborative team with a diverse workforce distributed globally, known for being creative, low-ego, and team-spirited.
Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against quality metrics.
Define the qualification bar: build the acceptance test suite, benchmark methodology, and quality thresholds to formalize tribal knowledge.
Guide providers through technical onboarding: work with their engineers to remediate gaps and bring clusters onto the network cleanly.
Andromeda Cluster provides early-stage startups access to scaled AI infrastructure that was once reserved for hyperscalers. It is a small, high-growth team at the center of the AI infrastructure boom, founded by Nat Friedman and Daniel Gross.