Lead investigation and resolution of complex infrastructure, networking, and platform incidents.
Provide technical leadership for Kubernetes platform operations and drive automation initiatives.
Mentor engineers and develop operational standards, runbooks, and best practices.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Serving enterprises like Adobe, PayPal, and Volkswagen, Mirantis is committed to open standards and freedom from lock-in.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
Oversee cluster reliability, GPU fleet operations, and incident response.
Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.
Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.
Lead the design and operation of GPU infrastructure for AI workloads.
Manage Kubernetes-based environments and optimize for AI training and inference.
Define operational standards, implement monitoring, and collaborate with AI engineering teams.
ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.
Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.
Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
Drive reliability, monitoring, automation, and incident response for AI infrastructure.
Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.
Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.
Design, operate, and improve reliable infrastructure for AI training and inference workloads.
Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.
Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.
They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.
Diagnose and resolve complex production issues across Linux, Kubernetes, networking, storage, and GPU systems.
Act as a senior escalation point for critical incidents, collaborating with engineering teams on root cause analysis.
Develop tools and automation in Python, Bash, or Go to improve troubleshooting efficiency and observability.
The partner company provides advanced AI and cloud infrastructure solutions, supporting large-scale distributed computing and AI workloads. They operate in a fast-moving, collaborative environment with highly skilled engineering teams focused on cutting-edge technology and operational excellence.
Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.
Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.
Build and maintain the Shadeform GPU platform and automated AI infrastructure services.
Work on novel solutions to GPU market challenges including provisioning, orchestration, and virtualization.
Own the systems that turn fragmented GPU capacity into a reliable, production-ready platform.
Shadeform provides a unified platform for deploying and managing GPU infrastructure across cloud providers, neoclouds, and data centers. They are a remote-first startup focused on innovation and working with the latest AI technologies.
Lead customers in designing and optimizing GPU-based solutions on Vultr's platform.
Collaborate with cross-functional teams to bring AI, ML, and GPU workloads into production.
Educate customers on the value of Vultr's cloud infrastructure and expand their possibilities.
Vultr provides high-performance cloud infrastructure solutions globally, making them easy to use, affordable, and locally accessible. It is a privately-held company with over a decade of self-funding, hundreds of thousands of customers across 185 countries, and a culture that emphasizes comprehensive benefits and employee growth.
Lead cross-team incident triage for high-impact customer outages, coordinating Engineering, Product, and Customer Experience response and contributing to root cause analysis.
Develop and maintain observability for cloud-hosted customer deployments by building and refining system monitors, dashboards, and alerting.
Serve as the senior escalation point for complex support cases in EMEA, working cases that involve deep platform internals and unusual failure modes.
Dragos defends industrial organizations that provide modern civilization necessities like water, electricity, and safe working environments. As a market leader in ICS/OT Cybersecurity, we operate globally with a remote-first culture and are looking for mission-oriented teammates who value authenticity, transparency, and trust.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Participate in sales meetings, provide architectural recommendations, and build proof-of-concept solutions for onboarding high-spending customers.
Troubleshoot and resolve complex technical issues using code analysis, scripting, and log analysis.
Create and maintain technical documentation and deliver training sessions, webinars, and demos.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We are a small, remote-first team that takes ownership seriously, moves fast, and ships work relied on by more than a million developers daily.
Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.
Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.
Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
Lead incident response and encode every fix as a new runbook and regression test.
Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.
Operate and maintain Linux-based infrastructure, deploy and scale Kubernetes clusters, and implement automation with Ansible and GitOps.
Design networking architecture, build observability stacks, and lead incident response across the platform.
Manage virtualization layers and collaborate with development teams to optimize resource utilization and system availability.
Pragmatike develops cutting-edge solutions in Cloud Computing, focusing on ambitious projects with a culture of collaboration and innovation. The team is passionate and collaborative, working in a dynamic and flexible environment to shape tomorrow's technologies.
Own high-severity technical escalations from intake through resolution or engineering handoff.
Partner with Support and Product/Engineering to close the gap between customer problems and engineering fixes.
Monitor patterns across escalations to catch systemic issues and translate into product improvements.
Tailscale builds software that makes it easy to securely interconnect people and devices. Founded in 2019, the company is fully distributed and backed by Accel, CRV, and others.
Build and operate the infrastructure behind AI-powered products, improving reliability, security, scalability, and cost efficiency.
Write code, automate infrastructure, investigate production issues, and design systems that reduce operational complexity.
Take ownership of unfamiliar systems, identify highest-leverage improvements, and balance immediate production needs with long-term platform investments.
Zencoder builds and orchestrates AI agents that ship real work across code, research, and operations. It is a growing platform where people and agents collaborate, with a high-caliber team and a culture that values individual contributors.