Source Job

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Kubernetes

20 jobs similar to Engineering Manager

Jobs ranked by similarity.

Canada

  • Lead the design and operation of GPU infrastructure for AI workloads.
  • Manage Kubernetes-based environments and optimize for AI training and inference.
  • Define operational standards, implement monitoring, and collaborate with AI engineering teams.

ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.

$150,000–$160,000/yr

  • Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
  • Oversee cluster reliability, GPU fleet operations, and incident response.
  • Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.

Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.

EMEA

  • Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
  • Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
  • Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.

They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.

  • Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
  • Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
  • Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.

Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.

Global Unlimited PTO

  • Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
  • Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
  • Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.

Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

$225,000–$325,000/yr
Global Unlimited PTO

  • Lead core infrastructure and SRE teams to ensure the highest reliability and performance for massive GPU computing demands.
  • Oversee HPC networking and distributed storage engine innovation to support massive multi-node AI workloads.
  • Build and scale a high-output engineering org while partnering cross-functionally with product and GTM leadership.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. As a small, remote-first team that closed a $100M Series A, we move fast and take ownership seriously.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

$153,351–$206,481/yr
Canada

  • Lead a team to build, scale, and optimize the ML infrastructure powering drug discovery.
  • Collaborate with ML engineering, data science, and research teams to deliver scalable solutions.
  • Mentor and coach team members in MLOps, distributed computing, and infrastructure engineering.

Recursion is a clinical-stage TechBio company decoding biology to develop medicines. With a focus on AI and machine learning, the company fosters a culture of bold integrity and cross-functional collaboration.

Global

  • Lead a high-leverage remote team of four infrastructure engineers, driving the evolution toward a scalable zero-toil platform.
  • Guide the team through an AI-driven engineering approach to reduce manual work and achieve zero-touch, scalable infrastructure.
  • Prepare and execute the strategy for CI/CD and artifact distribution systems to scale during a quality surge without increasing engineering toil.

Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex business processes. Trusted by over 700 organizations worldwide, including 9 of top 10 US banks, Camunda is a fully remote and global company with 150+ engineers across 20+ teams, and is transforming into an AI-first organization.

US

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.
  • Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.

We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.

Global

  • Own and optimize CI/CD pipelines, Kubernetes deployment, and infrastructure for model serving and inference.
  • Build telemetry, observability, and alerting to catch real problems and reduce noise.
  • Eliminate toil through thoughtful automation and improve developer and agent productivity.

Obvious is building an AI-native workspace that serves as an operating system for work, putting co-intelligence at the center. They are a small, talent-dense team with founders and leaders from top tech companies.

$110,000–$220,000/yr
Global Unlimited PTO

  • Lead a product engineering team to deliver customer-facing features across Runpod’s console, APIs, and developer workflows.
  • Own technical architecture, execution, quality, and launch for a roadmap area, partnering with Product and Design.
  • Build and grow a strong team through hiring, mentoring, and fostering a culture of ownership and speed.

Runpod provides cutting-edge cloud infrastructure for full-stack AI and machine learning applications. Founded in 2022, it is a rapidly growing, well-funded remote-first company with a global team.

US

  • Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
  • Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.

Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.

Europe

  • Lead investigation and resolution of complex infrastructure, networking, and platform incidents.
  • Provide technical leadership for Kubernetes platform operations and drive automation initiatives.
  • Mentor engineers and develop operational standards, runbooks, and best practices.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. Serving enterprises like Adobe, PayPal, and Volkswagen, Mirantis is committed to open standards and freedom from lock-in.

France Global

  • Lead a distributed team of Python engineers building infrastructure automation solutions for large-scale cloud and bare-metal environments.
  • Drive technical strategy, team growth, and engineering culture while contributing hands-on to codebases.
  • Collaborate with product, operations, and engineering stakeholders to align delivery with organizational goals.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective and fair processes. It operates as a distributed remote team, fostering a culture of autonomy and high expectations.

US Unlimited PTO

  • Design and maintain scalable ML infrastructure including data pipelines, training workflows, and model deployment systems.
  • Own end-to-end ML lifecycle operations, ensuring reliable delivery of models into production at scale.
  • Implement monitoring, telemetry, and feedback loops for ML models running across large-scale device fleets.

Our partner company develops ML systems for connected hardware products used by customers worldwide. They operate in a fast-paced, product-driven environment with a collaborative and technically ambitious culture focused on real-world ML impact.

$150,000–$200,000/yr
Global Unlimited PTO

  • Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
  • Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
  • Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.

Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.