Source Job

$133,940–$178,340/yr
Canada

  • Lead a team of platform engineers to build and operate ML training and serving infrastructure, including GPU and low-latency serving.
  • Combine strong people leadership with technical judgment in ML infrastructure, partnering with senior ICs and cross-functional teams.
  • Drive delivery, operational health, and evaluate modern ML tooling to support company-wide ML priorities.

Machine Learning Engineering Management Software Engineering Distributed Systems Infrastructure

20 jobs similar to Engineering Manager, Machine Learning Platform

Jobs ranked by similarity.

$200,000–$230,000/yr
US Canada Unlimited PTO

  • Lead Data + ML Platform strategy and execution, including MLOps and data engineering.
  • Build and scale the engineering team, coaching engineers and fostering a culture of ownership.
  • Partner cross-functionally with Product, Hardware, and Data Science to drive technical direction and business outcomes.

Inspiren offers the most complete and connected ecosystem in senior living, founded by a former Green Beret turned cardiothoracic nurse. The company is building an integrated platform with smart sensors and analytics to improve care outcomes.

$153,351–$206,481/yr
Canada

  • Lead a team to build, scale, and optimize the ML infrastructure powering drug discovery.
  • Collaborate with ML engineering, data science, and research teams to deliver scalable solutions.
  • Mentor and coach team members in MLOps, distributed computing, and infrastructure engineering.

Recursion is a clinical-stage TechBio company decoding biology to develop medicines. With a focus on AI and machine learning, the company fosters a culture of bold integrity and cross-functional collaboration.

$138,500–$225,500/yr
US 16w maternity 16w paternity

  • Design, train, and ship ML systems for governance and security like anomaly detection and trust scoring.
  • Build data pipelines, model serving, evaluation frameworks, and feedback loops.
  • Set technical direction, own architecture, and help recruit and mentor as the team grows.

Docker provides developer tooling trusted by over 20 million monthly users and billions of container pulls. They are a globally distributed, remote-first team building tools for software delivery.

India

  • Collaborate with data scientists and engineers to build scalable ML pipelines, troubleshoot infrastructure issues from Linux to Kubernetes, and optimize model performance.
  • Drive high engineering standards, design on-premises MLOps solutions, and maintain tools for deployment and monitoring.
  • Refine CI/CD workflows, incorporate ML model training and evaluation into testing, and ensure seamless handover between research and production.

Learneo is a platform of builder-driven businesses, including Course Hero, CliffsNotes, LitCharts, Quillbot, Symbolab, and Scribbr, focused on supercharging productivity and learning. The company supports high-growth businesses with centralized corporate operations and has a virtual-first culture with employees across multiple countries.

US Unlimited PTO

  • Design and maintain scalable ML infrastructure including data pipelines, training workflows, and model deployment systems.
  • Own end-to-end ML lifecycle operations, ensuring reliable delivery of models into production at scale.
  • Implement monitoring, telemetry, and feedback loops for ML models running across large-scale device fleets.

Our partner company develops ML systems for connected hardware products used by customers worldwide. They operate in a fast-paced, product-driven environment with a collaborative and technically ambitious culture focused on real-world ML impact.

US

  • Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
  • Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
  • Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.

Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

EMEA

  • Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
  • Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
  • Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.

They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.

$173,000–$233,000/yr
Canada

  • Build and ship AI agents, APIs, and applications on Affirm's internal platform, owning the full lifecycle from architecture to production.
  • Turn messy business requirements from People Operations stakeholders into production systems, integrating with tools like Workday and Notion.
  • Design reliability infrastructure for multi-model LLM services, including structured output validation and quality controls.

Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. The People Tech & Analytics team builds and owns the data, AI, and technology infrastructure for Affirm's People function, running like a product engineering group embedded in HR.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

$143,764–$184,404/yr
UK

  • Lead the ML strategy for Workforce Management, delivering production systems for forecasting, routing, and optimisation.
  • Coach and develop a growing team of Machine Learning Scientists embedded in product squads.
  • Work cross-functionally with Product, Engineering, Data, and Operations leaders to solve ambiguous operational problems.

Monzo is a digital bank on a mission to make money work for everyone, offering personal and business bank accounts, savings, investments, and pension consolidation. With over 10 years of growth in the UK and a focus on financial education and award-winning customer service, Monzo has a vibrant culture and a strong commitment to diversity and inclusion.

India

  • Collaborate with data scientists and software engineers to build scalable data pipelines and ML deployment systems.
  • Troubleshoot issues across the ML infrastructure stack, from Linux and Docker to Kubernetes and model serving.
  • Drive high engineering standards through code reviews, testing, and CI/CD enhancements.

Quillbot helps students and professionals strengthen their writing with AI-powered tools. We serve over 56 million users globally and foster a collaborative, virtual-first culture.

US

  • Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
  • Architect scalable inference systems and collaborate closely with application engineering.

This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.

Global Unlimited PTO

  • Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
  • Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
  • Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.

Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.

Australia

  • Design, build, and ship ML models that power content generation and quality eval scoring for Canva's generated element and template library.
  • Own the full ML lifecycle — from data pipelines and training through to deployment, monitoring, and iteration.
  • Partner with Content Engine, CORE AI Research, AI Media, and Discovery teams to align ML work with the broader content strategy.

Canva is redefining how the world experiences design with its intuitive design platform. We serve hundreds of millions of users globally and foster a culture of flexibility, inclusion, and innovation.

$220,000–$280,000/yr
US Unlimited PTO

  • Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
  • Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
  • Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.

Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.

France 5w PTO

  • Own ML models across their full lifecycle from data pipelines to deployment and monitoring, ensuring reliable performance.
  • Run and improve the ML platform including GitOps CI/CD, monitoring serving endpoints, and defining SLOs.
  • Collaborate with risk, operational, and product teams to turn ML into business value across the organization.

Alma provides installment and deferred payment solutions to help merchants boost sales and customer loyalty, without encouraging bad debt. With over 25,000 merchants, 10 million consumers, 380+ employees, and over €100M ARR, they are a Next40 member scaling rapidly across Europe.

United States

  • Lead engineering deployment, scaling, and operations of AI compute clusters with GPU fleets and bare metal environments.
  • Drive reliability, monitoring, automation, and incident response for AI infrastructure.
  • Collaborate with AI/ML, networking, and product teams to align infrastructure with business needs.

Our partner is a fast-growing cloud environment focused on building large-scale AI infrastructure. They seek a senior leader to manage engineering operations for advanced AI compute clusters.

Italy

  • Lead and develop a high-performing team of MLOps engineers, fostering technical excellence and collaboration.
  • Define and execute the MLOps roadmap, aligning infrastructure initiatives with research, engineering, and product goals.
  • Design and maintain scalable ML infrastructure including automated training pipelines, CI/CD, and model serving platforms.

Our partner is a company focused on cutting-edge machine learning infrastructure for large-scale AI systems. They foster an inclusive, mission-driven culture with international collaboration and value innovation, diversity, and continuous learning.

Canada

  • Lead a talented engineering team building advanced AI-powered systems for scientific discovery.
  • Translate open-ended research objectives into clear technical strategies and scalable solutions.
  • Balance innovation speed with reliability, security, and operational excellence.

The partner company is building advanced AI-powered systems for scientific discovery. It offers a collaborative environment where engineers work alongside researchers and technical experts.