Source Job

$153,351–$206,481/yr
Canada

  • Lead a team to build, scale, and optimize the ML infrastructure powering drug discovery.
  • Collaborate with ML engineering, data science, and research teams to deliver scalable solutions.
  • Mentor and coach team members in MLOps, distributed computing, and infrastructure engineering.

Python PyTorch Kubernetes MLOps GCP

20 jobs similar to Engineering Manager - Machine Learning

Jobs ranked by similarity.

India

  • Collaborate with data scientists and engineers to build scalable ML pipelines, troubleshoot infrastructure issues from Linux to Kubernetes, and optimize model performance.
  • Drive high engineering standards, design on-premises MLOps solutions, and maintain tools for deployment and monitoring.
  • Refine CI/CD workflows, incorporate ML model training and evaluation into testing, and ensure seamless handover between research and production.

Learneo is a platform of builder-driven businesses, including Course Hero, CliffsNotes, LitCharts, Quillbot, Symbolab, and Scribbr, focused on supercharging productivity and learning. The company supports high-growth businesses with centralized corporate operations and has a virtual-first culture with employees across multiple countries.

$220,000–$280,000/yr
US Unlimited PTO

  • Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
  • Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
  • Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.

Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.

US Unlimited PTO

  • Design and maintain scalable ML infrastructure including data pipelines, training workflows, and model deployment systems.
  • Own end-to-end ML lifecycle operations, ensuring reliable delivery of models into production at scale.
  • Implement monitoring, telemetry, and feedback loops for ML models running across large-scale device fleets.

Our partner company develops ML systems for connected hardware products used by customers worldwide. They operate in a fast-paced, product-driven environment with a collaborative and technically ambitious culture focused on real-world ML impact.

$200,000–$230,000/yr
US Canada Unlimited PTO

  • Lead Data + ML Platform strategy and execution, including MLOps and data engineering.
  • Build and scale the engineering team, coaching engineers and fostering a culture of ownership.
  • Partner cross-functionally with Product, Hardware, and Data Science to drive technical direction and business outcomes.

Inspiren offers the most complete and connected ecosystem in senior living, founded by a former Green Beret turned cardiothoracic nurse. The company is building an integrated platform with smart sensors and analytics to improve care outcomes.

India

  • Collaborate with data scientists and software engineers to build scalable data pipelines and ML deployment systems.
  • Troubleshoot issues across the ML infrastructure stack, from Linux and Docker to Kubernetes and model serving.
  • Drive high engineering standards through code reviews, testing, and CI/CD enhancements.

Quillbot helps students and professionals strengthen their writing with AI-powered tools. We serve over 56 million users globally and foster a collaborative, virtual-first culture.

EMEA

  • Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
  • Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
  • Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.

They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.

$161,000–$273,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Lead the design and operation of production machine learning systems for batch and online use cases with a focus on reliability and scalability.
  • Build and improve ML lifecycle infrastructure including training pipelines, inference workflows, monitoring, and automation.
  • Partner with cross-functional teams to translate business problems into ML solutions and guide prototypes to robust production systems.

Included Health is a healthcare company delivering integrated virtual care and navigation, aiming to raise the standard of healthcare for everyone. They are a remote-first organization offering comprehensive benefits and fostering a culture of inclusion.

US

  • Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
  • Architect scalable inference systems and collaborate closely with application engineering.

This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.

$185,000–$200,000/yr
US Unlimited PTO 16w maternity 4w paternity

  • Build and operate the ML lifecycle platform, including tooling for experiment tracking, model registry, and versioned pipelines.
  • Own CI/CD and deployment for ML workloads, building automated pipelines from notebook to production.
  • Make models observable and reliable in production with monitoring for latency, drift, data quality, and cost signals.

dv01 provides a data analytics platform for the structured finance market, offering transparency into investment performance and risk for lenders and Wall Street investors. With over 400 clients and coverage of over 100 million loans, dv01 is a data-first company with a diverse and innovative culture.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

US

  • Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
  • Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.

Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.

Canada

  • Investigate novel techniques combining class leading heuristics with optimization and ML.
  • Translate real world Supply Chain Management use cases into mathematical models.
  • Lead the design and implementation of mathematical models and ML systems.

Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains with an AI-infused platform. With over 2000 employees worldwide, they foster a culture of innovation, collaboration, and social responsibility.

Global Unlimited PTO

  • Lead and scale the Forward Deployed Engineering and Technical Support teams, defining engagement models and operating standards.
  • Own the FDE engagement lifecycle from technical discovery to deployment guidance, ensuring customer value.
  • Drive operational discipline across support tools and partner with Sales, Product, and Engineering on roadmap alignment.

Runpod is the AI Developer Cloud. More than one million developers use the platform to experiment, train, deploy, and scale AI, and we are a small, remote-first team that has processed over 20 billion inference requests and closed a $100M Series A.

  • Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
  • Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
  • Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.

Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.

Australia

  • Design, build, and ship ML models that power content generation and quality eval scoring for Canva's generated element and template library.
  • Own the full ML lifecycle — from data pipelines and training through to deployment, monitoring, and iteration.
  • Partner with Content Engine, CORE AI Research, AI Media, and Discovery teams to align ML work with the broader content strategy.

Canva is redefining how the world experiences design with its intuitive design platform. We serve hundreds of millions of users globally and foster a culture of flexibility, inclusion, and innovation.

Global 4w PTO

  • Own the ML serving API and deploy models to production with CI/CD and infrastructure as code.
  • Build monitoring, alerting, and reliability for NBA models and LLM agents.
  • Drive architectural decisions and mentor engineers on MLOps patterns.

Clutch is a vertical SaaS company backed by Andreessen Horowitz, revolutionizing how credit unions engage with members via fintech lending software. The company is small and ambitious, with a lean data team of five that values pragmatism and fast shipping.

Italy

  • Lead and develop a high-performing team of MLOps engineers, fostering technical excellence and collaboration.
  • Define and execute the MLOps roadmap, aligning infrastructure initiatives with research, engineering, and product goals.
  • Design and maintain scalable ML infrastructure including automated training pipelines, CI/CD, and model serving platforms.

Our partner is a company focused on cutting-edge machine learning infrastructure for large-scale AI systems. They foster an inclusive, mission-driven culture with international collaboration and value innovation, diversity, and continuous learning.

$150,000–$160,000/yr

  • Lead engineers deploying, operating, and optimizing AI compute clusters at scale.
  • Oversee cluster reliability, GPU fleet operations, and incident response.
  • Coordinate with cross-functional teams to ensure cluster capabilities meet AI workloads.

Vultr makes high-performance cloud infrastructure easy to use, affordable, and locally accessible for global enterprises and AI innovators. It is the world’s largest privately-held cloud infrastructure company, trusted by hundreds of thousands of customers across 185 countries.

US

  • Build, ship, and own product features end-to-end using cutting-edge AI/ML techniques.
  • Apply classical ML and LLM-based approaches like RAG, prompt engineering, and fine-tuning to enhance the audit and risk platform.
  • Collaborate with cross-functional teams in an Agile environment to deliver scalable, production-quality code.

Optro is a leading audit, risk, ESG, and InfoSec platform trusted by over 50% of the Fortune 500. The company has been named one of the 500 fastest-growing tech companies in North America for seven consecutive years, fostering a culture of innovation and collaboration.

Brazil

  • You will experiment with emerging technologies and contribute to building new models and systems.
  • You will implement prototypes in Python and focus on delivering solutions to production.
  • You will partner with the platform engineering team to streamline MLOps workflows and maintain high code quality.

Verve creates a more efficient and privacy-focused way to buy and monetize advertising by fusing data, media, and technology. With 30 offices globally, they serve top advertisers and publishers and foster a collaborative, fun culture.