Source Job

UK Germany Netherlands Ireland Spain Poland Bulgaria Unlimited PTO

  • Build and own the model serving infrastructure, real-time inference, feature retrieval, and the latency budget that governs both.
  • Build the deployment path for data scientists to ship models, including bring-your-own-model support.
  • Own models in production: monitoring, drift detection, retraining, incident response, and the on-call rotation.

Python Kubernetes GCP CI/CD Machine Learning

20 jobs similar to Machine Learning Engineer

Jobs ranked by similarity.

Global

  • Own and optimize CI/CD pipelines, Kubernetes deployment, and infrastructure for model serving and inference.
  • Build telemetry, observability, and alerting to catch real problems and reduce noise.
  • Eliminate toil through thoughtful automation and improve developer and agent productivity.

Obvious is building an AI-native workspace that serves as an operating system for work, putting co-intelligence at the center. They are a small, talent-dense team with founders and leaders from top tech companies.

Canada

  • Lead a team of platform engineers to build and operate ML training and serving infrastructure, including GPU and low-latency serving.
  • Combine strong people leadership with technical judgment in ML infrastructure, partnering with senior ICs and cross-functional teams.
  • Drive delivery, operational health, and evaluate modern ML tooling to support company-wide ML priorities.

Affirm is reinventing credit to make it more honest and friendly, offering buy now pay later solutions without hidden fees. It is a remote-first company with a strong engineering culture, prioritizing people and providing competitive benefits.

$153,351–$206,481/yr
Canada

  • Lead a team to build, scale, and optimize the ML infrastructure powering drug discovery.
  • Collaborate with ML engineering, data science, and research teams to deliver scalable solutions.
  • Mentor and coach team members in MLOps, distributed computing, and infrastructure engineering.

Recursion is a clinical-stage TechBio company decoding biology to develop medicines. With a focus on AI and machine learning, the company fosters a culture of bold integrity and cross-functional collaboration.

$149,200–$214,500/yr
US

  • Architect, design, build, deploy, and maintain Model Serving infrastructure using industry-standard AI tools.
  • Own projects that scale model serving and data processing services to handle 10x traffic.
  • Collaborate closely with MLE and Data Science teams to distill feedback and execute on strategy.

Abnormal AI protects the humans behind the world's most critical organizations from AI-powered cybercrime. Over 4,500 enterprises trust their behavioral AI platform, fostering a culture of innovation and security.

$140,000–$150,000/yr
United States Canada

  • Design and maintain ML model productionization infrastructure for high-visibility product features.
  • Collaborate with data science to streamline model training, validation, and deployment.
  • Implement robust monitoring and alerting for model performance, drift, and data quality.

The Athletic is a sports media company powered by one of the largest global newsrooms in sports, delivering in-depth coverage of professional and college teams across North America and Europe. With over 500 full-time staff, they foster a collaborative culture focused on high-quality journalism and data-driven innovation.

Europe

  • Design, deploy, and optimize production-grade machine learning systems for the full ML lifecycle.
  • Build scalable MLOps platforms, CI/CD workflows, and model serving infrastructure.
  • Collaborate with engineering teams to improve platform scalability, security, and operational best practices.

They build scalable MLOps infrastructure for enterprise AI solutions. They foster a collaborative, remote-first culture focused on innovation and professional growth.

$160,000–$190,000/yr
US

  • Build and improve ML components across data, training, evaluation, and inference.
  • Implement evaluation and testing to understand model behavior.
  • Debug model issues, performance problems, and production incidents.

This company builds core ML components for large-scale production systems. They emphasize real-world learning, iteration, and collaboration with senior engineers.

US

  • Design and develop machine learning infrastructure, tooling, and models to deliver world-class experiences.
  • Help product teams understand the data lifecycle and experimental nature of machine learning.
  • Build internal platforms to enable teams to incorporate AI into customer-facing products.

Weave provides AI-powered features and data services, handling data for hundreds of millions of people daily. The company fosters a culture of autonomy, collaboration, and innovation with cross-functional agile teams.

US

  • Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
  • Architect scalable inference systems and collaborate closely with application engineering.

This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

LATAM

  • Develop, deploy, and maintain production-grade Machine Learning models in cloud environments.
  • Build, maintain, and optimize data pipelines and monitor model performance post-deployment.
  • Partner with engineering, data, and business stakeholders to translate goals into scalable ML solutions.

In All Media is a global technology and design firm building impactful digital solutions through remote, distributed teams across LATAM. They partner with international clients across industries, providing long-term technical expertise and team augmentation.

$161,000–$273,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Lead the design and operation of production machine learning systems for batch and online use cases with a focus on reliability and scalability.
  • Build and improve ML lifecycle infrastructure including training pipelines, inference workflows, monitoring, and automation.
  • Partner with cross-functional teams to translate business problems into ML solutions and guide prototypes to robust production systems.

Included Health is a healthcare company delivering integrated virtual care and navigation, aiming to raise the standard of healthcare for everyone. They are a remote-first organization offering comprehensive benefits and fostering a culture of inclusion.

$138,500–$225,500/yr
US 16w maternity 16w paternity

  • Design, train, and ship ML systems for governance and security like anomaly detection and trust scoring.
  • Build data pipelines, model serving, evaluation frameworks, and feedback loops.
  • Set technical direction, own architecture, and help recruit and mentor as the team grows.

Docker provides developer tooling trusted by over 20 million monthly users and billions of container pulls. They are a globally distributed, remote-first team building tools for software delivery.

North America

  • Build and deploy production code to support customer AI inference workloads on Tenstorrent's hardware and software stack.
  • Debug and optimize across the full inference stack, from serving layer to kernel dispatch, and translate customer issues into actionable requirements.
  • Operate Kubernetes and observability tools to manage multi-node AI clusters and ensure reliability.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. Their diverse team of technologists has developed a high-performance RISC-V CPU from scratch, and they value collaboration, curiosity, and a commitment to solving hard problems.

$165,000–$330,000/yr
US Unlimited PTO

  • Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
  • Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
  • Work across product, software development, performance engineering, and customer-facing implementations.

Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.

$137,000–$252,000/yr
US Canada Unlimited PTO

  • Partner with product, data science, and engineering teams to design, develop, and deploy machine learning and optimization solutions for ads delivery.
  • Train, deploy, and scale ML models as microservices, working with infrastructure engineers to integrate model outputs into ads systems.
  • Establish unified logging, alerting, and monitoring for model inference, system latency, data drift, and concept drift.

Life360's mission is to keep people close to the ones they love through a mobile app, Tile tracking devices, and Pet GPS tracker. The company serves approximately 97.8 million monthly active users across 180+ countries and has a remote-first culture with over 500 employees.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

$154,000–$200,000/yr
US

  • Design, build, and deploy machine learning models for cybersecurity use cases like threat detection and risk scoring.
  • Own the full model lifecycle from data preparation to production deployment, working closely with engineering and product teams.
  • Build preprocessing and feature engineering pipelines, and monitor model performance with continuous improvement.

SpyCloud transforms recaptured darknet data to disrupt cybercrime. With over 250 employees, it is home to cybersecurity experts protecting businesses and consumers from stolen identity data.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.

US Unlimited PTO

  • Resolve complex escalations as the final authority, using code-level debugging and architectural investigation.
  • Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution.
  • Own end-to-end P1 resolution and deliver clear, actionable post-incident analysis.

TensorWave delivers a versatile cloud platform for AI compute at scale, eliminating infrastructure barriers. The company fosters a culture of innovation and reliability, empowering builders to focus on breakthrough AI.