Source Job

Europe

  • Design, deploy, and optimize production-grade machine learning systems for the full ML lifecycle.
  • Build scalable MLOps platforms, CI/CD workflows, and model serving infrastructure.
  • Collaborate with engineering teams to improve platform scalability, security, and operational best practices.

Python Kubernetes MLflow AWS PyTorch

20 jobs similar to MLOps Engineer

Jobs ranked by similarity.

Italy

  • Lead and develop a high-performing team of MLOps engineers, fostering technical excellence and collaboration.
  • Define and execute the MLOps roadmap, aligning infrastructure initiatives with research, engineering, and product goals.
  • Design and maintain scalable ML infrastructure including automated training pipelines, CI/CD, and model serving platforms.

Our partner is a company focused on cutting-edge machine learning infrastructure for large-scale AI systems. They foster an inclusive, mission-driven culture with international collaboration and value innovation, diversity, and continuous learning.

India

  • Collaborate with data scientists and engineers to build scalable ML pipelines, troubleshoot infrastructure issues from Linux to Kubernetes, and optimize model performance.
  • Drive high engineering standards, design on-premises MLOps solutions, and maintain tools for deployment and monitoring.
  • Refine CI/CD workflows, incorporate ML model training and evaluation into testing, and ensure seamless handover between research and production.

Learneo is a platform of builder-driven businesses, including Course Hero, CliffsNotes, LitCharts, Quillbot, Symbolab, and Scribbr, focused on supercharging productivity and learning. The company supports high-growth businesses with centralized corporate operations and has a virtual-first culture with employees across multiple countries.

Global 4w PTO

  • Own the ML serving API and deploy models to production with CI/CD and infrastructure as code.
  • Build monitoring, alerting, and reliability for NBA models and LLM agents.
  • Drive architectural decisions and mentor engineers on MLOps patterns.

Clutch is a vertical SaaS company backed by Andreessen Horowitz, revolutionizing how credit unions engage with members via fintech lending software. The company is small and ambitious, with a lean data team of five that values pragmatism and fast shipping.

Europe

  • Design, build, and automate enterprise-grade AWS SageMaker environments to support scalable machine learning initiatives.
  • Develop and implement DevOps automation for SageMaker Unified Studio and related cloud infrastructure.
  • Build and optimize CI/CD pipelines for deploying custom Docker images, kernels, and machine learning workloads.

This role is listed on behalf of a partner company that builds and optimizes enterprise-scale machine learning infrastructure. They operate with a collaborative international team and modern engineering practices.

India

  • Collaborate with data scientists and software engineers to build scalable data pipelines and ML deployment systems.
  • Troubleshoot issues across the ML infrastructure stack, from Linux and Docker to Kubernetes and model serving.
  • Drive high engineering standards through code reviews, testing, and CI/CD enhancements.

Quillbot helps students and professionals strengthen their writing with AI-powered tools. We serve over 56 million users globally and foster a collaborative, virtual-first culture.

US Unlimited PTO

  • Design and maintain scalable ML infrastructure including data pipelines, training workflows, and model deployment systems.
  • Own end-to-end ML lifecycle operations, ensuring reliable delivery of models into production at scale.
  • Implement monitoring, telemetry, and feedback loops for ML models running across large-scale device fleets.

Our partner company develops ML systems for connected hardware products used by customers worldwide. They operate in a fast-paced, product-driven environment with a collaborative and technically ambitious culture focused on real-world ML impact.

$220,000–$280,000/yr
US Unlimited PTO

  • Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
  • Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
  • Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.

Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.

$161,000–$273,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Lead the design and operation of production machine learning systems for batch and online use cases with a focus on reliability and scalability.
  • Build and improve ML lifecycle infrastructure including training pipelines, inference workflows, monitoring, and automation.
  • Partner with cross-functional teams to translate business problems into ML solutions and guide prototypes to robust production systems.

Included Health is a healthcare company delivering integrated virtual care and navigation, aiming to raise the standard of healthcare for everyone. They are a remote-first organization offering comprehensive benefits and fostering a culture of inclusion.

Europe

  • Build and scale massive distributed compute and storage systems for AI training.
  • Architect multi-cluster orchestration and optimize workload placement across global regions.
  • Design future-proof storage formats and implement metadata systems for exabyte-scale growth.

Mistral provides full-stack AI solutions, from frontier models to developer tools. They are a dynamic, collaborative team with a diverse workforce, passionate about innovation and low-ego teamwork.

Slovakia

  • Design and run Kubernetes environments optimized for AI inference, retrieval, and agent execution in secure settings.
  • Deploy and operate open-source model stacks, model gateways, vector databases, and supporting platform components.
  • Build reproducible platform automation using Infrastructure as Code and GitOps approaches for stable, auditable delivery.

Deutsche Telekom IT Solutions Slovakia provides innovative information and communication technology services. It has grown to become the second largest employer in eastern Slovakia with over 3900 employees, focusing on continuous transformation and improvement.

India

  • Design, build, and maintain scalable machine learning infrastructure on AWS, including training and deployment pipelines.
  • Develop and deploy ML models for recommendation systems, fraud detection, credit risk, and personalization use cases.
  • Implement monitoring, logging, and alerting systems to ensure model performance, stability, and reliability in production.

Our partner is a fast-growing, innovation-driven company where machine learning and AI systems directly power large-scale fintech and commerce experiences. They foster a highly dynamic environment with strong emphasis on experimentation, rapid iteration, and measurable business impact.

Brazil

  • Build end-to-end Machine Learning pipelines, including data preparation, code optimization, automated model training, and prediction workflows.
  • Provide technical leadership in MLOps processes, establishing best practices for development, deployment, and maintenance.
  • Design and implement cloud-based architectures for AI and data solutions using platforms such as GCP, AWS, or Azure.

A technology company specializing in data-driven solutions for consumers, leveraging AI and cloud technologies to transform data into insights. The company fosters a dynamic, innovation-driven culture with a remote-first approach, collaborating across multidisciplinary teams.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

Argentina

  • Design, build, and optimize data workflows for Machine Learning and GenAI solutions in cloud environments.
  • Develop and deploy Machine Learning and Generative AI models using AWS SageMaker.
  • Create and manage data and model pipelines to improve the efficiency of AI and machine learning systems.

Netrix Global provides the people, processes, and technology to run and scale modern data-driven businesses. It is a top system integrator with a culture focused on ownership, teamwork, and respect.

Portugal

  • Design, train, and evaluate machine learning models to address business problems.
  • Build and maintain data pipelines and infrastructure for model development and deployment.
  • Deploy ML models into production and monitor performance, reliability, and drift.

Critical Software delivers software solutions and consulting in complex, business-critical environments across industries like aerospace, defense, and healthcare. They are a Benefit Corporation committed to positive impact and an equal opportunity employer.

Brazil

  • You will experiment with emerging technologies and contribute to building new models and systems.
  • You will implement prototypes in Python and focus on delivering solutions to production.
  • You will partner with the platform engineering team to streamline MLOps workflows and maintain high code quality.

Verve creates a more efficient and privacy-focused way to buy and monetize advertising by fusing data, media, and technology. With 30 offices globally, they serve top advertisers and publishers and foster a collaborative, fun culture.

EMEA

  • Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
  • Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
  • Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.

They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.

Brazil

  • Lead the design and development of scalable AI solutions, from experimentation to production deployment.
  • Define AI engineering standards and best practices, influencing architecture decisions across teams.
  • Collaborate with cross-functional stakeholders to integrate generative AI and LLM capabilities into products.

The company is at the forefront of AI-driven product development, focusing on building scalable and intelligent systems. It fosters a culture of innovation and technical excellence, with a remote team and a commitment to engineering leadership.

France 5w PTO

  • Own ML models across their full lifecycle from data pipelines to deployment and monitoring, ensuring reliable performance.
  • Run and improve the ML platform including GitOps CI/CD, monitoring serving endpoints, and defining SLOs.
  • Collaborate with risk, operational, and product teams to turn ML into business value across the organization.

Alma provides installment and deferred payment solutions to help merchants boost sales and customer loyalty, without encouraging bad debt. With over 25,000 merchants, 10 million consumers, 380+ employees, and over €100M ARR, they are a Next40 member scaling rapidly across Europe.

$190,000–$225,000/yr
US

  • Build and maintain end-to-end deployment pipelines for AI-powered applications, including artifact builds, environment promotion, rollback, and observability hooks.
  • Stand up and operate the runtime and lifecycle infrastructure for production agents, including deployment, versioning, monitoring, rate-limiting, and retirement.
  • Design and build the shared developer harness that every AI-powered service uses: prompt management, model routing, retries, tracing, eval hooks, and policy enforcement.

RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company offers innovative solutions with integrated intelligence on a single enterprise platform, connecting the pharmacy ecosystem.