Source Job

$150,500–$173,000/yr
US

  • Design, deploy, and maintain scalable ML infrastructure for model training, batch processing, and real-time inference.
  • Build and manage cloud-based infrastructure with AWS and Snowflake using Infrastructure-as-Code practices.
  • Develop CI/CD pipelines, automation frameworks, and monitoring for ML systems to improve reliability and governance.

Python AWS Docker Kubernetes Terraform

20 jobs similar to Senior Machine Learning Ops Engineer

Jobs ranked by similarity.

India

  • Collaborate with data scientists and engineers to build scalable ML pipelines, troubleshoot infrastructure issues from Linux to Kubernetes, and optimize model performance.
  • Drive high engineering standards, design on-premises MLOps solutions, and maintain tools for deployment and monitoring.
  • Refine CI/CD workflows, incorporate ML model training and evaluation into testing, and ensure seamless handover between research and production.

Learneo is a platform of builder-driven businesses, including Course Hero, CliffsNotes, LitCharts, Quillbot, Symbolab, and Scribbr, focused on supercharging productivity and learning. The company supports high-growth businesses with centralized corporate operations and has a virtual-first culture with employees across multiple countries.

US Unlimited PTO

  • Design and maintain scalable ML infrastructure including data pipelines, training workflows, and model deployment systems.
  • Own end-to-end ML lifecycle operations, ensuring reliable delivery of models into production at scale.
  • Implement monitoring, telemetry, and feedback loops for ML models running across large-scale device fleets.

Our partner company develops ML systems for connected hardware products used by customers worldwide. They operate in a fast-paced, product-driven environment with a collaborative and technically ambitious culture focused on real-world ML impact.

Italy

  • Lead and develop a high-performing team of MLOps engineers, fostering technical excellence and collaboration.
  • Define and execute the MLOps roadmap, aligning infrastructure initiatives with research, engineering, and product goals.
  • Design and maintain scalable ML infrastructure including automated training pipelines, CI/CD, and model serving platforms.

Our partner is a company focused on cutting-edge machine learning infrastructure for large-scale AI systems. They foster an inclusive, mission-driven culture with international collaboration and value innovation, diversity, and continuous learning.

India

  • Collaborate with data scientists and software engineers to build scalable data pipelines and ML deployment systems.
  • Troubleshoot issues across the ML infrastructure stack, from Linux and Docker to Kubernetes and model serving.
  • Drive high engineering standards through code reviews, testing, and CI/CD enhancements.

Quillbot helps students and professionals strengthen their writing with AI-powered tools. We serve over 56 million users globally and foster a collaborative, virtual-first culture.

Europe

  • Design, deploy, and optimize production-grade machine learning systems for the full ML lifecycle.
  • Build scalable MLOps platforms, CI/CD workflows, and model serving infrastructure.
  • Collaborate with engineering teams to improve platform scalability, security, and operational best practices.

They build scalable MLOps infrastructure for enterprise AI solutions. They foster a collaborative, remote-first culture focused on innovation and professional growth.

$150,000–$195,000/yr
US Unlimited PTO

  • Design, secure, and scale modern cloud infrastructure to improve reliability and developer experiences.
  • Lead initiatives across infrastructure security, compliance, CI/CD, observability, and cloud operations.
  • Collaborate with engineering, IT, and business teams to solve complex security challenges and drive continuous improvement.

The company is a technology organization focused on building secure, scalable cloud infrastructure. It fosters a remote-first culture with an emphasis on innovation, ownership, and collaboration.

$206,261–$330,017/yr
US

  • Design, build, deploy, and optimize machine learning models that process large volumes of complex, unstructured data.
  • Develop and maintain scalable ML pipelines capable of supporting millions of documents and diverse customer requirements.
  • Lead technical initiatives from early experimentation through production implementation and ongoing improvement.

The company develops intelligent automation solutions using AI and machine learning for large-scale document processing. It offers a highly autonomous engineering environment with a focus on continuous improvement and innovation.

$190,000–$225,000/yr
US

  • Build and maintain end-to-end deployment pipelines for AI-powered applications, including artifact builds, environment promotion, rollback, and observability hooks.
  • Stand up and operate the runtime and lifecycle infrastructure for production agents, including deployment, versioning, monitoring, rate-limiting, and retirement.
  • Design and build the shared developer harness that every AI-powered service uses: prompt management, model routing, retries, tracing, eval hooks, and policy enforcement.

RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, the company offers innovative solutions with integrated intelligence on a single enterprise platform, connecting the pharmacy ecosystem.

Europe

  • Design, build, and automate enterprise-grade AWS SageMaker environments to support scalable machine learning initiatives.
  • Develop and implement DevOps automation for SageMaker Unified Studio and related cloud infrastructure.
  • Build and optimize CI/CD pipelines for deploying custom Docker images, kernels, and machine learning workloads.

This role is listed on behalf of a partner company that builds and optimizes enterprise-scale machine learning infrastructure. They operate with a collaborative international team and modern engineering practices.

$161,000–$273,000/yr
US Unlimited PTO 12w maternity 12w paternity

  • Lead the design and operation of production machine learning systems for batch and online use cases with a focus on reliability and scalability.
  • Build and improve ML lifecycle infrastructure including training pipelines, inference workflows, monitoring, and automation.
  • Partner with cross-functional teams to translate business problems into ML solutions and guide prototypes to robust production systems.

Included Health is a healthcare company delivering integrated virtual care and navigation, aiming to raise the standard of healthcare for everyone. They are a remote-first organization offering comprehensive benefits and fostering a culture of inclusion.

$220,000–$280,000/yr
US Unlimited PTO

  • Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
  • Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
  • Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.

Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.

$122,000–$129,200/yr
US Unlimited PTO 16w maternity 4w paternity

  • Build, operate, and scale AWS infrastructure supporting a microservices platform.
  • Develop and maintain CI/CD pipelines to make shipping fast, reliable, and safe.
  • Partner with product engineering teams, translating platform concepts into practical guidance.

Storyblocks is a content company that provides video, audio, images, and creative tools through a subscription model, empowering storytellers. They are a remote-first company with a culture driven by data, empowered communication, and ownership, recognized as a top workplace.

$160,000–$208,000/yr
US Unlimited PTO

  • Design, build, and optimize reliable infrastructure for healthcare technology.
  • Improve scalability, reliability, and performance across distributed systems.
  • Collaborate with engineers and data professionals to shape modern infrastructure practices.

This company provides innovative healthcare technology solutions. It fosters a remote-first culture with a focus on engineering excellence and collaboration.

Brazil

  • Build end-to-end Machine Learning pipelines, including data preparation, code optimization, automated model training, and prediction workflows.
  • Provide technical leadership in MLOps processes, establishing best practices for development, deployment, and maintenance.
  • Design and implement cloud-based architectures for AI and data solutions using platforms such as GCP, AWS, or Azure.

A technology company specializing in data-driven solutions for consumers, leveraging AI and cloud technologies to transform data into insights. The company fosters a dynamic, innovation-driven culture with a remote-first approach, collaborating across multidisciplinary teams.

$150,000–$200,000/yr
US Unlimited PTO

  • Design, build, and maintain scalable cloud infrastructure on AWS using Infrastructure as Code best practices.
  • Own and evolve the Atmos-based IaC framework and manage CI/CD pipelines with GitHub Actions.
  • Collaborate with engineers and scientists to support containerized and geospatial workloads.

Vibrant Planet develops a cloud-based AI platform to manage wildfire risk and modernize land management. They are a small, high-impact team backed by climate leaders, working on pressing climate challenges.

Argentina

  • Design, build, and optimize data workflows for Machine Learning and GenAI solutions in cloud environments.
  • Develop and deploy Machine Learning and Generative AI models using AWS SageMaker.
  • Create and manage data and model pipelines to improve the efficiency of AI and machine learning systems.

Netrix Global provides the people, processes, and technology to run and scale modern data-driven businesses. It is a top system integrator with a culture focused on ownership, teamwork, and respect.

India

  • Design, build, and maintain scalable machine learning infrastructure on AWS, including training and deployment pipelines.
  • Develop and deploy ML models for recommendation systems, fraud detection, credit risk, and personalization use cases.
  • Implement monitoring, logging, and alerting systems to ensure model performance, stability, and reliability in production.

Our partner is a fast-growing, innovation-driven company where machine learning and AI systems directly power large-scale fintech and commerce experiences. They foster a highly dynamic environment with strong emphasis on experimentation, rapid iteration, and measurable business impact.

$145,000–$175,000/yr
US 6w PTO

  • Leading design and implementation of robust, scalable, secure cloud-native solutions on AWS.
  • Developing and maintaining infrastructure-as-code for managing infrastructure across numerous Azure and AWS accounts.
  • Maintaining and optimizing CI/CD automation pipelines for rapid and reliable software deployments.

Element 84 is a woman-owned small business that develops geospatial data processing pipelines and builds software for public, private, and non-profit sectors. The company fosters a culture of curiosity and respect, supporting a large remote workforce with a flexible work schedule.

Unlimited PTO

  • Own and evolve cloud infrastructure and CI/CD pipelines, driving projects from inception to deployment.
  • Gain deep knowledge of the backend stack (Java Spring Boot) and optimize system reliability and scalability.
  • Actively adopt AI tools to enhance productivity, automate infrastructure provisioning, and ensure security.

Archy is a Series B vertical SaaS company that provides AI-powered software to revolutionize dental practice management. The company has a growing, collaborative team with a remote-friendly culture.

$184,558–$232,617/yr
US

  • Develop highly scalable and reliable data systems on AWS cloud platform for data-centric products.
  • Collaborate with Data Science teams to incorporate Machine Learning algorithms using Python and data engineering tools.
  • Lead and mentor team members, contribute to Agile cycles, and ensure timely delivery of documented software.

Experian is a global data and technology company that powers opportunities for people and businesses worldwide. With 22,500 employees across 32 countries, they foster a culture of innovation and data-driven solutions.