Source Job

$176,000–$195,000/yr
Canada

  • Lead the technical vision and roadmap for the ML platform.
  • Partner with Product and Engineering leadership to align investments.
  • Establish MLOps practices and optimize large-scale model training and serving.

Python Java Kubernetes AWS MLflow

20 jobs similar to Principal ML System Engineer

Jobs ranked by similarity.

$195,000–$217,000/yr
US

  • Define the technical vision and strategy for the machine learning platform supporting ML and generative AI development.
  • Set reference architectures and standards for scalable data and ML pipelines, MLOps practices, and production reliability.
  • Provide technical leadership and mentorship across engineering teams to raise the bar for ML systems.

PointClickCare provides cloud-based healthcare software solutions. With a large engineering team that values collaboration and innovation, they foster a culture of technical excellence and mentorship.

$150,500–$173,000/yr
US

  • Design, deploy, and maintain scalable ML infrastructure for model training, batch processing, and real-time inference.
  • Build and manage cloud-based infrastructure with AWS and Snowflake using Infrastructure-as-Code practices.
  • Develop CI/CD pipelines, automation frameworks, and monitoring for ML systems to improve reliability and governance.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They operate with a team-oriented culture and offer remote work flexibility, focusing on efficient, unbiased recruitment.

Europe

  • Design, deploy, and optimize production-grade machine learning systems for the full ML lifecycle.
  • Build scalable MLOps platforms, CI/CD workflows, and model serving infrastructure.
  • Collaborate with engineering teams to improve platform scalability, security, and operational best practices.

They build scalable MLOps infrastructure for enterprise AI solutions. They foster a collaborative, remote-first culture focused on innovation and professional growth.

Canada

  • Design and build scalable backend services for ML feature management, storage, and serving.
  • Own technical initiatives from design through implementation, balancing trade-offs and navigating ambiguity.
  • Collaborate with cross-functional teams to define solutions that align with business objectives and improve platform reliability.

The company is a technology firm focused on building scalable machine learning platforms and intelligent decision-making systems. It operates in a remote-first environment with a collaborative culture and a commitment to engineering excellence.

$100,000–$130,000/yr
Canada

  • Design, develop, and deploy foundational AI and ML models, building robust, scalable pipelines for advanced analytics.
  • Champion top-tier coding, testing, and MLOps practices, navigating ambiguity to refine pipelines and elevate workflows.
  • Partner with cross-functional stakeholders to convert strategic needs into technical specs and embed ML features into live applications.

Wave helps small businesses thrive so the heart of our communities beats stronger. They create an environment buzzing with creative energy and inspiration, valuing boldness, quick learning, and generous knowledge sharing.

Brazil

  • Build end-to-end Machine Learning pipelines, including data preparation, code optimization, automated model training, and prediction workflows.
  • Provide technical leadership in MLOps processes, establishing best practices for development, deployment, and maintenance.
  • Design and implement cloud-based architectures for AI and data solutions using platforms such as GCP, AWS, or Azure.

A technology company specializing in data-driven solutions for consumers, leveraging AI and cloud technologies to transform data into insights. The company fosters a dynamic, innovation-driven culture with a remote-first approach, collaborating across multidisciplinary teams.

$80,000–$120,000/yr
Global Unlimited PTO

  • Design, build, and operate high-load distributed backend services powering the company's ML infrastructure.
  • Take end-to-end ownership of core ML services and data pipelines from design to deployment and continuous improvement.
  • Partner with ML and product teams to understand their needs and turn them into reliable, reusable platform capabilities.

Constructor is an AI-first e-commerce search and discovery platform that helps shoppers find products and enables brands to drive revenue. The company is fully remote, diverse, and values ownership and collaboration.

US Unlimited PTO

  • Design, build, and operate scalable backend services, APIs, and data pipelines for ML-driven personalization.
  • Improve reliability, performance, and observability of production ML systems, including model versioning and safe rollout.
  • Collaborate with data scientists and product engineers to translate business needs into robust technical solutions.

Hungryroot uses AI to build a consumer-centric food and wellness company, acting as a personal assistant for healthy living by recommending and delivering healthy groceries, recipes, and supplements. They are a distributed team across 28+ US states with a remote-first culture emphasizing collaboration, flexibility, and an annual company retreat.

$140,000–$150,000/yr
United States Canada

  • Design and maintain ML model productionization infrastructure for high-visibility product features.
  • Collaborate with data science to streamline model training, validation, and deployment.
  • Implement robust monitoring and alerting for model performance, drift, and data quality.

The Athletic is a sports media company powered by one of the largest global newsrooms in sports, delivering in-depth coverage of professional and college teams across North America and Europe. With over 500 full-time staff, they foster a collaborative culture focused on high-quality journalism and data-driven innovation.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

Global

  • Own end-to-end delivery quality for major engagements, translating ambiguous client needs into practical execution plans.
  • Lead solution architecture and technical decision-making, making pragmatic tradeoffs between speed, quality, and client value.
  • Build and ship production AI/ML systems using Python, ML frameworks, and cloud-native infrastructure while mentoring other engineers.

Eliza is a technology services company and Advanced-tier OpenAI partner that helps organizations build and deploy AI solutions, from generative AI to predictive analytics. They are a collaborative, mission-driven team focused on real-world AI impact.

$161,000–$221,500/yr
United States

  • Drive AI platform architecture and execute the long-term roadmap for machine learning and generative AI.
  • Lead AI infrastructure vision including end-to-end training, fine-tuning, and low-latency inference platforms.
  • Set engineering excellence standards for the full AI/ML SDLC and mentor senior engineers.

Lyra Health is a leading provider of evidence-based mental health care, serving over 20 million people globally. The company operates with a focus on transformative care and has delivered over 15 million sessions, employing a large team that includes engineers, data scientists, and clinical professionals.

$147,900–$203,000/yr
US 6w PTO

  • Lead the development and improvement of MLOps platform capabilities supporting machine learning workflows across Oura.
  • Drive design and unification of workflows and tooling for reliable training, orchestration, and deployment of ML systems.
  • Partner with data scientists and engineers to improve the end-to-end ML lifecycle, including governance and maintainability.

Oura empowers people to own their inner potential through its award-winning Oura Ring and app, providing daily health insights. As a quickly growing company, we focus on team well-being and helping people live healthier, happier lives.

$206,261–$330,017/yr
US

  • Design, build, deploy, and optimize machine learning models that process large volumes of complex, unstructured data.
  • Develop and maintain scalable ML pipelines capable of supporting millions of documents and diverse customer requirements.
  • Lead technical initiatives from early experimentation through production implementation and ongoing improvement.

The company develops intelligent automation solutions using AI and machine learning for large-scale document processing. It offers a highly autonomous engineering environment with a focus on continuous improvement and innovation.

Canada

  • Lead a team of platform engineers to build and operate ML training and serving infrastructure, including GPU and low-latency serving.
  • Combine strong people leadership with technical judgment in ML infrastructure, partnering with senior ICs and cross-functional teams.
  • Drive delivery, operational health, and evaluate modern ML tooling to support company-wide ML priorities.

Affirm is reinventing credit to make it more honest and friendly, offering buy now pay later solutions without hidden fees. It is a remote-first company with a strong engineering culture, prioritizing people and providing competitive benefits.

US

  • Design, build, deploy, and maintain scalable machine learning and AI systems in GCP to solve complex business problems and create new value.
  • Partner across teams to understand problem definitions, data needs, and solution approaches, and support model deployment with MLOps practices.
  • Take ownership of production issues, perform root cause analysis, and improve documentation and quality assurance processes for ML systems.

General Mills makes food the world loves, operating across 100+ markets. As a large company, it prioritizes being a force for good and fosters a culture of bold thinking and big hearts, where employees challenge each other and grow together.

US

  • Lead a high-performing AI and data science team to develop and deploy AI-driven security systems.
  • Define and execute the technical roadmap for AI, ML, and data systems, ensuring robust performance and scalability.
  • Manage the entire lifecycle of data pipelines and AI models, integrating with platforms like AWS Bedrock and OpenAI.

Bugcrowd is a crowdsourced security platform that unites organizations with a global network of hackers to identify vulnerabilities. The company is backed by prominent investors and fosters a collaborative, inclusive culture with a diverse team of professionals.

$138,500–$225,500/yr
US 16w maternity 16w paternity

  • Design, train, and ship ML systems for governance and security like anomaly detection and trust scoring.
  • Build data pipelines, model serving, evaluation frameworks, and feedback loops.
  • Set technical direction, own architecture, and help recruit and mentor as the team grows.

Docker provides developer tooling trusted by over 20 million monthly users and billions of container pulls. They are a globally distributed, remote-first team building tools for software delivery.

$127,750–$160,600/yr
United States Canada

  • Design, build, and operate the online feature store serving ML features to production models with low latency and high reliability.
  • Build and maintain data pipelines (batch and streaming) that compute, validate, and publish features from source systems.
  • Act as a technical leader for feature store infrastructure, partnering with Data Science and Engineering teams to drive quality and scalability.

Forward Financing is a financial technology company on a mission to unlock capital for small businesses across America. The company has provided over $3.5 billion in funding to 71,000+ small businesses and is recognized as a Best Place to Work, with a culture focused on employee investment and customer experience.

Global

  • Design, build, and deploy production ML and LLM-based systems (RAG, agentic workflows, fine-tuning, embeddings) for enterprise clients.
  • Own technical delivery end-to-end: from architecture and prototyping to deployment, monitoring, and iteration.
  • Mentor and support other ML engineers on the team with code reviews, technical guidance, and knowledge sharing.

TensorOps is a boutique AI consultancy that bridges strategy and execution, designing and shipping production-grade AI systems for enterprise clients. We are a 100% remote team of 11+ people, partnering with unicorns and NASDAQ-listed companies, and have a culture of autonomy, open communication, and continuous learning.