Source Job

US Unlimited PTO

  • Lead and mentor a team of software engineers building scalable machine learning infrastructure.
  • Define and own the multi-quarter technical strategy for data generation, orchestration, and distributed training.
  • Represent the Scalable ML team in org-level planning and architecture decisions, resolving cross-team tradeoffs on shared infrastructure.

Python Machine Learning Distributed Systems Cloud Infrastructure PyTorch

20 jobs similar to Senior Engineering Manager - Scalable Machine Learning

Jobs ranked by similarity.

Global

  • Design and implement cutting-edge machine learning algorithms and model architectures.
  • Collaborate with product, engineering, and analytics teams to integrate AI features across Atlassian products.
  • Conduct rigorous experimentation, model evaluation, and mentor emerging ML engineers.

Atlassian creates team collaboration software like Jira, Confluence, and Trello, aimed at unleashing the potential of every team. They are a distributed-first company with thousands of employees, emphasizing diversity, inclusion, and flexible work arrangements.

US

  • Design high-performance training platform components including orchestration, observability, and performance tuning.
  • Deliver end-to-end ML pipelines from dataset curation to training, validation, and deployment.
  • Drive technical direction across ML Platform, Infrastructure, Autonomy, and Safety teams.

Stack is developing revolutionary AI and advanced autonomous systems for safer, more reliable, and efficient operations, focusing on autonomous trucking. The Stack team brings decades of experience in real-world systems and is committed to a culture of inclusion, entrepreneurship, and innovation.

$171,063–$269,075/yr
US

  • Drive development and implementation of advanced machine-learning algorithms for business-critical applications.
  • Build scalable forecasting models, including deep-learning-based approaches, to predict company top-line metrics.
  • Lead the end-to-end ML lifecycle from problem definition through deployment and iteration.

Atlassian builds software products that help teams all over the planet collaborate and work together. With a distributed-first culture, Atlassians have flexibility in where they work, and the company is motivated by unleashing the potential of every team.

$170,170–$286,000/yr
North America

  • Design and maintain reliable, low-latency ML APIs to integrate Safety AI model outputs into cloud applications.
  • Build scalable data pipelines for continuous model iteration, backtesting, and online evaluation.
  • Optimize model artifacts for production and monitor rollout health, ensuring predictable failure modes.

Samsara builds a Connected Operations Cloud that helps physical operations use IoT data to improve safety, efficiency, and sustainability. Samsara is a recently public company with an employee-led remote culture and a long-term focus.

$230,000–$400,000/yr
Global

  • Lead the team in roadmapping and executing machine learning features and products.
  • Improve team pace through technical expertise and pragmatic tradeoffs.
  • Drive recruiting, reliability, and growth of team members.

Hightouch helps companies sync data into their SaaS systems to automate and improve operations. They are a remote-first company with hundreds of customers and a culture focused on engineering excellence.

US

  • Develop advanced ML models and agentic workflows to accelerate model development.
  • Use AI-assisted tools like Claude and Cursor to investigate model behavior and automate analysis.
  • Set technical direction, mentor engineers, and raise the bar for modeling rigor.

Reddit is a community of communities, built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.

$211,000–$249,000/yr
US

  • Build and own training pipelines for ML models, including data prep, fine-tuning, and experiment tracking.
  • Develop inference and serving layer with model gateway, routing, caching, and cost optimization.
  • Own release path for model-layer artifacts with canary deployments and monitoring for regressions.

Horizon3 is a fast-growing cybersecurity company that provides autonomous pentesting through its NodeZero platform. The team is a fusion of former special operations cyber operators, startup engineers, and cybersecurity practitioners, fostering a culture of respect, collaboration, and ownership.

North America

  • Design and build systems that improve the efficiency of ML training and inference workloads.
  • Develop tooling and benchmarking frameworks to help ML engineers debug, profile, optimize, and monitor model performance.
  • Partner with ML teams to optimize distributed training, reduce costs, and drive technical strategy for scalability and reliability.

Reddit is a community of communities where users share interests and authentic conversations. With 100,000+ communities and ~130 million daily visitors, it has a flexible-first culture supporting remote work.

$230,000–$322,000/yr
US

  • Lead the development and optimization of machine learning models to detect and prevent AI security risks like prompt injection and jailbreaks.
  • Build reproducible training and evaluation pipelines on Reddit's ML platform, partnering with platform engineers to improve performance and reliability.
  • Set the technical vision and multi-quarter modeling roadmap, mentoring engineers and establishing best practices for responsible ML development.

Reddit is a community of communities, built on shared interests and authentic conversations, with 100,000+ active communities and 130 million daily active visitors. It is one of the internet's largest sources of information, fostering a culture of openness and trust.

US Unlimited PTO

  • Build the foundations of the EdgeRunner Research organization, including data pipelines, model evaluation, and efficient parallelization.
  • Own codebases in areas like distributed training, quantization, compression, or compute cluster management.
  • Collaborate with a team of self-starters who operate independently and translate business objectives into technical solutions.

EdgeRunner AI builds state-of-the-art AI models for the tactical edge, enabling warfighters to make faster decisions and interact with robotic platforms. As a Series-A startup, we move quickly, fostering a culture of initiative, ownership, and comfort with ambiguity.

Spain

  • Design and develop large-scale platforms for LLM training and AI workloads.
  • Tackle distributed-systems challenges including intelligent job scheduling and resource optimization.
  • Collaborate with international teams to build production-ready AI infrastructure.

This role is with a partner company, an AI-focused R&D team building infrastructure for large language models. They are a fast-moving, highly technical team with a collaborative and innovative culture.

Canada

  • Architect, design, build, deploy, and maintain Model Serving infrastructure for a world-class Detection Engine.
  • Own projects that scale model serving and data processing to handle 10x traffic, including real-time streaming pipelines and online feature serving.
  • Collaborate with MLE and Data Science teams to build the ML Training platform, improving MLE velocity and model precision and recall.

Abnormal protects the humans behind the world's most critical organizations from AI-powered cybercrime. 4,500+ enterprises trust our behavioral AI platform, and we foster a culture of innovation and impact.

US Unlimited PTO

  • Build the component layer around our layout synthesis engine, including API contracts, services, and evaluation gates.
  • Turn model retraining into a one-command job with built-in benchmarks and readable results for the whole team.
  • Own latency and cost budgets for learned capabilities and partner with infrastructure engineers on MLOps and deployment.

Higharc is a VC-backed startup that is changing how new homes are designed and built using spatial AI and generative floor plan technology. The company is fully remote, has raised over $175M, and values flexibility, collaboration, and asynchronous deep work.

India

  • Lead the design, implementation, deployment, and operation of complex platform capabilities supporting the machine-learning lifecycle.
  • Build reusable platform components, standards, and automation while improving reliability, scalability, and observability across ML systems.
  • Partner with Data Science, Product, Risk, Fraud, and Engineering teams to translate ambiguous business needs into practical technical solutions.

Oportun is a mission-driven financial services company that helps members build a better financial future through intelligent borrowing, savings, and budgeting. The company has provided over $22.7 billion in responsible credit and its culture emphasizes speed, high standards, and the thoughtful use of AI.

$97,600–$139,000/yr
United States Canada

  • Build and maintain core infrastructure for Quora's ML platform, ensuring high availability, scalability, and performance.
  • Build and improve distributed systems serving ML models in production, from Large Recommendation Models to Large Language Models.
  • Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models.

Quora's mission is to grow the world's collective intelligence through two platforms: Quora for global knowledge sharing and Poe for AI agent collaboration. We are a remote-first company with passionate, collaborative, and high-performing global teams, rooted in transparency and experimentation.

US

  • Design and build autonomous AI agents using modern agentic frameworks to analyze infrastructure and make intelligent decisions.
  • Develop and deploy ML models and MLOps pipelines that learn from infrastructure patterns to optimize resource policies and scaling.
  • Own AI systems end-to-end, from architecture to production, ensuring performance, reliability, and cost-effectiveness.

ScaleOps is redefining autonomous cloud and AI infrastructure by freeing DevOps teams from manual resource management, reducing cloud costs by up to 80%. Backed by over $210M from leading VCs, the company is trusted by Fortune 100 companies and fosters an innovative, mission-driven culture.

Global

  • Design, build, and deploy production ML and LLM-based systems for enterprise clients.
  • Own technical delivery end-to-end from architecture to deployment and iteration.
  • Mentor and support other ML engineers through code reviews and technical guidance.

TensorOps is a boutique AI consultancy that designs and ships production-grade AI systems for enterprise clients. The company has shipped AI systems impacting 200M+ end users daily, partnered with 11 unicorns, and operates fully remotely with a supportive, fast-growing culture.

Canada Unlimited PTO

  • Write and ship production AI code daily as an active contributor.
  • Build the retrieval, graph, and inference layers on top of our data platform.
  • Set patterns other teams build against and turn company goals into shipped systems.

Acquia empowers brands to create digital customer experiences using its Drupal-based Digital Experience Platform. It is a Great Place to Work-Certified company with thousands of global organizations as clients.

UK

  • Build tooling for capturing and processing data from agents and humans at significant scale.
  • Solve hard problems around compute, orchestration, scaling, security, and reliability.
  • Help develop approaches for training, benchmarking, and evaluating AI agents.

Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.

US

  • Lead the strategy and direction of a high-impact Machine Learning and AI function, shaping innovation and production capabilities.
  • Build and develop a high-performing ML organization while fostering a culture of accountability, inclusion, and technical excellence.
  • Partner across Engineering, Product, and Research to ensure responsible, secure, and production-ready AI systems involving sensitive data.

The company is a mission-driven technology organization applying advanced machine learning and AI to complex challenges with meaningful social impact, including keeping children safe. It fosters a remote-first, collaborative, and inclusive culture with multidisciplinary teams.