Source Job

US

  • Design high-performance training platform components including orchestration, observability, and performance tuning.
  • Deliver end-to-end ML pipelines from dataset curation to training, validation, and deployment.
  • Drive technical direction across ML Platform, Infrastructure, Autonomy, and Safety teams.

Python C++ PyTorch

20 jobs similar to Staff Software Engineer, ML Training Infrastructure

Jobs ranked by similarity.

$170,170–$286,000/yr
North America

  • Design and maintain reliable, low-latency ML APIs to integrate Safety AI model outputs into cloud applications.
  • Build scalable data pipelines for continuous model iteration, backtesting, and online evaluation.
  • Optimize model artifacts for production and monitor rollout health, ensuring predictable failure modes.

Samsara builds a Connected Operations Cloud that helps physical operations use IoT data to improve safety, efficiency, and sustainability. Samsara is a recently public company with an employee-led remote culture and a long-term focus.

North America

  • Design and build systems that improve the efficiency of ML training and inference workloads.
  • Develop tooling and benchmarking frameworks to help ML engineers debug, profile, optimize, and monitor model performance.
  • Partner with ML teams to optimize distributed training, reduce costs, and drive technical strategy for scalability and reliability.

Reddit is a community of communities where users share interests and authentic conversations. With 100,000+ communities and ~130 million daily visitors, it has a flexible-first culture supporting remote work.

Europe

  • Own end-to-end technical strategy for major areas of the Voyager SDK’s application layer, including model deployment and pipeline development.
  • Set technical direction for image processing, create reusable architectures, and partner with the compiler team to resolve compilation issues.
  • Represent the Applications team’s capabilities to business stakeholders, mentor engineers, and drive documentation standards.

Axelera AI is a deep-tech company creating a next-generation AI platform to advance humanity. With $450 million raised and 250+ employees including 60+ PhDs, they have a strong business pipeline exceeding $100 million and are committed to innovation.

$97,600–$139,000/yr
United States Canada

  • Build and maintain core infrastructure for Quora's ML platform, ensuring high availability, scalability, and performance.
  • Build and improve distributed systems serving ML models in production, from Large Recommendation Models to Large Language Models.
  • Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models.

Quora's mission is to grow the world's collective intelligence through two platforms: Quora for global knowledge sharing and Poe for AI agent collaboration. We are a remote-first company with passionate, collaborative, and high-performing global teams, rooted in transparency and experimentation.

US

  • Develop and train ML models for learned behavior systems using imitation and reinforcement learning.
  • Write production-quality ML code for training, evaluation, and inference in the autonomy stack.
  • Collaborate with teams to test and integrate learned behavior models across diverse driving environments.

Torc is an autonomous vehicle technology company that develops software for automated trucks, with a goal to transform freight movement. Now a part of the Daimler family, the company fosters a collaborative, energetic, and team-focused culture.

US

  • Develop advanced ML models and agentic workflows to accelerate model development.
  • Use AI-assisted tools like Claude and Cursor to investigate model behavior and automate analysis.
  • Set technical direction, mentor engineers, and raise the bar for modeling rigor.

Reddit is a community of communities, built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.

US Unlimited PTO

  • Build the component layer around our layout synthesis engine, including API contracts, services, and evaluation gates.
  • Turn model retraining into a one-command job with built-in benchmarks and readable results for the whole team.
  • Own latency and cost budgets for learned capabilities and partner with infrastructure engineers on MLOps and deployment.

Higharc is a VC-backed startup that is changing how new homes are designed and built using spatial AI and generative floor plan technology. The company is fully remote, has raised over $175M, and values flexibility, collaboration, and asynchronous deep work.

US

  • Partner with cross-functional stakeholders to define a vision for state-of-the-art AV data introspection software.
  • Deliver a pragmatic implementation roadmap while setting a high bar for engineering excellence.
  • Drive technical discussions and evangelize solutions to maximize business impact across the company.

Stack AV develops revolutionary AI and advanced autonomous systems for the trucking industry to enhance safety and efficiency. The team has decades of experience deploying real-world systems and is committed to innovation, safety, and inclusion.

Canada

  • Architect, design, build, deploy, and maintain Model Serving infrastructure for a world-class Detection Engine.
  • Own projects that scale model serving and data processing to handle 10x traffic, including real-time streaming pipelines and online feature serving.
  • Collaborate with MLE and Data Science teams to build the ML Training platform, improving MLE velocity and model precision and recall.

Abnormal protects the humans behind the world's most critical organizations from AI-powered cybercrime. 4,500+ enterprises trust our behavioral AI platform, and we foster a culture of innovation and impact.

UK

  • Build tooling for capturing and processing data from agents and humans at significant scale.
  • Solve hard problems around compute, orchestration, scaling, security, and reliability.
  • Help develop approaches for training, benchmarking, and evaluating AI agents.

Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.

$180,000–$270,000/yr
US

  • Build and operate production ML systems for recommendations, ranking, and personalized action selection across Stripe's Growth Platform.
  • Improve contextual bandit and policy-learning approaches, including exploration, reward design, and adaptive user feedback.
  • Develop reliable data and feature pipelines, reusable tooling, and online experiments to drive measurable business outcomes.

Stripe is a financial infrastructure platform for businesses of all sizes, enabling payments and revenue growth. With millions of companies as customers, Stripe fosters a culture of passion, grit, and integrity while working to increase the GDP of the internet.

Spain

  • Design and develop large-scale platforms for LLM training and AI workloads.
  • Tackle distributed-systems challenges including intelligent job scheduling and resource optimization.
  • Collaborate with international teams to build production-ready AI infrastructure.

This role is with a partner company, an AI-focused R&D team building infrastructure for large language models. They are a fast-moving, highly technical team with a collaborative and innovative culture.

Global

  • Build and improve the inference layer of the Gcore Inference platform, integrating frameworks like vLLM and TensorRT-LLM.
  • Bring new language and multimodal models into production, optimizing latency, throughput, and cost efficiency.
  • Debug performance issues across model code, GPU execution, and Kubernetes, collaborating with cross-functional teams.

Gcore is a global provider of AI, cloud, network, and security infrastructure and software. They are a team of 550+ professionals with a collaborative culture and partnerships with Intel, NVIDIA, Dell, and Equinix.

AI Engineer

Cyvl
US

  • Build and improve computer vision and 3D perception models that detect and assess infrastructure from LiDAR and imagery.
  • Develop LLM-powered products like agentic workflows, MCP servers, and AI copilots for customer tools.
  • Own models end-to-end, from data and training through evaluation, deployment, and monitoring in production.

Cyvl is a Physical AI company building purpose-built sensors, computer vision, and AI to turn every drive into current data for infrastructure. We're a fast-moving Boston startup with 500+ cities using our platform, and a culture of ownership, intensity, and care.

$180,000–$250,000/yr
Global

  • Work directly with leading AI labs and enterprises to define research goals and technical requirements.
  • Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
  • Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.

Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.

$230,000–$330,000/yr
US Unlimited PTO

  • Own Merlin's foundation and world-model work, including architecture selection, post-training, and capability roadmap.
  • Lead and mentor a small team of world-model engineers, setting the technical bar and review culture.
  • Design model interface to the autonomy stack with structured, schema-constrained plan outputs and build evaluation harnesses.

Merlin is a publicly traded aerospace and defense company building a non-human pilot for full-stack aircraft autonomy. Headquartered in Boston, it is expanding its organization to accelerate the deployment of its autonomy platform.

Southeast Asia

  • Serve as the technical bridge between Tenstorrent and customers across Southeast Asia, guiding evaluations and deployments of AI platforms.
  • Troubleshoot hardware and software issues, and translate complex AI/ML concepts for both technical and non-technical audiences.
  • Build trust with customers and internal teams through proactive problem solving and regional travel.

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. They value collaboration, curiosity, and a commitment to solving hard problems.

Global

  • Automate quality control for training data produced by companies using the platform's infrastructure.
  • Build QC systems grounded in human judgment, define quality standards, and design experiments.
  • Partner with data vendors to debug quality issues and feed learnings back into infrastructure tools.

The company is a fast-growing AI infrastructure startup focused on reinforcement learning environments and post-training data. With a roughly 15-person engineering team of published researchers and experienced AI practitioners, the culture is fast-paced, unstructured, and research-driven.

India

  • Design, build, and ship production services, APIs, and user-facing interfaces.
  • Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
  • Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.

Motive empowers people who run physical operations with tools to make their work safer, more productive, and more profitable. Serving nearly 100,000 customers across industries, the company values a diverse and inclusive workplace.

$189,507–$274,604/yr
Global

  • Develop and improve ads ranking models, including prediction objectives, feature interactions, user-history modeling, and calibration.
  • Take end-to-end ownership of machine learning systems from data pipelines to production integration.
  • Evaluate and apply advances in deep learning and recommendation modeling to improve ads ranking within production constraints.

Quora operates two knowledge-sharing platforms: Quora, a global Q&A platform, and Poe, a platform for interacting with AI language models. They have a remote-first culture with passionate, collaborative, and high-performing global teams focused on transparency and experimentation.