Source Job

Global

  • Design and scale production ML systems for LLM-based applications.
  • Build training and evaluation pipelines for continuous model improvement.
  • Fine-tune foundation models using modern adaptation techniques such as LoRA, QLoRA, SFT and DPO.

Python PyTorch

20 jobs similar to Technical Lead, Machine Learning

Jobs ranked by similarity.

US

  • Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
  • Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.

Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.

US

  • Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
  • Architect scalable inference systems and collaborate closely with application engineering.

This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.

United States

  • Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment.
  • Implement evaluation pipelines covering performance, robustness, safety, and bias.
  • Own production deployment including GPU optimization, memory efficiency, latency reduction, and scaling policies.

US

  • Develop and operate production-ready AI and machine learning systems for enterprise-scale products.
  • Build and optimize LLM-powered applications, RAG pipelines, and intelligent agents.
  • Implement software engineering best practices for AI development including CI/CD and testing.

Our partner is building enterprise-grade AI solutions that deliver measurable business impact. They offer a remote-friendly work environment with a collaborative engineering culture focused on innovation, quality, and continuous learning.

  • Build and ship AI features end-to-end, from model to system to user experience.
  • Design and iterate on prompts, tools, memory, and agent workflows for real-world reliability.
  • Debug full-stack issues and optimize for latency, cost, and production performance.

A1 builds a proactive smart assistant for everyday users, bringing intelligence to conversations, errands, organizing, and workflows with minimal prompting. The team is small, world-class, and focuses on rapid iteration and shipping high-quality AI products.

$216,700–$303,400/yr
US

  • Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
  • Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
  • Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.

Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.

United States

  • Lead technical discovery with foundation model labs, frontier AI teams, and large enterprises to understand model objectives and constraints.
  • Design end-to-end solutions across the post-training stack including SFT data curation, RLHF/DPO pipelines, custom benchmarks, and LLM-as-judge systems.
  • Author technical proposals, run workshops and POCs, and serve as ongoing technical advisor during delivery.

Innodata is a global data engineering company focused on enabling responsible AI advancement through data, evaluation frameworks, and human expertise. With a 36+ year legacy, they provide high-quality data solutions to foundation model labs, hyperscalers, and enterprise AI teams.

US

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.
  • Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.

We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.

Canada

  • Design, develop, and deploy production-grade AI-powered backend systems.
  • Integrate large language models and machine learning models into scalable architectures.
  • Optimize system performance and implement strong testing practices.

Our partner company is building advanced AI-powered systems to create meaningful customer value. The team operates in a high-autonomy, fast-moving environment focused on production-ready AI solutions.

India

  • Research and implement state-of-the-art techniques to accelerate AI inference: quantization, sparsity, distillation, speculative decoding, and caching.
  • Partner closely with hardware and compiler teams to ensure algorithmic improvements translate to real gains on custom silicon.
  • Build profiling tools and comprehensive benchmarking frameworks to measure model quality and efficiency.

EnCharge AI is building the next generation AI platform using novel in-memory-computing architecture. The team consists of experienced AI researchers, silicon & systems engineers, and architects backed by leading investors.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

Canada

  • Optimize machine learning inference systems for latency, throughput, and cost-efficiency.
  • Profile and troubleshoot GPU/CPU bottlenecks, implement advanced techniques like quantization and speculative decoding.
  • Collaborate with research and engineering teams to productionize new models and improve inference infrastructure.

The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.

France

  • Design and implement advanced knowledge distillation pipelines, including teacher-student approaches and multi-teacher architectures.
  • Run large-scale machine learning experiments to optimize model quality, latency, efficiency, and cost trade-offs.
  • Collaborate with research teams to transform emerging distillation techniques into reliable production-ready implementations.

Our partner is an innovative company focused on advancing the efficiency and scalability of next-generation machine learning systems. They offer a remote-friendly work environment with an async-first culture and a small, senior team combining research expertise and engineering excellence.

Global

  • Lead the design and development of production AI systems powering insurance workflow automation.
  • Architect AI orchestration layers connecting LLMs, backend services, and business workflows.
  • Own end-to-end AI system design including inference pipelines, routing, caching, and fallback strategies.

BJAK is Southeast Asia's largest digital insurance platform, building AI-powered products that simplify insurance and financial services for millions of users. They use AI and automation to transform real-world insurance workflows and foster a high-ownership culture with small, empowered teams.

Canada

  • Design, develop, and deploy production-grade AI-powered backend systems integrating LLMs and traditional ML models.
  • Integrate and optimize vector databases for RAG pipelines, and write clean, well-structured Python code.
  • Debug complex cross-layer issues and collaborate with product and engineering teams to deliver cohesive solutions.

We are a fast-growing product company integrating cutting-edge AI capabilities into our core offering to deliver exceptional value to customers. Our small, fast-moving team works on practical, real-world AI applications with high autonomy.

EMEA

  • Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
  • Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
  • Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.

They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.

$220,000–$280,000/yr
US Unlimited PTO

  • Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
  • Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
  • Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.

Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.

$160,000–$190,000/yr
US

  • Build and improve ML components across data, training, evaluation, and inference.
  • Implement evaluation and testing to understand model behavior.
  • Debug model issues, performance problems, and production incidents.

This company builds core ML components for large-scale production systems. They emphasize real-world learning, iteration, and collaboration with senior engineers.

$190,000–$230,000/yr
US

  • Lead the design and development of our production inference platform, defining the technical roadmap for inference infrastructure, model serving, and runtime optimization.
  • Build and operate scalable, cost-effective systems for serving large language models in production, optimizing latency, throughput, GPU utilization, and memory efficiency.
  • Partner with ML engineers to productionize new models and inference techniques, establish benchmarking methodologies, and make key architectural decisions.

Syllo is on a mission to transform litigation with a unified platform that enables lawyers to safely harness AI. Since going to market, they have gained diverse enterprise customers including big law firms and corporations, and are quickly expanding.

UK Poland 5w PTO

  • Train, evaluate, and iterate on ML models for customer feedback, including custom fine-tuning pipelines.
  • Build and maintain LLM-powered features like retrieval pipelines and insight agents.
  • Design and run robust evaluation frameworks to measure model performance.

Chattermill helps large brands like Uber, Amazon, and Wise put customers at the center using AI. They offer a flexible, trust-based culture with a choice-first environment.