Source Job

US UK Unlimited PTO

  • Design and optimize training and post-training pipelines for large language models.
  • Improve model quality through supervised fine-tuning, reinforcement learning, and evaluation.
  • Build PyTorch-based training infrastructure and optimize distributed training across multi-GPU environments.

PyTorch Python

20 jobs similar to Senior Research Engineer, LLM Training & Post-Training

Jobs ranked by similarity.

Global

  • Own the full post-training pipeline from data curation to deployment.
  • Advance techniques across the post-training stack including SFT, RLHF, DPO, and reward modeling.
  • Build personalization and customization capabilities for user adaptation.

Black Forest Labs is a research lab behind foundational generative AI technologies like Stable Diffusion and FLUX, powering tools used by millions worldwide. They are a fast-growing, distributed team with a culture of research excellence, open science, and low ego.

US

  • Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
  • Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
  • Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.

Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.

Global 6w PTO 26w maternity 26w paternity

  • Conduct cutting-edge machine learning research, building and training large language models.
  • Focus on research projects aimed at expanding the frontier of knowledge in language modelling and associated areas.
  • Disseminate your research results through publications, datasets, and code.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products for enterprise AI systems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, and more.

US

  • Design, train, and improve large-scale machine learning models for recommendation and personalization, leveraging modern deep learning architectures.
  • Own end-to-end delivery of major ML system components, from problem framing to production rollout, with cross-functional partners.
  • Optimize distributed training, model efficiency, and online inference performance while ensuring low-latency, high-throughput production systems.

Reddit is a community of communities built on shared interests, passion, and trust, hosting authentic conversations. With over 100,000 active communities and about 130 million daily active users, Reddit is one of the internet's largest sources of information.

US

  • Fine-tune and optimize Large Language Models for healthcare and life science applications.
  • Develop and enhance RAG pipelines to improve AI-powered information retrieval.
  • Collaborate with cross-functional teams to deploy deep learning solutions for complex healthcare challenges.

The company develops AI-powered healthcare solutions using Large Language Models. They foster a highly collaborative and innovation-driven environment with a diverse international team.

Global 6w PTO 26w maternity 26w paternity

  • Develop, prototype, and deploy techniques to improve LLM inference efficiency in production.
  • Explore and ship breakthroughs across model architecture, decoding, and software/hardware co-design.
  • Optimize performance without compromising model quality.

Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products for real-world business problems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, New York City, Montreal, Seoul, and Paris.

$216,700–$303,400/yr
US

  • Design, develop, and deploy ML models, including large language models, for various NLP tasks.
  • Collaborate with cross-functional teams to gather requirements, define architectures, and iterate on model development.
  • Stay up-to-date with latest research and contribute to best practices for responsible ML development.

Reddit is a community of communities built on shared interests, passion, and trust, hosting the most open and authentic conversations on the internet. With over 100,000 active communities and approximately 126 million daily active users, it is one of the internet's largest sources of information, fostering a culture of authenticity and community.

Global

  • Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
  • Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
  • Benchmark performance by conducting comparative testing between model outputs based on technical metrics.

Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.

UK Poland 5w PTO

  • Train, evaluate, and iterate on ML models for customer feedback, including custom fine-tuning pipelines.
  • Build and maintain LLM-powered features like retrieval pipelines and insight agents.
  • Design and run robust evaluation frameworks to measure model performance.

Chattermill helps large brands like Uber, Amazon, and Wise put customers at the center using AI. They offer a flexible, trust-based culture with a choice-first environment.

Global 5w PTO

  • Contribute to the entire development cycle of cutting-edge large deep learning models, from dataset preparation to deployment.
  • Collaborate with engineers and researchers to translate research into practical applications, wearing many hats.
  • Work at an exciting moment to build something from the ground up with a kind and collaborative team.

Reka builds useful multimodal AI to empower organizations, as a globally distributed foundation model startup headquartered in San Francisco Bay Area. The team, including contributors from Google DeepMind and FAIR, embraces a remote-first culture and collaborates globally.

$200,000–$350,000/yr
US

  • Develop and deploy machine learning and AI systems.
  • Work with LLMs, generative AI, and modern ML frameworks.
  • Optimize model performance, latency, and cost.

A fast-growing technology company building critical infrastructure that powers high-volume, real-time business operations across multiple systems and platforms. It is a collaborative, fast-moving environment where engineers have meaningful influence on architecture and product direction.

Global 6w PTO 26w maternity 26w paternity

  • Design and write high-performing scalable software for training models.
  • Develop new tools to support and accelerate research and LLM training.
  • Collaborate with engineering teams and scientific teams to implement experiments on cluster and data infrastructure.

Cohere is a security-first enterprise AI company building cutting-edge foundation AI models and end-to-end products for real-world business problems. The company is a global team of researchers, engineers, and designers passionate about AI, headquartered in Toronto with offices worldwide.

Global

  • Design and scale production ML systems for LLM-based applications.
  • Build training and evaluation pipelines for continuous model improvement.
  • Fine-tune foundation models using modern adaptation techniques such as LoRA, QLoRA, SFT and DPO.

A1 is a new AI venture building the next generation of AI-native productivity applications, starting with an email agent that uses autonomous AI. Backed by an initial $100M investment, the company is a small, high-talent founding engineering team focused on solving challenging AI infrastructure problems.

France

  • Design and implement advanced knowledge distillation pipelines, including teacher-student approaches and multi-teacher architectures.
  • Run large-scale machine learning experiments to optimize model quality, latency, efficiency, and cost trade-offs.
  • Collaborate with research teams to transform emerging distillation techniques into reliable production-ready implementations.

Our partner is an innovative company focused on advancing the efficiency and scalability of next-generation machine learning systems. They offer a remote-friendly work environment with an async-first culture and a small, senior team combining research expertise and engineering excellence.

US Unlimited PTO

  • Architect and build AI platforms, working across distributed teams to deliver GenAI products like Abacus Agents, MCP, and RAG systems.
  • Develop ML/AI pipelines for data preprocessing, model training, and evaluation, enabling rapid experimentation and self-healing systems.
  • Mentor software developers, participate in architectural discussions, and communicate technical direction to align teams and stakeholders.

Abacus Insights transforms healthcare data for health plans to improve outcomes, reduce waste, and enhance member and provider experiences. Backed by $100M from top investors, the company is a collaborative, curiosity-driven team of bold individuals embracing AI and automation to innovate in the healthcare industry.

  • Build and ship AI features end-to-end, from model to system to user experience.
  • Design and iterate on prompts, tools, memory, and agent workflows for real-world reliability.
  • Debug full-stack issues and optimize for latency, cost, and production performance.

A1 builds a proactive smart assistant for everyday users, bringing intelligence to conversations, errands, organizing, and workflows with minimal prompting. The team is small, world-class, and focuses on rapid iteration and shipping high-quality AI products.

Global

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
  • Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
  • Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.

Global

  • Deploy LLMs into production across GPU infrastructure, owning the full pipeline from customer query to served response.
  • Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
  • Apply quantization, batching, caching, and routing to optimize latency and cost at scale.

vCluster Labs is a venture-backed tech startup pioneering Kubernetes virtualization for the AI era, enabling AI Cloud providers and AI factories to operate GPU infrastructure with hyperscaler-like experiences. We raised over $30M from top-tier VCs like Khosla Ventures, are in a hyper-growth phase, and maintain a remote-first, distributed global team with headquarters in San Francisco.

Global

  • Design, build, and deploy production ML and LLM-based systems (RAG, agentic workflows, fine-tuning, embeddings) for enterprise clients.
  • Own technical delivery end-to-end: from architecture and prototyping to deployment, monitoring, and iteration.
  • Mentor and support other ML engineers on the team with code reviews, technical guidance, and knowledge sharing.

TensorOps is a boutique AI consultancy that bridges strategy and execution, designing and shipping production-grade AI systems for enterprise clients. We are a 100% remote team of 11+ people, partnering with unicorns and NASDAQ-listed companies, and have a culture of autonomy, open communication, and continuous learning.

Global

  • Taking ML or LLM proof-of-concept to production for large enterprises.
  • Designing and hardening data and training pipelines for enterprise ML systems.
  • Building LLM and RAG systems with retrieval quality, evaluation and cost control.

Janea Systems (USA) is a dynamic team of the best & brightest software engineering specialists and solutions innovators from around the world. From kernel to cloud, we provide high-impact software development services to Fortune 500 companies.