Source Job

Global

  • Own the full post-training pipeline from data curation to deployment.
  • Advance techniques across the post-training stack including SFT, RLHF, DPO, and reward modeling.
  • Build personalization and customization capabilities for user adaptation.

PyTorch Reinforcement Learning

20 jobs similar to Staff / Senior IC Post-Training Engineer

Jobs ranked by similarity.

US

  • Design and evaluate reinforcement learning systems for agentic AI workflows, including RL environments, reward models, and post-training pipelines for LLM-based agents.
  • Develop simulation environments, reward functions, and evaluation frameworks for enterprise workflows.
  • Collaborate with researchers to translate research into practical enterprise solutions, with opportunities to publish and present findings.

Centific is a frontier AI data foundry that curates diverse, high-quality data to empower clients with safe, scalable AI deployment. Their team includes over 150 PhDs and data scientists, along with 4,000 AI practitioners and engineers, fostering a culture of innovation and excellence.

India

  • Research and implement state-of-the-art techniques to accelerate AI inference: quantization, sparsity, distillation, speculative decoding, and caching.
  • Partner closely with hardware and compiler teams to ensure algorithmic improvements translate to real gains on custom silicon.
  • Build profiling tools and comprehensive benchmarking frameworks to measure model quality and efficiency.

EnCharge AI is building the next generation AI platform using novel in-memory-computing architecture. The team consists of experienced AI researchers, silicon & systems engineers, and architects backed by leading investors.

  • Build and ship AI features end-to-end, from model to system to user experience.
  • Design and iterate on prompts, tools, memory, and agent workflows for real-world reliability.
  • Debug full-stack issues and optimize for latency, cost, and production performance.

A1 builds a proactive smart assistant for everyday users, bringing intelligence to conversations, errands, organizing, and workflows with minimal prompting. The team is small, world-class, and focuses on rapid iteration and shipping high-quality AI products.

United States

  • Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment.
  • Implement evaluation pipelines covering performance, robustness, safety, and bias.
  • Own production deployment including GPU optimization, memory efficiency, latency reduction, and scaling policies.

France

  • Design and implement advanced knowledge distillation pipelines, including teacher-student approaches and multi-teacher architectures.
  • Run large-scale machine learning experiments to optimize model quality, latency, efficiency, and cost trade-offs.
  • Collaborate with research teams to transform emerging distillation techniques into reliable production-ready implementations.

Our partner is an innovative company focused on advancing the efficiency and scalability of next-generation machine learning systems. They offer a remote-friendly work environment with an async-first culture and a small, senior team combining research expertise and engineering excellence.

$216,700–$303,400/yr
US

  • Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
  • Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
  • Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.

Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.

United States

  • Lead technical discovery with foundation model labs, frontier AI teams, and large enterprises to understand model objectives and constraints.
  • Design end-to-end solutions across the post-training stack including SFT data curation, RLHF/DPO pipelines, custom benchmarks, and LLM-as-judge systems.
  • Author technical proposals, run workshops and POCs, and serve as ongoing technical advisor during delivery.

Innodata is a global data engineering company focused on enabling responsible AI advancement through data, evaluation frameworks, and human expertise. With a 36+ year legacy, they provide high-quality data solutions to foundation model labs, hyperscalers, and enterprise AI teams.

Global

  • Work directly with top AI labs to design custom data pipelines and integrations for training and evaluation data.
  • Build technical infrastructure for advanced quality control workflows, including model-in-the-loop and human-in-the-loop systems.
  • Ship fast and iterate constantly, translating ambiguous research needs into high-leverage technical systems.

Surge AI is a platform that powers the most advanced AI models in partnership with leading AI labs like Anthropic, Google, Microsoft, and Meta. Founded by engineers and researchers, the company is profitable from day one without venture funding.

Global

  • Develop innovative solutions to enhance annotation efficiency using the latest AI advancements.
  • Create intuitive interfaces for data exploration and visualization, including tools for RLHF processes.
  • Design and implement systems for large-scale data processing to support AI model training and evaluation.

Surge AI builds a platform that powers advanced AI models by combining elite human expertise with cutting-edge tools for scalable oversight. They are a small, profitable, and dynamic team founded by engineers and researchers.

  • Conduct end-to-end research and development of vision-language models, including training, evaluation, optimization, and deployment.
  • Design and implement advanced post-training methodologies such as supervised fine-tuning, knowledge distillation, and reinforcement learning from human feedback.
  • Build, curate, filter, and maintain high-quality multimodal datasets tailored to domain-specific applications.

Jobgether is an AI-powered job matching platform that connects candidates with partner companies. They prioritize objective, fair review of applications using AI, and operate with a focus on efficiency and data privacy.

UK Poland 5w PTO

  • Train, evaluate, and iterate on ML models for customer feedback, including custom fine-tuning pipelines.
  • Build and maintain LLM-powered features like retrieval pipelines and insight agents.
  • Design and run robust evaluation frameworks to measure model performance.

Chattermill helps large brands like Uber, Amazon, and Wise put customers at the center using AI. They offer a flexible, trust-based culture with a choice-first environment.

US

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.
  • Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.

We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.

US Unlimited PTO

  • Develop and refine features for deep learning models using large-scale customer and behavioral datasets.
  • Implement model architecture changes informed by recent academic research from venues like NeurIPS.
  • Optimize model training pipelines for efficiency and scalability while collaborating with client teams.

OpenTeams builds AI that empowers, offering energy-efficient and cost-effective models with a commitment to open source. The company values freedom, teamwork, accountability, and quality, and reinvests 3% of profits into the open-source community.

US

  • Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
  • Architect scalable inference systems and collaborate closely with application engineering.

This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.

Global

  • Own the agent layer orchestrating specialist agents for investor research, outreach, and conversations.
  • Build and scale the matching engine using embeddings, ranking, and feedback loops.
  • Ship production code end to end with Python, Postgres, and AI tooling.

Paires is a fundraising platform that pairs founders with investors using AI agents for outreach and relationship management. It is a small, senior, flat team that is profitable and self-funded.

AI Developer

Aphex
Global 4w PTO

  • Build and implement AI features by selecting the right model and approach for each use case.
  • Evaluate and monitor AI feature performance in production to ensure accuracy and reliability.
  • Improve and iterate on prompts and implementations based on how features behave in the wild.

Aphex is a construction execution platform that replaces traditional spreadsheets with collaborative tools for delivery teams. They are a remote-first company with a growing engineering team in the Philippines, serving major contractors on multi-billion dollar projects.

US

  • Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
  • Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.

Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.

US

  • Own the model strategy and technical direction for advanced generative and multimodal vision systems.
  • Drive model benchmarking, ablation studies, and the path from pre-trained baselines to production-ready capability.
  • Mentor other scientists and act as the senior technical owner for the model domain.

HERE Technologies is a location data and technology platform company that empowers customers to achieve better outcomes. As a global organization, we have a diverse team and foster a culture of innovation, opportunity, and inclusion.

LATAM North America EMEA

  • Spend your first weeks in the operator's seat, learning the customer's job from the inside before writing any code.
  • Ship production GenAI/LLM systems that move business unit metrics, not just complete scope.
  • Work embedded in small, senior teams alongside Principal Architects, owning the outcome from start to finish.

Provectus is a Premier AWS partner and an Anthropic Strategic Partner at the forefront of applied AI, helping enterprises turn Claude, agentic systems, and their own data into measurable business outcomes. With offices in North America, LATAM, and EMEA, we partner with clients worldwide and our team holds 100+ AWS certifications and is Claude Code certified.

Germany

  • Design, build, and operate multi-agent workflows and tool-enabled agents for resilient production pipelines.
  • Architect and maintain end-to-end RAG systems covering document ingestion, chunking, embedding, retrieval, and answer synthesis.
  • Define and own evaluation frameworks for generative outputs, including automated metrics, LLM-as-judge, and hallucination detection.

Mitratech builds world-class products that simplify operations in Legal, Risk, Compliance, and HR functions. With over 35 years of experience, we serve 20,000 client companies globally, including 30% of the Fortune 500, and foster a diverse, inclusive culture centered on individual excellence and learning.