Source Job

$110,301–$171,885/yr
UK Unlimited PTO

  • Design and implement model-based RL agents and planning controllers for real industrial systems.
  • Develop learned world models and training pipelines to make control agents reliable and safe.
  • Work closely with researchers and engineers to translate cutting-edge RL research into production outcomes.

Python PyTorch Reinforcement Learning

8 jobs similar to Senior AI Research Scientist (Model-based RL)

Jobs ranked by similarity.

Global 6w PTO 26w maternity 26w paternity

  • Build new RL environments targeting different agentic capabilities and industry areas.
  • Train and evaluate agents in those environments, improving both agents and environments.
  • Work across modeling and product to identify performance gaps and automate measurement.

Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products. It is a global team of researchers, engineers, and designers headquartered in Toronto with offices worldwide.

  • Build sandboxed environments that wrap real Niural workflows, stateful across episodes and seeded for reproducibility.
  • Design programmatic verifiers from known ground truth, scoring trajectories to prevent reward hacking.
  • Train agents via RL loops, curriculum schedules, and fine-tuning, then publish negative results internally.

Niural is a global Payroll, Employer of Record (EOR), Agent of Record (AOR), and Contractor Management platform that empowers businesses in the digital economy. We are a team building foundational internet infrastructure with a focus on speed, ownership, and ambition.

$120–$120/hr
Global

  • Design and implement tool-based environments for knowledge work simulations.
  • Write clean and modular code to support reinforcement learning training pipelines.
  • Troubleshoot and refine environment mechanics based on testing feedback.

Terac is building the world's largest pool of vetted human experts for AI. They are a platform for researchers and AI labs to recruit, screen, and pay study participants across industries and skill sets.

US Unlimited PTO

  • Collaborate with product and engineering teams to translate product objectives into autonomous agent-based solutions.
  • Design and build new agent data types, pipelines, and frameworks to coordinate reasoning, function calling, and actions.
  • Develop and optimize autonomous agents leveraging LLMs, planning algorithms, and multi-step reasoning approaches.

PointClickCare is a leading health tech company that helps providers deliver exceptional care. As a founder-led, privately held company with over 30,000 provider organizations and 400+ integrated partners, they are recognized by Forbes as a top private cloud company and honored as one of Canada's Most Admired Corporate Cultures, offering flexibility and growth opportunities.

$100,000–$300,000/yr
US

  • Build the software backbone for autonomous foundation models, including multimodal data pipelines and training workflows.
  • Implement and iterate on LLM, VLM, and VLA architectures, owning model code paths and inference runners.
  • Deliver production-grade serving tooling for low-latency operation and own systems from architecture to iteration.

We're building the next generation of ground transportation with advanced physical AI to simplify freight challenges. Our stealth team, founded by engineers who scaled autonomous driving, is developing a new vehicle platform and focuses on creating reliable real-time autonomous systems.

Global

  • Set and evolve the research direction for A1’s core intelligence, including context representation, memory, reasoning, planning, and orchestration.
  • Define evaluation frameworks that measure real-world usefulness, robustness, safety, and long-term behavior.
  • Own alignment, safety, and guardrail strategy as first-class product concerns.

A1 builds a proactive smart assistant for everyday users to bring intelligence to conversations, errands, organizing, and workflows. We are a small, high-talent-density team focused on shipping high-quality work and learning at rapid speed.

US UK Unlimited PTO

  • Design and optimize training and post-training pipelines for large language models.
  • Improve model quality through supervised fine-tuning, reinforcement learning, and evaluation.
  • Build PyTorch-based training infrastructure and optimize distributed training across multi-GPU environments.

Lightning AI builds an end-to-end platform for developing, training, and deploying AI systems, founded in 2019. They are a global company with offices in New York, San Francisco, Seattle, and London, backed by major venture capital firms, and foster a builder culture that values urgency, ownership, and open communication.

Global

  • Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
  • Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
  • Benchmark performance by conducting comparative testing between model outputs based on technical metrics.

Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.