Source Job

Global

  • Build the Learning Gym, a sandboxed environment for training and evaluating AI agents on real Niural workflows with ground-truth reward signals.
  • Engineer verifiers, run reinforcement learning loops, and establish honest baselines to prove agent improvements generalize.
  • Write first-author research papers and internal technical reports, and represent the work externally through preprints and talks.

Python Reinforcement Learning Machine Learning Docker

20 jobs similar to AI Research Engineer

Jobs ranked by similarity.

India

  • Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
  • Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
  • Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.

Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.

India

  • Build RL environments, agentic systems, LLM pipelines, and evaluation frameworks for real-world AI use cases.
  • Develop benchmarks and evaluation harnesses to assess model quality across accuracy, safety, latency, and cost.
  • Conduct fine-tuning and model experiments, deploy self-hosted models, and document reproducible methodologies.

The employer is an organization focused on applied AI research, building practical and reusable AI systems for real-world use cases. It values curiosity, accountability, innovation, collaboration, and continuous learning in a remote environment.

$130,000–$200,000/yr
US

  • Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
  • Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
  • Turn research ideas into working prototypes and production-facing capabilities.

This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.

$130,000–$200,000/yr
US

  • Develop and improve AI methods, algorithms, and systems across the lifecycle of reliable AI agents.
  • Create new approaches for simulating, evaluating, and optimizing agent behavior in real-world settings.
  • Turn research ideas into working systems, prototypes, and production-facing capabilities.

The company builds infrastructure that helps enterprises make AI agents more reliable in production. It is a technical and research-focused team working on AI methods and practical systems.

AI Engineer

Sibill
$64,800–$86,400/yr
Italy

  • Build and ship production AI systems for the core finance workflows.
  • Own evals: metrics, test datasets, and monitoring of quality, cost, and latency.
  • Iterate end-to-end on AI products, from data pipelines to deployment.

Sibill is an AI-powered platform that unifies invoicing, payments, treasury, and accounting for SMEs and accountants in Italy. Backed by €12M from Creandum, the company serves over 5,000 businesses and 300+ firms with a team of over 130 people.

Global

  • Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
  • Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
  • Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.

Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.

Global

  • Build LLM-based agents on the platform's scaffolding, integrating tool calls, internal APIs, and guardrails.
  • Take agents to production on AWS with containers, CI/CD, secrets, permissions, and security controls.
  • Define and run evals, monitor with Langfuse, and document runbooks for independent operation.

Muttdata builds innovative Data Products and Machine Learning solutions to help companies solve complex business challenges. It is a fast-growing, remote-first startup that values collaboration, continuous learning, and a positive, ownership-driven culture.

North America 3w PTO

  • Develop, configure, deploy, and optimize AI agents using Cresta's AI platform and tools.
  • Build AI agent integrations with external systems such as APIs, databases, and CRMs to ensure seamless workflow integration.
  • Collaborate with customers and internal stakeholders to gather technical requirements and translate business needs into AI agent solutions.

Cresta builds a unified AI platform that combines conversational AI agents, real-time human agent augmentation, and conversation intelligence to drive revenue and efficiency for customer experience channels. Born from Stanford AI Lab, the company has raised over $270 million from top investors and is led by AI industry veterans.

North America Europe

  • Build tools for understanding traces, evaluating agent behavior, and turning results into improvements.
  • Run experiments on Mastra's primitives to find where they fall short and improve configuration and behavior.
  • Help turn findings into capabilities developers can use in their own applications.

Mastra is the open-source TypeScript framework for building AI agents. We're a small, fully remote team backed by Y Combinator and Spark Capital with $35M raised.

Europe 6w PTO

  • Own agent quality end-to-end for Bliro's chat and voice agents, including prompting, context, and model evaluation.
  • Build eval sets, LLM judges, and benchmarks to measure quality, latency, and cost with statistical rigor.
  • Establish a repeatable process for improving agent quality and set the long-term technology roadmap.

Bliro builds an AI assistant that handles desk work for field sales teams, with over one million customer touchpoints documented. Backed by leading investors and trusted by German Mittelstand companies, it's a small, high-performing team that values ownership and continuous improvement.

Global Unlimited PTO

  • Create and publish coding benchmarks that challenge state-of-the-art frontier agents.
  • Shape the frontier of coding agents and publish impactful studies of interesting findings.
  • Help top AI labs improve their models for coding through benchmarks, open-source tools, and research.

Handshake's mission is to organize expert human knowledge to advance the AI economy, working directly with frontier labs on data, evaluations, and post-training challenges. It is a rapidly growing startup with a small, high-ownership technical team from top AI companies.

$86,400–$151,200/yr
Europe

  • Transform an existing AI agent product into scalable, reliable infrastructure.
  • Design and operate the core agent runtime, including orchestration, memory, and multi-agent workflows.
  • Work directly with the founder to set architecture, engineering standards, and development culture.

An early-stage AI infrastructure startup transforming a working AI agent product into scalable infrastructure. As the first engineering hire, you'll work directly with the founder to shape engineering culture and standards from the ground up.

$225,000–$250,000/yr
US

  • Design and implement state-of-the-art ML models and training pipelines for robotics.
  • Develop efficient data/training strategies and evaluation frameworks for rapid experimentation.
  • Collaborate with engineering to optimize training infrastructure and deployment.

We're revolutionizing real-world automation by making robotic systems accessible to everyone. Our AI-powered platform brings software automation to physical spaces, and we're a small startup team working across the stack to solve customer problems.

US Unlimited PTO

  • Build the component layer around our layout synthesis engine, including API contracts, services, and evaluation gates.
  • Turn model retraining into a one-command job with built-in benchmarks and readable results for the whole team.
  • Own latency and cost budgets for learned capabilities and partner with infrastructure engineers on MLOps and deployment.

Higharc is a VC-backed startup that is changing how new homes are designed and built using spatial AI and generative floor plan technology. The company is fully remote, has raised over $175M, and values flexibility, collaboration, and asynchronous deep work.

Global

  • Automate quality control for training data produced by companies using the platform's infrastructure.
  • Build QC systems grounded in human judgment, define quality standards, and design experiments.
  • Partner with data vendors to debug quality issues and feed learnings back into infrastructure tools.

The company is a fast-growing AI infrastructure startup focused on reinforcement learning environments and post-training data. With a roughly 15-person engineering team of published researchers and experienced AI practitioners, the culture is fast-paced, unstructured, and research-driven.

$150,000–$220,000/yr
US Unlimited PTO

  • Build production-grade agentic workflows to automate fan issue triage, seller risk monitoring, and marketplace optimization.
  • Coach non-technical teammates to use AI tools and create their own automations.
  • Own the full lifecycle of built solutions, from design to production, and scale successes across operations teams.

Gametime makes it easy for people to discover and access live experiences, with platforms supporting over 60,000 events across the US and Canada. They are a mid-sized company focused on reimagining the event ticket industry, with a culture that emphasizes inclusivity and innovation.

$105,000–$147,000/yr
US

  • Build and ship features of AI/ML and LLM-powered systems with guidance from senior engineers.
  • Implement and maintain AI/ML and AI agent pipelines from data ingestion through model deployment.
  • Contribute to LLM-powered features such as prompts, evaluations, retrieval, and tool integrations, while documenting experiments clearly.

Robots & Pencils designs AI systems for a human world, pairing engineering with creativity. Teams average fifteen-plus years of experience and value ownership, craft, direct feedback, and continuous learning.

$129,211–$178,250/yr
Canada

  • Lead the design and delivery of AI/ML systems, defining architecture and establishing engineering standards for agentic workflows.
  • Build and maintain RAG pipelines, agentic workflows, and production system prompts with eval and observability.
  • Mentor engineers and take end-to-end ownership of features, including debugging and hardening production systems.

Robots and Pencils designs AI systems for a human world, pairing engineering with creativity to ship production-ready AI in 30 to 45 days. They are a company where every role contributes directly to integrating AI into enterprise operations, with a culture that values ownership, craft, and measurable impact.

Latin America

  • Build production Agentforce agents including topics, actions, grounding, guardrails, and evaluation harnesses.
  • Implement actions across Flow, Apex, prompt templates, and external APIs, grounding answers in Data 360.
  • Design guardrails against prompt injection and data leakage, and run evaluation and regression tests before changes.

AspenView Technology Partners builds high-performing nearshore IT teams for North American clients, focusing on innovation and efficiency. They are a people-first, purpose-driven company that believes great culture drives great outcomes.

Europe US

  • Maintain and evolve platform foundations, tools, and processes for agentic development to keep n8n operating as an AI-native engineering organization.
  • Build and integrate agentic workflows across Linear, GitHub, CI, and developer review processes while ensuring safety and security.
  • Define quality metrics, benchmarks, and adoption signals to measure agent output and improve engineering productivity.

n8n is an open workflow orchestration platform built for the new era of AI, giving technical teams the freedom of code with the speed of no-code. Since 2019, the company has grown to over 260 employees across Europe and the US, with a community of 650,000+ developers, 190K+ GitHub stars, and a $5.2bn valuation.