Source Job

Unlimited PTO

  • Own the research roadmap for recursive self-improvement and ship what works into production.
  • Design agent strategies, exploration policies, and evaluations for autonomous LLM research systems.
  • Partner with applied scientists and build production-quality experiments end to end.

Python LLM Agents Reinforcement Learning

20 jobs similar to Staff+ Research Engineer — Recursive Self-Improvement Lead

Jobs ranked by similarity.

India

  • Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
  • Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
  • Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.

Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.

Europe 6w PTO

  • Own agent quality end-to-end for Bliro's chat and voice agents, including prompting, context, and model evaluation.
  • Build eval sets, LLM judges, and benchmarks to measure quality, latency, and cost with statistical rigor.
  • Establish a repeatable process for improving agent quality and set the long-term technology roadmap.

Bliro builds an AI assistant that handles desk work for field sales teams, with over one million customer touchpoints documented. Backed by leading investors and trusted by German Mittelstand companies, it's a small, high-performing team that values ownership and continuous improvement.

$250,000–$300,000/yr
US Canada Unlimited PTO

  • Own the build out of new agents, skills, and platform capability for teams across TLDR.
  • Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
  • Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.

TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.

Global

  • Design and implement autonomous agents and multi-agent orchestration systems for complex, open-ended tasks.\n- Build and optimize production-grade LLM applications with a focus on reliability, observability, and low-latency performance.\n- Architect advanced RAG pipelines and vector database strategies to provide agents with accurate, real-time context.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective evaluation. The company operates with a distributed, global-first team culture, emphasizing fairness and innovation in recruitment.

Global

  • Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
  • Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
  • Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.

Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.

India

  • Build RL environments, agentic systems, LLM pipelines, and evaluation frameworks for real-world AI use cases.
  • Develop benchmarks and evaluation harnesses to assess model quality across accuracy, safety, latency, and cost.
  • Conduct fine-tuning and model experiments, deploy self-hosted models, and document reproducible methodologies.

The employer is an organization focused on applied AI research, building practical and reusable AI systems for real-world use cases. It values curiosity, accountability, innovation, collaboration, and continuous learning in a remote environment.

$180,000–$250,000/yr
Global

  • Work directly with leading AI labs and enterprises to define research goals and technical requirements.
  • Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
  • Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.

Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.

$200,000–$350,000/yr
US Unlimited PTO

  • Architect, build, and optimize high-performance production LLM systems while maintaining a strong personal technical presence on the team.
  • Spearhead strategic technological changes and champion code refactoring efforts to keep the core codebase cutting-edge and performant.
  • Lead technical story breakdowns, architectural design, and mentor engineers across the department.

Appian provides AI automation for mission-critical work, automating complex processes in large enterprises and governments. With over 25 years of experience, the company is known for its reliability and scale, and fosters an inclusive culture with employee-led affinity groups.

$130,000–$200,000/yr
US

  • Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
  • Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
  • Turn research ideas into working prototypes and production-facing capabilities.

This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.

$150,000–$220,000/yr
US Unlimited PTO

  • Build production-grade agentic workflows to automate fan issue triage, seller risk monitoring, and marketplace optimization.
  • Coach non-technical teammates to use AI tools and create their own automations.
  • Own the full lifecycle of built solutions, from design to production, and scale successes across operations teams.

Gametime makes it easy for people to discover and access live experiences, with platforms supporting over 60,000 events across the US and Canada. They are a mid-sized company focused on reimagining the event ticket industry, with a culture that emphasizes inclusivity and innovation.

US

  • Design and deploy AI agents that automate sales, operations, and support workflows.
  • Own technical delivery end-to-end, from architecture to production and post-launch improvements.
  • Work directly with clients to understand workflows and explain technical tradeoffs.

Stello is an AI transformation firm that partners with mid-market and enterprise companies to automate repetitive work with AI agents. They work across sales, operations, and customer support, deploying solutions inside existing client tools.

$0–$135,000/yr
US

  • Design and build AI agents and automation to solve real problems across engineering, product, and delivery.
  • Partner with stakeholders to identify high-leverage opportunities and deliver end-to-end solutions.
  • Stay current with LLM and agentic frameworks to drive innovation in healthcare technology.

HealtheDGE provides AI-powered operational infrastructure for health insurance companies, helping them modernize operations. The company is experiencing strong market momentum and invests in its people, offering a collaborative culture focused on innovation.

Europe Unlimited PTO

  • Lead a small engineering team as a hands-on player-coach, setting technical direction and owning delivery.
  • Ship AI-powered features to production while raising standards across observability, testing, security, and architecture.
  • Partner with Product and leadership to plan, prioritize, and follow through on complex technical work.

Usersnap is building AI-driven capabilities into its product, from automated setup to intelligent surveys and reporting. It is part of saas.group, a remote portfolio of 380+ SaaS professionals across 50+ countries, with a high-trust, flexible culture.

US Unlimited PTO

  • Lead the engineering strategy and execution for evaluations of AI agents, owning the core evaluation platform.
  • Design scalable evaluation infrastructure, APIs, workflows, and production systems across software categories.
  • Mentor and develop a team of engineers as the technical authority on agentic evaluation.

This company focuses on building credible, scalable evaluations of AI agents from software vendors. It operates as a fully remote, inclusive team with a flexible culture and a focus on professional growth.

Canada Unlimited PTO

  • Design and deliver production-grade agents that investigate, reason, and act on live observability data.
  • Own agent work from rough prototype to production, including evals.
  • Extend the agentic workspace with new canvas capabilities, MCP server, and skills.

Honeycomb provides an observability platform that helps engineers understand and debug complex systems. It is a fully distributed company of over 200 people, with a culture that values impact, autonomy, and inclusivity.

Global

  • Design agent architectures for planning, reasoning, tool use, and memory integration.
  • Improve reliability on long-running tasks with failure recovery and evaluation systems.
  • Build loops for agents to improve with real use, balancing quality, latency, and cost.

Adaption builds AI systems that evolve in real-time, making them flexible and personalized. They are a global-first team focused on talent density and collaboration.

Global Unlimited PTO

  • Create and publish coding benchmarks that challenge state-of-the-art frontier agents.
  • Shape the frontier of coding agents and publish impactful studies of interesting findings.
  • Help top AI labs improve their models for coding through benchmarks, open-source tools, and research.

Handshake's mission is to organize expert human knowledge to advance the AI economy, working directly with frontier labs on data, evaluations, and post-training challenges. It is a rapidly growing startup with a small, high-ownership technical team from top AI companies.

Europe US

  • Maintain and evolve platform foundations, tools, and processes for agentic development to keep n8n operating as an AI-native engineering organization.
  • Build and integrate agentic workflows across Linear, GitHub, CI, and developer review processes while ensuring safety and security.
  • Define quality metrics, benchmarks, and adoption signals to measure agent output and improve engineering productivity.

n8n is an open workflow orchestration platform built for the new era of AI, giving technical teams the freedom of code with the speed of no-code. Since 2019, the company has grown to over 260 employees across Europe and the US, with a community of 650,000+ developers, 190K+ GitHub stars, and a $5.2bn valuation.

$170,000–$170,000/yr
Global 7w PTO

  • Lead the engineering foundation behind reliable, measurable, and scalable AI-powered features.
  • Build evaluation frameworks, observability tooling, and diagnostic infrastructure for AI agents in production.
  • Combine hands-on software engineering with technical leadership and people management.

The company builds AI-powered products with a focus on reliability, scalability, and evaluation infrastructure. It operates as a fully remote, globally distributed team with an ownership-oriented culture.

$150,000–$190,000/yr
Global Unlimited PTO

  • Monitor production health across every graph, catching errors and silent failures proactively.
  • Triage incoming issues, fix small things directly, and route larger problems to the right owner.
  • Track cost, latency, usage, and build business metrics to show the agent's ROI.

LangChain builds the foundation for intelligent agent engineering, helping developers move from prototypes to production-ready AI agents. With over 100M monthly open source downloads and backing from top VCs, the company is at a stage where all team members have meaningful impact.