Source Job

North America Unlimited PTO

  • Study how engineers and customers use Archie in the field, identify where it succeeds or fails, and turn observations into actionable evaluations.
  • Recreate real-world engineering tasks and failure modes in repeatable environments for development teams.
  • Build and refine evaluation methodologies, including automated judges and human evaluation processes, to align with expert judgment.

Agentic AI

20 jobs similar to AI Evaluations Engineer

Jobs ranked by similarity.

US

  • Set the technical direction for AI engineering across the team: agent architecture patterns, evaluation methodology, deployment and monitoring strategies
  • Design the AI platform layer, including shared agent frameworks, tool integrations, and evaluation infrastructure
  • Work directly with clients on the most complex engagements, identifying new problem domains and ensuring production quality

Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. With over 1,500 firms in 60 countries managing nearly $10 trillion in assets, Addepar fosters a culture of ownership, collaboration, and innovation.

US Unlimited PTO

  • Lead the engineering strategy and execution for evaluations of AI agents, owning the core evaluation platform.
  • Design scalable evaluation infrastructure, APIs, workflows, and production systems across software categories.
  • Mentor and develop a team of engineers as the technical authority on agentic evaluation.

This company focuses on building credible, scalable evaluations of AI agents from software vendors. It operates as a fully remote, inclusive team with a flexible culture and a focus on professional growth.

India

  • Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
  • Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
  • Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.

Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.

UK

  • Build tooling for capturing and processing data from agents and humans at significant scale.
  • Solve hard problems around compute, orchestration, scaling, security, and reliability.
  • Help develop approaches for training, benchmarking, and evaluating AI agents.

Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.

US

  • Design and deploy AI agents that automate sales, operations, and support workflows.
  • Own technical delivery end-to-end, from architecture to production and post-launch improvements.
  • Work directly with clients to understand workflows and explain technical tradeoffs.

Stello is an AI transformation firm that partners with mid-market and enterprise companies to automate repetitive work with AI agents. They work across sales, operations, and customer support, deploying solutions inside existing client tools.

Global

  • Design agent architectures for planning, reasoning, tool use, and memory integration.
  • Improve reliability on long-running tasks with failure recovery and evaluation systems.
  • Build loops for agents to improve with real use, balancing quality, latency, and cost.

Adaption builds AI systems that evolve in real-time, making them flexible and personalized. They are a global-first team focused on talent density and collaboration.

US

  • Evaluate software engineering tasks for technical accuracy, realism, and reproducibility.
  • Investigate codebases, tests, and integration issues to identify technical weaknesses.
  • Provide clear, actionable feedback that directly improves AI training and evaluation workflows.

Jobgether is an AI-powered job platform that connects candidates to roles through objective, skill-based matching. It focuses on remote and freelance opportunities, with a data-driven recruitment process and a global candidate pool.

Global

  • Embed with customers to understand workflows, constraints, and success measures, then turn operational problems into technical plans.
  • Design, build, and deploy AI agents that integrate with customer tools, data, and business processes from prototype to production.
  • Measure real-world impact, iterate based on evidence, and share reusable patterns across deployments.

DehazeLabs transforms complex enterprise workflows into AI agents that operate in real-world environments. The company is a growing startup focused on applied AI, with a culture that blends engineering, product, and customer delivery.

North America

  • Design and build production-grade AI agentic systems for web and chat experiences.
  • Architect agent workflows involving reasoning, tool use, retrieval, guardrails, and monitoring.
  • Build LLM evaluation systems and diagnose agent performance issues across prompts and tools.

Netomi is an agentic AI platform for enterprise customer experience, enabling automation for global brands. Backed by Y Combinator and Index Ventures, the company drives efficiency and higher quality experiences.

India

  • Build RL environments, agentic systems, LLM pipelines, and evaluation frameworks for real-world AI use cases.
  • Develop benchmarks and evaluation harnesses to assess model quality across accuracy, safety, latency, and cost.
  • Conduct fine-tuning and model experiments, deploy self-hosted models, and document reproducible methodologies.

The employer is an organization focused on applied AI research, building practical and reusable AI systems for real-world use cases. It values curiosity, accountability, innovation, collaboration, and continuous learning in a remote environment.

Global

  • Design and improve prompts for classification and extraction tasks, running structured evaluation cycles for accuracy and iteration.
  • Build datasets for testing and validation, analyze outputs from real usage (logs, SQLite), and write Python scripts for testing and scoring.
  • Identify failure patterns, propose improvements, validate system behavior in iOS simulations, and document findings clearly.

PiggyBank Ventures is an incubator building AI-first consumer applications, currently developing Nestora, an AI-powered mobile app for intelligent communication management. The team includes repeat founders and operators with experience building and scaling products, from top-tier universities and technical programs.

$111,000–$180,000/yr
US

  • Own the hardest customer problems end-to-end, driving ambiguous challenges from prototype to production.
  • Design and deploy AI-powered workflows with appropriate controls for regulated environments.
  • Serve as a trusted technical advisor and contribute to a compounding knowledge base.

Workiva provides an AI-powered platform that unifies finance, risk, and sustainability on a single secure foundation. The company values meaningful challenges, collaborative teams, and helping organizations turn uncertainty into advantage.

Latin America

  • Build production Agentforce agents including topics, actions, grounding, guardrails, and evaluation harnesses.
  • Implement actions across Flow, Apex, prompt templates, and external APIs, grounding answers in Data 360.
  • Design guardrails against prompt injection and data leakage, and run evaluation and regression tests before changes.

AspenView Technology Partners builds high-performing nearshore IT teams for North American clients, focusing on innovation and efficiency. They are a people-first, purpose-driven company that believes great culture drives great outcomes.

$230,000–$330,000/yr
US Unlimited PTO

  • Own Merlin's foundation and world-model work, including architecture selection, post-training, and capability roadmap.
  • Lead and mentor a small team of world-model engineers, setting the technical bar and review culture.
  • Design model interface to the autonomy stack with structured, schema-constrained plan outputs and build evaluation harnesses.

Merlin is a publicly traded aerospace and defense company building a non-human pilot for full-stack aircraft autonomy. Headquartered in Boston, it is expanding its organization to accelerate the deployment of its autonomy platform.

$88,000–$180,000/yr
US Unlimited PTO

  • Direct the agent array on production workstreams by decomposing problems into tasks and integrating agent output into shipped software.
  • Review agent-generated pull requests at volume and depth, identifying correctness, security, and accessibility defects.
  • Author evaluation suites that make quality measurable using eval-driven development and own end-to-end quality within a FedRAMP-authorized environment.

Granicus provides cloud-based solutions for government communications, website design, meeting management, and records management, serving over 5,500 agencies and 300 million citizens. With a globally distributed team and a culture of transparency and inclusion, Granicus has been recognized on the GovTech 100 list for the past 5 years.

US

  • Test and evaluate AI model responses across diverse topics and conversation types.
  • Write and refine system prompts to shape model behavior and personality.
  • Apply detailed rubrics consistently to assess model performance and document findings.

We are a team dedicated to evaluating and improving advanced AI models and products. Our fast-moving environment emphasizes quality, independent judgment, and attention to detail.

Canada Unlimited PTO

  • Design and deliver production-grade agents that investigate, reason, and act on live observability data.
  • Own agent work from rough prototype to production, including evals.
  • Extend the agentic workspace with new canvas capabilities, MCP server, and skills.

Honeycomb provides an observability platform that helps engineers understand and debug complex systems. It is a fully distributed company of over 200 people, with a culture that values impact, autonomy, and inclusivity.

$86,400–$151,200/yr
Europe

  • Transform an existing AI agent product into scalable, reliable infrastructure.
  • Design and operate the core agent runtime, including orchestration, memory, and multi-agent workflows.
  • Work directly with the founder to set architecture, engineering standards, and development culture.

An early-stage AI infrastructure startup transforming a working AI agent product into scalable infrastructure. As the first engineering hire, you'll work directly with the founder to shape engineering culture and standards from the ground up.

United States

  • Evaluate AI model responses across diverse topics and use cases against established quality standards.
  • Write, test, and refine system prompts to influence model behavior and performance.
  • Identify and document model failures and provide clear, actionable feedback to product teams.

This company specializes in evaluating and improving advanced AI systems through content evaluation and prompt engineering. It operates as a partner company managing applications for this role, focusing on quality and consistency in AI behavior.

US 4w PTO

  • Own the technical direction of a product area, defining how teams specify, delegate, and verify agentic work.
  • Set standards for specification quality, verification, and security architecture to keep AI-generated work safe at scale.
  • Mentor engineers across teams and drive large-scale initiatives combining human and agentic contributors.

BambooHR builds a people intelligence platform that transforms HR for small and mid-sized businesses. The company is a values-driven market leader with a culture that champions growth, flexibility, and meaningful work.