Source Job

Global

  • Review real user interaction traces with an AI shopping assistant
  • Identify logical failures, inaccuracies, or poor recommendations in the text
  • Create structured rubrics and verifiers to judge response quality

Data Evaluation Quality Assurance Analytical Skills E-commerce AI Training

20 jobs similar to AI Evaluators: Assessing A Shopping Assistant

Jobs ranked by similarity.

US

  • Review real user interactions with an AI shopping assistant and identify flaws in accuracy and usefulness.
  • Analyze response quality from an e-commerce perspective, considering product recommendations and user needs.
  • Create structured rubrics and verifiers for consistent evaluation of future AI responses.

The partner company is developing an AI-powered digital shopping assistant and seeks evaluators to assess and improve its responses. The team size and culture are not specified.

US

  • Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
  • Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
  • Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.

The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.

Canada

  • Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
  • Identify factual, formatting, visual, and structural issues in professional deliverables.
  • Provide clear, structured feedback to enhance AI output quality and consistency.

This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.

  • Translate a complex workflow into a demanding AI prompt designed to expose model limitations.
  • Test your prompt in ChatGPT, refine it until the AI fails, and write a grading rubric for others to use.
  • Submit your prompt, failure notes, rubric, and a screen recording of your thought process.

Terac builds the world's largest pool of vetted human experts for AI research and evaluation. They are a growing platform used by AI labs and researchers to recruit, screen, and pay study participants globally.

$15–$15/hr
US

  • Review, evaluate, and validate AI-generated content and data according to project guidelines.
  • Compare AI outputs and identify the most accurate, relevant, or high-quality results.
  • Annotate, categorize, or label text, audio, images, video, or other data with clear feedback.

Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train advanced AI systems. They offer limitless flexibility and growth for a diverse community in 100+ countries.

Global

  • Evaluate LLM responses for accuracy, clarity, and completeness.
  • Fact-check technical claims using authoritative references.
  • Validate code and outputs, and annotate model performance.

Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.

UK

  • Evaluate AI-generated support replies for accuracy and brand voice consistency.
  • Simulate realistic customer interactions to test AI handling of edge cases.
  • Audit AI-generated FAQ articles and knowledge bases for technical correctness.

Prolific is building the world's largest pool of quality human data for AI development. They connect researchers with paid participants to gather diverse behavioral data, and are at the forefront of ethical AI data collection.

Canada

  • Evaluate AI-generated documents and presentations against quality standards.
  • Apply humanities expertise to identify inaccuracies and cultural issues.
  • Provide structured feedback to improve AI model performance.

A partner company is seeking a humanities evaluator to assess AI-generated content for accuracy and quality. The company emphasizes cultural awareness and critical thinking in a remote, asynchronous work environment.

$45–$55/hr
US

  • Evaluate user requests and AI model responses against detailed customer policies with precise reasoning.
  • Distinguish subtle differences in context and intent to classify ambiguous cases accurately.
  • Participate in calibration discussions and contribute to improving evaluation frameworks.

Handshake powers a platform connecting 25 million job seekers with employers and educational institutions. Through Handshake AI, they provide data to frontier AI labs, having grown to a ~$1B run rate and paying over 30K individuals monthly.

US

  • Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
  • Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
  • Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.

Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.

US

  • Evaluate AI-generated content for quality, accuracy, and cultural relevance
  • Apply Castilian Spanish expertise to assess response appropriateness for Spain
  • Provide structured feedback and document decisions to improve AI performance

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use objective, data-driven recruitment processes and prioritize privacy and fairness.

$15–$15/hr
United States

  • Rating and assessing the performance of AI models based on their output or behavior.
  • Labeling and categorizing content to train machine learning models.
  • Generating prompts, responses, and summaries to improve language model reasoning.

Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, the company focuses on enabling responsible AI advancement.

Brazil

  • Accurately annotate and evaluate text, video, and geographic data following detailed guidelines.
  • Conduct research and review your work to ensure high standards of consistency and quality.
  • Collaborate with an international team and provide feedback to improve annotation processes.

They are an AI data company that improves the accuracy and performance of generative AI models through data annotation. They operate with a fully remote, collaborative team culture.

US Unlimited PTO

  • Evaluate AI systems at a scale only possible by combining thousands of vetted experts with model graders.
  • Innovate at the frontier of QA by shaping industry standards for validating agentic AI and large language models.
  • Collaborate with global market leaders to architect AI quality blueprints and drive high-impact consultative visibility.

Testlio provides a fully managed crowdsourced testing platform powered by proprietary intelligence technology, LeoCore. They are a female-founded, fully remote company with an inclusive culture, half of their team identifying as women, and are growing profitably.

US

  • Support participants in virtual HR workshops by diagnosing and improving generative AI outputs.
  • Provide tool-agnostic guidance to help attendees refine prompts and achieve practical results.
  • Ensure responsible AI use and escalate issues as needed while fostering productive learning.

This company is a leading technology firm specializing in internet-related services and products, including search, cloud computing, and AI. It is a large global organization with a culture of innovation and collaboration.

$200,000–$275,000/yr
Global

  • Shape Dia's AI experience by defining the right behavior for users and finding the best path to it.
  • Ship every week with hands-on work in prompts, code, and data that reaches users continuously.
  • Improve systematically by grounding work in evals that make quality trackable and drive faster iteration.

The Browser Company is building a better way to use the internet. They are a remote-first, distributed team of close to 100 people who are passionate about building great products.

  • Write detailed outlines of your regular workflows, focusing on one critical task performed at least weekly.
  • Provide structured evaluation tasks and nuanced feedback to train AI models.
  • Complete paid tasks remotely on a freelance basis, with most tasks requiring one hour of uninterrupted work.

Prolific builds the world's largest pool of quality human data for AI development. With over 35,000 AI developers and organizations using its platform, it focuses on ethically sourced behavioral data from paid participants.

Canada

  • Evaluate AI-generated spreadsheets against quality standards and domain-specific rubrics.
  • Identify calculation errors, formatting issues, and inconsistencies in workbooks.
  • Provide structured, actionable feedback to improve AI output accuracy and usability.

The partner company focuses on evaluating AI-generated spreadsheets and workbooks. It offers a remote, asynchronous work environment with flexible scheduling.

Portugal

  • Evaluate AI-generated responses for logical consistency and business accuracy across various management scenarios.
  • Assess AI models' understanding of business strategy, operations, and organizational behavior.
  • Provide structured feedback to improve AI reasoning and decision-making in realistic business contexts.

Our partner is a technology company that develops advanced AI systems and seeks freelance professionals to train AI models in business contexts. The company offers flexible remote work and is looking for independent contractors with strong business knowledge.

US

  • Evaluate financial documents and reports to verify accuracy and provide AI training data.
  • Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
  • Validate AI outputs against professional financial standards and provide expert feedback.

Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.