Review real user interaction traces with an AI shopping assistant
Identify logical failures, inaccuracies, or poor recommendations in the text
Create structured rubrics and verifiers to judge response quality
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply careful judgment to ensure high-quality results aligned with task objectives.
LILT provides multilingual AI and human-verified services to enterprises, governments, and AI developers. The company has a global community of linguists and language professionals committed to innovation and excellence.
Analyze video content and annotate sequences with concise, objective summaries using simple English.
Segment videos into action-based sequences and label only visually verifiable actions.
Review AI-generated captions for accuracy and apply quality feedback to improve precision.
Jobgether operates an AI-powered recruitment platform that matches candidates with job opportunities. They emphasize objective, fair review processes and work with partner companies to manage hiring.
Work remotely with a commitment to reside in Central African Republic.
Appen specializes in providing high-quality training data for AI and machine learning models. They operate a global platform with a large crowd of independent contractors, focusing on flexible task-based work.
Evaluate AI-generated responses for relevance, accuracy, and personalization using personalized prompts and data from connected Google applications.
Identify subtle issues such as incorrect assumptions, irrelevant recommendations, inconsistencies, and inappropriate personalization.
Provide clear, detailed, and structured feedback to support improvements to AI models and personalization systems.
Our partner company is seeking an AI Response Quality Evaluator to improve AI-generated responses. This is a project-based contract role with a remote, independent working environment and a duration of up to 16 weeks.
Evaluate AI-generated content across humanities, arts, and culture domains for accuracy and quality.
Provide structured feedback and identify issues in AI outputs.
Work remotely on a flexible schedule contributing to AI system improvement.
The company specializes in developing and improving AI systems through expert human evaluation. They operate as a remote, flexible organization that values specialized domain knowledge and critical analysis.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Review real user interactions with an AI shopping assistant and identify flaws in accuracy and usefulness.
Analyze response quality from an e-commerce perspective, considering product recommendations and user needs.
Create structured rubrics and verifiers for consistent evaluation of future AI responses.
The partner company is developing an AI-powered digital shopping assistant and seeks evaluators to assess and improve its responses. The team size and culture are not specified.
Review robot manipulation video annotations for accuracy across multiple camera views.
Provide constructive feedback to annotators and identify recurring errors.
Maintain a QC log and participate in calibration sessions to align standards.
Welo Data, part of Welocalize, is a global AI data company with over 500,000 contributors. They provide high-quality, ethical data to train advanced AI systems and emphasize flexibility and growth for their contributors.
Label, annotate, and evaluate German-language content including photos, graphics, and videos for linguistic and cultural accuracy.
Evaluate AI-generated content against Canva's quality bar for German users to shape language experiences.
Build and contribute to German-specific datasets to support the internationalization of Canva AI features.
Canva is a design platform redefining how the world experiences design. It is a global company with a large user base, known for its innovative culture and focus on AI-powered features.
Listen to recorded conversations between users and AI voice agents to evaluate response quality.
Rate each agent turn on two scales: content helpfulness and prosody naturalness.
Write short, specific justifications for each score given.
Welo Data is an AI services company that provides evaluation and data services for AI voice agents. They are a global organization with a freelance workforce, focusing on quality assessment of AI interactions.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.
Evaluate AI-generated Azerbaijani text for naturalness and authenticity.
Compare text snippets and provide quality control on cultural nuance.
Rate AI-generated text and tag data on tone and naturalness.
Prolific is building the largest pool of quality human data in the world, with over 35,000 AI developers and researchers using its platform. They connect researchers with paid participants to collect ethically sourced human behavioral data and feedback.
Evaluate conversations between users and AI voice agents by listening to recorded interactions.
Rate each AI agent turn on two 1-5 scales: content quality and prosody naturalness.
Provide clear, specific written justifications for each rating given.
Welo Data provides AI services, specializing in data evaluation and human feedback for AI systems. They are an established company with a global community of freelancers and a focus on remote, collaborative work.
Evaluate AI-generated design outputs across various formats, including visual assets, layouts, websites, and games.
Assess designs based on clarity, functionality, accessibility, and effectiveness, not just visual appeal.
Write detailed annotations and critiques to distinguish between mediocre, good, and exceptional design.
Our partner is a company developing advanced AI systems. They seek design professionals to evaluate and improve AI-generated design outputs, and offer remote work opportunities with flexible schedules.
Review and approve or reject user-generated content based on platform standards, documenting moderation decisions.
Conduct safety and compliance reviews on in-house content and workflows, including datasets and AI characters.
Moderate video, image, and text content, taking action on flagged violations and categorizing content for appropriate visibility.
EverAI is building the world's largest AI companionship platform, driven by a proprietary moderation system called EverGuard. With 50 million users in two years and a fully remote team of about 100, the company is a fast-growing, category-creating AI company.
Perform data annotation and classification tasks for AI training.
Complete tasks like sentiment analysis and categorization.
Work on a pay-per-task basis with flexible hours.
Appen is a leading provider of high-quality data for machine learning and artificial intelligence. They offer flexible, gig-based opportunities with a large community of contributors worldwide.
Own end-to-end quality and delivery for Human Data projects, reviewing work and managing performance.
Coach and develop AI Tutors, conduct regular reviews, and implement process improvements.
Collaborate with Human Data Managers and Engineering to translate model requirements into labeling strategies.
Our partner is a company that produces and evaluates high-quality human data to support the development of advanced AI systems. They are part of a collaborative, engineering-focused environment with a distributed team, emphasizing initiative and ownership.
Evaluate AI-generated slides, spreadsheets, and documents for real-world usability and professional quality.
Assess outputs for accuracy, clarity, relevance, and alignment with data science standards.
Provide structured written feedback to help improve AI systems and their outputs.
A partner company is seeking a Data Science Expert to evaluate AI-generated work. The company focuses on improving AI systems and operates with a flexible, remote team.