Review real user interaction traces with an AI shopping assistant
Identify logical failures, inaccuracies, or poor recommendations in the text
Create structured rubrics and verifiers to judge response quality
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate user requests and AI model responses against detailed customer policies with precise reasoning.
Distinguish subtle differences in context and intent to classify ambiguous cases accurately.
Participate in calibration discussions and contribute to improving evaluation frameworks.
Handshake powers a platform connecting 25 million job seekers with employers and educational institutions. Through Handshake AI, they provide data to frontier AI labs, having grown to a ~$1B run rate and paying over 30K individuals monthly.
Evaluate search results and AI-generated content for quality and relevance.
Conduct online research to verify information and support rating decisions.
Provide feedback and document edge cases to improve AI systems.
TELUS Digital AI is a global AI community of over 1 million contributors helping clients collect, enhance, and train data to build better AI models. They offer flexible remote work and a diverse, inclusive culture.
Review real user interactions with an AI shopping assistant and identify flaws in accuracy and usefulness.
Analyze response quality from an e-commerce perspective, considering product recommendations and user needs.
Create structured rubrics and verifiers for consistent evaluation of future AI responses.
The partner company is developing an AI-powered digital shopping assistant and seeks evaluators to assess and improve its responses. The team size and culture are not specified.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply careful judgment to ensure high-quality results aligned with task objectives.
LILT provides multilingual AI and human-verified services to enterprises, governments, and AI developers. The company has a global community of linguists and language professionals committed to innovation and excellence.
Evaluate and score AI-generated customer interactions across Foundational, Experiential, and Operational dimensions using complex rubrics.
Benchmark informational and transactional customer queries against authoritative business sources to ensure accuracy.
Participate in dual-review processes and daily calibration audits to maintain inter-rater agreement and quality standards.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. The company has a 36+ year legacy delivering high-quality data and outstanding outcomes for customers.
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.
Evaluate AI model performance through real-time, voice-based conversations by roleplaying assigned scenarios with two different models.
Compare model responses across five defined dimensions, identify error clusters, and select the stronger performer with a detailed rationale.
Maintain consistent conversational turns and voice recording to ensure fair, objective comparisons.
Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, they deliver high-quality data and outcomes for AI builders and adopters.
Evaluate search results and AI-generated content for quality, relevance, accuracy, and usefulness.
Conduct online research to verify information and support rating decisions.
Apply rating guidelines consistently and participate in training and calibration sessions.
TELUS Digital AI & Data Solutions partners with a diverse and vibrant community to help our customers enhance their AI and machine learning models. Our global AI community includes over 1 million contributors across 500+ languages and dialects, offering flexible remote and onsite opportunities.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Label, annotate, and evaluate German-language content including photos, graphics, and videos for linguistic and cultural accuracy.
Evaluate AI-generated content against Canva's quality bar for German users to shape language experiences.
Build and contribute to German-specific datasets to support the internationalization of Canva AI features.
Canva is a design platform redefining how the world experiences design. It is a global company with a large user base, known for its innovative culture and focus on AI-powered features.
Evaluate AI-generated work products in real estate, hospitality, and events using quality rubrics.
Identify factual, aesthetic, and presentation errors and provide actionable feedback.
Apply industry expertise to distinguish realistic, commercially sound work from generic AI content.
The company develops AI systems and evaluates their outputs for quality. They seek experienced industry professionals for flexible remote contract work.
Evaluate AI-generated slides, spreadsheets, and documents for real-world usability and professional quality.
Assess outputs for accuracy, clarity, relevance, and alignment with data science standards.
Provide structured written feedback to help improve AI systems and their outputs.
A partner company is seeking a Data Science Expert to evaluate AI-generated work. The company focuses on improving AI systems and operates with a flexible, remote team.
Support participants in virtual HR workshops by diagnosing and improving generative AI outputs.
Provide tool-agnostic guidance to help attendees refine prompts and achieve practical results.
Ensure responsible AI use and escalate issues as needed while fostering productive learning.
This company is a leading technology firm specializing in internet-related services and products, including search, cloud computing, and AI. It is a large global organization with a culture of innovation and collaboration.
Perform side-by-side comparisons of AI-generated responses and evaluate them for factual accuracy, relevance, and overall quality.
Apply deep Korean expertise to assess language, terminology, tone, and cultural context specific to Korea.
Complete evaluations within established time and productivity expectations while maintaining consistent judgment.
Blueprint is a technology solutions firm that helps organizations turn complex challenges into meaningful outcomes across AI, cloud, data, and emerging technology. The company has teams across the United States and a culture built on high standards, ownership, and mutual support.
Own and extend the offline evaluation suite for AI products, building datasets and metrics.
Build online quality dashboards and close the production feedback loop by mining failure patterns.
Translate numbers into clear decisions for Product and domain experts.
Finom is a European tech startup developing an all-in-one financial B2B platform integrating banking, accounting, and invoicing for entrepreneurs. With over €115 million in Series C funding and a team dedicated to innovation, they foster a start-up culture that values bold ideas and swift implementation.
Evaluate AI-generated market research and competitive intelligence artifacts for accuracy, rigor, and quality.
Apply structured rubrics to assess deliverables and identify factual inaccuracies and analytical gaps.
Provide clear, actionable written feedback to support evaluation decisions and improve AI training.
The partner company specializes in evaluating AI-generated market research and competitive intelligence content. They hire independent contractors for flexible remote engagements.
Quality-check ocean freight contracts and rates against source documents.
Collaborate with international vendor partners to resolve discrepancies.
Use AI tools to streamline QA workflows and improve efficiency.
The company specializes in ocean freight data quality assurance. It is a remote-first, international team with a feedback-driven culture that values ownership and continuous improvement.
Design, iterate, and maintain system prompts and instruction sets for AI agents in enrollment, learning, and student support.
Build and maintain evaluation frameworks to measure agent accuracy, tone, and task completion using observability tools like Langfuse.
Collaborate with product managers, learning designers, and engineers to translate educational goals into reliable agent behavior.
Noodle is higher education's leading strategy, services, and technology partner. We are a dynamic team focused on transforming education through innovative technology solutions.