Review real user interactions with an AI shopping assistant and identify flaws in accuracy and usefulness.
Analyze response quality from an e-commerce perspective, considering product recommendations and user needs.
Create structured rubrics and verifiers for consistent evaluation of future AI responses.
The partner company is developing an AI-powered digital shopping assistant and seeks evaluators to assess and improve its responses. The team size and culture are not specified.
Conduct voice-mode conversations with AI chatbots in Javanese (Indonesia).
Evaluate and rate the quality of AI voice interactions based on detailed guidelines.
Complete all evaluations in English and maintain confidentiality.
This company is a partner organization that develops conversational AI systems. They are looking for Javanese speakers to evaluate AI voice modes on a project basis.
Review real user interaction traces with an AI shopping assistant
Identify logical failures, inaccuracies, or poor recommendations in the text
Create structured rubrics and verifiers to judge response quality
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate and score AI-generated customer interactions across Foundational, Experiential, and Operational dimensions using complex rubrics.
Benchmark informational and transactional customer queries against authoritative business sources to ensure accuracy.
Participate in dual-review processes and daily calibration audits to maintain inter-rater agreement and quality standards.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. The company has a 36+ year legacy delivering high-quality data and outstanding outcomes for customers.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Evaluate AI systems at a scale only possible by combining thousands of vetted experts with model graders.
Innovate at the frontier of QA by shaping industry standards for validating agentic AI and large language models.
Collaborate with global market leaders to architect AI quality blueprints and drive high-impact consultative visibility.
Testlio provides a fully managed crowdsourced testing platform powered by proprietary intelligence technology, LeoCore. They are a female-founded, fully remote company with an inclusive culture, half of their team identifying as women, and are growing profitably.
Comparing text and voice snippets to assess quality and authenticity.
Listening to AI-generated audio and rating how natural the voice sounds.
Identifying where AI tone or pronunciation feels unnatural or culturally mismatched.
Prolific builds the world's largest pool of quality human data, connecting AI developers and researchers with paid study participants. Over 35,000 organizations use our platform for ethically sourced human behavioral data to improve AI.
You will design the architecture for specialized subagents operating within live customer conversations, including technical QA, product expertise, and objection handling.
You will build routing and delegation systems that determine when to answer directly, invoke a subagent, or escalate to a human.
You will master the dialogue platform, train AI agents via prompting and fine-tuning, and document workflows to educate the team.
1mind builds autonomous customer experience software that deploys AI-powered 'Superhumans' to engage, demo, onboard, and support customers across the entire buying journey. The company offers a remote-first, fast-moving culture with ownership, autonomy, and impact from day one.
Record high-quality Latvian voice content for AI training.
Evaluate AI-generated speech for accuracy and naturalness.
Provide detailed feedback to improve voice models.
Our partner company develops next-generation AI voice technologies. As a freelance contractor, you will work remotely to create high-quality training data for AI voice models.
Evaluate user requests and AI model responses against detailed customer policies with precise reasoning.
Distinguish subtle differences in context and intent to classify ambiguous cases accurately.
Participate in calibration discussions and contribute to improving evaluation frameworks.
Handshake powers a platform connecting 25 million job seekers with employers and educational institutions. Through Handshake AI, they provide data to frontier AI labs, having grown to a ~$1B run rate and paying over 30K individuals monthly.
Join a brief, remote audio session with an AI moderator
Listen to automated prompts and respond out loud with your honest opinions
Provide candid feedback based on your personal preferences
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Own an AI-native product end to end, from customer problem to launch criteria to production performance.
Deploy forward with some of the largest IT organizations, building in their environment and analyzing AI Agent execution logs.
Define quality standards for autonomous agents, own outcome metrics, and run design partnerships with enterprises.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better with an AI-native culture. They are a large enterprise company focused on putting AI to work for people through their platform.
Record scripted training material in Welsh with clear pronunciation and natural delivery.
Evaluate AI-generated speech for linguistic accuracy, fluency, naturalness, and expression.
Collaborate with project teams to refine prompts and voice design guidelines.
Our partner is developing next-generation AI voice technologies. The project offers a flexible freelance environment where your linguistic and performance expertise can have a global impact.
Record scripted speech in Icelandic with clear pronunciation and natural delivery.
Evaluate AI-generated speech for linguistic accuracy, naturalness, and performance quality.
Provide expert feedback on pronunciation, pacing, tone, and regional variation.
The partner company is developing next-generation AI voice and speech technologies. This is a remote contract project with a team of voice professionals and AI experts.
Design and build voice AI systems for real-time, conversational user experiences.
Develop and maintain AI orchestration layers, including agent workflows, state management, and system handoffs.
Build data pipelines that capture, transform, and surface insights from patient and member interactions.
Blooming Health is a mission-driven, venture-backed health tech company transforming the social care landscape by building an intelligent platform that helps people navigate and access resources for holistic well-being. They are a fast-moving, mission-focused team working to improve health equity.
Translate a complex workflow into a demanding AI prompt designed to expose model limitations.
Test your prompt in ChatGPT, refine it until the AI fails, and write a grading rubric for others to use.
Submit your prompt, failure notes, rubric, and a screen recording of your thought process.
Terac builds the world's largest pool of vetted human experts for AI research and evaluation. They are a growing platform used by AI labs and researchers to recruit, screen, and pay study participants globally.
Evaluate AI-generated legal research and analysis for accuracy, relevance, and completeness.
Verify legal citations, authorities, and reasoning to identify errors and weaknesses.
Develop objective evaluation criteria and provide structured feedback to improve AI legal content.
Jobgether uses an AI-powered matching process to connect top-fitting candidates with hiring companies. They prioritize objective and fair review, sharing shortlists directly with employers for final decisions.
Support participants in virtual HR workshops by diagnosing and improving generative AI outputs.
Provide tool-agnostic guidance to help attendees refine prompts and achieve practical results.
Ensure responsible AI use and escalate issues as needed while fostering productive learning.
This company is a leading technology firm specializing in internet-related services and products, including search, cloud computing, and AI. It is a large global organization with a culture of innovation and collaboration.