Listen carefully to recorded conversations between a person and an AI voice agent.
Rate each agent turn on two 1–5 scales: Content (helpfulness/relevance) and Prosody (naturalness/expressiveness).
Write short, specific justifications for each score given.
Welo Data provides AI services, specializing in evaluating and improving AI voice agents through detailed audio rating. As a growing community of freelancers, they emphasize attention to detail and clear communication for project-based work.
Evaluate conversations between users and AI voice agents by listening to recorded interactions.
Rate each AI agent turn on two 1-5 scales: content quality and prosody naturalness.
Provide clear, specific written justifications for each rating given.
Welo Data provides AI services, specializing in data evaluation and human feedback for AI systems. They are an established company with a global community of freelancers and a focus on remote, collaborative work.
Listen to recorded conversations between users and AI voice agents to evaluate response quality.
Rate each agent turn on two scales: content helpfulness and prosody naturalness.
Write short, specific justifications for each score given.
Welo Data is an AI services company that provides evaluation and data services for AI voice agents. They are a global organization with a freelance workforce, focusing on quality assessment of AI interactions.
Listen carefully to recorded conversations between a person and an AI voice agent.
Rate each of the agent's turns on two independent 1–5 scales for content and prosody.
Write a short, specific, and clear justification for each score given.
Welo Data provides AI services, specializing in evaluating and improving AI voice agents. They are a growing community of freelancers focused on data quality and human-in-the-loop tasks.
Evaluate AI model performance through real-time, voice-based conversations by roleplaying assigned scenarios with two different models.
Compare model responses across five defined dimensions, identify error clusters, and select the stronger performer with a detailed rationale.
Maintain consistent conversational turns and voice recording to ensure fair, objective comparisons.
Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, they deliver high-quality data and outcomes for AI builders and adopters.
Audit 30% of production output to ensure rating quality and consistency during the pilot phase.
Identify and flag miscalibration or inconsistent scoring across raters.
Provide clear, actionable feedback and corrections to raters to maintain alignment.
Welo Data provides AI services, focusing on data annotation and quality assurance for machine learning models. They are a global freelance community with a collaborative, independent culture.
Audit 30% of production output to ensure rating quality and consistency during the initial phase.
Identify miscalibration or inconsistent scoring across raters and flag issues early.
Provide clear, actionable feedback and corrections to raters to maintain alignment.
Welo Data provides AI services and data solutions for machine learning and artificial intelligence projects. They are a growing community of freelancers and contractors working on various AI-related tasks.
Audit 30% of production output to identify and flag miscalibration or inconsistent scoring across raters.
Provide clear, actionable feedback and corrections to raters to ensure rating quality and consistency.
Complete onboarding and training ahead of the production pool's rollout to maintain alignment.
Welo Data provides AI services, specializing in data annotation and quality assurance for machine learning projects. They are a growing company focused on building a community of freelance raters and testers, offering flexible remote work opportunities.
Audit 30% of production output during the initial pilot phase to ensure rating quality and consistency.
Identify and flag miscalibration or inconsistent scoring across raters, providing clear feedback.
Complete onboarding and training ahead of the production pool's rollout while working independently.
Welo Data is an AI services company that provides data annotation and quality assurance for AI projects. They are a growing community of freelancers and contractors, fostering a remote work culture with a focus on accuracy and consistency.
Audit 30% of production output during the pilot phase to ensure rating quality and consistency.
Identify and flag miscalibration or inconsistent scoring across raters.
Provide clear, actionable feedback and corrections to raters.
Welo Data provides AI services, including data annotation and rating for machine learning projects. It is a community of freelancers working on various AI-related tasks.
Audit 30% of production output during the pilot phase to ensure rating quality and consistency.
Identify and flag miscalibration or inconsistent scoring across raters.
Provide clear, actionable feedback and corrections to raters.
Welo Data provides AI services and general application support. They are a freelance-oriented company focused on data quality and AI training projects.
Evaluate AI-generated responses for relevance, accuracy, and personalization using personalized prompts and data from connected Google applications.
Identify subtle issues such as incorrect assumptions, irrelevant recommendations, inconsistencies, and inappropriate personalization.
Provide clear, detailed, and structured feedback to support improvements to AI models and personalization systems.
Our partner company is seeking an AI Response Quality Evaluator to improve AI-generated responses. This is a project-based contract role with a remote, independent working environment and a duration of up to 16 weeks.
Conduct voice-mode conversations with AI chatbots in Javanese (Indonesia).
Evaluate and rate the quality of AI voice interactions based on detailed guidelines.
Complete all evaluations in English and maintain confidentiality.
This company is a partner organization that develops conversational AI systems. They are looking for Javanese speakers to evaluate AI voice modes on a project basis.
Define the persona and voice system for Deepgram voice agents, ensuring coherence across use cases.
Design turn-taking, barge-in, and repair behaviors with ML and Engineering, tuning for real-time interaction.
Build conversational-quality evals and developer-facing guidance that establish best practices for voice agents.
Deepgram is the leading platform for the Voice AI economy, providing real-time APIs for speech-to-text and text-to-speech, enabling organizations to build production-grade voice agents. Backed by a recent Series C and processed over 50,000 years of audio, Deepgram is a growth-stage company with an AI-first culture where innovation and rapid adaptation are key.
Review real user interaction traces with an AI shopping assistant
Identify logical failures, inaccuracies, or poor recommendations in the text
Create structured rubrics and verifiers to judge response quality
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate search results and AI-generated content for quality and relevance.
Conduct online research to verify information and support rating decisions.
Provide feedback and document edge cases to improve AI systems.
TELUS Digital AI is a global AI community of over 1 million contributors helping clients collect, enhance, and train data to build better AI models. They offer flexible remote work and a diverse, inclusive culture.
Evaluate user requests and AI model responses against detailed customer policies with precise reasoning.
Distinguish subtle differences in context and intent to classify ambiguous cases accurately.
Participate in calibration discussions and contribute to improving evaluation frameworks.
Handshake powers a platform connecting 25 million job seekers with employers and educational institutions. Through Handshake AI, they provide data to frontier AI labs, having grown to a ~$1B run rate and paying over 30K individuals monthly.
Comparing text and voice snippets to assess quality and authenticity.
Listening to AI-generated audio and rating how natural the voice sounds.
Identifying where AI tone or pronunciation feels unnatural or culturally mismatched.
Prolific builds the world's largest pool of quality human data, connecting AI developers and researchers with paid study participants. Over 35,000 organizations use our platform for ethically sourced human behavioral data to improve AI.
Review text or media samples based on provided project guidelines
Apply accurate labels and categorizations to diverse data sets
Evaluate AI-generated responses for clarity, safety, and factual accuracy
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.