Follow detailed project guidelines to evaluate AI voice and audio applications.
Record audio outputs and test application behavior for instruction adherence.
Identify bugs, inconsistencies, and naturalness issues affecting system quality.
The company is an enterprise client that requires Arabic-speaking Mac users to test application-level AI voice and audio evaluation tasks. The project is short-term, taking up to 4-5 hours, with flexible scheduling and potential for additional work.
Engage in conversations with a real-time speech-to-speech AI model
Evaluate performance on speech recognition, audio quality, conversation flow, and content accuracy
Provide accurate ratings based on project guidelines within specified timelines
Appen is a global leader in AI training data and crowd-sourced solutions. They work with a large community of independent contractors to improve AI systems through human evaluation.
Perform side-by-side comparisons of AI-generated responses.
Evaluate outputs for accuracy, relevance, clarity, instruction-following, and overall quality.
Assess AI responses across general-purpose questions and answers, web search results, file-based and image-based responses, content-generation tasks, and single-turn and multi-turn conversations.
We are a technology solutions firm headquartered in Bellevue, Washington, with teams across the United States. We help organizations turn complex challenges into meaningful outcomes by connecting strategy and execution across AI, cloud, data, product development, and emerging technology.
Record natural, conversational customer-service style speech from your home recording setup.
Submit two audition recordings: one unscripted and one scripted customer-service sample.
Complete two supervised recording sessions of approximately 4 hours each.
Welo Data provides AI services, specializing in voice data collection for conversational AI applications. The company operates as a freelance platform with a community of contributors.
Evaluate AI-generated content for quality, accuracy, and cultural relevance
Apply Castilian Spanish expertise to assess response appropriateness for Spain
Provide structured feedback and document decisions to improve AI performance
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use objective, data-driven recruitment processes and prioritize privacy and fairness.
Facilitate live voice recording sessions for a multilingual conversational AI project in Spanish.
Provide real-time coaching and technical quality monitoring to ensure natural, authentic recordings.
Complete calibration and support sessions independently, resolving issues to minimize re-records.
Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train the world’s most advanced AI systems. We build smarter, more human AI with a diverse community in 100+ countries.
Lead and support a team of approximately 10 Quality Control Reviewers, ensuring consistent quality standards across the Spanish (Spain) locale.
Act as the primary quality point of contact for the Spanish (Spain) locale, resolving quality questions and escalating complex issues when necessary.
Conduct calibration sessions and provide structured coaching to improve reviewer performance.
Welo Data provides data quality services for AI training programs, focusing on linguistic quality and annotation. The team is dedicated to improving AI systems through rigorous quality assurance and collaboration.
Evaluate AI-generated responses in Dutch for language quality, customer experience, technical accuracy, and JSON structure.
Participate in real-time conversational scenarios with AI systems and review transcripts.
Capture session artifacts and provide written justifications using structured rubrics.
An enterprise client is hiring contract experts to evaluate AI-generated audio outputs for quality and technical accuracy. The project offers flexible scheduling and remote work, with a focus on language fluency and coding skills.
Side-by-side evaluation of text and voice snippets to assess quality and authenticity.
Listening to AI-generated audio and rating how natural the voice sounds.
Identifying cultural mismatches in tone, pronunciation, and intonation.
Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. They focus on ethical data collection and diverse human perspectives to improve AI.
Evaluate AI-generated text and audio in Catalan for accuracy and natural flow.
Provide corrections and constructive feedback on grammar, tone, and cultural context.
Complete approximately 10 hours of asynchronous tasks each week via our online platform.
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Curate and annotate multilingual audio data to train AI for voice interactions and speech recognition.
Ensure high-quality voice recordings and accurate transcriptions across diverse languages and accents.
Collaborate with technical staff to improve annotation tools and audio workflows.
SpaceXAI creates AI systems to understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence.
Evaluate AI-generated text and voice snippets in Marathi for quality and authenticity.
Listen to audio clips and rate how natural the AI voice sounds.
Provide feedback on tone, pronunciation, and cultural context.
Prolific is an AI data platform that connects researchers with a global pool of participants to gather high-quality, ethically sourced human data. With over 35,000 AI developers and organizations using the platform, Prolific is building the largest pool of quality human data to train AI models.
Listen to two audio recordings and evaluate which is better.
Follow project guidelines to make consistent choices.
Complete 20-23 cases per hour remotely.
CrowdGen by Appen is an AI data company that improves AI systems through human feedback. They offer flexible, project-based remote work for independent contractors.