Lead a team of Quality Control Reviewers to ensure consistent quality standards for the Japanese locale.
Review and audit feedback to ensure alignment with project guidelines.
Act as the primary quality point of contact for the Japanese locale.
Welo Data provides AI training data and quality services to improve AI systems. They are a remote-first company focused on language quality and team collaboration.
Evaluate prompts and AI-generated outputs for accuracy, cultural appropriateness, and brand alignment.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply local cultural insight and consistent evaluation guidelines to ensure high-quality AI training.
Lilt provides multilingual AI and human-verified services to enterprises, governments, and AI developers. They foster a global community of linguists and subject matter experts working on cutting-edge AI and language technology.
Review scientific papers and LLM-generated graphical abstracts for accuracy.
Fact-check and identify inaccuracies in AI-generated summaries.
Use domain expertise to verify that technical concepts are correctly represented.
Prolific is building the largest pool of quality human data for AI development, serving over 35,000 developers. They focus on ethical data collection and aim to integrate human perspectives into AI.
Provide native-level Canadian French language vetting and QA for AI data projects.
Annotate and review AI outputs for grammatical accuracy, cultural context, and naturalness.
Develop educational resources and feedback documentation to improve AI alignment.
We are an AI training company that focuses on language alignment and data annotation for AI systems. Our remote team values linguistic precision and cultural nuance.
Translate, edit, and proofread certified legal content from English (en-US) to Japanese (ja-JP) preserving legal meaning.
Work in XTM Cloud, following style guides, terminology, and strict turnaround SLAs in a multi-vendor environment.
Perform terminology research, implement LQA feedback, and flag linguistic or legal-risk issues proactively.
Welo is a global language services company providing translation, localization, and interpretation solutions. They operate as a multi-vendor platform serving tech companies and other industries, with a focus on certified legal content.
Design realistic Real Estate and Leasing scenarios in Urdu or English grounded in Indian operational contexts.
Adapt structured evaluation rubrics for customer interaction, operational problem-solving, leasing workflows, and sales support tasks.
Review AI and human-generated responses for factual accuracy, service quality, policy compliance, and operational realism.
LILT is an AI and language technology company that provides multilingual AI and human-verified services to enterprises, governments, and AI developers. The company fosters a supportive, innovative global community of linguists and subject matter experts who work on diverse projects.
Evaluate AI-generated research on environmental topics for scientific accuracy and nuance.
Fact-check environmental claims, emissions factors, and legislative requirements.
Assess sustainability logic and annotate geospatial data to improve model outputs.
Prolific builds the world's largest pool of quality human data for AI development. With over 35,000 AI developers, researchers, and organizations using the platform, they focus on ethically sourced human behavioral data to train and evaluate AI models.
Train and evaluate cutting-edge AI models by completing language tasks in Italian/Spanish/French/German/Dutch.
Judge the performance of AI in performing Italian prompts and improve its capabilities.
Analyze, edit, and write in target languages with strong attention to detail for up to one hour per task.
Prolific is building the largest pool of quality human data in the world for AI training. Over 35,000 AI developers, researchers, and organizations use the platform to gather data from paid participants with diverse experiences, skills, and knowledge.
Provide linguistic QA and develop alignment resources for AI data projects.
Review, annotate, and test AI outputs for grammatical accuracy and cultural context.
Act as a primary quality check to identify and correct subtle errors in Czech language.
We specialize in AI data projects, focusing on linguistic quality and cultural relevance. We are a global project that operates remotely with freelance contractors.
Evaluate AI-generated scientific responses for accuracy and reasoning in biology.
Fact-check technical claims from public databases like PubMed and NCBI.
Assess experimental logic and annotate errors in biological sequences or protocols.
Prolific builds the largest pool of quality human data for AI development. With over 35,000 AI developers and researchers using the platform, it connects experts to train and evaluate AI models through ethical, paid participation.
Evaluate AI-generated text and voice snippets in Punjabi for naturalness and authenticity.
Assess audio clips for cultural and tonal accuracy of AI speech.
Provide feedback on linguistic nuance and quality of AI outputs.
Prolific builds the biggest pool of quality human data in the world, serving over 35,000 AI developers and researchers. The company connects researchers with paid study participants from diverse backgrounds to gather high-quality, ethically sourced behavioral data.
Evaluate AI-generated text and audio in Catalan for accuracy and natural flow.
Provide corrections and constructive feedback on grammar, tone, and cultural context.
Complete approximately 10 hours of asynchronous tasks each week via our online platform.
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate the Japanese app and web experience across core user journeys to identify language, terminology, and cultural issues.
Update translations directly in the translation management system and document issues that require broader product or design changes.
Collaborate with the Localization team to prioritize improvements and ensure consistency in Japanese terminology.
Whatnot is the largest live shopping platform in North America and Europe, enabling users to buy, sell, and discover items through live video. They are a remote co-located team with hubs across multiple countries, recently named the #1 Best Startup Employer in America by Forbes.
Evaluate AI-generated responses for accuracy, grammar, and cultural relevance.
Identify issues and provide refined, high-quality rewritten responses.
Create natural prompts and responses in Hindi to improve conversational datasets.
Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train the world’s most advanced AI systems. They're building smarter, more human AI with a diverse community in 100+ countries.
Own the design and defense of frontier model evaluations across reasoning, coding, agents, tool use, and multi-modal.
Build benchmark packages with expert-verified ground truth, multi-model headroom results, and rigorous QC.
Recruit, calibrate, and review a pool of subject-matter experts in coding, agentic/tool-use, and STEM/reasoning.
Anyone AI measures frontier model capability through expert-verified evaluation packages. The company operates as a remote team with a focus on rigorous benchmarking and lab collaboration.
Review English source documents alongside two machine-generated Assamese translations, evaluating accuracy, fluency, and overall quality.
Select the preferred translation and provide a clear written justification for your assessment.
Complete assigned samples independently within established timelines, adhering strictly to project and client guidelines.
Welo Data provides AI operations and data generation services for leading technology clients. They are a global contributor community offering project-based opportunities with flexible, remote work.
Evaluate financial documents and verify information accuracy for AI model training.
Respond to prompts with financial expertise to help AI understand complex fiscal concepts.
Provide data validation and expert feedback to bridge human financial knowledge and machine learning.
Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers and researchers. They connect researchers with a global pool of participants for ethically sourced human behavioral data, focusing on integrating diverse human perspectives into AI development.
Side-by-side evaluation of text and voice snippets to assess quality and authenticity.
Naturalness assessment of AI-generated audio to rate how natural the voice sounds.
Quality control to identify where the AI's tone or pronunciation feels unnatural or culturally mismatched.
Prolific builds the largest pool of quality human data for AI development. Over 35,000 AI developers, researchers, and organizations use Prolific to gather data from paid study participants.
Create realistic, domain-specific tasks in Spanish/English reflecting Mexico-specific finance practices.
Adapt and apply clear scoring rubrics to evaluate AI-generated and human responses.
Review and score submissions for accuracy, regulatory alignment, and professional quality.
Lilt's mission is to make the world's information available to everyone, no matter the language they speak. As a global community of linguists and experts, they deliver multilingual AI and human-verified services to Enterprises, Governments, and AI Developers worldwide.
Write prompts in Luxembourgish to test AI models and evaluate responses using a structured rating guide.
Identify areas for improvement in AI responses and provide accurate English translations.
Work remotely on a flexible schedule with competitive pay.
CrowdGen by Appen is a platform that connects freelancers to help improve AI models through data annotation and evaluation. It is part of a large global company with a diverse community of independent contractors working remotely.