Evaluate AI-generated text and voice snippets in Marathi for quality and authenticity.
Listen to audio clips and rate how natural the AI voice sounds.
Provide feedback on tone, pronunciation, and cultural context.
Prolific is an AI data platform that connects researchers with a global pool of participants to gather high-quality, ethically sourced human data. With over 35,000 AI developers and organizations using the platform, Prolific is building the largest pool of quality human data to train AI models.
Perform side-by-side comparisons of AI-generated responses.
Evaluate outputs for accuracy, relevance, clarity, instruction-following, and overall quality.
Assess AI responses across general-purpose questions and answers, web search results, file-based and image-based responses, content-generation tasks, and single-turn and multi-turn conversations.
We are a technology solutions firm headquartered in Bellevue, Washington, with teams across the United States. We help organizations turn complex challenges into meaningful outcomes by connecting strategy and execution across AI, cloud, data, product development, and emerging technology.
Rating and assessing the performance of AI models based on their output or behavior.
Labeling and categorizing content to train machine learning models.
Generating prompts, responses, and summaries to improve language model reasoning.
Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, the company focuses on enabling responsible AI advancement.
Create realistic, domain-specific tasks in the target language reflecting local finance practices.
Adapt and apply scoring rubrics to evaluate AI-generated and human responses.
Review and score submissions for accuracy, regulatory alignment, and professional quality.
LILT is a company that makes the world's information available to everyone, no matter the language they speak, through multilingual AI and human-verified services. They serve Enterprises, Governments, and AI Developers globally, with a community of linguists and experts focused on innovation and excellence.
Evaluate AI-generated content for quality, accuracy, and cultural relevance
Apply Castilian Spanish expertise to assess response appropriateness for Spain
Provide structured feedback and document decisions to improve AI performance
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use objective, data-driven recruitment processes and prioritize privacy and fairness.
Review search results and evaluate their relevance to user queries
Answer true/false questions about content quality
Rate search results based on guidelines to improve AI systems
Welo Data provides AI services and data validation to improve search engine and AI systems. They are a remote-first company with a focus on quality and support for their contractors.
Contribute to shaping safer, smarter AI by joining a global network of linguists and culturally aware contributors.
Work on flexible, remote projects in annotation, evaluation, and prompt creation, always on your terms.
Get first access to projects that match your skills, from short tasks to multi-week assignments.
Welo Data, part of Welocalize, is a global AI data company with a network of over 500,000 contributors. They build smarter, more human AI by offering flexible, remote projects to a diverse community in over 100 countries, emphasizing growth and work-life balance.
Dive into a cutting-edge AI benchmarking project focused on highly specific professional domains like software engineering, healthcare, and finance.
Design realistic scenarios in English or Korean, adapt evaluation rubrics, and review AI responses for accuracy and cultural appropriateness.
Contribute to gold-standard solutions that reflect best practices across your target locale and domain.
Lilt is a company that makes the world's information available to everyone, regardless of language, through multilingual AI and human-verified services for Enterprises, Governments, and AI Developers. The company has a global community of linguists and subject matter experts who thrive on innovation and excellence.
Evaluate AI-generated documents and presentations against quality standards.
Apply humanities expertise to identify inaccuracies and cultural issues.
Provide structured feedback to improve AI model performance.
A partner company is seeking a humanities evaluator to assess AI-generated content for accuracy and quality. The company emphasizes cultural awareness and critical thinking in a remote, asynchronous work environment.
Evaluate AI systems at a scale only possible by combining thousands of vetted experts with model graders.
Innovate at the frontier of QA by shaping industry standards for validating agentic AI and large language models.
Collaborate with global market leaders to architect AI quality blueprints and drive high-impact consultative visibility.
Testlio provides a fully managed crowdsourced testing platform powered by proprietary intelligence technology, LeoCore. They are a female-founded, fully remote company with an inclusive culture, half of their team identifying as women, and are growing profitably.
Lead the coordination of learning design projects, ensuring timelines, responsibilities, and deliverables are met.
Serve as the main point of contact for stakeholders, building positive relationships and aligning expectations.
Support people management, assign projects based on skills fit, and contribute to process improvements.
INFUSE is a demand generation company providing B2B solutions to help clients deliver audience, buyers, and account engagement. With a global team across more than 60 countries, the company has been recognized on the Inc. 5000 list and received over 60 industry awards, including Inc. Best Workplaces.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply local insight into tone, symbolism, visual cues, and market fit to deliver culturally relevant content.
LILT is an AI company that makes the world's information available to everyone, no matter the language they speak. They work with a global community of linguists and subject matter experts to deliver multilingual AI and human-verified services to Enterprises, Governments, and AI Developers.
Audit multiple-choice math questions for technical accuracy and correctness.
Review grading rubrics to ensure clear explanation of correct mathematical reasoning.
Propose edge cases and test conditions to strengthen assessment logic.
Terac is building the world's largest pool of vetted human experts for AI researchers and labs. They are a growing platform used to recruit, screen, and pay study participants across industries, languages, and skill sets.
Define and enforce quality standards for motion and static creative through review and coaching, while owning the team's delivery speed from brief to launch.
Own the team's production stack, evaluate and train on AI-assisted creative workflows, and build reusable systems to scale output without scaling headcount.
Translate performance data (CTR, CVR, CPA) into concrete creative direction, partnering with the user acquisition team and owning localization QA.
Ruby Labs is a leading tech company that creates and operates innovative consumer products across health, education, and entertainment industries. They are a fast-growing, ambitious team that values high standards, speed, and results, offering significant upside for strong operators.
Define and articulate design vision for Flex's most critical product initiatives, ensuring alignment with business goals and user needs.
Own the design process end-to-end on flagship initiatives, from research through prototyping and final execution.
Collaborate with cross-functional teams and advocate for design thinking at all levels.
Flex is a growth-stage FinTech company that enables renters to pay rent on a flexible schedule. The company has a dynamic team with employees across the US, Australia, Canada, and South America, and is focused on building an inclusive culture.
Complete AI training tasks such as analyzing, editing, and writing in Korean.
Evaluate and judge AI performance on Korean prompts.
Help improve cutting-edge AI models with your expertise.
Prolific builds the largest pool of quality human data for AI development. Over 35,000 developers and researchers use Prolific to gather diverse data from paid participants.
You will complete AI training tasks such as analyzing, editing, and writing in Japanese.
You will judge the performance of AI in performing Japanese prompts.
You will improve cutting-edge AI models.
We are building the biggest pool of quality human data in the world. Over 35,000 AI developers, researchers, and organizations use our platform to gather data from paid study participants.
Design conversational speech systems using Agentic AI and NLU for customer service automation.
Lead requirement gathering and client design sessions for contact center transformations.
Analyze interaction data to identify automation opportunities and drive data-driven experience improvements.
NeuraFlash, Part of Accenture, is a trusted leader in AI, AWS, and Salesforce innovation, redefining business through AI and technologies like Agentforce. They foster a culture of trust, flexibility, and collaboration, with over half of employees working remotely and a focus on celebrating achievements.
Evaluate and score AI-generated customer interactions across Foundational, Experiential, and Operational dimensions using complex rubrics.
Benchmark informational and transactional customer queries against authoritative business sources to ensure accuracy.
Participate in dual-review processes and daily calibration audits to maintain inter-rater agreement and quality standards.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. The company has a 36+ year legacy delivering high-quality data and outstanding outcomes for customers.