Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.
Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.
Evaluate LLM responses for accuracy, clarity, and completeness.
Fact-check technical claims using authoritative references.
Validate code and outputs, and annotate model performance.
Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.
Walk us through your current professional workflows and AI tool usage.
Discuss specific pain points or limitations you encounter regularly.
Brainstorm and describe future capabilities you would like to see developed.
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Construct domain-specific materials based on provided prompts.
Provide structured feedback on legal reasoning workflows.
Terac is building the world's largest pool of vetted human experts for AI. They work with researchers, AI labs, and product teams to recruit, screen, and pay study participants across industries.
Write detailed outlines of your regular workflows, focusing on one critical task performed at least weekly.
Provide structured evaluation tasks and nuanced feedback to train AI models.
Complete paid tasks remotely on a freelance basis, with most tasks requiring one hour of uninterrupted work.
Prolific builds the world's largest pool of quality human data for AI development. With over 35,000 AI developers and organizations using its platform, it focuses on ethically sourced behavioral data from paid participants.
Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
Identify factual, formatting, visual, and structural issues in professional deliverables.
Provide clear, structured feedback to enhance AI output quality and consistency.
This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.
Navigate through a typical Salesforce Marketing Cloud workflow while recording your screen and narrating your step-by-step actions.
Explain the reasoning behind your campaign setup choices and identify any usability challenges or platform workarounds you encounter.
Complete the asynchronous screen recording session remotely from your own desktop.
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Create original graduate-to-PhD-level academic problems in your field
Write rigorous, step-by-step solutions with exact and verifiable answers
Review AI-generated responses to identify specific reasoning errors
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Apply deep subject-matter expertise to AI model evaluation and large language model projects.
Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
Collaborate with AI research teams to improve training datasets and evaluation methodologies.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.
Create concise 4–6-minute narrated screen-recording lessons demonstrating practical generative AI workflows for college-level learners.
Use tools such as ChatGPT, Claude, Gemini, or Copilot to demonstrate real-world professional applications.
Produce high-quality screen captures with clear voice narration and perform light video editing as needed.
This partner company focuses on creating accessible, practical educational content around real-world generative AI applications. The team works asynchronously with independent contractors, offering flexible project-based collaboration.
You will design the architecture for specialized subagents operating within live customer conversations, including technical QA, product expertise, and objection handling.
You will build routing and delegation systems that determine when to answer directly, invoke a subagent, or escalate to a human.
You will master the dialogue platform, train AI agents via prompting and fine-tuning, and document workflows to educate the team.
1mind builds autonomous customer experience software that deploys AI-powered 'Superhumans' to engage, demo, onboard, and support customers across the entire buying journey. The company offers a remote-first, fast-moving culture with ownership, autonomy, and impact from day one.
Write detailed, honest accounts of complex travel agent workflows for AI training.
Choose from real tasks like planning multi-destination itineraries or managing group bookings.
Work fully remote on a freelance basis with flexible hours and competitive pay.
Prolific is building the largest pool of quality human data in the world. Over 35,000 AI developers, researchers, and organizations use the platform to collect high-quality data from diverse participants.
Complete a survey about your generative AI daily habits and attitudes.
Become part of the AI insights community for additional earning opportunities.
Receive a $10 one-time payment for your participation.
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate model-generated content across multiple modalities including text, images, audio, and video.
Apply defined quality rubrics such as factuality, consistency, and aesthetics to assess outputs.
Conduct independent research on unfamiliar topics to make well-supported evaluation judgments.
Our client is a global technology company that helps businesses build, train, and manage AI systems. They offer flexible, remote work with variable workload and a focus on high-quality model evaluation.
Review AI-generated clinical responses for accuracy, safety, and reasoning quality.
Compare multiple model answers and select or justify the best response.
Write improved exemplars and structured feedback to enhance AI learning.
Prolific builds the world's largest pool of quality human data for AI development, used by over 35,000 researchers and organizations. They focus on ethically sourced human behavioral data to improve AI models, with a culture of innovation and global reach.
You will rate and assess the performance of AI models based on their output or behavior.
You will label elements of content and assign predefined categories to generate training data.
You will create prompts, summaries, and evaluate relevance to improve AI system understanding.
Innodata (Nasdaq: INOD) is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. The company has a 36+ year legacy of delivering high-quality data and outstanding outcomes for customers.
Evaluate and rank AI-generated scientific explanations based on accuracy and logic.
Review scientific papers alongside AI-generated abstracts to identify inaccuracies.
Verify AI-generated data against source documentation for material properties and formulas.
Prolific builds the largest pool of quality human data for AI training, connecting researchers with a global participant network. We serve over 35,000 AI developers and organizations, focusing on ethical data collection.
Participate in a paid research study on everyday Salesforce Sales Cloud workflows.
Record a brief screen capture video walking through a standard sales workflow and narrating your process.
Earn $70 per completed task for sharing your expertise.
Terac builds the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate financial documents and reports to verify accuracy and provide AI training data.
Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
Validate AI outputs against professional financial standards and provide expert feedback.
Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.
Evaluate AI systems at a scale only possible by combining thousands of vetted experts with model graders.
Innovate at the frontier of QA by shaping industry standards for validating agentic AI and large language models.
Collaborate with global market leaders to architect AI quality blueprints and drive high-impact consultative visibility.
Testlio provides a fully managed crowdsourced testing platform powered by proprietary intelligence technology, LeoCore. They are a female-founded, fully remote company with an inclusive culture, half of their team identifying as women, and are growing profitably.