Write detailed, honest accounts of complex travel agent workflows for AI training.
Choose from real tasks like planning multi-destination itineraries or managing group bookings.
Work fully remote on a freelance basis with flexible hours and competitive pay.
Prolific is building the largest pool of quality human data in the world. Over 35,000 AI developers, researchers, and organizations use the platform to collect high-quality data from diverse participants.
Evaluate AI-generated content across humanities, arts, and culture domains for accuracy and quality.
Provide structured feedback and identify issues in AI outputs.
Work remotely on a flexible schedule contributing to AI system improvement.
The company specializes in developing and improving AI systems through expert human evaluation. They operate as a remote, flexible organization that values specialized domain knowledge and critical analysis.
Evaluate AI-generated documents, spreadsheets, and presentations against domain-specific quality standards.
Assess outputs for factual accuracy, procurement relevance, completeness, clarity, and practical applicability.
Provide structured written feedback identifying issues and opportunities for improvement.
This partner company specializes in evaluating AI-generated work products for public-sector procurement and RFI responses. The company size and culture are not specified in the posting, but the engagement is flexible and fully remote.
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Evaluate AI-generated market research and competitive intelligence artifacts for accuracy, rigor, and quality.
Apply structured rubrics to assess deliverables and identify factual inaccuracies and analytical gaps.
Provide clear, actionable written feedback to support evaluation decisions and improve AI training.
The partner company specializes in evaluating AI-generated market research and competitive intelligence content. They hire independent contractors for flexible remote engagements.
Evaluate AI-generated documents, spreadsheets, and presentations for privacy and regulatory compliance accuracy.
Apply domain expertise to assess outputs against defined quality standards and provide actionable feedback.
Contribute to improving AI systems by providing expert judgment and structured feedback.
The company is a partner of Jobgether, offering remote opportunities for privacy and compliance professionals to evaluate AI-generated content. The company values autonomy and flexibility, providing independent contractor arrangements.
Write detailed outlines of your regular workflows, focusing on one critical task performed at least weekly.
Provide structured evaluation tasks and nuanced feedback to train AI models.
Complete paid tasks remotely on a freelance basis, with most tasks requiring one hour of uninterrupted work.
Prolific builds the world's largest pool of quality human data for AI development. With over 35,000 AI developers and organizations using its platform, it focuses on ethically sourced behavioral data from paid participants.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.
Review, evaluate, and validate AI-generated content and data according to project guidelines.
Compare AI outputs and identify the most accurate, relevant, or high-quality results.
Annotate, categorize, or label text, audio, images, video, or other data with clear feedback.
Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train advanced AI systems. They offer limitless flexibility and growth for a diverse community in 100+ countries.
Support participants in virtual HR workshops by diagnosing and improving generative AI outputs.
Provide tool-agnostic guidance to help attendees refine prompts and achieve practical results.
Ensure responsible AI use and escalate issues as needed while fostering productive learning.
This company is a leading technology firm specializing in internet-related services and products, including search, cloud computing, and AI. It is a large global organization with a culture of innovation and collaboration.
Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.
Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.
Evaluate AI-generated responses for accuracy, grammar, and cultural relevance.
Create natural prompts and responses in Filipino to improve conversational datasets.
Collaborate with global teams to help improve AI language models.
Welo Data, part of Welocalize, is a global AI data company that delivers high-quality, ethical data to train advanced AI systems. They have a community of over 500,000 contributors across 100+ countries, focusing on building smarter, more human AI with limitless opportunities for growth.
Review real user interactions with an AI shopping assistant and identify flaws in accuracy and usefulness.
Analyze response quality from an e-commerce perspective, considering product recommendations and user needs.
Create structured rubrics and verifiers for consistent evaluation of future AI responses.
The partner company is developing an AI-powered digital shopping assistant and seeks evaluators to assess and improve its responses. The team size and culture are not specified.
Review AI-generated financial article drafts for accuracy, relevance, and clarity.
Fact-check financial data and market developments using reliable research methods.
Write concise, investor-focused commentary of 150–300 words explaining the significance of news events.
They are a company focused on financial journalism and AI-assisted content production. They foster a collaborative, fast-paced environment with an emphasis on accuracy and analytical thinking.
Review real user interaction traces with an AI shopping assistant
Identify logical failures, inaccuracies, or poor recommendations in the text
Create structured rubrics and verifiers to judge response quality
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Apply expertise in incident management and SRE to evaluate AI-generated documents, spreadsheets, and slide decks for technical accuracy and operational rigor.
Assess outputs against real-world reliability practices, identifying factual, technical, and reasoning errors.
Provide clear, structured written feedback and collaborate asynchronously with a research team to refine evaluation approaches.
This partner company focuses on AI evaluation and development, seeking experienced professionals to assess AI-generated work products. They offer flexible remote work and independent contractor engagements with weekly payments.
Evaluate model outputs in humanities fields for factual accuracy, logical coherence, and ideological bias.
Create exemplary responses and datasets emphasizing intellectual honesty and thorough source evaluation.
Collaborate with engineering teams to design evaluation tasks and define desired model behavior.
SpaceXAI creates AI systems to understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence with a flat organizational structure.
Assess technical accuracy, realism, and reproducibility of AI-assisted developer workflow tasks.
Provide actionable feedback on IDE integration faults, AI-assisted coding inefficiencies, and logic errors.
Apply deep knowledge of modern developer tooling, including AI coding assistants and telemetry/trace analysis.
We are sourcing experienced technical specialists to audit tasks used in training and evaluating AI systems. We focus on ensuring technical rigor and accuracy in developer workflow tasks, operating as a remote freelance project.