Review text or media samples based on provided project guidelines
Apply accurate labels and categorizations to diverse data sets
Evaluate AI-generated responses for clarity, safety, and factual accuracy
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate AI-generated content across humanities, arts, and culture domains for accuracy and quality.
Provide structured feedback and identify issues in AI outputs.
Work remotely on a flexible schedule contributing to AI system improvement.
The company specializes in developing and improving AI systems through expert human evaluation. They operate as a remote, flexible organization that values specialized domain knowledge and critical analysis.
Evaluate AI-generated slides, spreadsheets, and documents for real-world usability and professional quality.
Assess outputs for accuracy, clarity, relevance, and alignment with data science standards.
Provide structured written feedback to help improve AI systems and their outputs.
A partner company is seeking a Data Science Expert to evaluate AI-generated work. The company focuses on improving AI systems and operates with a flexible, remote team.
Evaluate model outputs in humanities fields for factual accuracy, logical coherence, and ideological bias.
Create exemplary responses and datasets emphasizing intellectual honesty and thorough source evaluation.
Collaborate with engineering teams to design evaluation tasks and define desired model behavior.
SpaceXAI creates AI systems to understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence with a flat organizational structure.
Own the build out of new agents, skills, and platform capability for teams across TLDR.
Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.
TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.
Evaluate AI-generated design outputs across various formats, including visual assets, layouts, websites, and games.
Assess designs based on clarity, functionality, accessibility, and effectiveness, not just visual appeal.
Write detailed annotations and critiques to distinguish between mediocre, good, and exceptional design.
Our partner is a company developing advanced AI systems. They seek design professionals to evaluate and improve AI-generated design outputs, and offer remote work opportunities with flexible schedules.
Evaluate AI-generated work products in real estate, hospitality, and events using quality rubrics.
Identify factual, aesthetic, and presentation errors and provide actionable feedback.
Apply industry expertise to distinguish realistic, commercially sound work from generic AI content.
The company develops AI systems and evaluates their outputs for quality. They seek experienced industry professionals for flexible remote contract work.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.
Label, annotate, and evaluate German-language content including photos, graphics, and videos for linguistic and cultural accuracy.
Evaluate AI-generated content against Canva's quality bar for German users to shape language experiences.
Build and contribute to German-specific datasets to support the internationalization of Canva AI features.
Canva is a design platform redefining how the world experiences design. It is a global company with a large user base, known for its innovative culture and focus on AI-powered features.
Listen to recorded conversations between users and AI voice agents to evaluate response quality.
Rate each agent turn on two scales: content helpfulness and prosody naturalness.
Write short, specific justifications for each score given.
Welo Data is an AI services company that provides evaluation and data services for AI voice agents. They are a global organization with a freelance workforce, focusing on quality assessment of AI interactions.
Evaluate AI-generated Azerbaijani text for naturalness and authenticity.
Compare text snippets and provide quality control on cultural nuance.
Rate AI-generated text and tag data on tone and naturalness.
Prolific is building the largest pool of quality human data in the world, with over 35,000 AI developers and researchers using its platform. They connect researchers with paid participants to collect ethically sourced human behavioral data and feedback.
Evaluate AI model performance through real-time, voice-based conversations by roleplaying assigned scenarios with two different models.
Compare model responses across five defined dimensions, identify error clusters, and select the stronger performer with a detailed rationale.
Maintain consistent conversational turns and voice recording to ensure fair, objective comparisons.
Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, they deliver high-quality data and outcomes for AI builders and adopters.
Evaluate AI-generated Icelandic text for naturalness and authenticity.
Compare side-by-side text snippets to assess quality.
Provide feedback on tone, register, and cultural context.
Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. They focus on ethically sourced human behavioral data to improve AI systems.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply careful judgment to ensure high-quality results aligned with task objectives.
LILT provides multilingual AI and human-verified services to enterprises, governments, and AI developers. The company has a global community of linguists and language professionals committed to innovation and excellence.
Evaluate and edit AI-generated translations between German and English for accuracy and fluency.
Translate content between German and English while maintaining cultural relevance and context.
Annotate translation errors and provide feedback to improve AI model performance.
This partner company is seeking a German translator for an AI training project. It offers a flexible, remote contract opportunity for language professionals to improve AI translation quality.
Evaluate search results and AI-generated content for quality, relevance, accuracy, and usefulness.
Conduct online research to verify information and support rating decisions.
Apply rating guidelines consistently and participate in training and calibration sessions.
TELUS Digital AI & Data Solutions partners with a diverse and vibrant community to help our customers enhance their AI and machine learning models. Our global AI community includes over 1 million contributors across 500+ languages and dialects, offering flexible remote and onsite opportunities.
Evaluate AI-generated coding interactions end to end for correctness and engineering judgment.
Assess whether outputs reflect strong engineering taste and provide clear, opinionated feedback.
Help define what great looks like for AI coding tools like Codex, Claude Code, and Cursor.
G2i Inc. is a technology staffing company that connects software engineers with remote contract opportunities. The company values engineering excellence and provides flexible, ongoing projects for senior-level developers.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.