Review AI-generated responses against source images and quality guidelines.
Identify issues like hallucinations, missing details, or policy violations.
Provide structured feedback to improve model performance and output quality.
Jobgether uses AI-powered matching to connect candidates with partner companies. They focus on efficient, objective hiring processes and operate as a platform for remote opportunities.
Evaluate and assess AI model outputs based on predefined quality, accuracy, relevance, and behavioral guidelines.
Annotate, classify, and label text, images, or audio to support AI model training.
Create prompts and generate high-quality responses to improve language model reasoning capabilities.
Jobgether uses an AI-powered matching process to connect candidates with hiring companies. They focus on efficient, fair recruitment and handle data privacy in compliance with GDPR.
Coordinate assigned project workstreams to ensure on-time delivery against execution standards.
Partner with client teams to validate quote assumptions, analyze datasets for insights, and provide data-backed findings.
Maintain version control and translate QA findings into targeted improvements to reduce recurring annotation errors.
Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, large language models, computer vision, and speech recognition. With a global crowd of over 1 million contributors in more than 200 countries, the company fosters a culture of innovation, collaboration, and excellence.
Provide native-level Canadian French language vetting and QA for AI data projects.
Annotate and review AI outputs for grammatical accuracy, cultural context, and naturalness.
Develop educational resources and feedback documentation to improve AI alignment.
We are an AI training company that focuses on language alignment and data annotation for AI systems. Our remote team values linguistic precision and cultural nuance.
Evaluate simulated advertiser-AI conversations for technical accuracy and campaign structure.
Fact-check platform strategies against correct hierarchy and full-funnel metrics.
Write clear, actionable feedback to correct errors and improve AI model performance.
RWS specializes in AI training data and language services, providing data annotation and evaluation solutions to improve AI model reliability. The company fosters a culture of diversity and inclusion, operating as a global employer with a focus on equal opportunity.
Contribute to AI training by ranking responses or providing creative prompts.
Participate in behavioral experiments, user research, and academic studies.
Complete surveys, interviews, and feedback sessions on various topics.
Prolific builds the world's largest pool of quality human data, serving over 35,000 AI developers and researchers. A fast-growing company, it fosters a culture of innovation and diversity, connecting participants with academic and applied research studies.
Evaluate prompts and AI-generated outputs for accuracy, cultural appropriateness, and brand alignment.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply local cultural insight and consistent evaluation guidelines to ensure high-quality AI training.
Lilt provides multilingual AI and human-verified services to enterprises, governments, and AI developers. They foster a global community of linguists and subject matter experts working on cutting-edge AI and language technology.
Design and deliver high-quality AI evaluation data initiatives, from proposals through pilot execution and production readiness.
Recruit and manage subject-matter experts across technical domains, ensuring rigorous quality control frameworks.
Act as key interface with AI lab partners, converting pilots into scaled production engagements.
Jobgether uses AI-powered matching to connect candidates with roles quickly and fairly. They are a remote-first company that shares top-fitting candidates with hiring partners.
Own the design and defense of frontier model evaluations across reasoning, coding, agents, tool use, and multi-modal.
Build benchmark packages with expert-verified ground truth, multi-model headroom results, and rigorous QC.
Recruit, calibrate, and review a pool of subject-matter experts in coding, agentic/tool-use, and STEM/reasoning.
Anyone AI measures frontier model capability through expert-verified evaluation packages. The company operates as a remote team with a focus on rigorous benchmarking and lab collaboration.
Perform data collection, evaluation, and annotation for AI training.
Conduct pairwise comparisons and counting tasks.
Tag and label objects across audio, video, images, or collected data.
RWS provides AI training data services. They are a global company with a focus on diversity, equity, and inclusion, offering flexible remote work opportunities.
Evaluate search results and AI-generated content for quality, relevance, accuracy, and usefulness.
Apply rating guidelines consistently while maintaining high quality standards and meeting productivity expectations.
Participate in training, calibration sessions, and ongoing quality reviews to improve AI model performance.
TELUS Digital enriches data for better AI via human intelligence. They empower generative AI, computer vision, and NLP models with a skilled team and a managed AI community of over one million contributors.
Evaluate AI-generated text and voice snippets in Punjabi for naturalness and authenticity.
Assess audio clips for cultural and tonal accuracy of AI speech.
Provide feedback on linguistic nuance and quality of AI outputs.
Prolific builds the biggest pool of quality human data in the world, serving over 35,000 AI developers and researchers. The company connects researchers with paid study participants from diverse backgrounds to gather high-quality, ethically sourced behavioral data.
Research data collection strategies and design high-impact data slices that uncover model failure modes.
Model annotator behavior and design experiments to optimize instruction clarity and reward signal reliability.
Develop metrics and frameworks for evaluating dataset quality, diversity, and impact on downstream model alignment.
Surge AI builds a platform that powers the most powerful AI models in partnership with companies like Anthropic, Google, Microsoft, and Meta. They are a profitable, bootstrapped company focused on human intelligence and data quality.
Evaluate AI quality across the advisor stack, including pre-call briefs, in-call guidance, and post-call outputs.
Iterate inside ORA by refining prompts, updating knowledge base entries, and tweaking skills to close the loop on issues.
Surface trends and drive continuous improvement by tagging conversations, logging issues, and recommending prioritized improvements.
HighLevel is an AI-powered business operating system that gives agencies, entrepreneurs and SMBs the infrastructure to build, automate and scale. With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership.
Annotate and review multimedia data (video, images, metadata) using defined labeling rules and guidelines.
Perform self-QA, track recurring issues, and contribute to guideline improvements with clear documentation.
Collaborate with stakeholders to meet throughput and quality targets while participating in calibration sessions.
Welo Data provides AI services, specializing in data annotation and quality review for multimedia content. As a growing remote team, we offer freelance opportunities for detail-oriented professionals to enhance global AI systems.
Collect, evaluate, and annotate diverse data to improve AI-generated content in Italian.
Perform pairwise comparisons, counting tasks, and object tagging across audio, video, images, and text.
Work remotely on a flexible, part-time schedule with a long-term contract.
RWS provides technology-enabled language, content management, and intellectual property services. It is a large global company that values diversity and equal opportunity, offering flexible remote work.
Review scientific papers and LLM-generated graphical abstracts for accuracy.
Fact-check and identify inaccuracies in AI-generated summaries.
Use domain expertise to verify that technical concepts are correctly represented.
Prolific is building the largest pool of quality human data for AI development, serving over 35,000 developers. They focus on ethical data collection and aim to integrate human perspectives into AI.