You will rate and assess the performance of AI models based on their output or behavior.
You will label elements of content and assign predefined categories to generate training data.
You will create prompts, summaries, and evaluate relevance to improve AI system understanding.
Innodata (Nasdaq: INOD) is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. The company has a 36+ year legacy of delivering high-quality data and outstanding outcomes for customers.
Evaluate and assess AI model outputs based on predefined quality, accuracy, relevance, and behavioral guidelines.
Annotate, classify, and label text, images, or audio to support AI model training.
Create prompts and generate high-quality responses to improve language model reasoning capabilities.
Jobgether uses an AI-powered matching process to connect candidates with hiring companies. They focus on efficient, fair recruitment and handle data privacy in compliance with GDPR.
Review, evaluate, and annotate AI-generated content across text, images, audio, and video.
Perform quality checks to ensure accuracy, consistency, and compliance with project guidelines.
Identify edge cases and inconsistencies, contribute to high-quality dataset development, and participate in calibration activities.
Welo Data, part of Welocalize, is a global AI data company with over 500,000 contributors that provides high-quality, ethical data for training advanced AI systems. The company supports a diverse, global community across 100+ countries and offers project-based freelance opportunities with flexibility and growth potential.
Compare and rank AI-generated responses for accuracy, logic, and safety.
Review CS research papers alongside AI summaries to ensure scientific integrity.
Fact-check technical data and code for logical flaws and inaccuracies.
Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers and researchers. They connect researchers with paid participants to gather high-quality, ethically sourced behavioral data for AI development.
Evaluate prompts and AI-generated outputs for accuracy, cultural appropriateness, and brand alignment.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply local cultural insight and consistent evaluation guidelines to ensure high-quality AI training.
Lilt provides multilingual AI and human-verified services to enterprises, governments, and AI developers. They foster a global community of linguists and subject matter experts working on cutting-edge AI and language technology.
Evaluate model-generated content across multiple modalities including text, images, audio, and video.
Apply defined quality rubrics such as factuality, consistency, and aesthetics to assess outputs.
Conduct independent research on unfamiliar topics to make well-supported evaluation judgments.
Our client is a global technology company that helps businesses build, train, and manage AI systems. They offer flexible, remote work with variable workload and a focus on high-quality model evaluation.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply local insight into tone, symbolism, visual cues, and market fit to deliver culturally relevant content.
LILT is an AI company that makes the world's information available to everyone, no matter the language they speak. They work with a global community of linguists and subject matter experts to deliver multilingual AI and human-verified services to Enterprises, Governments, and AI Developers.
Evaluate AI-generated French responses, rate them, and flag cultural issues.
Rewrite weak responses into clear, natural Canadian French.
Create original French prompts and example responses to expand training data.
We are a global AI data company that delivers high-quality, ethical data to train the world's most advanced AI systems. With over 500,000 contributors, we offer flexible, remote project-based opportunities with a supportive global community.
Label and evaluate photos, graphics, videos, stickers, and designs in Dutch for linguistic accuracy and cultural appropriateness.
Assess AI-generated content against Canva's quality bar for Dutch users to shape localised AI experiences.
Build and contribute to Dutch-specific datasets and deliver labelled assets on time across varied task types.
Canva is redefining how the world experiences design, empowering users to create visual content. The company has a global team and supports flexible, remote-friendly work, with a focus on collaboration and innovation.
Write prompts in Luxembourgish to test AI models and evaluate responses using a structured rating guide.
Identify areas for improvement in AI responses and provide accurate English translations.
Work remotely on a flexible schedule with competitive pay.
CrowdGen by Appen is a platform that connects freelancers to help improve AI models through data annotation and evaluation. It is part of a large global company with a diverse community of independent contractors working remotely.
Review search results and evaluate their relevance to user queries
Answer true/false questions about content quality
Rate search results based on guidelines to improve AI systems
Welo Data provides AI services and data validation to improve search engine and AI systems. They are a remote-first company with a focus on quality and support for their contractors.
Contribute to building smarter, more accurate AI by annotating, evaluating, and creating prompts for language models.
Work flexibly on your own terms with remote projects that fit your schedule and skills.
Be part of a global community of linguists and tech enthusiasts shaping the future of AI.
We are a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train the world's most advanced AI systems. We are building a diverse community in 100+ countries, offering flexible remote work and limitless opportunities for growth.
Contribute to AI training through annotation, evaluation, and prompt creation tasks.
Work flexibly on remote projects that match your skills and availability.
Be part of a global community shaping safer, smarter AI.
Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train the world’s most advanced AI systems. We build smarter AI through the power of human contribution, offering limitless opportunities for our global community to grow, contribute, and work on their terms.
Review AI-generated responses against source images and quality guidelines.
Identify issues like hallucinations, missing details, or policy violations.
Provide structured feedback to improve model performance and output quality.
Jobgether uses AI-powered matching to connect candidates with partner companies. They focus on efficient, objective hiring processes and operate as a platform for remote opportunities.
Evaluate AI responses using a structured rating guide.
Translate prompts and evaluations into English.
CrowdGen by Appen provides AI training data services to improve machine learning models. They operate a global community of independent contractors and emphasize flexible, remote work.
Design, build, and maintain automated AI evaluation pipelines for production LLM applications.
Develop prompt engineering strategies and evaluate model performance using quantitative methods.
Analyze production AI behavior with Python, SQL, and statistical techniques to identify improvement opportunities.
GovWorx provides an AI-powered platform, CommsCoach, that supports 9-1-1 and emergency communications centers by automating quality assurance, training, and real-time call evaluation. The company is a growing technology team focused on public safety, collaborating across AI, engineering, product, and data science.
Review domain archives to understand subject matter and extract key facts.
Create accurate question-and-answer pairs covering various complexity levels.
Ensure answers are traceable, unambiguous, and consistent with approved source content.
Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With over 36 years of experience, the company delivers high-quality data and outcomes for Generative AI builders.
Develop and execute a content strategy that tells Surge's story across web, blog, and campaigns
Turn complex technical ideas into sharp, memorable narratives that resonate
Collaborate with leadership to shape how Surge shows up across the AI landscape
Surge AI is a platform that powers the most advanced AI models by combining elite human expertise with cutting-edge tools for scalable oversight. The company was founded by engineers and researchers, has been profitable from day one without venture funding, and partners with leading AI companies like OpenAI and Anthropic.
Evaluate simulated advertiser-AI conversations for technical accuracy and campaign structure.
Fact-check platform strategies against correct hierarchy and full-funnel metrics.
Write clear, actionable feedback to correct errors and improve AI model performance.
RWS specializes in AI training data and language services, providing data annotation and evaluation solutions to improve AI model reliability. The company fosters a culture of diversity and inclusion, operating as a global employer with a focus on equal opportunity.
Evaluate financial documents and reports to verify accuracy and provide AI training data.
Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
Validate AI outputs against professional financial standards and provide expert feedback.
Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.