Evaluate AI-generated responses for relevance, accuracy, and personalization using personalized prompts and data from connected Google applications.
Identify subtle issues such as incorrect assumptions, irrelevant recommendations, inconsistencies, and inappropriate personalization.
Provide clear, detailed, and structured feedback to support improvements to AI models and personalization systems.
Our partner company is seeking an AI Response Quality Evaluator to improve AI-generated responses. This is a project-based contract role with a remote, independent working environment and a duration of up to 16 weeks.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply careful judgment to ensure high-quality results aligned with task objectives.
LILT provides multilingual AI and human-verified services to enterprises, governments, and AI developers. The company has a global community of linguists and language professionals committed to innovation and excellence.
Edit and rewrite AI-generated content for client websites and third-party publications.
Verify facts, research unfamiliar topics, and flag unsupported claims or weak sourcing.
Optimize content for search intent, SEO, GEO, and readability across brands and industries.
Interdependence is rebuilding the PR industry with a proprietary platform that analyzes 300,000+ stories daily across a network of 250,000 journalists. The 100+ person team, named one of America's Best PR Agencies, works with brands and leaders across consumer, healthcare, tech, B2B, travel and entertainment.
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.
Label, annotate, and evaluate German-language content including photos, graphics, and videos for linguistic and cultural accuracy.
Evaluate AI-generated content against Canva's quality bar for German users to shape language experiences.
Build and contribute to German-specific datasets to support the internationalization of Canva AI features.
Canva is a design platform redefining how the world experiences design. It is a global company with a large user base, known for its innovative culture and focus on AI-powered features.
Evaluate AI-generated Macedonian text for naturalness and cultural authenticity.
Compare text side-by-side to assess quality and nuance.
Provide feedback on tone, register, and word choice to improve AI language models.
Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. We connect a global community of participants with researchers to ethically source human behavior and feedback for AI development.
Edit and enhance AI-generated content for client websites, blogs, and marketing campaigns.
Research and fact-check claims to ensure accurate, credible, publication-ready material.
Optimize content for SEO, GEO, and AI-driven search while collaborating with teams.
Jobgether is an AI-powered platform that connects candidates with hiring companies. This position is with a partner communications and marketing firm, known for a collaborative, entrepreneurial culture and career development opportunities.
Evaluate AI-generated content across humanities, arts, and culture domains for accuracy and quality.
Provide structured feedback and identify issues in AI outputs.
Work remotely on a flexible schedule contributing to AI system improvement.
The company specializes in developing and improving AI systems through expert human evaluation. They operate as a remote, flexible organization that values specialized domain knowledge and critical analysis.
Evaluate and label AI model outputs to improve performance and alignment with project guidelines.
Create prompts, rewrite text, and generate training data for large language models.
Work on flexible, remote, project-based tasks while helping shape the future of AI.
Innodata is a global data engineering company that enables the responsible advancement of AI by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company delivers high-quality data and outstanding outcomes for customers.
Write original, human-authored text in your domain of expertise following provided prompts.
Ensure all content is entirely original with no use of AI or translation tools.
Attest to the originality of your submission and adhere to confidentiality terms.
LILT provides multilingual AI and human-verified services to enterprises, governments, and AI developers worldwide. The company fosters a global community of linguists and subject matter experts, emphasizing innovation, flexibility, and collaboration.
Evaluate AI-generated slides, spreadsheets, and documents for real-world usability and professional quality.
Assess outputs for accuracy, clarity, relevance, and alignment with data science standards.
Provide structured written feedback to help improve AI systems and their outputs.
A partner company is seeking a Data Science Expert to evaluate AI-generated work. The company focuses on improving AI systems and operates with a flexible, remote team.
Review and approve or reject user-generated content based on platform standards, documenting moderation decisions.
Conduct safety and compliance reviews on in-house content and workflows, including datasets and AI characters.
Moderate video, image, and text content, taking action on flagged violations and categorizing content for appropriate visibility.
EverAI is building the world's largest AI companionship platform, driven by a proprietary moderation system called EverGuard. With 50 million users in two years and a fully remote team of about 100, the company is a fast-growing, category-creating AI company.
Evaluate AI-generated Icelandic text for naturalness and authenticity.
Compare side-by-side text snippets to assess quality.
Provide feedback on tone, register, and cultural context.
Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. They focus on ethically sourced human behavioral data to improve AI systems.
Evaluate search results and AI-generated content for quality, relevance, accuracy, and usefulness.
Conduct online research to verify information and support rating decisions.
Apply rating guidelines consistently and participate in training and calibration sessions.
TELUS Digital AI & Data Solutions partners with a diverse and vibrant community to help our customers enhance their AI and machine learning models. Our global AI community includes over 1 million contributors across 500+ languages and dialects, offering flexible remote and onsite opportunities.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Evaluate AI-generated Azerbaijani text for naturalness and authenticity.
Compare text snippets and provide quality control on cultural nuance.
Rate AI-generated text and tag data on tone and naturalness.
Prolific is building the largest pool of quality human data in the world, with over 35,000 AI developers and researchers using its platform. They connect researchers with paid participants to collect ethically sourced human behavioral data and feedback.
Review text or media samples based on provided project guidelines
Apply accurate labels and categorizations to diverse data sets
Evaluate AI-generated responses for clarity, safety, and factual accuracy
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate model outputs in humanities fields for factual accuracy, logical coherence, and ideological bias.
Create exemplary responses and datasets emphasizing intellectual honesty and thorough source evaluation.
Collaborate with engineering teams to design evaluation tasks and define desired model behavior.
SpaceXAI creates AI systems to understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence with a flat organizational structure.
Write and edit content for websites, blogs, social media, email campaigns, and marketing materials.
Create engaging copy for product features, landing pages, campaigns, and promotional materials.
Research topics related to careers, job searching, recruitment, technology, and AI.
CareerSwift is a technology company focused on making the job search process simpler, faster, and more effective. The company is growing and offers a collaborative and supportive team environment.