Apply deep subject-matter expertise to AI model evaluation and large language model projects.
Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
Collaborate with AI research teams to improve training datasets and evaluation methodologies.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.
Evaluate financial documents and reports to verify accuracy and provide AI training data.
Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
Validate AI outputs against professional financial standards and provide expert feedback.
Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.
Evaluate AI-generated documents and presentations against quality standards.
Apply humanities expertise to identify inaccuracies and cultural issues.
Provide structured feedback to improve AI model performance.
A partner company is seeking a humanities evaluator to assess AI-generated content for accuracy and quality. The company emphasizes cultural awareness and critical thinking in a remote, asynchronous work environment.
Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
Identify factual, formatting, visual, and structural issues in professional deliverables.
Provide clear, structured feedback to enhance AI output quality and consistency.
This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.
Evaluate LLM responses for accuracy, clarity, and completeness.
Fact-check technical claims using authoritative references.
Validate code and outputs, and annotate model performance.
Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.
Review real user interaction traces with an AI shopping assistant
Identify logical failures, inaccuracies, or poor recommendations in the text
Create structured rubrics and verifiers to judge response quality
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate AI-generated content against domain-specific quality rubrics in humanities, arts, and culture.
Review documents, spreadsheets, and presentations for accuracy, relevance, clarity, and overall quality.
Provide structured feedback and collaborate with AI research teams to improve model outputs.
A partner company is seeking subject-matter experts to evaluate AI-generated content across humanities, arts, and culture. The company offers a flexible, remote contract environment, with no details on team size provided.
Create original graduate-to-PhD-level academic problems in your field
Write rigorous, step-by-step solutions with exact and verifiable answers
Review AI-generated responses to identify specific reasoning errors
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Rating and assessing the performance of AI models based on their output or behavior.
Labeling and categorizing content to train machine learning models.
Generating prompts, responses, and summaries to improve language model reasoning.
Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, the company focuses on enabling responsible AI advancement.
Write detailed outlines of your regular workflows, focusing on one critical task performed at least weekly.
Provide structured evaluation tasks and nuanced feedback to train AI models.
Complete paid tasks remotely on a freelance basis, with most tasks requiring one hour of uninterrupted work.
Prolific builds the world's largest pool of quality human data for AI development. With over 35,000 AI developers and organizations using its platform, it focuses on ethically sourced behavioral data from paid participants.
Create original, human-authored content in specialized domains like law, business, and philosophy for AI training.
Work independently from creative prompts to produce accurate, original written samples without AI tools.
Collaborate with a global community of writers and experts to improve language technology and AI evaluation.
Jobgether uses AI-powered matching to connect candidates with hiring companies, streamlining the application process. It operates as a platform that shares shortlisted candidates with employers, who manage final decisions and next steps.
Evaluate AI-generated legal research and analysis for accuracy, relevance, and completeness.
Verify legal citations, authorities, and reasoning to identify errors and weaknesses.
Develop objective evaluation criteria and provide structured feedback to improve AI legal content.
Jobgether uses an AI-powered matching process to connect top-fitting candidates with hiring companies. They prioritize objective and fair review, sharing shortlists directly with employers for final decisions.
Evaluate AI-generated documents, spreadsheets, and presentations for privacy and regulatory compliance accuracy.
Apply domain expertise to assess outputs against defined quality standards and provide actionable feedback.
Contribute to improving AI systems by providing expert judgment and structured feedback.
The company is a partner of Jobgether, offering remote opportunities for privacy and compliance professionals to evaluate AI-generated content. The company values autonomy and flexibility, providing independent contractor arrangements.
Evaluate AI-generated responses for logical consistency and business accuracy across various management scenarios.
Assess AI models' understanding of business strategy, operations, and organizational behavior.
Provide structured feedback to improve AI reasoning and decision-making in realistic business contexts.
Our partner is a technology company that develops advanced AI systems and seeks freelance professionals to train AI models in business contexts. The company offers flexible remote work and is looking for independent contractors with strong business knowledge.
Train and evaluate AI models on complex financial institutions, regulatory, and risk-management topics.
Assess model responses for factual accuracy, financial logic, and regulatory interpretation.
Produce clear error traces and structured feedback to improve AI performance.
This company specializes in AI training projects for financial institutions, offering flexible freelance opportunities. It operates with a small remote team and provides a fully remote work environment.
Conduct red-team evaluations to identify jailbreaks, prompt injections, and misuse scenarios in conversational AI models.
Develop creative adversarial prompts and scenarios to systematically probe model behavior and uncover weaknesses.
Generate high-quality human evaluation data by annotating failures and classifying vulnerabilities.
The partner company is a technology organization focused on AI safety and responsible AI development. They work with a remote, asynchronous team to improve the robustness of conversational AI systems.
Review AI-generated clinical responses for accuracy, safety, and reasoning quality.
Compare multiple model answers and select or justify the best response.
Write improved exemplars and structured feedback to enhance AI learning.
Prolific builds the world's largest pool of quality human data for AI development, used by over 35,000 researchers and organizations. They focus on ethically sourced human behavioral data to improve AI models, with a culture of innovation and global reach.
You will design the architecture for specialized subagents operating within live customer conversations, including technical QA, product expertise, and objection handling.
You will build routing and delegation systems that determine when to answer directly, invoke a subagent, or escalate to a human.
You will master the dialogue platform, train AI agents via prompting and fine-tuning, and document workflows to educate the team.
1mind builds autonomous customer experience software that deploys AI-powered 'Superhumans' to engage, demo, onboard, and support customers across the entire buying journey. The company offers a remote-first, fast-moving culture with ownership, autonomy, and impact from day one.
Design and apply evaluation criteria for consulting deliverables such as market analyses and financial models.
Assess AI-generated and human work, providing evidence-based scores and justifications.
Work independently in a remote, asynchronous environment to improve AI model reasoning.
They are an AI-focused organization improving the quality of AI outputs through expert evaluation. The work is fully remote and asynchronous, with an emphasis on independent problem-solving and collaboration with senior reviewers.
Source, vet, and onboard specialized annotators and performers across dozens of skill profiles and languages.
Scale sourcing channels using AI workflows, automation, and LinkedIn Recruiter.
Build relationships with experts and collaborate with delivery leads to meet customer needs.
HumanSignal is a human data partner for AI companies, creating and annotating real-world datasets and running the open-source platform Label Studio. They serve advanced ML and AI teams with data collection and multi-step workflows, though company size and culture details are not specified.