Source Job

$35–$75/hr
Global

  • Evaluate model outputs in humanities fields for factual accuracy, logical coherence, and ideological bias.
  • Create exemplary responses and datasets emphasizing intellectual honesty and thorough source evaluation.
  • Collaborate with engineering teams to design evaluation tasks and define desired model behavior.

Linguistics Analytical Writing AI Evaluation

20 jobs similar to AI Tutor - Humanities

Jobs ranked by similarity.

Canada

  • Evaluate AI-generated documents and presentations against quality standards.
  • Apply humanities expertise to identify inaccuracies and cultural issues.
  • Provide structured feedback to improve AI model performance.

A partner company is seeking a humanities evaluator to assess AI-generated content for accuracy and quality. The company emphasizes cultural awareness and critical thinking in a remote, asynchronous work environment.

India

  • Evaluate AI-generated content against domain-specific quality rubrics in humanities, arts, and culture.
  • Review documents, spreadsheets, and presentations for accuracy, relevance, clarity, and overall quality.
  • Provide structured feedback and collaborate with AI research teams to improve model outputs.

A partner company is seeking subject-matter experts to evaluate AI-generated content across humanities, arts, and culture. The company offers a flexible, remote contract environment, with no details on team size provided.

US

  • Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
  • Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
  • Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.

The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.

Canada

  • Apply deep subject-matter expertise to AI model evaluation and large language model projects.
  • Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
  • Collaborate with AI research teams to improve training datasets and evaluation methodologies.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.

Global

  • Evaluate LLM responses for accuracy, clarity, and completeness.
  • Fact-check technical claims using authoritative references.
  • Validate code and outputs, and annotate model performance.

Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.

US

  • Evaluate AI-generated content related to job descriptions, compensation benchmarking, and employee policy interpretation.
  • Create realistic HR workflow scenarios such as performance management cycles, benefits enrollment, and employee investigations.
  • Annotate and label data across offer letter generation, leave administration, and workforce analytics.

Prolific builds the largest pool of quality human data in the world, used by over 35,000 AI developers, researchers, and organizations. The platform connects researchers with paid participants to gather ethically sourced behavioral data, placing itself at the forefront of AI innovation.

US

  • Review AI-generated financial article drafts for accuracy, relevance, and clarity.
  • Fact-check financial data and market developments using reliable research methods.
  • Write concise, investor-focused commentary of 150–300 words explaining the significance of news events.

They are a company focused on financial journalism and AI-assisted content production. They foster a collaborative, fast-paced environment with an emphasis on accuracy and analytical thinking.

Global

  • Evaluate and refine design outputs across formats to train AI models.
  • Write precise annotations critiquing design quality from mediocre to excellent.
  • Collaborate with engineering teams to set high standards for design quality.

SpaceXAI creates AI systems to understand the universe and aid humanity. Our team is small, highly motivated, and focused on engineering excellence with a flat structure.

US

  • Review scientific papers alongside LLM-generated graphical abstracts.
  • Fact-check AI outputs for scientific accuracy and integrity.
  • Verify technical concepts using your neuroscience expertise.

Prolific is building the biggest pool of quality human data in the world, connecting AI developers, researchers, and organizations with paid study participants. With over 35,000 AI developers and researchers using the platform, it enables flexible, ethical data collection for AI training.

  • Write detailed outlines of your regular workflows, focusing on one critical task performed at least weekly.
  • Provide structured evaluation tasks and nuanced feedback to train AI models.
  • Complete paid tasks remotely on a freelance basis, with most tasks requiring one hour of uninterrupted work.

Prolific builds the world's largest pool of quality human data for AI development. With over 35,000 AI developers and organizations using its platform, it focuses on ethically sourced behavioral data from paid participants.

Global

  • Contribute to AI model training and evaluation in your area of expertise, including writing, reviewing, and assessing responses.
  • Evaluate AI-generated outputs for accuracy, logic, and nuance, and provide actionable feedback for model improvement.
  • Apply PhD-level judgment to real-world tasks in social sciences, humanities, arts, or linguistics.

Welo Data, part of Welocalize, is a global AI data company that delivers high-quality, ethical data to train advanced AI systems. With a community of over 500,000 contributors in 100+ countries, they emphasize flexibility, growth, and support for their contributors.

Canada

  • Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
  • Identify factual, formatting, visual, and structural issues in professional deliverables.
  • Provide clear, structured feedback to enhance AI output quality and consistency.

This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.

US

  • Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
  • Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
  • Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.

Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.

Canada

  • Evaluate AI-generated media, journalism, and communications content against quality rubrics.
  • Review outputs for factual accuracy, relevance, clarity, tone, and structure.
  • Provide structured, actionable feedback to improve AI model performance.

Jobgether is an AI-powered job matching platform that connects professionals with remote opportunities. It uses technology to streamline recruitment and provide a flexible, remote-first work environment.

US

  • Evaluate AI-generated legal research and analysis for accuracy, relevance, and completeness.
  • Verify legal citations, authorities, and reasoning to identify errors and weaknesses.
  • Develop objective evaluation criteria and provide structured feedback to improve AI legal content.

Jobgether uses an AI-powered matching process to connect top-fitting candidates with hiring companies. They prioritize objective and fair review, sharing shortlists directly with employers for final decisions.

UK

  • Review AI-generated responses to psychological scenarios and behavioral prompts.
  • Complete AI training tasks such as analyzing, editing, and writing psychology-related content.
  • Compare model answers and select or justify the best response to improve AI understanding.

Prolific is building the world's largest pool of quality human data for AI training. They connect over 35,000 researchers and developers with paid participants, offering a flexible, remote platform for ethical data collection.

US

  • Review, write, and publish hundreds of AI-assisted articles per month, fact-checking and adding Foolish context.
  • Edit AI-generated content for accuracy, clarity, and substance in the content management system.
  • Provide feedback to prompt engineers to improve AI outputs based on your editing observations.

The Motley Fool is a purpose-driven financial services company on a mission to make the world smarter, happier, and richer. We are a fast-moving, collaborative team that values high-quality work, curiosity, and initiative.

$9–$20/hr
US

  • Write detailed, honest accounts of complex travel agent workflows for AI training.
  • Choose from real tasks like planning multi-destination itineraries or managing group bookings.
  • Work fully remote on a freelance basis with flexible hours and competitive pay.

Prolific is building the largest pool of quality human data in the world. Over 35,000 AI developers, researchers, and organizations use the platform to collect high-quality data from diverse participants.

Canada

  • Conduct red-team evaluations to identify jailbreaks, prompt injections, and misuse scenarios in conversational AI models.
  • Develop creative adversarial prompts and scenarios to systematically probe model behavior and uncover weaknesses.
  • Generate high-quality human evaluation data by annotating failures and classifying vulnerabilities.

The partner company is a technology organization focused on AI safety and responsible AI development. They work with a remote, asynchronous team to improve the robustness of conversational AI systems.

Global

  • Complete AI training tasks such as analyzing, editing, and writing in Korean.
  • Evaluate and judge AI performance on Korean prompts.
  • Help improve cutting-edge AI models with your expertise.

Prolific builds the largest pool of quality human data for AI development. Over 35,000 developers and researchers use Prolific to gather diverse data from paid participants.