Source Job

Canada

  • Evaluate AI-generated documents and presentations against quality standards.
  • Apply humanities expertise to identify inaccuracies and cultural issues.
  • Provide structured feedback to improve AI model performance.

Critical Thinking Analytical Skills Written Communication Cultural Awareness Microsoft Office

20 jobs similar to Humanities / Arts / Culture Evaluator

Jobs ranked by similarity.

India

  • Evaluate AI-generated content against domain-specific quality rubrics in humanities, arts, and culture.
  • Review documents, spreadsheets, and presentations for accuracy, relevance, clarity, and overall quality.
  • Provide structured feedback and collaborate with AI research teams to improve model outputs.

A partner company is seeking subject-matter experts to evaluate AI-generated content across humanities, arts, and culture. The company offers a flexible, remote contract environment, with no details on team size provided.

United States

  • Review search results and evaluate their relevance to user queries
  • Answer true/false questions about content quality
  • Rate search results based on guidelines to improve AI systems

Welo Data provides AI services and data validation to improve search engine and AI systems. They are a remote-first company with a focus on quality and support for their contractors.

Canada

  • Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
  • Identify factual, formatting, visual, and structural issues in professional deliverables.
  • Provide clear, structured feedback to enhance AI output quality and consistency.

This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.

UK

  • Evaluate AI-generated French responses, rate them, and flag cultural issues.
  • Rewrite weak responses into clear, natural Canadian French.
  • Create original French prompts and example responses to expand training data.

We are a global AI data company that delivers high-quality, ethical data to train the world's most advanced AI systems. With over 500,000 contributors, we offer flexible, remote project-based opportunities with a supportive global community.

US

  • Evaluate AI-generated content for quality, accuracy, and cultural relevance
  • Apply Castilian Spanish expertise to assess response appropriateness for Spain
  • Provide structured feedback and document decisions to improve AI performance

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use objective, data-driven recruitment processes and prioritize privacy and fairness.

$15–$15/hr
United States

  • Rating and assessing the performance of AI models based on their output or behavior.
  • Labeling and categorizing content to train machine learning models.
  • Generating prompts, responses, and summaries to improve language model reasoning.

Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, the company focuses on enabling responsible AI advancement.

  • Evaluate prompts and AI-generated outputs for accuracy, cultural appropriateness, and brand alignment.
  • Review and correct text, analyze multimedia content, and contribute voice recordings.
  • Apply local cultural insight and consistent evaluation guidelines to ensure high-quality AI training.

Lilt provides multilingual AI and human-verified services to enterprises, governments, and AI developers. They foster a global community of linguists and subject matter experts working on cutting-edge AI and language technology.

Canada

  • Evaluate AI-generated media, journalism, and communications content against quality rubrics.
  • Review outputs for factual accuracy, relevance, clarity, tone, and structure.
  • Provide structured, actionable feedback to improve AI model performance.

Jobgether is an AI-powered job matching platform that connects professionals with remote opportunities. It uses technology to streamline recruitment and provide a flexible, remote-first work environment.

Global

  • Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
  • Review and correct text, analyze multimedia content, and contribute voice recordings.
  • Apply local insight into tone, symbolism, visual cues, and market fit to deliver culturally relevant content.

LILT is an AI company that makes the world's information available to everyone, no matter the language they speak. They work with a global community of linguists and subject matter experts to deliver multilingual AI and human-verified services to Enterprises, Governments, and AI Developers.

Canada

  • Evaluate AI-generated finance documents against quality standards and identify issues.
  • Review outputs for accuracy, consistency, and formatting errors.
  • Provide structured, actionable feedback to improve AI content.

The partner company develops AI systems and seeks evaluators to improve output quality. The work is remote and asynchronous, with a focus on independent contributions.

Netherlands

  • Label and evaluate photos, graphics, videos, stickers, and designs in Dutch for linguistic accuracy and cultural appropriateness.
  • Assess AI-generated content against Canva's quality bar for Dutch users to shape localised AI experiences.
  • Build and contribute to Dutch-specific datasets and deliver labelled assets on time across varied task types.

Canva is redefining how the world experiences design, empowering users to create visual content. The company has a global team and supports flexible, remote-friendly work, with a focus on collaboration and innovation.

$35–$35/hr
Germany UK

  • Evaluate AI-generated responses for accuracy, grammar, and cultural relevance in German.
  • Create natural prompts and responses in German to improve conversational datasets.
  • Collaborate with global teams to help refine AI language models.

Welo Data, part of Welocalize, is a global AI data company that provides high-quality, ethical data to train advanced AI systems. With over 500,000 contributors worldwide, they focus on building smarter, more human AI through a diverse, global community.

$24–$24/hr
Germany

  • Provide expert rating support and guidance, enhancing the skills of other raters on AI-powered advertising systems.
  • Analyze datasets, identify patterns, and use KPIs to drive data-informed decisions for quality improvements.
  • Collaborate with clients and teams to conduct test rounds, ensure quality benchmarks, and maintain rating proficiency.

Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train advanced AI systems. The company values limitless flexibility and growth, offering a supportive global community for its contributors.

Global

  • Provide native-level Canadian French language vetting and QA for AI data projects.
  • Annotate and review AI outputs for grammatical accuracy, cultural context, and naturalness.
  • Develop educational resources and feedback documentation to improve AI alignment.

We are an AI training company that focuses on language alignment and data annotation for AI systems. Our remote team values linguistic precision and cultural nuance.

Canada

  • Create original, human-authored content in specialized domains like law, business, and philosophy for AI training.
  • Work independently from creative prompts to produce accurate, original written samples without AI tools.
  • Collaborate with a global community of writers and experts to improve language technology and AI evaluation.

Jobgether uses AI-powered matching to connect candidates with hiring companies, streamlining the application process. It operates as a platform that shares shortlisted candidates with employers, who manage final decisions and next steps.

India

  • Label, annotate, and evaluate Hindi content for linguistic quality, accuracy, and cultural appropriateness.
  • Build and contribute to Hindi-specific datasets to support internationalisation of AI features.
  • Review and refine labels based on feedback to maintain consistency across task types.

Canva is a design platform that redefines how the world experiences design. We have a global team that supports remote collaboration and values diverse skills and backgrounds.

US

  • Evaluate simulated advertiser-AI conversations for technical accuracy and campaign structure.
  • Fact-check platform strategies against correct hierarchy and full-funnel metrics.
  • Write clear, actionable feedback to correct errors and improve AI model performance.

RWS specializes in AI training data and language services, providing data annotation and evaluation solutions to improve AI model reliability. The company fosters a culture of diversity and inclusion, operating as a global employer with a focus on equal opportunity.

US

  • Evaluate financial documents and reports to verify accuracy and provide AI training data.
  • Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
  • Validate AI outputs against professional financial standards and provide expert feedback.

Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.

Global

  • Evaluate model-generated content across multiple modalities including text, images, audio, and video.
  • Apply defined quality rubrics such as factuality, consistency, and aesthetics to assess outputs.
  • Conduct independent research on unfamiliar topics to make well-supported evaluation judgments.

Our client is a global technology company that helps businesses build, train, and manage AI systems. They offer flexible, remote work with variable workload and a focus on high-quality model evaluation.

Global

  • Evaluate AI-generated text and audio in Catalan for accuracy and natural flow.
  • Provide corrections and constructive feedback on grammar, tone, and cultural context.
  • Complete approximately 10 hours of asynchronous tasks each week via our online platform.

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.