Source Job

Canada

  • Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
  • Identify factual, formatting, visual, and structural issues in professional deliverables.
  • Provide clear, structured feedback to enhance AI output quality and consistency.

Quality Assurance Microsoft Office Google Workspace Attention To Detail

20 jobs similar to Document / Deck Production QA Evaluator

Jobs ranked by similarity.

Canada

  • Evaluate AI-generated spreadsheets against quality standards and domain-specific rubrics.
  • Identify calculation errors, formatting issues, and inconsistencies in workbooks.
  • Provide structured, actionable feedback to improve AI output accuracy and usability.

The partner company focuses on evaluating AI-generated spreadsheets and workbooks. It offers a remote, asynchronous work environment with flexible scheduling.

Canada

  • Evaluate AI-generated finance documents against quality standards and identify issues.
  • Review outputs for accuracy, consistency, and formatting errors.
  • Provide structured, actionable feedback to improve AI content.

The partner company develops AI systems and seeks evaluators to improve output quality. The work is remote and asynchronous, with a focus on independent contributions.

$4–$5/hr
Global

  • Review AI-generated responses against source images and quality guidelines.
  • Identify issues like hallucinations, missing details, or policy violations.
  • Provide structured feedback to improve model performance and output quality.

Jobgether uses AI-powered matching to connect candidates with partner companies. They focus on efficient, objective hiring processes and operate as a platform for remote opportunities.

Canada

  • Evaluate AI-generated legal and business documents against quality standards and apply professional judgment.
  • Review contracts, diligence materials, redlines for accuracy, consistency, and completeness.
  • Provide clear, structured feedback to improve AI-generated legal content.

Canada

  • Evaluate AI-generated documents and presentations against quality standards.
  • Apply humanities expertise to identify inaccuracies and cultural issues.
  • Provide structured feedback to improve AI model performance.

A partner company is seeking a humanities evaluator to assess AI-generated content for accuracy and quality. The company emphasizes cultural awareness and critical thinking in a remote, asynchronous work environment.

India

  • Evaluate AI-generated content against domain-specific quality rubrics in humanities, arts, and culture.
  • Review documents, spreadsheets, and presentations for accuracy, relevance, clarity, and overall quality.
  • Provide structured feedback and collaborate with AI research teams to improve model outputs.

A partner company is seeking subject-matter experts to evaluate AI-generated content across humanities, arts, and culture. The company offers a flexible, remote contract environment, with no details on team size provided.

$22–$22/hr
Canada

  • Evaluate and rank model outputs, stress-test models for failure modes, and create high-quality datasets with detailed rubrics.
  • Annotate and correct multimodal data, maintain consistency through calibration exercises, and adapt to evolving task types.
  • Report on model performance trends and provide clear feedback to cross-functional partners on model successes and failures.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for real-world business problems. It is a global technology company with offices in Toronto, San Francisco, London, New York, Montreal, Seoul, Germany, and Paris, staffed by a team of passionate researchers, engineers, and designers.

United States

  • Review search results and evaluate their relevance to user queries
  • Answer true/false questions about content quality
  • Rate search results based on guidelines to improve AI systems

Welo Data provides AI services and data validation to improve search engine and AI systems. They are a remote-first company with a focus on quality and support for their contractors.

Canada

  • Evaluate AI-generated media, journalism, and communications content against quality rubrics.
  • Review outputs for factual accuracy, relevance, clarity, tone, and structure.
  • Provide structured, actionable feedback to improve AI model performance.

Jobgether is an AI-powered job matching platform that connects professionals with remote opportunities. It uses technology to streamline recruitment and provide a flexible, remote-first work environment.

UK

  • Compare and rank AI-generated responses for accuracy, logic, and safety.
  • Review CS research papers alongside AI summaries to ensure scientific integrity.
  • Fact-check technical data and code for logical flaws and inaccuracies.

Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers and researchers. They connect researchers with paid participants to gather high-quality, ethically sourced behavioral data for AI development.

US

  • Review, evaluate, and annotate AI-generated content across text, images, audio, and video.
  • Perform quality checks to ensure accuracy, consistency, and compliance with project guidelines.
  • Identify edge cases and inconsistencies, contribute to high-quality dataset development, and participate in calibration activities.

Welo Data, part of Welocalize, is a global AI data company with over 500,000 contributors that provides high-quality, ethical data for training advanced AI systems. The company supports a diverse, global community across 100+ countries and offers project-based freelance opportunities with flexibility and growth potential.

Canada

  • Apply deep subject-matter expertise to AI model evaluation and large language model projects.
  • Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
  • Collaborate with AI research teams to improve training datasets and evaluation methodologies.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.

India

  • Evaluate AI-generated customer support materials against quality rubrics.
  • Review documents, spreadsheets, and presentations for accuracy and consistency.
  • Provide structured feedback to improve AI system performance.

The company is a partner firm specializing in AI system development and evaluation. It operates remotely with a focus on professional expertise and flexible contract work.

Global

  • Evaluate model-generated content across multiple modalities including text, images, audio, and video.
  • Apply defined quality rubrics such as factuality, consistency, and aesthetics to assess outputs.
  • Conduct independent research on unfamiliar topics to make well-supported evaluation judgments.

Our client is a global technology company that helps businesses build, train, and manage AI systems. They offer flexible, remote work with variable workload and a focus on high-quality model evaluation.

Global

  • Evaluate LLM responses for accuracy, clarity, and completeness.
  • Fact-check technical claims using authoritative references.
  • Validate code and outputs, and annotate model performance.

Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.

US

  • Evaluate AI-generated content for quality, accuracy, and cultural relevance
  • Apply Castilian Spanish expertise to assess response appropriateness for Spain
  • Provide structured feedback and document decisions to improve AI performance

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use objective, data-driven recruitment processes and prioritize privacy and fairness.

$15–$15/hr
United States

  • Rating and assessing the performance of AI models based on their output or behavior.
  • Labeling and categorizing content to train machine learning models.
  • Generating prompts, responses, and summaries to improve language model reasoning.

Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, the company focuses on enabling responsible AI advancement.

Global

  • Provide native-level Canadian French language vetting and QA for AI data projects.
  • Annotate and review AI outputs for grammatical accuracy, cultural context, and naturalness.
  • Develop educational resources and feedback documentation to improve AI alignment.

We are an AI training company that focuses on language alignment and data annotation for AI systems. Our remote team values linguistic precision and cultural nuance.

US

  • Evaluate financial documents and reports to verify accuracy and provide AI training data.
  • Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
  • Validate AI outputs against professional financial standards and provide expert feedback.

Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.

Brazil

  • Review and validate AI training data for accuracy and consistency.
  • Provide constructive feedback to improve data quality.
  • Ensure compliance with project guidelines and escalate issues.

They are a partner company working on innovative AI data projects. The team is global and focused on improving AI accuracy through quality assurance.