Source Job

Canada

  • Evaluate AI-generated spreadsheets against quality standards and domain-specific rubrics.
  • Identify calculation errors, formatting issues, and inconsistencies in workbooks.
  • Provide structured, actionable feedback to improve AI output accuracy and usability.

Microsoft Office Google Workspace Data Integrity

17 jobs similar to Spreadsheet QA / Workbook Maintenance Evaluator

Jobs ranked by similarity.

Canada

  • Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
  • Identify factual, formatting, visual, and structural issues in professional deliverables.
  • Provide clear, structured feedback to enhance AI output quality and consistency.

This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.

Canada

  • Evaluate AI-generated finance documents against quality standards and identify issues.
  • Review outputs for accuracy, consistency, and formatting errors.
  • Provide structured, actionable feedback to improve AI content.

The partner company develops AI systems and seeks evaluators to improve output quality. The work is remote and asynchronous, with a focus on independent contributions.

$4–$5/hr
Global

  • Review AI-generated responses against source images and quality guidelines.
  • Identify issues like hallucinations, missing details, or policy violations.
  • Provide structured feedback to improve model performance and output quality.

Jobgether uses AI-powered matching to connect candidates with partner companies. They focus on efficient, objective hiring processes and operate as a platform for remote opportunities.

United States

  • Review search results and evaluate their relevance to user queries
  • Answer true/false questions about content quality
  • Rate search results based on guidelines to improve AI systems

Welo Data provides AI services and data validation to improve search engine and AI systems. They are a remote-first company with a focus on quality and support for their contractors.

Canada

  • Evaluate AI-generated legal and business documents against quality standards and apply professional judgment.
  • Review contracts, diligence materials, redlines for accuracy, consistency, and completeness.
  • Provide clear, structured feedback to improve AI-generated legal content.

Canada

  • Evaluate AI-generated documents and presentations against quality standards.
  • Apply humanities expertise to identify inaccuracies and cultural issues.
  • Provide structured feedback to improve AI model performance.

A partner company is seeking a humanities evaluator to assess AI-generated content for accuracy and quality. The company emphasizes cultural awareness and critical thinking in a remote, asynchronous work environment.

Japan

  • Review rater worksheets for compliance with formatting and guidelines.
  • Verify that responses are logically correct and aligned with assigned scenarios.
  • Provide clear, concise, actionable feedback and request corrections where needed.

Welo Data provides AI services including data annotation and evaluation. They work with a freelance team and focus on quality and accuracy in their projects.

India

  • Evaluate AI-generated content against domain-specific quality rubrics in humanities, arts, and culture.
  • Review documents, spreadsheets, and presentations for accuracy, relevance, clarity, and overall quality.
  • Provide structured feedback and collaborate with AI research teams to improve model outputs.

A partner company is seeking subject-matter experts to evaluate AI-generated content across humanities, arts, and culture. The company offers a flexible, remote contract environment, with no details on team size provided.

India

  • Evaluate AI-generated customer support materials against quality rubrics.
  • Review documents, spreadsheets, and presentations for accuracy and consistency.
  • Provide structured feedback to improve AI system performance.

The company is a partner firm specializing in AI system development and evaluation. It operates remotely with a focus on professional expertise and flexible contract work.

$22–$22/hr
Canada

  • Evaluate and rank model outputs, stress-test models for failure modes, and create high-quality datasets with detailed rubrics.
  • Annotate and correct multimodal data, maintain consistency through calibration exercises, and adapt to evolving task types.
  • Report on model performance trends and provide clear feedback to cross-functional partners on model successes and failures.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for real-world business problems. It is a global technology company with offices in Toronto, San Francisco, London, New York, Montreal, Seoul, Germany, and Paris, staffed by a team of passionate researchers, engineers, and designers.

Ireland

  • Listen to two audio recordings and compare them to determine which is better.
  • Follow provided evaluation guidelines to make consistent judgments.
  • Complete approximately 20-23 cases per hour with flexible remote work.

Appen leverages human feedback to train AI speech models. It is a large global company that connects independent contractors to AI projects.

Sweden

  • Review completed tasks from trainers, including questions, images, and golden answers, to ensure accuracy and consistency.
  • Independently verify golden answers based on images and flag errors, inconsistencies, or ambiguities.
  • Provide clear feedback and track error patterns to maintain high data quality and escalate unclear cases.

Welo Data, part of Welocalize, is a global AI data company with over 500,000 contributors delivering high-quality, ethical data to train advanced AI systems. They operate in 100+ countries, offering flexible remote work and growth opportunities.

Global

  • Evaluate LLM responses for accuracy, clarity, and completeness.
  • Fact-check technical claims using authoritative references.
  • Validate code and outputs, and annotate model performance.

Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.

Global

  • Evaluate AI-generated data entries and records for accuracy, completeness, and consistency.
  • Simulate realistic data entry scenarios to test AI handling of messy inputs and edge cases.
  • Audit AI datasets for errors in categorization, labeling, and field mapping, ensuring quality standards.

Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers, researchers, and organizations. We connect researchers with a global pool of participants to collect ethically sourced human behavioral data, fostering a culture of flexibility and remote work.

  • Listen to two audio recordings and evaluate which is better.
  • Follow project guidelines to make consistent choices.
  • Complete 20-23 cases per hour remotely.

CrowdGen by Appen is an AI data company that improves AI systems through human feedback. They offer flexible, project-based remote work for independent contractors.

Canada

  • Evaluate AI-generated program management and implementation planning deliverables against quality standards.
  • Identify gaps, risks, and inconsistencies in project plans, schedules, and presentations.
  • Provide structured written feedback to improve accuracy and practical feasibility of AI outputs.

Our partner specializes in evaluating AI-generated program management and implementation planning artifacts. They offer a fully remote, asynchronous work environment where professionals apply their expertise to improve AI outputs.

Italy

  • Review search queries and evaluate personalized place recommendations based on your activity history.
  • Rate the relevance and usefulness of suggested places according to project guidelines.
  • Complete tasks accurately while following dynamic project schedules and requirements.

Welo Data provides AI services, including data annotation and evaluation for machine learning models. They operate with a global network of remote freelancers and prioritize accuracy and critical thinking.