Source Job

Canada

  • Apply deep subject-matter expertise to AI model evaluation and large language model projects.
  • Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
  • Collaborate with AI research teams to improve training datasets and evaluation methodologies.

AI Research Analytical Problem-solving

20 jobs similar to STEM PhDs and Technical Domain Experts

Jobs ranked by similarity.

Canada

  • Conduct red-team evaluations to identify jailbreaks, prompt injections, and misuse scenarios in conversational AI models.
  • Develop creative adversarial prompts and scenarios to systematically probe model behavior and uncover weaknesses.
  • Generate high-quality human evaluation data by annotating failures and classifying vulnerabilities.

The partner company is a technology organization focused on AI safety and responsible AI development. They work with a remote, asynchronous team to improve the robustness of conversational AI systems.

Global

  • Evaluate LLM responses for accuracy, clarity, and completeness.
  • Fact-check technical claims using authoritative references.
  • Validate code and outputs, and annotate model performance.

Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.

UK

  • Compare and rank AI-generated responses for accuracy, logic, and safety.
  • Review CS research papers alongside AI summaries to ensure scientific integrity.
  • Fact-check technical data and code for logical flaws and inaccuracies.

Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers and researchers. They connect researchers with paid participants to gather high-quality, ethically sourced behavioral data for AI development.

Canada

  • Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
  • Identify factual, formatting, visual, and structural issues in professional deliverables.
  • Provide clear, structured feedback to enhance AI output quality and consistency.

This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.

Global

  • You will rate and assess the performance of AI models based on their output or behavior.
  • You will label elements of content and assign predefined categories to generate training data.
  • You will create prompts, summaries, and evaluate relevance to improve AI system understanding.

Innodata (Nasdaq: INOD) is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. The company has a 36+ year legacy of delivering high-quality data and outstanding outcomes for customers.

$22–$22/hr
Canada

  • Evaluate and rank model outputs, stress-test models for failure modes, and create high-quality datasets with detailed rubrics.
  • Annotate and correct multimodal data, maintain consistency through calibration exercises, and adapt to evolving task types.
  • Report on model performance trends and provide clear feedback to cross-functional partners on model successes and failures.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for real-world business problems. It is a global technology company with offices in Toronto, San Francisco, London, New York, Montreal, Seoul, Germany, and Paris, staffed by a team of passionate researchers, engineers, and designers.

US

  • Evaluate financial documents and reports to verify accuracy and provide AI training data.
  • Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
  • Validate AI outputs against professional financial standards and provide expert feedback.

Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.

$80–$150/hr
UK

  • Review AI-generated responses to clinical scenarios for accuracy and safety.
  • Compare and justify the best responses among multiple model answers.
  • Write improved exemplars and structured feedback to enhance AI model learning.

Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. It is a platform that connects researchers with a global participant pool for ethically sourced human data.

  • Create original graduate-to-PhD-level academic problems in your field
  • Write rigorous, step-by-step solutions with exact and verifiable answers
  • Review AI-generated responses to identify specific reasoning errors

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

Global

  • Evaluate model-generated content across multiple modalities including text, images, audio, and video.
  • Apply defined quality rubrics such as factuality, consistency, and aesthetics to assess outputs.
  • Conduct independent research on unfamiliar topics to make well-supported evaluation judgments.

Our client is a global technology company that helps businesses build, train, and manage AI systems. They offer flexible, remote work with variable workload and a focus on high-quality model evaluation.

Global

  • Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
  • Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
  • Benchmark performance by conducting comparative testing between model outputs based on technical metrics.

Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.

$15–$15/hr
United States

  • Rating and assessing the performance of AI models based on their output or behavior.
  • Labeling and categorizing content to train machine learning models.
  • Generating prompts, responses, and summaries to improve language model reasoning.

Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, the company focuses on enabling responsible AI advancement.

Canada

  • Own the roadmap for AI-powered development frameworks including agent workflows and evaluation systems.
  • Expand AI-enabled workflows across teams to improve efficiency and business outcomes.
  • Productize internal AI solutions by developing connector layers, skill libraries, and onboarding experiences.

A company focused on building AI-powered platforms to transform team productivity. The company operates remotely and values innovation and collaboration.

$100–$150/hr
Canada

  • Design precise grading criteria for pre-sales, solutions engineering, and technical sales deliverables.
  • Evaluate AI-generated and human-produced work samples against established criteria with detailed justifications.
  • Assess technical discovery, solution design, demonstrations, and proof-of-concept work for quality and effectiveness.

The partner company focuses on evaluating AI-generated and human-created sales engineering deliverables. The remote team is collaborative, consisting of experienced professionals and senior reviewers, fostering a culture of expert feedback and calibration.

$120,000–$170,000/yr
US

  • Design, build, and maintain automated AI evaluation pipelines for production LLM applications.
  • Develop prompt engineering strategies and evaluate model performance using quantitative methods.
  • Analyze production AI behavior with Python, SQL, and statistical techniques to identify improvement opportunities.

GovWorx provides an AI-powered platform, CommsCoach, that supports 9-1-1 and emergency communications centers by automating quality assurance, training, and real-time call evaluation. The company is a growing technology team focused on public safety, collaborating across AI, engineering, product, and data science.

Canada

  • Evaluate AI-generated documents and presentations against quality standards.
  • Apply humanities expertise to identify inaccuracies and cultural issues.
  • Provide structured feedback to improve AI model performance.

A partner company is seeking a humanities evaluator to assess AI-generated content for accuracy and quality. The company emphasizes cultural awareness and critical thinking in a remote, asynchronous work environment.

$75–$100/hr
United States

  • Review domain archives to understand subject matter and extract key facts.
  • Create accurate question-and-answer pairs covering various complexity levels.
  • Ensure answers are traceable, unambiguous, and consistent with approved source content.

Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With over 36 years of experience, the company delivers high-quality data and outcomes for Generative AI builders.

Canada

  • Conduct hands-on adversarial testing across AI models, applications, and data pipelines to identify vulnerabilities.
  • Perform advanced red-team assessments including jailbreaks, guardrail bypass, and prompt injection analysis.
  • Produce detailed vulnerability reports and collaborate with AI safety and engineering teams to improve security.

This company specializes in AI security research and adversarial machine learning. It is a remote-first organization with a collaborative culture focused on technical impact and professional growth.

$150–$220/hr
India

  • Design and apply evaluation criteria for consulting deliverables such as market analyses and financial models.
  • Assess AI-generated and human work, providing evidence-based scores and justifications.
  • Work independently in a remote, asynchronous environment to improve AI model reasoning.

They are an AI-focused organization improving the quality of AI outputs through expert evaluation. The work is fully remote and asynchronous, with an emphasis on independent problem-solving and collaboration with senior reviewers.

$150–$300/hr
Global

  • Develop difficult, novel tasks for models that challenge growing time horizons.
  • Conduct quality assurance to ensure tasks are solvable and appropriately scoped.
  • Baseline and score tasks within your domain of expertise for AI or human performance.

METR is a nonprofit research organization developing scientific methods to assess AI capabilities, risks, and mitigations, focusing on catastrophic AI risk evaluations. It is a mission-driven, tight-knit team with a low-ego, collaborative culture committed to high-quality, trustworthy science.