Source Job

India

  • Create original graduate- and PhD-level problems within your STEM expertise.
  • Develop rigorous, step-by-step solutions and review AI-generated responses for accuracy.
  • Work fully asynchronously through an online platform to contribute to AI reasoning research.

Problem Solving Technical Writing Critical Thinking Analytical Skills

20 jobs similar to STEM PhDs and Postdocs: Problem Creation for AI Reasoning

Jobs ranked by similarity.

$40–$40/hr
Global

  • Design advanced bioinformatics problems that challenge frontier AI models and require multi-step computational reasoning.
  • Create deterministic problems with exactly one verifiable correct answer and complete, verified solutions.
  • Develop tasks involving sequence analysis, genomics, and other computational biology workflows using Python or R.

Anyone AI creates high-quality STEM training data for frontier AI models, used directly in training pipelines at leading AI labs. The company is a contractor-focused organization with a lean team of domain experts.

US

  • Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
  • Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
  • Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.

The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.

Global

  • Design advanced mechanical engineering problems for frontier AI training and evaluation.
  • Write complete, verified solutions with clearly documented reasoning.
  • Develop problems that test engineering analysis and multi-step reasoning across mechanics, thermodynamics, and fluid dynamics.

We create high-quality STEM training data for frontier AI models, used directly in training and evaluation pipelines at leading AI labs. As a small, focused team, we value technical precision and rigorous problem-solving.

India

  • Analyze, evaluate, and review diverse datasets to support AI system training and improvement.
  • Assess AI-generated content for accuracy, relevance, consistency, and quality, providing actionable feedback.
  • Work independently in a remote, digital-first environment, managing multiple tasks and deadlines.

$40–$40/hr
Global

  • Design advanced civil engineering problems for frontier AI training and evaluation.
  • Create deterministic problems with exactly one verifiable correct answer and complete solutions.
  • Develop problems involving structural analysis, geotechnical engineering, hydraulics, transportation, or infrastructure systems.

We create high-quality STEM training data for frontier AI models. We are a remote-first company seeking experts in civil engineering to design rigorous problems for AI training.

Global

  • Design advanced electrical engineering problems for frontier AI training and evaluation.
  • Create deterministic problems with exactly one verifiable correct answer and complete solutions.
  • Develop problems covering circuits, electronics, signal processing, control systems, electromagnetics, power systems, or communications.

Anyone AI creates high-quality STEM training data for frontier AI models. They are a small, expert team dedicated to improving AI reasoning in technical domains.

India

  • Evaluate AI-generated content across humanities, arts, and culture domains for accuracy and quality.
  • Provide structured feedback and identify issues in AI outputs.
  • Work remotely on a flexible schedule contributing to AI system improvement.

The company specializes in developing and improving AI systems through expert human evaluation. They operate as a remote, flexible organization that values specialized domain knowledge and critical analysis.

Brazil

  • Design advanced and original mechanical engineering problems for AI training and evaluation.
  • Develop rigorous deterministic problems with verifiable solutions.
  • Work across domains including mechanics, thermodynamics, fluid mechanics, heat transfer, dynamics, and mechanical systems.

A partner company is seeking a Mechanical Engineering Expert to design engineering problems for AI training. The company offers a flexible, remote contract role with a part-time commitment of 20 hours per week.

Global

  • Review and analyze complex biological data sets or technical literature.
  • Provide detailed written feedback on emerging concepts in genetics and biochemistry.
  • Participate in occasional remote interviews to explain your analytical reasoning.

Terac is building the world's largest pool of vetted human experts for AI. They work with researchers, AI labs, and product teams to recruit, screen, and pay study participants.

India

  • Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
  • Assess visual and aesthetic quality including layout, formatting, and readability.
  • Provide clear, structured written feedback to identify issues and improve AI outputs.

Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.

India

  • Investigate user-reported technical issues end-to-end using logs, telemetry, and database queries.
  • Collaborate closely with Engineering and Product teams to identify root causes and improve product quality.
  • Communicate with users in a precise, empathetic manner to deliver a consistently positive experience.

Our partner is a fast-paced SaaS company focused on AI-powered solutions, offering customer support and technical troubleshooting. The company operates 100% remotely in India with a culture emphasizing collaboration between support, engineering, and product teams.

Global

  • Create and review realistic professional services scenarios in Nepali or English for AI benchmarking in Indian corporate contexts.
  • Adapt evaluation rubrics for analytical reasoning, technical problem-solving, and project coordination tasks.
  • Review AI and human-generated responses for factual accuracy, professional standards, and operational realism.

LILT provides multilingual AI and human-verified services to enterprises and governments worldwide. The company fosters a global, innovative community of linguists and subject matter experts dedicated to advancing human knowledge.

US

  • Evaluate software engineering tasks for technical accuracy, realism, and reproducibility.
  • Investigate codebases, tests, and integration issues to identify technical weaknesses.
  • Provide clear, actionable feedback that directly improves AI training and evaluation workflows.

Jobgether is an AI-powered job platform that connects candidates to roles through objective, skill-based matching. It focuses on remote and freelance opportunities, with a data-driven recruitment process and a global candidate pool.

$130,000–$200,000/yr
US

  • Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
  • Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
  • Turn research ideas into working prototypes and production-facing capabilities.

This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.

$15–$15/hr
Global

  • Evaluate and label AI model outputs to improve performance and alignment with project guidelines.
  • Create prompts, rewrite text, and generate training data for large language models.
  • Work on flexible, remote, project-based tasks while helping shape the future of AI.

Innodata is a global data engineering company that enables the responsible advancement of AI by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company delivers high-quality data and outstanding outcomes for customers.

$100–$150/hr
Canada

  • Evaluate AI-generated slides, spreadsheets, and documents for real-world usability and professional quality.
  • Assess outputs for accuracy, clarity, relevance, and alignment with data science standards.
  • Provide structured written feedback to help improve AI systems and their outputs.

A partner company is seeking a Data Science Expert to evaluate AI-generated work. The company focuses on improving AI systems and operates with a flexible, remote team.

Global

  • Assess software engineering tasks for technical accuracy, realism, and reproducibility.
  • Provide actionable feedback on codebase integration issues and logic errors.
  • Ensure AI training workflows are rigorous and practically applicable.

Project World Wide sources experienced technical specialists for AI training task auditing. This freelance contract opportunity focuses on ensuring technical rigor and accuracy in AI workflows.

$35–$75/hr
Global

  • Evaluate model outputs in humanities fields for factual accuracy, logical coherence, and ideological bias.
  • Create exemplary responses and datasets emphasizing intellectual honesty and thorough source evaluation.
  • Collaborate with engineering teams to design evaluation tasks and define desired model behavior.

SpaceXAI creates AI systems to understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence with a flat organizational structure.

US

  • Define end-to-end requirements for AI capabilities, from model behavior to user experience.
  • Translate model capabilities and technical constraints into product decisions.
  • Work closely with ML and engineering teams on system design and iteration.

They develop AI-powered email applications and work at the intersection of user needs and model capability. The company is fast-moving and collaborative with a strong focus on technical ownership and innovation.

India

  • Evaluate search engine results and online content against defined quality, relevance, and accuracy guidelines.
  • Review text, images, audio, and video content and provide structured feedback.
  • Work independently in a flexible, remote environment while maintaining high standards of quality.

A partner company is seeking Search Engine Evaluators to assess search results and digital content for AI improvement. The role offers flexible, remote work with no previous experience required.