Author original, technically rigorous free-response engineering problems in Materials Science and Engineering.
Design challenging questions that reflect realistic industrial scenarios and engineering constraints.
Develop expert-level solutions, explanations, and reasoning grounded in established engineering principles.
The partner company is focused on developing and evaluating frontier AI models through expert-level engineering problems. They operate as a remote, contract-based organization with a collaborative culture.
Evaluate AI-generated chemistry content for factual accuracy and structural integrity.
Fact-check chemical reactions, balancing equations, and synthetic routes.
Audit technical documentation including lab protocols and safety summaries.
Prolific is building the biggest pool of quality human data in the world. Over 35,000 AI developers, researchers, and organizations use Prolific to gather data from paid study participants.
Compare and rank AI-generated responses for accuracy, logic, and safety.
Review CS research papers alongside AI summaries to ensure scientific integrity.
Fact-check technical data and code for logical flaws and inaccuracies.
Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers and researchers. They connect researchers with paid participants to gather high-quality, ethically sourced behavioral data for AI development.
Evaluate AI-generated research on environmental topics for scientific accuracy and nuance.
Fact-check environmental claims, emissions factors, and legislative requirements.
Assess sustainability logic and annotate geospatial data to improve model outputs.
Prolific builds the world's largest pool of quality human data for AI development. With over 35,000 AI developers, researchers, and organizations using the platform, they focus on ethically sourced human behavioral data to train and evaluate AI models.
Create original graduate-to-PhD-level academic problems in your field
Write rigorous, step-by-step solutions with exact and verifiable answers
Review AI-generated responses to identify specific reasoning errors
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Fact-check AI outputs for scientific accuracy and integrity.
Verify technical concepts using your neuroscience expertise.
Prolific is building the biggest pool of quality human data in the world, connecting AI developers, researchers, and organizations with paid study participants. With over 35,000 AI developers and researchers using the platform, it enables flexible, ethical data collection for AI training.
Design difficult chemistry problems reflecting real scientific workflows.
Create deterministic tasks with one correct answer and full verified solutions.
Develop reasoning-intensive and computationally grounded problems using Python.
Anyone AI creates high-quality STEM training data for frontier AI models used by leading AI labs. The company is a remote-first, small team offering part-time contract work.
Design difficult chemistry problems that reflect real scientific workflows
Create deterministic tasks with one correct answer and full verified solutions
Develop reasoning-intensive and computationally grounded problems
They create high-quality STEM training data for frontier AI models that is directly used in training and evaluation workflows at leading AI labs. The company is a small team of contractors, and they value technical rigor and clear documentation.
Evaluate LLM responses for accuracy, clarity, and completeness.
Fact-check technical claims using authoritative references.
Validate code and outputs, and annotate model performance.
Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.
Design advanced biology problems that challenge frontier AI systems in molecular biology, genetics, or computational biology.
Create deterministic tasks with exactly one correct answer and submit complete, verified solutions.
Use Python and bioinformatics tools to build problems involving experimental reasoning and computational analysis.
We create high-quality STEM training data for frontier AI models used by leading AI labs. We are a team of experts working to improve model reasoning in scientific domains.
Evaluate financial documents and reports to verify accuracy and provide AI training data.
Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
Validate AI outputs against professional financial standards and provide expert feedback.
Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.
Review AI-generated responses to psychological scenarios and provide feedback.
Complete AI training tasks including analyzing, editing, and writing psychology-related content.
Compare model answers and select the best response, writing improved exemplars.
Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. They focus on ethically sourced human behavioural data to improve AI models.
Apply deep subject-matter expertise to AI model evaluation and large language model projects.
Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
Collaborate with AI research teams to improve training datasets and evaluation methodologies.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.
Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
Identify factual, formatting, visual, and structural issues in professional deliverables.
Provide clear, structured feedback to enhance AI output quality and consistency.
This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.
Review AI-generated responses to clinical scenarios for accuracy and safety.
Compare and justify the best responses among multiple model answers.
Write improved exemplars and structured feedback to enhance AI model learning.
Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. It is a platform that connects researchers with a global participant pool for ethically sourced human data.
Review AI-generated responses to psychological scenarios and behavioral prompts.
Complete AI training tasks such as analyzing, editing, and writing psychology-related content.
Compare model answers and select or justify the best response to improve AI understanding.
Prolific is building the world's largest pool of quality human data for AI training. They connect over 35,000 researchers and developers with paid participants, offering a flexible, remote platform for ethical data collection.
Complete a paid vetting survey and screening call to qualify for the study.
Design 3D-printable objects using open-source CAD software and capture workflow to train AI models.
Write and execute Python validation scripts and complete tasks across multiple complexity levels.
Prolific builds the largest pool of quality human data for AI developers and researchers. With over 35,000 users, they connect researchers with participants to gather data for AI training.
Create realistic, domain-specific tasks in the target language reflecting local natural sciences practices.
Adapt and apply clear scoring rubrics to evaluate AI-generated and human responses.
Review and score submissions for accuracy, regulatory alignment, and professional quality.
LILT provides multilingual AI and human-verified services to Enterprises, Governments, and AI Developers. They have a global community of linguists and subject matter experts focused on innovation and excellence.
Evaluate AI-generated content related to HR and talent acquisition tasks.
Create realistic HR workflow scenarios for performance management, benefits, and investigations.
Annotate and label data for offer letters, leave administration, and workforce analytics.
Prolific is building the world's largest pool of quality human data for AI development. They are trusted by over 35,000 AI developers and researchers and focus on ethical, high-quality data collection.