Develop difficult, novel tasks for models that challenge growing time horizons.
Conduct quality assurance to ensure tasks are solvable and appropriately scoped.
Baseline and score tasks within your domain of expertise for AI or human performance.
METR is a nonprofit research organization developing scientific methods to assess AI capabilities, risks, and mitigations, focusing on catastrophic AI risk evaluations. It is a mission-driven, tight-knit team with a low-ego, collaborative culture committed to high-quality, trustworthy science.
Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.
Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.
Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
Identify factual, formatting, visual, and structural issues in professional deliverables.
Provide clear, structured feedback to enhance AI output quality and consistency.
This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.
Evaluate AWS Serverless and IaC tasks for technical accuracy and reliability.
Provide clear, actionable feedback on deployment pipeline errors and architecture issues.
Ensure tasks are realistic, reproducible, and supported by robust tests.
We source experienced technical specialists to audit tasks used to train AI systems. Our project ensures that AI training workflows are technically rigorous and accurate, with a focus on high-quality deliverables.
Develop and evaluate AI training data for LLM and AI agent platforms.
Create coding tasks and write reference-quality solutions for evaluation.
Critically assess AI-generated code for correctness, security, and maintainability.
Toloka is a leading expert human data platform for AI agents and LLMs, providing high-quality training data. The company focuses on improving AI models through human feedback and structured evaluation.
Evaluate Kubernetes tasks for technical accuracy, realism, and reproducibility.
Provide clear feedback on orchestration issues, configuration bugs, or logic errors.
Utilize your deep Kubernetes expertise to audit complex technical scenarios.
Greenhouse is a hiring platform that powers recruitment for modern companies. They are an established firm with a distributed team and a culture focused on innovation and flexibility.
Write detailed outlines of your regular workflows, focusing on one critical task performed at least weekly.
Provide structured evaluation tasks and nuanced feedback to train AI models.
Complete paid tasks remotely on a freelance basis, with most tasks requiring one hour of uninterrupted work.
Prolific builds the world's largest pool of quality human data for AI development. With over 35,000 AI developers and organizations using its platform, it focuses on ethically sourced behavioral data from paid participants.
Evaluate AI model responses on software engineering tasks using TypeScript.
Challenge AI systems across algorithms, data structures, and development practices.
Provide structured feedback to improve model reasoning and code quality.
You'll work on a cutting-edge AI training project where your TypeScript expertise directly improves advanced language models. The project is a flexible freelance opportunity with a focus on software engineering and coding challenges, though the size and culture of the hiring company are not specified.
Assess technical accuracy and reproducibility of AWS Trainium/NKI tasks.
Provide actionable feedback on kernel execution bugs and logic errors.
Evaluate hardware acceleration inefficiencies and compilation issues.
We source experienced technical specialists to audit AI training tasks and evaluation workflows. The project is remote and freelance, with a focus on technical accuracy and efficiency.
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.
Evaluate AI coding-agent interactions for technical accuracy and engineering judgment.
Assess explanations and reasoning to ensure they genuinely help developers.
Provide structured feedback to improve AI-assisted development experience.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective AI-based review. The role is posted on behalf of a partner company that develops AI coding tools.
Evaluate AI-generated spreadsheets against quality standards and domain-specific rubrics.
Identify calculation errors, formatting issues, and inconsistencies in workbooks.
Provide structured, actionable feedback to improve AI output accuracy and usability.
The partner company focuses on evaluating AI-generated spreadsheets and workbooks. It offers a remote, asynchronous work environment with flexible scheduling.
Design precise grading criteria for pre-sales, solutions engineering, and technical sales deliverables.
Evaluate AI-generated and human-produced work samples against established criteria with detailed justifications.
Assess technical discovery, solution design, demonstrations, and proof-of-concept work for quality and effectiveness.
The partner company focuses on evaluating AI-generated and human-created sales engineering deliverables. The remote team is collaborative, consisting of experienced professionals and senior reviewers, fostering a culture of expert feedback and calibration.
Evaluate LLM responses for accuracy, clarity, and completeness.
Fact-check technical claims using authoritative references.
Validate code and outputs, and annotate model performance.
Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.
Assess technical accuracy, realism, and reproducibility of AI-assisted developer workflow tasks.
Provide actionable feedback on IDE integration faults, AI-assisted coding inefficiencies, and logic errors.
Apply deep knowledge of modern developer tooling, including AI coding assistants and telemetry/trace analysis.
We are sourcing experienced technical specialists to audit tasks used in training and evaluating AI systems. We focus on ensuring technical rigor and accuracy in developer workflow tasks, operating as a remote freelance project.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Design and build technical interview questions, including coding exercises, debugging challenges, and real-world engineering scenarios.
Write and test interview content across multiple programming languages, check solutions for correctness, and use AI-assisted tools to speed up creation.
Conduct peer reviews, update content based on performance data, and partner with Content Architecture to improve processes.
Karat is the world's largest interviewing company, providing a powerful system for technical leaders at companies like PayPal, Atlassian, and Citi. They are a growing, remote-first team that values inclusivity and diversity, and they are committed to preventing barriers to success.
Conduct structured 60-minute technical interviews across back-end, full-stack, front-end, and system design domains.
Submit thorough written evaluations of each candidate's coding approach, technical knowledge, and communication skills.
Maintain Karat's standards for structured, inclusive, bias-mitigated interviewing throughout the engagement.
Karat is a technical interviewing platform that connects experienced engineers to conduct structured interviews with job candidates. It operates a distributed community of interview engineers and focuses on inclusive, bias-mitigated hiring.
Commit to a consistent weekly availability block of 40 hours for conducting interviews.
Conduct structured 60-minute technical interviews across back-end, full-stack, front-end, and system design.
Submit thorough written evaluations of each candidate's coding approach, technical knowledge, and communication skills.
Karat provides a platform for structured technical interviews, enabling companies to assess candidates effectively. They have a large community of interview engineers and emphasize inclusive, bias-mitigated interviewing practices.