Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.
Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.
Audit multiple-choice options and correct answers for technical accuracy, eliminating ambiguous distractors.
Verify coding question prompts and grading rubrics, and write additional edge test cases.
Format and return the final corrected exam in a valid JSON structure.
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate Kubernetes tasks for technical accuracy, realism, and reproducibility.
Provide clear feedback on orchestration issues, configuration bugs, or logic errors.
Utilize your deep Kubernetes expertise to audit complex technical scenarios.
Greenhouse is a hiring platform that powers recruitment for modern companies. They are an established firm with a distributed team and a culture focused on innovation and flexibility.
Audit multiple-choice and coding questions for technical accuracy in a FastAPI microservices exam.
Verify code prompts, evaluate grading rubrics, and add edge test cases.
Submit a corrected JSON file with your expert modifications.
Terac builds the world's largest pool of vetted human experts for AI. Researchers and AI labs use Terac to recruit, screen, and pay study participants across many industries and languages.
Evaluate AI coding-agent interactions for technical accuracy and engineering judgment.
Assess explanations and reasoning to ensure they genuinely help developers.
Provide structured feedback to improve AI-assisted development experience.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective AI-based review. The role is posted on behalf of a partner company that develops AI coding tools.
Assess software engineering tasks for technical accuracy, realism, and reproducibility.
Provide actionable feedback on codebase integration issues and logic errors.
Ensure AI training workflows are rigorous and practically applicable.
Project World Wide sources experienced technical specialists for AI training task auditing. This freelance contract opportunity focuses on ensuring technical rigor and accuracy in AI workflows.
Evaluate generated images against prompts for adherence, composition, realism, and technical defects.
Compare images side by side and select the stronger one with clear, evidence-based rationale.
Classify images against customer content policies covering sexual content, violence, and real-person likeness.
Handshake is a career platform that helps 25 million job seekers connect with employers and 1,600 educational institutions. Handshake AI, a division, works with frontier AI labs on complex data at scale, having grown to a $1B run rate and paying over 30K individuals monthly.
Evaluate LLM responses for accuracy, clarity, and completeness.
Fact-check technical claims using authoritative references.
Validate code and outputs, and annotate model performance.
Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.
Develop and evaluate AI training data for LLM and AI agent platforms.
Create coding tasks and write reference-quality solutions for evaluation.
Critically assess AI-generated code for correctness, security, and maintainability.
Toloka is a leading expert human data platform for AI agents and LLMs, providing high-quality training data. The company focuses on improving AI models through human feedback and structured evaluation.
Support participants in virtual HR workshops by diagnosing and improving generative AI outputs.
Provide tool-agnostic guidance to help attendees refine prompts and achieve practical results.
Ensure responsible AI use and escalate issues as needed while fostering productive learning.
This company is a leading technology firm specializing in internet-related services and products, including search, cloud computing, and AI. It is a large global organization with a culture of innovation and collaboration.
Audit multiple-choice questions for technical correctness and unambiguous distractors
Verify the accuracy of coding prompts and their associated grading rubrics
Create and add edge test cases for API endpoint coding questions
Terac is building the world's largest pool of vetted human experts for AI. They recruit, screen, and pay study participants across industries, languages, and skill sets.
Create original graduate-to-PhD-level academic problems in your field
Write rigorous, step-by-step solutions with exact and verifiable answers
Review AI-generated responses to identify specific reasoning errors
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Design advanced bioinformatics problems that challenge frontier AI models and require multi-step computational reasoning.
Create deterministic problems with exactly one verifiable correct answer and complete, verified solutions.
Develop tasks involving sequence analysis, genomics, and other computational biology workflows using Python or R.
Anyone AI creates high-quality STEM training data for frontier AI models, used directly in training pipelines at leading AI labs. The company is a contractor-focused organization with a lean team of domain experts.
Own the Vault's end-to-end lifecycle in code: ingest, processing, outtake, and the tooling that runs it all.
Build durable software to replace operational ceilings, shipping to production from week two with Python, Postgres, AWS, React, and TypeScript.
Drive measurable SLAs, time-to-liquidity, and cost per item by instrumenting your own data and hunting bottlenecks.
Alt is unlocking the value of alternative assets, starting with the $5 B trading-card market. We let collectors buy, sell, vault, and finance their cards in one place and we are backed by leaders at Stripe, Coinbase, Seven Seven Six, and pro athletes like Tom Brady and Giannis Antetokounmpo.