Source Job

US

  • Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
  • Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
  • Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.

Python TypeScript JavaScript AI Coding Tools

20 jobs similar to AI Interaction Evaluator (Codex / Claude Code)

Jobs ranked by similarity.

US

  • Evaluate AI coding-agent interactions for technical accuracy and engineering judgment.
  • Assess explanations and reasoning to ensure they genuinely help developers.
  • Provide structured feedback to improve AI-assisted development experience.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective AI-based review. The role is posted on behalf of a partner company that develops AI coding tools.

Global

  • Evaluate LLM responses for accuracy, clarity, and completeness.
  • Fact-check technical claims using authoritative references.
  • Validate code and outputs, and annotate model performance.

Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.

Canada

  • Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
  • Identify factual, formatting, visual, and structural issues in professional deliverables.
  • Provide clear, structured feedback to enhance AI output quality and consistency.

This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.

US

  • Develop and evaluate AI training data for LLM and AI agent platforms.
  • Create coding tasks and write reference-quality solutions for evaluation.
  • Critically assess AI-generated code for correctness, security, and maintainability.

Toloka is a leading expert human data platform for AI agents and LLMs, providing high-quality training data. The company focuses on improving AI models through human feedback and structured evaluation.

Global

  • Evaluate AI-generated JavaScript and TypeScript code for correctness and best practices.
  • Audit step-by-step explanations provided by AI for complex algorithmic solutions.
  • Execute model-generated scripts to verify performance and identify inefficiencies.

Prolific builds the largest pool of quality human data for AI development. Over 35,000 AI developers and researchers use the platform to gather data from paid participants with diverse experiences.

Global

  • Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
  • Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
  • Benchmark performance by conducting comparative testing between model outputs based on technical metrics.

Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.

$150–$300/hr
Global

  • Develop difficult, novel tasks for models that challenge growing time horizons.
  • Conduct quality assurance to ensure tasks are solvable and appropriately scoped.
  • Baseline and score tasks within your domain of expertise for AI or human performance.

METR is a nonprofit research organization developing scientific methods to assess AI capabilities, risks, and mitigations, focusing on catastrophic AI risk evaluations. It is a mission-driven, tight-knit team with a low-ego, collaborative culture committed to high-quality, trustworthy science.

US

  • Design and develop secure AI evaluation environments using modern full-stack development practices.
  • Analyze real-world web application clones to identify bugs and workflow issues for improving evaluation quality.
  • Build automated testing frameworks to accurately assess AI-generated code submissions.

The company focuses on creating secure, reliable testing systems for AI-generated code, directly influencing next-generation AI evaluation. They operate as a remote, collaborative team of highly skilled technical professionals.

US

  • Evaluate AI-generated content for quality, accuracy, and cultural relevance
  • Apply Castilian Spanish expertise to assess response appropriateness for Spain
  • Provide structured feedback and document decisions to improve AI performance

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use objective, data-driven recruitment processes and prioritize privacy and fairness.

Canada

  • Apply deep subject-matter expertise to AI model evaluation and large language model projects.
  • Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
  • Collaborate with AI research teams to improve training datasets and evaluation methodologies.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.

United States

  • Review search results and evaluate their relevance to user queries
  • Answer true/false questions about content quality
  • Rate search results based on guidelines to improve AI systems

Welo Data provides AI services and data validation to improve search engine and AI systems. They are a remote-first company with a focus on quality and support for their contractors.

Ireland

  • Listen to two audio recordings and compare them to determine which is better.
  • Follow provided evaluation guidelines to make consistent judgments.
  • Complete approximately 20-23 cases per hour with flexible remote work.

Appen leverages human feedback to train AI speech models. It is a large global company that connects independent contractors to AI projects.

UK

  • Compare and rank AI-generated responses for accuracy, logic, and safety.
  • Review CS research papers alongside AI summaries to ensure scientific integrity.
  • Fact-check technical data and code for logical flaws and inaccuracies.

Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers and researchers. They connect researchers with paid participants to gather high-quality, ethically sourced behavioral data for AI development.

South Africa

  • Evaluate and improve AI model performance on complex infrastructure and platform engineering challenges.
  • Analyze system designs, assess code quality, and provide detailed feedback on architectural soundness and technical accuracy.
  • Create reproducible failure cases and communicate complex technical concepts to enhance AI reasoning capabilities.

Jobgether uses an AI-powered matching process to connect candidates with hiring companies quickly and fairly. As a platform, it facilitates remote freelance opportunities for technical professionals.

Canada

  • Evaluate AI-generated spreadsheets against quality standards and domain-specific rubrics.
  • Identify calculation errors, formatting issues, and inconsistencies in workbooks.
  • Provide structured, actionable feedback to improve AI output accuracy and usability.

The partner company focuses on evaluating AI-generated spreadsheets and workbooks. It offers a remote, asynchronous work environment with flexible scheduling.

Global

  • Review, analyze, and evaluate Lua code for accuracy, quality, and best practices.
  • Complete technical evaluations and scripting tasks involving Lua programming.
  • Identify bugs, performance issues, and contribute to AI model training via code evaluations.

An enterprise client partners with leading AI organizations to improve AI model quality by leveraging expert software developers. They are seeking experienced Lua developers to support AI model training through technical evaluations and code reviews.

Brazil

  • Test and evaluate AI chatbots and language models through structured conversations using assigned criteria.
  • Assess AI-generated responses for quality, relevance, safety, and linguistic accuracy.
  • Submit accurate deliverables such as written evaluations, ratings, and audio recordings within required timelines.

This company specializes in AI development and data evaluation, focusing on improving generative AI systems. The organization operates with a flexible, project-based team and values linguistic expertise.

Canada

  • Evaluate AI-generated legal and business documents against quality standards and apply professional judgment.
  • Review contracts, diligence materials, redlines for accuracy, consistency, and completeness.
  • Provide clear, structured feedback to improve AI-generated legal content.

Global

  • Lead agent optimization efforts to design tools for optimizing token spend and response time.
  • Work closely with the AI Foundations team on agentic development and ecosystem integrations.
  • Take end-to-end ownership of features, ensuring reliability and great developer experience.

Temporal is an open source programming model that simplifies code and makes applications more reliable. They are a growing, fully-remote company with values of curiosity, collaboration, and humility.

$60,000–$90,000/yr
Global

  • Act as the technical bridge between product and customer deployments, translating business requirements into robust implementations.
  • Lead end-to-end production rollouts of AI workflows, including integrations, glue code, and monitoring.
  • Build reusable templates and deployment playbooks to make customer setups scalable and repeatable.

This company develops an AI-powered voice assistant platform for businesses. It is a seed-stage startup with a small, collaborative team focused on rapid iteration and customer success.