Source Job

US

  • Evaluate software engineering tasks for technical accuracy, realism, and reproducibility.
  • Investigate codebases, tests, and integration issues to identify technical weaknesses.
  • Provide clear, actionable feedback that directly improves AI training and evaluation workflows.

Software Engineering Debugging Code Review AI Evaluation

20 jobs similar to SWE-Bench AI Task Auditor - Freelance AI Trainer Project

Jobs ranked by similarity.

Global

  • Assess software engineering tasks for technical accuracy, realism, and reproducibility.
  • Provide actionable feedback on codebase integration issues and logic errors.
  • Ensure AI training workflows are rigorous and practically applicable.

Project World Wide sources experienced technical specialists for AI training task auditing. This freelance contract opportunity focuses on ensuring technical rigor and accuracy in AI workflows.

  • Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
  • Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
  • Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.

Global

  • Evaluate AWS Serverless and IaC tasks for technical accuracy and reliability.
  • Provide clear, actionable feedback on deployment pipeline errors and architecture issues.
  • Ensure tasks are realistic, reproducible, and supported by robust tests.

We source experienced technical specialists to audit tasks used to train AI systems. Our project ensures that AI training workflows are technically rigorous and accurate, with a focus on high-quality deliverables.

Global

  • Evaluate Kubernetes tasks for technical accuracy, realism, and reproducibility.
  • Provide clear feedback on orchestration issues, configuration bugs, or logic errors.
  • Utilize your deep Kubernetes expertise to audit complex technical scenarios.

Greenhouse is a hiring platform that powers recruitment for modern companies. They are an established firm with a distributed team and a culture focused on innovation and flexibility.

Global

  • Assess technical accuracy and reproducibility of AWS Trainium/NKI tasks.
  • Provide actionable feedback on kernel execution bugs and logic errors.
  • Evaluate hardware acceleration inefficiencies and compilation issues.

We source experienced technical specialists to audit AI training tasks and evaluation workflows. The project is remote and freelance, with a focus on technical accuracy and efficiency.

Global

  • Review real user interaction traces with an AI shopping assistant
  • Identify logical failures, inaccuracies, or poor recommendations in the text
  • Create structured rubrics and verifiers to judge response quality

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

Global

  • Evaluate AI-generated coding interactions end to end for correctness and engineering judgment.
  • Assess whether outputs reflect strong engineering taste and provide clear, opinionated feedback.
  • Help define what great looks like for AI coding tools like Codex, Claude Code, and Cursor.

G2i Inc. is a technology staffing company that connects software engineers with remote contract opportunities. The company values engineering excellence and provides flexible, ongoing projects for senior-level developers.

Global

  • Review text or media samples based on provided project guidelines
  • Apply accurate labels and categorizations to diverse data sets
  • Evaluate AI-generated responses for clarity, safety, and factual accuracy

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

Global

  • Review coding tasks derived from real GitHub issues and pull requests to assess technical soundness and reproducibility.
  • Evaluate unit tests for correctness, coverage, and robustness, identifying flaky tests or missing dependencies.
  • Provide clear recommendations on whether tasks should be accepted, improved, or excluded.

Anyone AI is a company that focuses on AI-related projects, particularly in evaluating software engineering tasks. They are a smaller organization with a culture that values technical expertise and attention to detail.

US

  • Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
  • Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
  • Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.

The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.

Global

  • Evaluate GPU kernel tasks for technical accuracy, realism, solvability, reproducibility, and robust testing criteria.
  • Rigorously test and troubleshoot complex GPU programming scenarios to identify memory allocation bugs, execution bottlenecks, and parallel computing logic errors.
  • Review CUDA, Triton, and other GPU kernel implementations and provide clear, actionable technical feedback.

Our partner is a company specializing in AI training and evaluation, seeking experienced GPU kernel specialists to audit AI training tasks. The project is globally distributed and offers fully remote freelance work.

$88,000–$180,000/yr
US Unlimited PTO

  • Direct the agent array on production workstreams by decomposing problems into tasks and integrating agent output into shipped software.
  • Review agent-generated pull requests at volume and depth, identifying correctness, security, and accessibility defects.
  • Author evaluation suites that make quality measurable using eval-driven development and own end-to-end quality within a FedRAMP-authorized environment.

Granicus provides cloud-based solutions for government communications, website design, meeting management, and records management, serving over 5,500 agencies and 300 million citizens. With a globally distributed team and a culture of transparency and inclusion, Granicus has been recognized on the GovTech 100 list for the past 5 years.

US

  • Audit AWS Serverless and Infrastructure as Code tasks for technical accuracy and realism.
  • Evaluate deployment scenarios, architecture logic, and testing criteria.
  • Provide clear, actionable feedback to improve AI training and evaluation systems.

Jobgether uses AI-powered matching to connect professionals with freelance roles at partner companies. This project offers independent remote work with competitive hourly rates and flexible scheduling.

Global

  • Assess technical accuracy, realism, and reproducibility of AI-assisted developer workflow tasks.
  • Provide actionable feedback on IDE integration faults, AI-assisted coding inefficiencies, and logic errors.
  • Apply deep knowledge of modern developer tooling, including AI coding assistants and telemetry/trace analysis.

We are sourcing experienced technical specialists to audit tasks used in training and evaluating AI systems. We focus on ensuring technical rigor and accuracy in developer workflow tasks, operating as a remote freelance project.

Brazil

  • Design and engineer challenging benchmark tasks for evaluating coding agents in multilingual terminal environments.
  • Create authentic task environments using native language assets and identify model failure points.
  • Participate in rigorous quality assurance processes including calibration and audit of benchmark tasks.

The hiring company specializes in AI evaluation and multilingual language technology. They are a global team of engineers and linguists working on cutting-edge AI systems.

India

  • Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
  • Assess visual and aesthetic quality including layout, formatting, and readability.
  • Provide clear, structured written feedback to identify issues and improve AI outputs.

Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.

Global

  • Apply your expertise in software engineering to help train next-generation AI systems by creating reinforcement learning environments and solving complex software engineering problems.
  • Contribute expert-level code samples, debugging strategies, and development insights in languages such as Python, Java, Rust, Go, C++, or TypeScript.
  • Refactor and optimize code, review and validate peer contributions, and document technical reasoning to enhance AI training data quality.

The company is a rapidly growing, venture-backed AI company that combines world-class human expertise with advanced machine learning to build and improve cutting-edge AI models. It is backed by over $40 million in funding and has a rapidly expanding international network of experts.

US

  • Apply expertise in incident management and SRE to evaluate AI-generated documents, spreadsheets, and slide decks for technical accuracy and operational rigor.
  • Assess outputs against real-world reliability practices, identifying factual, technical, and reasoning errors.
  • Provide clear, structured written feedback and collaborate asynchronously with a research team to refine evaluation approaches.

This partner company focuses on AI evaluation and development, seeking experienced professionals to assess AI-generated work products. They offer flexible remote work and independent contractor engagements with weekly payments.

UK

  • Evaluate AI model responses on software engineering tasks using TypeScript.
  • Challenge AI systems across algorithms, data structures, and development practices.
  • Provide structured feedback to improve model reasoning and code quality.

You'll work on a cutting-edge AI training project where your TypeScript expertise directly improves advanced language models. The project is a flexible freelance opportunity with a focus on software engineering and coding challenges, though the size and culture of the hiring company are not specified.

Global

  • Design and build rigorous, verifiable Terminal-Bench tasks that test multilingual robustness in LLMs across prompt language effects and encoding edge cases.
  • Create realistic task environments with datasets and files in your native language, ensuring assets remain in the target language to genuinely measure multilingual handling.
  • Calibrate task difficulty by analyzing execution logs and participate in a 4-layer human quality control process to ensure benchmark integrity.

LILT is an AI and language technology company whose mission is to make the world's information available to everyone, regardless of language. They operate with a global community of linguists, engineers, and subject matter experts, fostering a culture of innovation and excellence.