Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.
Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.
Develop and evaluate AI training data for LLM and AI agent platforms.
Create coding tasks and write reference-quality solutions for evaluation.
Critically assess AI-generated code for correctness, security, and maintainability.
Toloka is a leading expert human data platform for AI agents and LLMs, providing high-quality training data. The company focuses on improving AI models through human feedback and structured evaluation.
Evaluate AI-generated JavaScript and TypeScript code for correctness and best practices.
Audit step-by-step explanations provided by AI for complex algorithmic solutions.
Execute model-generated scripts to verify performance and identify inefficiencies.
Prolific builds the largest pool of quality human data for AI development. Over 35,000 AI developers and researchers use the platform to gather data from paid participants with diverse experiences.
Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
Benchmark performance by conducting comparative testing between model outputs based on technical metrics.
Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.
Design and develop secure AI evaluation environments using modern full-stack development practices.
Analyze real-world web application clones to identify bugs and workflow issues for improving evaluation quality.
Build automated testing frameworks to accurately assess AI-generated code submissions.
The company focuses on creating secure, reliable testing systems for AI-generated code, directly influencing next-generation AI evaluation. They operate as a remote, collaborative team of highly skilled technical professionals.
Evaluate and improve AI model performance on complex infrastructure and platform engineering challenges.
Analyze system designs, assess code quality, and provide detailed feedback on architectural soundness and technical accuracy.
Create reproducible failure cases and communicate complex technical concepts to enhance AI reasoning capabilities.
Jobgether uses an AI-powered matching process to connect candidates with hiring companies quickly and fairly. As a platform, it facilitates remote freelance opportunities for technical professionals.
Design and implement core application components and services, setting standards for code quality and architecture. - Utilize AI-assisted development tools like Claude Code and Cursor to accelerate prototyping and refactoring. - Collaborate cross-functionally with product and engineering teams to deliver production-ready features in a healthcare SaaS environment.
CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, the company fosters a culture of impact, inclusion, and flexibility, and has been recognized as a best place to work over 10 times.
Evaluate LLM responses for accuracy, clarity, and completeness.
Fact-check technical claims using authoritative references.
Validate code and outputs, and annotate model performance.
Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.
Design and implement full-stack applications, AI agents, and platform components for rapid GenAI agent development, validation, and deployment.
Build developer tooling, CI/CD, and observability for safe, fast iteration with evals, canaries, and rollout/rollback systems.
Apply current LLM patterns like RAG, retrieval, routing, and tool-use to deliver measurable customer value and improved trust/safety metrics.
Wolters Kluwer provides expert software and information solutions that professionals in healthcare, legal, tax, and compliance industries rely on for critical decisions. With over 21,000 employees worldwide and 2025 annual revenues of €6.1 billion, they operate in over 40 countries with a culture valuing diversity and inclusion.
Develop difficult, novel tasks for models that challenge growing time horizons.
Conduct quality assurance to ensure tasks are solvable and appropriately scoped.
Baseline and score tasks within your domain of expertise for AI or human performance.
METR is a nonprofit research organization developing scientific methods to assess AI capabilities, risks, and mitigations, focusing on catastrophic AI risk evaluations. It is a mission-driven, tight-knit team with a low-ego, collaborative culture committed to high-quality, trustworthy science.
Evaluate AI-generated contract review outputs for legal accuracy and practical usefulness.
Apply real-world legal judgment to determine if AI suggestions reflect experienced attorney workflows.
Provide structured feedback to improve redline quality and consistency across contract types.
TAC is an AI-native law firm that specializes in commercial contract review for startups, using AI-assisted workflows. Backed by top investors and led by alumni from major tech companies, TAC offers a fast-paced, collaborative culture.
Lead agent optimization efforts to design tools for optimizing token spend and response time.
Work closely with the AI Foundations team on agentic development and ecosystem integrations.
Take end-to-end ownership of features, ensuring reliability and great developer experience.
Temporal is an open source programming model that simplifies code and makes applications more reliable. They are a growing, fully-remote company with values of curiosity, collaboration, and humility.
Build and evolve the services that capture, ingest, and process meeting audio, video, metadata, and real-time events.
Design systems that remain reliable across changing network conditions, platform behaviour, permissions, media formats, and third-party APIs.
Develop tooling and observability that make complex recording failures easier to detect, diagnose, reproduce, and resolve.
Fellow is an AI meeting assistant that helps teams record, transcribe, summarise, and act on their meetings. We are a Series A company backed by Craft Ventures, iNovia Capital, and Felicis Ventures, and our culture values speed, collaboration, and continuous improvement.
Own the operational backbone of an AI-native development workflow, including environment setup, CI/CD, testing, documentation, and API prototyping.
Prepare dev environments, run test suites, review code, manage deployments, and handle lower-complexity coding tasks using AI tools.
Research and prototype third-party API integrations and maintain technical documentation for regulatory compliance.
SPHERE Technology is a fully capitalized HealthTech SaaS company building an AI-native platform from scratch, backed by a Miami-based family office with a long runway and zero burn-rate pressure. The company is a startup with a small, focused team, ambitious product roadmap, and a direct access culture with no layers or committees.
Build and ship MVPs fast, working hands-on with LLMs, APIs, and AI-assisted workflows.
Use tools like Copilot, Cursor, GPT, and Claude daily to ship AI-powered product inside high-growth startups.
Work fully remote with flexible schedules, part-time or side gig, with some US time overlap.
Futureproofing is a talent platform focused on embedding high-caliber engineers into startups building AI-driven products. They match intentionally, not through keyword filtering, and focus on output, ownership, and shipping speed in a developer-first model.
Deliver high-quality work end-to-end, from scoping and design through implementation and release, leveraging AI tools to accelerate velocity.
Partner closely with Engineering, Product Management, and Product Design to refine ambiguous work into clearly defined, valuable deliverables.
Maintain reliable systems by applying best practices for observability, quality, and incident response, and actively use AI tooling to detect and resolve issues faster.
Newsela is a leading education technology company dedicated to meaningful classroom learning for every student. We deliver integrated, AI-powered solutions designed to unlock student engagement, empower teachers, and drive meaningful learning outcomes, and we focus on making a difference in the lives of students and teachers.
Support Meridian engineering teams by building, testing, and maintaining AI-assisted development workflows for cloud software.
Implement and maintain small AI-assisted engineering utilities for repository indexing, code summarization, and documentation generation.
Test open-source coding models and document their strengths, limitations, and practical usage guidance.
Deutsche Telekom IT Solutions Slovakia provides innovative information and communication technology services. It has grown to become the second largest employer in eastern Slovakia with over 3900 employees, fostering a culture of continuous improvement and transformation.
Implement AI-enabled features and agent workflows with senior guidance.
Design, build, test, and deploy small services with CI/CD pipelines.
Participate in on-call for owned components; mentor junior team members.
Mitratech builds world-class products simplifying operations in Legal, Risk, Compliance, and HR. We are a close-knit global team of over 35 years, serving 20,000 companies including 30% of the Fortune 500, with a diverse and inclusive culture that supports individual excellence.
Apply deep subject-matter expertise to AI model evaluation and large language model projects.
Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
Collaborate with AI research teams to improve training datasets and evaluation methodologies.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.
Build asynchronous Python/TypeScript services with FastAPI, Pydantic, and asyncio.
Develop agent orchestration, subagent delegation, tool calling, and structured outputs.
Create behavioral evals and regression datasets using Langfuse.
We are a venture-backed defense-tech company building AI-native intelligent software for sovereign institutions and their affiliated organizations. We have offices across the U.S., Europe, and the Middle East, and we are expanding our engineering team with people who demonstrate strong ownership and sound judgment.