Design and engineer challenging benchmark tasks for evaluating coding agents in multilingual terminal environments.
Create authentic task environments using native language assets and identify model failure points.
Participate in rigorous quality assurance processes including calibration and audit of benchmark tasks.
The hiring company specializes in AI evaluation and multilingual language technology. They are a global team of engineers and linguists working on cutting-edge AI systems.
Design and build rigorous, verifiable Terminal-Bench tasks that test multilingual robustness in LLMs across prompt language effects and encoding edge cases.
Create realistic task environments with datasets and files in your native language, ensuring assets remain in the target language to genuinely measure multilingual handling.
Calibrate task difficulty by analyzing execution logs and participate in a 4-layer human quality control process to ensure benchmark integrity.
LILT is an AI and language technology company whose mission is to make the world's information available to everyone, regardless of language. They operate with a global community of linguists, engineers, and subject matter experts, fostering a culture of innovation and excellence.
Annotate text, image, audio, and video data to train large language models and AI systems.
Evaluate search results, ads, and chatbot outputs for factual accuracy, linguistic quality, and safety.
Provide detailed feedback and documentation to refine AI model reasoning and evaluation frameworks.
Our partner is a company focused on AI training and data annotation, working with a global network of linguists and tech enthusiasts. It values flexibility, transparency, and cultural accuracy in AI development.
Evaluate software engineering tasks for technical accuracy, realism, and reproducibility.
Investigate codebases, tests, and integration issues to identify technical weaknesses.
Provide clear, actionable feedback that directly improves AI training and evaluation workflows.
Jobgether is an AI-powered job platform that connects candidates to roles through objective, skill-based matching. It focuses on remote and freelance opportunities, with a data-driven recruitment process and a global candidate pool.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply careful judgment to ensure high-quality results aligned with task objectives.
LILT provides multilingual AI and human-verified services to enterprises, governments, and AI developers. The company has a global community of linguists and language professionals committed to innovation and excellence.
Annotate and label text, images, audio, or other content to support AI model training.
Evaluate AI-generated outputs such as search results and chatbot responses for quality and relevance.
Create and refine prompts while reviewing Greek linguistic and cultural accuracy.
This opportunity is a talent network for Greek-speaking contributors supporting AI model training through annotation, evaluation, and prompt creation. It is a global, remote community of independent contributors with flexible project-based work.
Apply your expertise in software engineering to help train next-generation AI systems by creating reinforcement learning environments and solving complex software engineering problems.
Contribute expert-level code samples, debugging strategies, and development insights in languages such as Python, Java, Rust, Go, C++, or TypeScript.
Refactor and optimize code, review and validate peer contributions, and document technical reasoning to enhance AI training data quality.
The company is a rapidly growing, venture-backed AI company that combines world-class human expertise with advanced machine learning to build and improve cutting-edge AI models. It is backed by over $40 million in funding and has a rapidly expanding international network of experts.
Record high-quality French speech samples for AI training datasets.
Evaluate AI-generated French audio for pronunciation, tone, and naturalness.
Provide detailed feedback and collaborate with technical teams to refine voice models.
The hiring partner is an innovative technology company focused on training cutting-edge conversational and expressive AI voice systems. They operate with a remote-first, autonomous freelance culture, collaborating with global talent to shape next-generation speech technology.
Label, annotate, and evaluate German-language content including photos, graphics, and videos for linguistic and cultural accuracy.
Evaluate AI-generated content against Canva's quality bar for German users to shape language experiences.
Build and contribute to German-specific datasets to support the internationalization of Canva AI features.
Canva is a design platform redefining how the world experiences design. It is a global company with a large user base, known for its innovative culture and focus on AI-powered features.
Annotate and label text, images, audio, or other data to improve AI systems.
Evaluate search results, advertisements, and chatbot responses for quality and accuracy.
Create, test, and refine prompts for large language models with Maltese-language insight.
Our partner is a global contributor network that provides flexible remote AI projects focused on annotation, evaluation, and prompt creation. The work is independent, project-based, and aims to improve the accuracy, relevance, and inclusivity of AI systems.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.
Collect data according to detailed project guidelines and specifications.
Ensure collected data meets the project's required standards and apply feedback for corrections.
Maintain consistent quality and productivity throughout the project.
Our enterprise client is a leading provider of high-quality, diverse datasets for AI model development. They are seeking skilled contributors to support AI training initiatives focused on improving AI-powered language and speech models.
Evaluate AI-generated Macedonian text for naturalness and cultural authenticity.
Compare text side-by-side to assess quality and nuance.
Provide feedback on tone, register, and word choice to improve AI language models.
Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. We connect a global community of participants with researchers to ethically source human behavior and feedback for AI development.
Design realistic scenarios in your target language or English grounded in operational contexts.
Adapt structured evaluation rubrics and review AI/human responses for accuracy, quality, and cultural appropriateness.
Contribute to gold-standard solutions reflecting best practices across target locale and domain.
LILT is an AI language company that provides multilingual AI and human-verified services to enterprises, governments, and AI developers. It operates with a global community of linguists and subject matter experts focused on innovation and excellence.
Review AI-generated business emails, reports, and strategies to ensure they are professionally sound.
Fact-check AI suggestions for marketing plans, budgets, and adherence to standard business rules.
Teach AI common office tasks such as meeting summaries, invoice drafting, and project timeline creation.
TELUS Digital operates a global AI community of over one million contributors who help collect, enhance, and train content for AI models. The community is diverse, flexible, and focused on shaping innovative AI technologies used by world-class brands.
Evaluate AI-generated responses for safety and bias against strict rubrics.
Classify harmful content categories like hate speech and self-harm.
Verify factual accuracy of model claims using external sources to prevent hallucinations.
TELUS Digital is a leading AI data company that trains models to be safe and accurate. They have a global community of over one million contributors and foster a collaborative, flexible culture.
Evaluate search results and AI-generated content for quality, relevance, accuracy, and usefulness.
Conduct online research to verify information and support rating decisions.
Apply rating guidelines consistently and participate in training and calibration sessions.
TELUS Digital AI & Data Solutions partners with a diverse and vibrant community to help our customers enhance their AI and machine learning models. Our global AI community includes over 1 million contributors across 500+ languages and dialects, offering flexible remote and onsite opportunities.
Create and review realistic professional services scenarios in Nepali or English for AI benchmarking in Indian corporate contexts.
Adapt evaluation rubrics for analytical reasoning, technical problem-solving, and project coordination tasks.
Review AI and human-generated responses for factual accuracy, professional standards, and operational realism.
LILT provides multilingual AI and human-verified services to enterprises and governments worldwide. The company fosters a global, innovative community of linguists and subject matter experts dedicated to advancing human knowledge.
Evaluate AI-generated coding interactions end to end for correctness and engineering judgment.
Assess whether outputs reflect strong engineering taste and provide clear, opinionated feedback.
Help define what great looks like for AI coding tools like Codex, Claude Code, and Cursor.
G2i Inc. is a technology staffing company that connects software engineers with remote contract opportunities. The company values engineering excellence and provides flexible, ongoing projects for senior-level developers.
Ship real features in production using coding agents like Claude Code, Cursor, and Copilot, setting intent and owning the outcome.
Keep the codebase and context lean by maintaining standards, architecture notes, and feeding SAST findings back for real patches.
Build team-level tooling, review AI-assisted pull requests, and share workflows that improve how everyone works with agents.
Insider One is a marketing and customer engagement platform that unifies data, personalization, and journey orchestration across channels. The company is powered by 1500+ employees across 30+ offices, and is recognized as a woman-founded, women-led B2B SaaS unicorn.