Design and build rigorous, verifiable Terminal-Bench tasks that test multilingual robustness in LLMs across prompt language effects and encoding edge cases.
Create realistic task environments with datasets and files in your native language, ensuring assets remain in the target language to genuinely measure multilingual handling.
Calibrate task difficulty by analyzing execution logs and participate in a 4-layer human quality control process to ensure benchmark integrity.
Evaluate and edit AI-generated translations between German and English for accuracy and fluency.
Translate content between German and English while maintaining cultural relevance and context.
Annotate translation errors and provide feedback to improve AI model performance.
This partner company is seeking a German translator for an AI training project. It offers a flexible, remote contract opportunity for language professionals to improve AI translation quality.
Develop and evaluate AI training data for LLM and AI agent platforms.
Create coding tasks and write reference-quality solutions for evaluation.
Critically assess AI-generated code for correctness, security, and maintainability.
Toloka is a leading expert human data platform for AI agents and LLMs, providing high-quality training data. The company focuses on improving AI models through human feedback and structured evaluation.
Contribute to shaping safer, smarter AI by joining a global network of linguists and culturally aware contributors.
Work on flexible, remote projects in annotation, evaluation, and prompt creation, always on your terms.
Get first access to projects that match your skills, from short tasks to multi-week assignments.
Welo Data, part of Welocalize, is a global AI data company with a network of over 500,000 contributors. They build smarter, more human AI by offering flexible, remote projects to a diverse community in over 100 countries, emphasizing growth and work-life balance.
Test and evaluate AI chatbots and language models through structured conversations using assigned criteria.
Assess AI-generated responses for quality, relevance, safety, and linguistic accuracy.
Submit accurate deliverables such as written evaluations, ratings, and audio recordings within required timelines.
This company specializes in AI development and data evaluation, focusing on improving generative AI systems. The organization operates with a flexible, project-based team and values linguistic expertise.
Review and evaluate AI-generated text for linguistic accuracy, grammatical correctness, and cultural appropriateness.
Identify issues and provide high-quality rewritten responses to improve language outputs.
Develop natural prompts and responses in the target language to enhance conversational datasets.
Welo Data is a global AI data company with over 500,000 contributors delivering high-quality, ethical data to train advanced AI systems. We are building smarter, more human AI with a diverse community in over 100 countries.
Translate content between English and Mongolian, preserving cultural nuance and clarity.
Edit and refine language content for accuracy, tone, and contextual relevance across digital platforms.
Collaborate with cross-functional teams to localize content and support language model quality improvements.
LILT provides multilingual AI and human-verified translation services to enterprises, governments, and AI developers. They foster a global community of linguists and experts dedicated to innovation and excellence.
Adapt, review, and refine scripts for AI-powered dubbing projects to ensure natural, accurate, and engaging localized dialogue.
Review AI-generated dubbing outputs, identifying linguistic, performance, and quality issues that require correction.
Collaborate with production and language teams to improve localization quality and achieve strong creative outcomes.
Our partner is a media and entertainment localization company that adapts film, television, and digital content for global audiences. They are building a global talent pool of freelancers and focus on integrating AI-powered dubbing technologies into their workflows.
Evaluate AI-generated responses for accuracy, grammar, and cultural relevance.
Create natural prompts and responses in Filipino to improve conversational datasets.
Collaborate with global teams to help improve AI language models.
Welo Data, part of Welocalize, is a global AI data company that delivers high-quality, ethical data to train advanced AI systems. They have a community of over 500,000 contributors across 100+ countries, focusing on building smarter, more human AI with limitless opportunities for growth.
Have natural, unscripted conversations in Mandarin over video calls with other participants.
Discuss everyday topics to help AI models learn natural speech patterns and cultural nuances.
Choose your own schedule with no minimum hours, and get paid per task.
Prolific is building the biggest pool of quality human data in the world, used by over 35,000 AI developers and researchers. The platform connects researchers with a global participant pool for ethically sourced data.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.
Edit and review machine-translated customer service content for accuracy and consistency.
Evaluate MT output quality and apply severity ratings following project guidelines.
Produce high-quality human translations as reference data for model training.
Welo Data specializes in AI services and data solutions. They operate with a freelance workforce and focus on improving AI-driven content through human expertise.
Record scripted training material in Welsh with clear pronunciation and natural delivery.
Evaluate AI-generated speech for linguistic accuracy, fluency, naturalness, and expression.
Collaborate with project teams to refine prompts and voice design guidelines.
Our partner is developing next-generation AI voice technologies. The project offers a flexible freelance environment where your linguistic and performance expertise can have a global impact.
Review short-form video clips to identify languages in audio and on-screen text.
Validate media content against your target locale for authenticity and cultural relevance.
Generate accurate ground truth labels using internal classification tools following strict guidelines.
RWS is a technology-enabled language services company that helps global organizations connect with their audiences. With a large team of linguists and AI specialists, RWS fosters an inclusive culture focused on innovation and quality.
Complete AI training tasks such as analyzing, editing, and writing in Korean.
Evaluate and judge AI performance on Korean prompts.
Help improve cutting-edge AI models with your expertise.
Prolific builds the largest pool of quality human data for AI development. Over 35,000 developers and researchers use Prolific to gather diverse data from paid participants.
Accurately annotate and evaluate text, video, and geographic data following detailed guidelines.
Conduct research and review your work to ensure high standards of consistency and quality.
Collaborate with an international team and provide feedback to improve annotation processes.
They are an AI data company that improves the accuracy and performance of generative AI models through data annotation. They operate with a fully remote, collaborative team culture.
Support Meridian engineering teams by building, testing, and maintaining AI-assisted development workflows for cloud software.
Implement and maintain small AI-assisted engineering utilities for repository indexing, code summarization, and documentation generation.
Test open-source coding models and document their strengths, limitations, and practical usage guidance.
Deutsche Telekom IT Solutions Slovakia provides innovative information and communication technology services. It has grown to become the second largest employer in eastern Slovakia with over 3900 employees, fostering a culture of continuous improvement and transformation.
Work directly with customer engineering teams to design and implement real-time applications built on LiveKit.
Design scalable architectures for voice AI, real-time media, and developer platform use cases.
Help customers navigate SDKs, APIs, infrastructure, and deployment strategies for production systems.
LiveKit builds the infrastructure layer for the agentic era of computing, providing developers with everything needed to build, test, deploy, scale, and observe AI agents in production. Founded in 2021, the company powers voice and agentic AI applications for OpenAI, Salesforce, Spotify, Meta, and tens of thousands of other developers, collectively facilitating billions of calls each year.
Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.
Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.