Create and review realistic professional services scenarios in Nepali or English for AI benchmarking in Indian corporate contexts.
Adapt evaluation rubrics for analytical reasoning, technical problem-solving, and project coordination tasks.
Review AI and human-generated responses for factual accuracy, professional standards, and operational realism.
LILT provides multilingual AI and human-verified services to enterprises and governments worldwide. The company fosters a global, innovative community of linguists and subject matter experts dedicated to advancing human knowledge.
Contribute to a cutting-edge AI benchmarking project by creating and reviewing high-quality, real-world natural sciences scenarios in Korean.
Adapt and apply clear scoring rubrics to evaluate AI-generated and human responses for accuracy and regulatory alignment.
Provide expert feedback and contribute to high-quality gold standard solutions for AI evaluation.
LILT is an AI and language technology company that makes the world's information available to everyone, regardless of language. They are a global community of linguists, subject matter experts, and language professionals working on cutting-edge AI projects, offering flexible independent contractor opportunities.
Design and build rigorous, verifiable Terminal-Bench tasks that test multilingual robustness in LLMs across prompt language effects and encoding edge cases.
Create realistic task environments with datasets and files in your native language, ensuring assets remain in the target language to genuinely measure multilingual handling.
Calibrate task difficulty by analyzing execution logs and participate in a 4-layer human quality control process to ensure benchmark integrity.
LILT is an AI and language technology company whose mission is to make the world's information available to everyone, regardless of language. They operate with a global community of linguists, engineers, and subject matter experts, fostering a culture of innovation and excellence.
Design and engineer challenging benchmark tasks for evaluating coding agents in multilingual terminal environments.
Create authentic task environments using native language assets and identify model failure points.
Participate in rigorous quality assurance processes including calibration and audit of benchmark tasks.
The hiring company specializes in AI evaluation and multilingual language technology. They are a global team of engineers and linguists working on cutting-edge AI systems.
Create realistic, domain-specific mathematics tasks in Arabic reflecting local practices.
Adapt and apply clear scoring rubrics to evaluate AI-generated and human responses.
Review submissions and provide expert feedback for high-quality gold standard solutions.
LILT provides multilingual AI and human-verified services to Enterprises, Governments, and AI Developers worldwide. They have a global community of linguists and subject matter experts who thrive on innovation and excellence.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply careful judgment to ensure high-quality results aligned with task objectives.
LILT provides multilingual AI and human-verified services to enterprises, governments, and AI developers. The company has a global community of linguists and language professionals committed to innovation and excellence.
Contribute to shaping safer, smarter AI by joining a global network of linguists and culturally aware contributors.
Work on flexible, remote projects in annotation, evaluation, and prompt creation, always on your terms.
Get first access to projects that match your skills, from short tasks to multi-week assignments.
Welo Data, part of Welocalize, is a global AI data company with a network of over 500,000 contributors. They build smarter, more human AI by offering flexible, remote projects to a diverse community in over 100 countries, emphasizing growth and work-life balance.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Design realistic healthcare and social assistance scenarios reflecting clinical and administrative settings in German-speaking locales.
Develop structured evaluation rubrics and review AI responses for medical correctness, operational feasibility, and patient safety.
Ensure cultural and contextual appropriateness of healthcare content, including hospital workflows and regulatory expectations.
LILT provides multilingual AI and human-verified services to Enterprises, Governments, and AI Developers worldwide. They have a global community of linguists and subject matter experts who collaborate on innovative projects advancing human knowledge.
Label, annotate, and evaluate German-language content including photos, graphics, and videos for linguistic and cultural accuracy.
Evaluate AI-generated content against Canva's quality bar for German users to shape language experiences.
Build and contribute to German-specific datasets to support the internationalization of Canva AI features.
Canva is a design platform redefining how the world experiences design. It is a global company with a large user base, known for its innovative culture and focus on AI-powered features.
Have natural, unscripted conversations in Mandarin over video calls with other participants.
Discuss everyday topics to help AI models learn natural speech patterns and cultural nuances.
Choose your own schedule with no minimum hours, and get paid per task.
Prolific is building the biggest pool of quality human data in the world, used by over 35,000 AI developers and researchers. The platform connects researchers with a global participant pool for ethically sourced data.
Contribute to AI model training and evaluation in your area of expertise, including writing, reviewing, and assessing responses.
Evaluate AI-generated outputs for accuracy, logic, and nuance, and provide actionable feedback for model improvement.
Apply PhD-level judgment to real-world tasks in social sciences, humanities, arts, or linguistics.
Welo Data, part of Welocalize, is a global AI data company that delivers high-quality, ethical data to train advanced AI systems. With a community of over 500,000 contributors in 100+ countries, they emphasize flexibility, growth, and support for their contributors.
Review, localize, fact-check, translate, and adapt sports content (mainly football) to resonate with Brazilian Portuguese audiences.
Conduct market-specific research on terminology trends and best localization practices for the target market.
Provide feedback on written content and ensure adherence to style guides and guidelines.
Welocalize is a global transformation partner that helps brands reach, engage, and grow international audiences through multilingual content transformation services. It has a network of over 400,000 in-country linguistic resources and a team spanning North America, Europe, and Asia.
Source, vet, and onboard specialized annotators and performers across dozens of skill profiles and languages.
Scale sourcing channels using AI workflows, automation, and LinkedIn Recruiter.
Build relationships with experts and collaborate with delivery leads to meet customer needs.
HumanSignal is a human data partner for AI companies, creating and annotating real-world datasets and running the open-source platform Label Studio. They serve advanced ML and AI teams with data collection and multi-step workflows, though company size and culture details are not specified.
Evaluate and score AI-generated customer interactions across Foundational, Experiential, and Operational dimensions using complex rubrics.
Benchmark informational and transactional customer queries against authoritative business sources to ensure accuracy.
Participate in dual-review processes and daily calibration audits to maintain inter-rater agreement and quality standards.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. The company has a 36+ year legacy delivering high-quality data and outstanding outcomes for customers.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.
Research technical profiles on LinkedIn based on precise criteria.
Compile lists of 20 to 30 profiles per mission.
Work 100% remotely without contacting candidates.
Steinhem Direct Search is a headhunting firm based in Switzerland, specializing in civil engineering and construction. The company is experienced and growing, with a focus on confidentiality.
Evaluate search results and AI-generated content for quality, relevance, accuracy, and usefulness.
Conduct online research to verify information and support rating decisions.
Apply rating guidelines consistently and participate in training and calibration sessions.
TELUS Digital AI & Data Solutions partners with a diverse and vibrant community to help our customers enhance their AI and machine learning models. Our global AI community includes over 1 million contributors across 500+ languages and dialects, offering flexible remote and onsite opportunities.
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.
Create and review realistic Wholesale Trade scenarios in Odia or English for AI benchmarking.
Adapt structured evaluation rubrics for commercial problem-solving, sales operations, and order management.
Review AI responses for accuracy and contribute to gold standard solutions for Indian B2B trade.
LILT is a leading provider of multilingual AI and human-verified services for enterprises, governments, and AI developers. The company fosters a flexible, global contractor community focused on innovation and excellence.