Design realistic scenarios in your target language or English grounded in operational contexts.
Adapt structured evaluation rubrics and review AI/human responses for accuracy, quality, and cultural appropriateness.
Contribute to gold-standard solutions reflecting best practices across target locale and domain.
LILT is an AI language company that provides multilingual AI and human-verified services to enterprises, governments, and AI developers. It operates with a global community of linguists and subject matter experts focused on innovation and excellence.
Review Saudi Arabic text against original text for naturalness, grammar, and accuracy.
Identify errors in spelling, numbers, symbols, and overall text quality.
Work 3-4 hours per day on a fully remote, project-based basis.
Appen provides high-quality training data for AI and machine learning projects. They are a large global company with a diverse community of contributors, offering flexible project-based work.
Review Judeo-Iraqi Arabic text against the original for naturalness, grammar, spelling, and formatting.
Identify and flag errors accurately according to project guidelines.
Ensure numbers, symbols, and special characters follow required conventions.
CrowdGen by Appen provides AI training data and language technology services. It operates a large platform connecting independent contractors worldwide for flexible, project-based work.
Audit multiple-choice math questions for technical accuracy and correctness.
Review grading rubrics to ensure clear explanation of correct mathematical reasoning.
Propose edge cases and test conditions to strengthen assessment logic.
Terac is building the world's largest pool of vetted human experts for AI researchers and labs. They are a growing platform used to recruit, screen, and pay study participants across industries, languages, and skill sets.
Translate and review patent-related content from English into Arabic, maintaining accuracy and readability for target audience.
Follow project instructions, including CAT tool usage (XTM/JIRA) and communicate promptly in English.
Ensure style adequacy and attention to detail while working remotely with immediate availability.
Welocalize is a trusted global transformation partner that helps brands reach international audiences through multilingual content services. With over 400,000 in-country linguistic resources, the company delivers translation and localization solutions for over 250 languages.
Design and build rigorous, verifiable Terminal-Bench tasks that test multilingual robustness in LLMs across prompt language effects and encoding edge cases.
Create realistic task environments with datasets and files in your native language, ensuring assets remain in the target language to genuinely measure multilingual handling.
Calibrate task difficulty by analyzing execution logs and participate in a 4-layer human quality control process to ensure benchmark integrity.
LILT is an AI and language technology company whose mission is to make the world's information available to everyone, regardless of language. They operate with a global community of linguists, engineers, and subject matter experts, fostering a culture of innovation and excellence.
Own the technical design and delivery of customer implementations across MENA, from kickoff through production cutover.
Lead technical discovery, proof-of-concept design, and evaluations for regional prospects, partnering with Account Executives.
Represent Deepgram in front of Arabic-speaking customers and at regional events, and bring customer feedback to Product and Research teams.
Deepgram is the leading platform for the Voice AI economy, providing real-time APIs for speech-to-text and text-to-speech. Over 200,000 developers and 1,300+ organizations use Deepgram, and the company is backed by a recent Series C with a culture that expects an AI-first mindset.
Contribute to a cutting-edge AI benchmarking project by creating and reviewing high-quality, real-world natural sciences scenarios in Korean.
Adapt and apply clear scoring rubrics to evaluate AI-generated and human responses for accuracy and regulatory alignment.
Provide expert feedback and contribute to high-quality gold standard solutions for AI evaluation.
LILT is an AI and language technology company that makes the world's information available to everyone, regardless of language. They are a global community of linguists, subject matter experts, and language professionals working on cutting-edge AI projects, offering flexible independent contractor opportunities.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Create and review realistic professional services scenarios in Nepali or English for AI benchmarking in Indian corporate contexts.
Adapt evaluation rubrics for analytical reasoning, technical problem-solving, and project coordination tasks.
Review AI and human-generated responses for factual accuracy, professional standards, and operational realism.
LILT provides multilingual AI and human-verified services to enterprises and governments worldwide. The company fosters a global, innovative community of linguists and subject matter experts dedicated to advancing human knowledge.
Review, localize, fact check, translate, and adapt football-related content for Arabic-speaking audiences.
Conduct market-specific research on local content, product, and terminology trends.
Provide feedback on written content to ensure quality and cultural resonance.
Welocalize is a global transformation partner that enables brands to reach and engage international audiences through multilingual content transformation services. They have a network of over 400,000 in-country linguistic resources and teams across North America, Europe, and Asia.
Design realistic healthcare and social assistance scenarios reflecting clinical and administrative settings in German-speaking locales.
Develop structured evaluation rubrics and review AI responses for medical correctness, operational feasibility, and patient safety.
Ensure cultural and contextual appropriateness of healthcare content, including hospital workflows and regulatory expectations.
LILT provides multilingual AI and human-verified services to Enterprises, Governments, and AI Developers worldwide. They have a global community of linguists and subject matter experts who collaborate on innovative projects advancing human knowledge.
Translate and review patent-related content from Arabic to English, ensuring original meaning is conveyed clearly and understandably to the target audience.
Follow instructions for translation process, CAT tool usage, and style adequacy as per client requirements.
Communicate effectively and respond promptly in English with the team.
Welocalize accelerates the global business journey by enabling brands and companies to reach, engage, and grow international audiences through multilingual content transformation services. With over 400,000 in-country linguistic resources and a global team across North America, Europe, and Asia, Welocalize fosters a culture of innovation and collaboration.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply careful judgment to ensure high-quality results aligned with task objectives.
LILT provides multilingual AI and human-verified services to enterprises, governments, and AI developers. The company has a global community of linguists and language professionals committed to innovation and excellence.
Design and engineer challenging benchmark tasks for evaluating coding agents in multilingual terminal environments.
Create authentic task environments using native language assets and identify model failure points.
Participate in rigorous quality assurance processes including calibration and audit of benchmark tasks.
The hiring company specializes in AI evaluation and multilingual language technology. They are a global team of engineers and linguists working on cutting-edge AI systems.
Review elementary math olympiad questions for technical accuracy
Identify and correct ambiguous wording in problem statements
Audit grading rubrics to ensure fair and logical point distribution
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Contribute to shaping safer, smarter AI by joining a global network of linguists and culturally aware contributors.
Work on flexible, remote projects in annotation, evaluation, and prompt creation, always on your terms.
Get first access to projects that match your skills, from short tasks to multi-week assignments.
Welo Data, part of Welocalize, is a global AI data company with a network of over 500,000 contributors. They build smarter, more human AI by offering flexible, remote projects to a diverse community in over 100 countries, emphasizing growth and work-life balance.
Evaluate model outputs in humanities fields for factual accuracy, logical coherence, and ideological bias.
Create exemplary responses and datasets emphasizing intellectual honesty and thorough source evaluation.
Collaborate with engineering teams to design evaluation tasks and define desired model behavior.
SpaceXAI creates AI systems to understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence with a flat organizational structure.