Evaluate and rank model outputs, stress-test models for failure modes, and create high-quality datasets with detailed rubrics.
Annotate and correct multimodal data, maintain consistency through calibration exercises, and adapt to evolving task types.
Report on model performance trends and provide clear feedback to cross-functional partners on model successes and failures.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for real-world business problems. It is a global technology company with offices in Toronto, San Francisco, London, New York, Montreal, Seoul, Germany, and Paris, staffed by a team of passionate researchers, engineers, and designers.
Evaluate and assess AI model outputs based on predefined quality, accuracy, relevance, and behavioral guidelines.
Annotate, classify, and label text, images, or audio to support AI model training.
Create prompts and generate high-quality responses to improve language model reasoning capabilities.
Jobgether uses an AI-powered matching process to connect candidates with hiring companies. They focus on efficient, fair recruitment and handle data privacy in compliance with GDPR.
Review AI-generated responses against source images and quality guidelines.
Identify issues like hallucinations, missing details, or policy violations.
Provide structured feedback to improve model performance and output quality.
Jobgether uses AI-powered matching to connect candidates with partner companies. They focus on efficient, objective hiring processes and operate as a platform for remote opportunities.
You will rate and assess the performance of AI models based on their output or behavior.
You will label elements of content and assign predefined categories to generate training data.
You will create prompts, summaries, and evaluate relevance to improve AI system understanding.
Innodata (Nasdaq: INOD) is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. The company has a 36+ year legacy of delivering high-quality data and outstanding outcomes for customers.
Rating and assessing the performance of AI models based on their output or behavior.
Labeling and categorizing content to train machine learning models.
Generating prompts, responses, and summaries to improve language model reasoning.
Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, the company focuses on enabling responsible AI advancement.
Review, evaluate, and annotate AI-generated content across text, images, audio, and video.
Perform quality checks to ensure accuracy, consistency, and compliance with project guidelines.
Identify edge cases and inconsistencies, contribute to high-quality dataset development, and participate in calibration activities.
Welo Data, part of Welocalize, is a global AI data company with over 500,000 contributors that provides high-quality, ethical data for training advanced AI systems. The company supports a diverse, global community across 100+ countries and offers project-based freelance opportunities with flexibility and growth potential.
Run the full eval pipeline end to end, reproducing results and pairing with senior engineers.
Build a judge calibration protocol to measure agreement and identify drift zones.
Extend benchmarks like GAIA and SWE-bench with new tasks targeting capability gaps.
Nous Research is an AI research lab that develops evaluation infrastructure for LLMs. They are a small, high-growth team valuing ownership and rapid shipping.
Review search results and evaluate their relevance to user queries
Answer true/false questions about content quality
Rate search results based on guidelines to improve AI systems
Welo Data provides AI services and data validation to improve search engine and AI systems. They are a remote-first company with a focus on quality and support for their contractors.
Own AI evaluation methods and operations for Figma's AI-powered experiences, defining quality dimensions and designing measurement frameworks.
Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars using human and automated approaches.
Produce clear readouts and dashboards to enable stakeholders to make confident shipping decisions.
Figma is on a mission to make design accessible to all, empowering teams to bring ideas to life through collaborative design and prototyping tools. With a growing team of passionate creatives and builders, Figma fosters a culture of growth and inclusivity.
Research data collection strategies and design high-impact data slices that uncover model failure modes.
Model annotator behavior and design experiments to optimize instruction clarity and reward signal reliability.
Develop metrics and frameworks for evaluating dataset quality, diversity, and impact on downstream model alignment.
Surge AI builds a platform that powers the most powerful AI models in partnership with companies like Anthropic, Google, Microsoft, and Meta. They are a profitable, bootstrapped company focused on human intelligence and data quality.
Evaluate AI quality across the advisor stack, including pre-call briefs, in-call guidance, and post-call outputs.
Iterate inside ORA by refining prompts, updating knowledge base entries, and tweaking skills to close the loop on issues.
Surface trends and drive continuous improvement by tagging conversations, logging issues, and recommending prioritized improvements.
HighLevel is an AI-powered business operating system that gives agencies, entrepreneurs and SMBs the infrastructure to build, automate and scale. With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership.
Lead technical discovery with foundation model labs, frontier AI teams, and large enterprises to understand model objectives and constraints.
Design end-to-end solutions across the post-training stack including SFT data curation, RLHF/DPO pipelines, custom benchmarks, and LLM-as-judge systems.
Author technical proposals, run workshops and POCs, and serve as ongoing technical advisor during delivery.
Innodata is a global data engineering company focused on enabling responsible AI advancement through data, evaluation frameworks, and human expertise. With a 36+ year legacy, they provide high-quality data solutions to foundation model labs, hyperscalers, and enterprise AI teams.
Evaluate prompts and AI-generated outputs for accuracy, cultural appropriateness, and brand alignment.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply local cultural insight and consistent evaluation guidelines to ensure high-quality AI training.
Lilt provides multilingual AI and human-verified services to enterprises, governments, and AI developers. They foster a global community of linguists and subject matter experts working on cutting-edge AI and language technology.
Evaluate prompts and AI-generated outputs for accuracy, clarity, and cultural appropriateness.
Review and correct text, analyze multimedia content, and contribute voice recordings.
Apply local insight into tone, symbolism, visual cues, and market fit to deliver culturally relevant content.
LILT is an AI company that makes the world's information available to everyone, no matter the language they speak. They work with a global community of linguists and subject matter experts to deliver multilingual AI and human-verified services to Enterprises, Governments, and AI Developers.