Review search results and evaluate their relevance to user queries
Answer true/false questions about content quality
Rate search results based on guidelines to improve AI systems
Welo Data provides AI services and data validation to improve search engine and AI systems. They are a remote-first company with a focus on quality and support for their contractors.
Evaluate AI-generated documents and presentations against quality standards.
Apply humanities expertise to identify inaccuracies and cultural issues.
Provide structured feedback to improve AI model performance.
A partner company is seeking a humanities evaluator to assess AI-generated content for accuracy and quality. The company emphasizes cultural awareness and critical thinking in a remote, asynchronous work environment.
Evaluate AI agent conversations for acute symptom triage, diabetes management, and travel health guidance.
Score and annotate agent performance using structured rubrics and provide actionable feedback.
Participate in calibration sessions and help refine test scenarios for clinical safety.
Hippocratic AI develops AI-driven clinical agents to support healthcare delivery. As a specialized AI company, it focuses on patient safety and clinical accuracy, engaging clinical experts to evaluate and refine its systems.
Evaluate AI-generated documents, spreadsheets, and presentations for privacy and regulatory compliance accuracy.
Apply domain expertise to assess outputs against defined quality standards and provide actionable feedback.
Contribute to improving AI systems by providing expert judgment and structured feedback.
The company is a partner of Jobgether, offering remote opportunities for privacy and compliance professionals to evaluate AI-generated content. The company values autonomy and flexibility, providing independent contractor arrangements.
Evaluate AI-generated content for quality, accuracy, and cultural relevance
Apply Castilian Spanish expertise to assess response appropriateness for Spain
Provide structured feedback and document decisions to improve AI performance
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use objective, data-driven recruitment processes and prioritize privacy and fairness.
Design and iterate on agent behavior across real, live workflows, including long-horizon, multi-turn agentic tasks.
Build and run evaluations against production conditions to measure performance, regressions, and edge cases.
Partner with Product and Engineering to ensure agents are steerable, trustworthy, and ready for scale.
Arcadia is a healthcare platform that transforms complex data into trusted intelligence for providers, payers, and life sciences organizations. Backed by Nordic Capital, the company has hundreds of organizational clients and a mission-driven culture focused on improving patient outcomes.
Evaluate model-generated content across multiple modalities including text, images, audio, and video.
Apply defined quality rubrics such as factuality, consistency, and aesthetics to assess outputs.
Conduct independent research on unfamiliar topics to make well-supported evaluation judgments.
Our client is a global technology company that helps businesses build, train, and manage AI systems. They offer flexible, remote work with variable workload and a focus on high-quality model evaluation.
Evaluate and score AI-generated customer interactions across Foundational, Experiential, and Operational dimensions using complex rubrics.
Benchmark informational and transactional customer queries against authoritative business sources to ensure accuracy.
Participate in dual-review processes and daily calibration audits to maintain inter-rater agreement and quality standards.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. The company has a 36+ year legacy delivering high-quality data and outstanding outcomes for customers.
You will design the architecture for specialized subagents operating within live customer conversations, including technical QA, product expertise, and objection handling.
You will build routing and delegation systems that determine when to answer directly, invoke a subagent, or escalate to a human.
You will master the dialogue platform, train AI agents via prompting and fine-tuning, and document workflows to educate the team.
1mind builds autonomous customer experience software that deploys AI-powered 'Superhumans' to engage, demo, onboard, and support customers across the entire buying journey. The company offers a remote-first, fast-moving culture with ownership, autonomy, and impact from day one.
Review, evaluate, and validate AI-generated content and data according to project guidelines.
Compare AI outputs and identify the most accurate, relevant, or high-quality results.
Annotate, categorize, or label text, audio, images, video, or other data with clear feedback.
Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train advanced AI systems. They offer limitless flexibility and growth for a diverse community in 100+ countries.
Evaluate AI-generated customer support materials against quality rubrics.
Review documents, spreadsheets, and presentations for accuracy and consistency.
Provide structured feedback to improve AI system performance.
The company is a partner firm specializing in AI system development and evaluation. It operates remotely with a focus on professional expertise and flexible contract work.
Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
Identify factual, formatting, visual, and structural issues in professional deliverables.
Provide clear, structured feedback to enhance AI output quality and consistency.
This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.
Support participants in virtual HR workshops by diagnosing and improving generative AI outputs.
Provide tool-agnostic guidance to help attendees refine prompts and achieve practical results.
Ensure responsible AI use and escalate issues as needed while fostering productive learning.
This company is a leading technology firm specializing in internet-related services and products, including search, cloud computing, and AI. It is a large global organization with a culture of innovation and collaboration.
Own escalated cases end-to-end, investigating issues across conversational AI flows, voice integrations, REST APIs, and core banking connectors.
Analyze escalated AI conversations, categorize root causes, and close the feedback loop by writing and refining knowledge base articles.
Lead incident response for P1/P2 issues, coordinate with Engineering and customer teams, and write post-mortems for customer leadership.
Interface.ai builds Generative AI-powered virtual assistants for banks and credit unions, automating customer service across voice and chat for 100+ financial institutions. The company is remote-first and emphasizes high ownership and AI innovation at the frontier of financial services.
Rating and assessing the performance of AI models based on their output or behavior.
Labeling and categorizing content to train machine learning models.
Generating prompts, responses, and summaries to improve language model reasoning.
Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, the company focuses on enabling responsible AI advancement.
Evaluate AI-generated content related to job descriptions, compensation benchmarking, and employee policy interpretation.
Create realistic HR workflow scenarios such as performance management cycles, benefits enrollment, and employee investigations.
Annotate and label data across offer letter generation, leave administration, and workforce analytics.
Prolific builds the largest pool of quality human data in the world, used by over 35,000 AI developers, researchers, and organizations. The platform connects researchers with paid participants to gather ethically sourced behavioral data, placing itself at the forefront of AI innovation.
Own a client's voice or chat agent end-to-end, from requirements through production and continuous iteration.
Engineer the agent's behavior by writing and maintaining instruction sets, and build automated evaluation gates.
Debug live conversations, work directly with US enterprise clients, and ship changes that impact real callers.
Rifa AI builds an AI agents platform for contact centers in regulated industries. They are a small team of passionate engineers with paying enterprise clients and growing revenue, backed by Seaborne Capital.
Evaluate AI-generated content against domain-specific quality rubrics in humanities, arts, and culture.
Review documents, spreadsheets, and presentations for accuracy, relevance, clarity, and overall quality.
Provide structured feedback and collaborate with AI research teams to improve model outputs.
A partner company is seeking subject-matter experts to evaluate AI-generated content across humanities, arts, and culture. The company offers a flexible, remote contract environment, with no details on team size provided.
Evaluate AI model options (commercial, open source, on-prem) and build a repeatable framework for cost, capability, and security tradeoffs.
Build and maintain cost models comparing API-based pricing vs. flat subscriptions for heavy AI users.
Support vendor evaluation projects for SMS and messaging, including compliance and pricing analysis.
Customer.io is a platform that helps over 9,000 companies send automated communications like emails, push notifications, and SMS using real-time behavioral data. The company values empathy, transparency, and responsibility, fostering an inclusive and bias-free work culture.