Evaluate AI agent conversations for acute symptom triage, diabetes management, and travel health guidance.
Score and annotate agent performance using structured rubrics and provide actionable feedback.
Participate in calibration sessions and help refine test scenarios for clinical safety.
Hippocratic AI develops AI-driven clinical agents to support healthcare delivery. As a specialized AI company, it focuses on patient safety and clinical accuracy, engaging clinical experts to evaluate and refine its systems.
Review AI-generated clinical responses for accuracy, safety, and reasoning quality.
Compare multiple model answers and select or justify the best response.
Write improved exemplars and structured feedback to enhance AI learning.
Prolific builds the world's largest pool of quality human data for AI development, used by over 35,000 researchers and organizations. They focus on ethically sourced human behavioral data to improve AI models, with a culture of innovation and global reach.
Design realistic clinical scenarios and instructions based on your day-to-day specialty work.
Assess clinical AI outputs for accuracy, flagging incorrect or harmful information.
Validate reference answers and grading criteria, providing rationale for disagreements.
Aptura develops frontier AI models by integrating clinical expertise from physicians to review and validate AI outputs in healthcare scenarios. As a remote contractor-focused organization, Aptura offers flexible, asynchronous work alongside clinicians and AI researchers.
Design and customize AI-powered clinical documentation templates for doctors.
Analyze clinical documents to understand documentation styles and improve workflows.
Collaborate with cross-functional teams to ensure high-quality template solutions.
A digital health company that uses AI-powered tools to help doctors focus on care, not paperwork. They are fast-growing with hundreds of clinics across multiple countries and emphasize high-quality customer support.
Review and label clinical data to ensure accuracy and consistency for AI training.
Provide clinical expertise on healthcare workflows and EHR documentation.
Collaborate with teams to evaluate AI outputs and improve model performance.
This company creates AI-powered healthcare solutions to advance patient care and clinical efficiency. It offers a remote, collaborative environment for experienced healthcare professionals.
Review conversation transcripts against clinical quality, accuracy, and safety standards.
Identify and flag clinical risk, safety concerns, and moments requiring human clinical judgment.
Document clear, structured feedback that our clinical and product teams can act on.
Limbic is the leading clinical AI company for mental healthcare, powering patient-facing AI that has supported over half a million patients across the UK and US. We are a world-class team that has published in top journals, won major awards, and advises governments on AI safety and the future of clinical agentic systems.
Ensure clinical safety and operational efficiency of AI agent deployments.
Drive clinician adoption and confidence in AI tools through training and support.
Identify new use cases and workflows where AI can transform care delivery.
Hippocratic AI is building safety-first generative intelligence for healthcare. The company emphasizes clinical safety, patient advocacy, and cross-functional collaboration in a remote-first environment.
Provide expert pathology services for an AI project focused on medical image interpretation.
Annotate medical images and review AI-generated interpretations for accuracy.
Evaluate the clinical relevance and quality of model outputs to improve AI performance.
Prolific builds the world's largest pool of quality human data, serving over 35,000 AI developers and researchers. It connects companies with a global participant base for ethically sourced behavioral data and feedback.
Define what good looks like for AI-augmented primary care, setting clinical standards and safety guardrails.
Contribute to provider and patient-facing AI applications, including building multi-agent architectures for acute and chronic care.
Lead or co-author clinical research publications in top-tier journals.
We are a health tech company on a mission to expand access to affordable, high-quality primary care for everyone. Since 2017, we have operated an AI-powered virtual primary care platform that has delivered over 700,000 patient visits, and we are a mission-driven team of talented colleagues.
Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
Identify factual, formatting, visual, and structural issues in professional deliverables.
Provide clear, structured feedback to enhance AI output quality and consistency.
This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.
Apply deep subject-matter expertise to AI model evaluation and large language model projects.
Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
Collaborate with AI research teams to improve training datasets and evaluation methodologies.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.
Directly supervise, coach, and develop the non-clinician care team, setting performance expectations and productivity targets.
Own clinician scheduling and coverage planning across a 24x7x365 operation, managing PTO, shift swaps, and real-time staffing decisions.
Drive patient queue management and throughput to meet SLAs, identifying bottlenecks and implementing AI-driven improvements.
Curai Health is an AI-powered virtual multispecialty primary care clinic that uses machine learning to deliver affordable and accessible healthcare. It is a remote-first company with a mission-driven team committed to improving health outcomes at scale.
Validate AI-generated benefits guidance against accuracy and coverage standards.
Design and pressure-test member scenarios to ensure the rubric reflects real member questions.
Provide subject-matter expertise on benefits plan design, claims operations, or regulatory compliance.
League is a healthcare experience platform that closes the gap between people and their health by identifying what each person needs to do next and getting it done. It is one of the fastest-growing technology companies in Canada, serving 70 million+ people, and fosters an AI-native, flexible, and remote-friendly culture.
Evaluate AI-generated responses to pharmaceutical scenarios for accuracy and safety.
Compare multiple model answers and select the best based on current pharmacy standards.
Write improved exemplars and structured feedback to help AI models learn.
Prolific is building the world's largest pool of quality human data, connecting AI developers and researchers with paid study participants. With over 35,000 AI developers and organizations using the platform, Prolific offers flexible, remote work opportunities for domain experts to train and evaluate AI models.
Shape what we build by bringing doctors' real needs, workflows, and habits into the product.
Ensure clinical quality by defining standards, evaluating AI outputs, and turning errors into improvements.
Keep products safe and scalable by collaborating with regulatory, engineering, and clinical teams.
Docplanner is a global healthcare platform connecting over 300,000 doctors with 100 million patients monthly across 13 markets. Noa, a part of Docplanner, builds AI to help doctors reduce administrative work and improve patient care.
Review AI-generated responses to psychological scenarios and behavioral prompts.
Complete AI training tasks such as analyzing, editing, and writing psychology-related content.
Compare model answers and select or justify the best response to improve AI understanding.
Prolific is building the world's largest pool of quality human data for AI training. They connect over 35,000 researchers and developers with paid participants, offering a flexible, remote platform for ethical data collection.
Conduct red-team evaluations to identify jailbreaks, prompt injections, and misuse scenarios in conversational AI models.
Develop creative adversarial prompts and scenarios to systematically probe model behavior and uncover weaknesses.
Generate high-quality human evaluation data by annotating failures and classifying vulnerabilities.
The partner company is a technology organization focused on AI safety and responsible AI development. They work with a remote, asynchronous team to improve the robustness of conversational AI systems.
Evaluate and rank model outputs, stress-test models for failure modes, and create high-quality datasets with detailed rubrics.
Annotate and correct multimodal data, maintain consistency through calibration exercises, and adapt to evolving task types.
Report on model performance trends and provide clear feedback to cross-functional partners on model successes and failures.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for real-world business problems. It is a global technology company with offices in Toronto, San Francisco, London, New York, Montreal, Seoul, Germany, and Paris, staffed by a team of passionate researchers, engineers, and designers.
Evaluate AI-generated documents, spreadsheets, and presentations for privacy and regulatory compliance accuracy.
Apply domain expertise to assess outputs against defined quality standards and provide actionable feedback.
Contribute to improving AI systems by providing expert judgment and structured feedback.
The company is a partner of Jobgether, offering remote opportunities for privacy and compliance professionals to evaluate AI-generated content. The company values autonomy and flexibility, providing independent contractor arrangements.