Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Design realistic clinical scenarios and instructions based on your day-to-day specialty work.
Assess clinical AI outputs for accuracy, flagging incorrect or harmful information.
Validate reference answers and grading criteria, providing rationale for disagreements.
Aptura develops frontier AI models by integrating clinical expertise from physicians to review and validate AI outputs in healthcare scenarios. As a remote contractor-focused organization, Aptura offers flexible, asynchronous work alongside clinicians and AI researchers.
Design advanced bioinformatics problems that challenge frontier AI models and require multi-step computational reasoning.
Create deterministic problems with exactly one verifiable correct answer and complete, verified solutions.
Develop tasks involving sequence analysis, genomics, and other computational biology workflows using Python or R.
Anyone AI creates high-quality STEM training data for frontier AI models, used directly in training pipelines at leading AI labs. The company is a contractor-focused organization with a lean team of domain experts.
Support the development and validation of AI agent rubrics, reviewing AI-generated health responses for safety and utility.
Provide expert clinical perspective to design care pathways and clinically-focused solutions.
Validate that League's capabilities, programs, and content reflect current evidence and standards of care.
League is a leading healthcare experience platform that closes the gap between what people need to do next and getting it done. With over 70 million people on its platform, League is a fast-growing technology company that values AI-native workflows and equal opportunity.
Review and analyze complex biological data sets or technical literature.
Provide detailed written feedback on emerging concepts in genetics and biochemistry.
Participate in occasional remote interviews to explain your analytical reasoning.
Terac is building the world's largest pool of vetted human experts for AI. They work with researchers, AI labs, and product teams to recruit, screen, and pay study participants.
Develop Clinical Evaluation Plans (CEPs) and Clinical Evaluation Reports (CERs) per EU MDR to support CE Mark submissions.
Interact with cross-functional teams to define strategies for clinical documentation and regulatory submissions.
Conduct literature searches, critically appraise scientific data, and write clinical summaries for products and surgical procedures.
Intuitive is a global leader in robotic-assisted surgery and minimally invasive care, developing technologies like the da Vinci surgical system and Ion to transform surgery. The company is a team of engineers, clinicians, and innovators united by a purpose to make surgery smarter, safer, and more human, with challenging but deeply meaningful work.
Work cross-functionally across strategy, engineering, and marketing to appraise and optimize clinical foundations.
Independently analyze clinical cases and think about clinical databases as an independent contributor.
Design and test digital experiments alongside an engineering team.
AKASA provides generative AI solutions for healthcare revenue cycle, helping health systems capture the full patient clinical journey. With over $205M in funding, they serve leading health systems like Cleveland Clinic and Duke, and have been certified as a Great Place to Work for 6 years.
We create high-quality STEM training data for frontier AI models. We are a remote-first company seeking experts in civil engineering to design rigorous problems for AI training.
Identify real cancer-biology datasets and convert them into rigorous, deterministically-graded tasks for AI agents.
Establish defensible reference analyses and scoring criteria across conclusion types from basic biology to clinical outcomes.
Anticipate plausible agent errors and test handling of experimental variability, incomplete data, and conflicting evidence.
Latch builds rigorous scientific benchmarks for AI agents in biology, partnering with frontier labs and pharma. The team emphasizes cross-domain integration, long-horizon reasoning, and a culture of respectful debate and flexible schedules.