Design advanced biology problems that challenge frontier AI systems in molecular biology, genetics, or computational biology.
Create deterministic tasks with exactly one correct answer and submit complete, verified solutions.
Use Python and bioinformatics tools to build problems involving experimental reasoning and computational analysis.
We create high-quality STEM training data for frontier AI models used by leading AI labs. We are a team of experts working to improve model reasoning in scientific domains.
Design advanced bioinformatics problems that challenge frontier AI models and require multi-step computational reasoning.
Create deterministic problems with exactly one verifiable correct answer and complete, verified solutions.
Develop tasks involving sequence analysis, genomics, and other computational biology workflows using Python or R.
Anyone AI creates high-quality STEM training data for frontier AI models, used directly in training pipelines at leading AI labs. The company is a contractor-focused organization with a lean team of domain experts.
Review and label biological and biotechnology-related content against defined policies and guidelines.
Evaluate technically complex or ambiguous biological exchanges and make consistent risk determinations.
Distinguish legitimate scientific research from potentially harmful or dual-use applications.
10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs and top tech companies. It is a specialized firm focused on adversarial red teaming, model evaluations, and intelligence collection for safe AI deployment.
Evaluate and rank AI-generated scientific explanations based on accuracy and logic.
Review scientific papers alongside AI-generated abstracts to identify inaccuracies.
Verify AI-generated data against source documentation for material properties and formulas.
Prolific builds the largest pool of quality human data for AI training, connecting researchers with a global participant network. We serve over 35,000 AI developers and organizations, focusing on ethical data collection.
Identify real cancer-biology datasets and convert them into rigorous, deterministically-graded tasks for AI agents.
Establish defensible reference analyses and scoring criteria across conclusion types from basic biology to clinical outcomes.
Anticipate plausible agent errors and test handling of experimental variability, incomplete data, and conflicting evidence.
Latch builds rigorous scientific benchmarks for AI agents in biology, partnering with frontier labs and pharma. The team emphasizes cross-domain integration, long-horizon reasoning, and a culture of respectful debate and flexible schedules.