Lead development of Dyno’s external-facing AI platforms (e.g., design.dynotx.com)
Develop field-relevant benchmarks for agentic tasks in biological design
Build and improve AI-driven workflows that integrate models, data systems, and experimental feedback
Dyno Therapeutics is a biotech company that develops AI-driven genetic technologies to transform patient lives. They are a high-energy, high-impact team of world-class scientists and engineers working at the intersection of AI and genetic medicine.
Evaluate AI-generated scientific responses for accuracy and reasoning in biology.
Fact-check technical claims from public databases like PubMed and NCBI.
Assess experimental logic and annotate errors in biological sequences or protocols.
Prolific builds the largest pool of quality human data for AI development. With over 35,000 AI developers and researchers using the platform, it connects experts to train and evaluate AI models through ethical, paid participation.
Lead analysis and interpretation of cancer's molecular signatures, guiding model development for biological changes.
Execute rigorous computational analyses on data from high-throughput molecular assays including sequencing and protein quantitation.
Partner with molecular biologists and development scientists to refine experiments and turn research models into products.
Freenome is a biotechnology company developing blood tests for early cancer detection. We are a growing organization dedicated to changing the landscape of cancer, with a collaborative and diverse team.
Own the design and defense of frontier model evaluations across reasoning, coding, agents, tool use, and multi-modal.
Build benchmark packages with expert-verified ground truth, multi-model headroom results, and rigorous QC.
Recruit, calibrate, and review a pool of subject-matter experts in coding, agentic/tool-use, and STEM/reasoning.
Anyone AI measures frontier model capability through expert-verified evaluation packages. The company operates as a remote team with a focus on rigorous benchmarking and lab collaboration.
Design, build, and maintain automated AI evaluation pipelines for production LLM applications.
Develop prompt engineering strategies and evaluate model performance using quantitative methods.
Analyze production AI behavior with Python, SQL, and statistical techniques to identify improvement opportunities.
GovWorx provides an AI-powered platform, CommsCoach, that supports 9-1-1 and emergency communications centers by automating quality assurance, training, and real-time call evaluation. The company is a growing technology team focused on public safety, collaborating across AI, engineering, product, and data science.
Build and ship AI agents, APIs, and applications on Affirm's internal platform, owning the full lifecycle from architecture to production.
Turn messy business requirements from People Operations stakeholders into production systems, integrating with tools like Workday and Notion.
Design reliability infrastructure for multi-model LLM services, including structured output validation and quality controls.
Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. The People Tech & Analytics team builds and owns the data, AI, and technology infrastructure for Affirm's People function, running like a product engineering group embedded in HR.
Build AI agents, workflows, and automations that solve real customer and operations problems.
Apply LLMs, retrieval, and tool calling to practical financial workflows.
Create AI-assisted experiences for onboarding, support, KYC, and risk review.
We are the leading insurance platform in Southeast Asia, helping people plan, save, and grow their money. With over 20 nationalities working remotely and from offices, we are expanding to offer spending, saving, investing, and more.
Act as a domain authority establishing benchmarks for AI-driven statistical results.
Provide knowledge sharing about day-to-day workflows and regulatory submissions.
Identify and correct discrepancies in data quality and generate reference artifacts.
Edison Scientific builds and commercializes AI agents for science. We are assembling a team of top researchers and engineers across AI and biology to build an AI scientist.
Run the full eval pipeline end to end, reproducing results and pairing with senior engineers.
Build a judge calibration protocol to measure agreement and identify drift zones.
Extend benchmarks like GAIA and SWE-bench with new tasks targeting capability gaps.
Nous Research is an AI research lab that develops evaluation infrastructure for LLMs. They are a small, high-growth team valuing ownership and rapid shipping.
Build agents with real authority over real workflows, not just Q&A toys.
Design typed, intent-routed tools over MCP for scale and performance.
Develop the evaluation harness to measure and improve agent quality.
Rootstock Software is building an AI-first ERP for manufacturing, replacing traditional UI with agent-driven interactions. They have a deep bench of senior engineers, real products and customers, and a culture of experimentation and mutual growth.
Build internal tools and coach-facing features on top of LLMs like Claude and OpenAI.
Turn unstructured conversational data into structured, reliable signals for daily prioritisation.
Own the full stack of what you ship: frontend, backend, model integration, and deployment.
Best10 delivers online, highly personalised diet and lifestyle coaching through WhatsApp and its own platform, working with medical insurers to reduce costs from high-risk members. The company fosters a builder culture where teams ship working AI solutions rapidly and take pride in systems that outlast them.
Build AI-powered product features integrated into real insurance workflows.
Design and optimize LLM-based interactions for customer and internal systems.
Improve reliability of AI outputs through guardrails, fallback logic, and validation layers.
BJAK is Southeast Asia's largest digital insurance platform, using AI to simplify insurance and financial services for millions of users. The company values technical excellence, speed of execution, and practical decision-making, with a global engineering team that works closely across product, design, and AI teams.
Build AI agents, workflows, and automations that solve real customer and operational problems.
Apply LLMs, retrieval, tool calling, and guardrails to practical financial workflows.
Create AI-assisted experiences for onboarding, support, KYC, risk review, and document handling.
BJAK is a mobile-first insurance platform that enables insurance to be accessible online in Southeast Asia, and is expanding to spending, saving, investing, and more. The company has a global team of over 20 nationalities working remotely and in offices, with a culture that values passion and building next-generation products.
Identify workflows where AI agents can achieve 10-100x efficiency gains, build the business case, and align with domain leadership.
Design and implement agentic workflows integrating CRM, ERP, ticketing, and other enterprise systems with human-in-the-loop checkpoints.
Own production agent performance, track KPIs, tune prompts and retrieval, and provide observability for continuous improvement.
Natera is a global leader in cell-free DNA testing, dedicated to oncology, women's health, and organ health. The team consists of highly dedicated professionals from world-class institutions, working in a challenging and collaborative culture.
Build internal tools and coach-facing features on top of LLMs to turn unstructured conversational data into structured signals.
Design and maintain prompt chains, RAG pipelines, and agent workflows, rapidly prototyping from idea to demo in days.
Wire AI into existing systems, own full stack, and sit with coaches to build for their reality.
We deliver highly personalised diet and lifestyle coaching at scale through WhatsApp and our own technology platform. We partner with leading health insurers and employers, and our culture is about using AI to remove administrative work and enhance the personal coach-member relationship.
Identify and design AI-powered agentic architecture to automate repetitive tasks in creative and technical workflows.
Define end-to-end architecture for AI frameworks, integrating with DCC tools and building reusable components.
Drive adoption, create documentation, and report automation impact to studio leadership.
DreamWorks Animation is a leading producer of animated films and series, known for high-quality, award-winning content. They foster a growth-minded community of artists, technologists, and innovators who value transparency, trust, and collaboration.
Design and own the datasets and evaluation specifications for financial-domain LLMs, vision-language models, and AI agents, focusing on unstructured and multimodal financial data.
Translate customer goals into concrete dataset specifications, taxonomies, rubrics, and acceptance criteria, ensuring domain validity and statistical defensibility.
Develop evaluation methodology beyond surface accuracy, covering numerical consistency, hallucination rates, refusal appropriateness, and fairness across customer segments.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. With over 36 years of experience, the company delivers high-quality data and solutions for Generative AI builders and adopters.
Own problem spaces end to end: write specs, acceptance criteria, and own the architecture.
Build AI into the product, e.g., turning free-text email replies into bookable quotes.
Make AI trustworthy with structured outputs, evals, and confidence-gated human review.
Cargo.one operates an AI-native operating system for freight, serving 30,000+ users across 172 countries with customers like Lufthansa Cargo and Kuehne+Nagel, backed by Index and Bessemer. The culture is positive, diverse, hard-working, feedback-heavy, and playful.
Research data collection strategies and design high-impact data slices that uncover model failure modes.
Model annotator behavior and design experiments to optimize instruction clarity and reward signal reliability.
Develop metrics and frameworks for evaluating dataset quality, diversity, and impact on downstream model alignment.
Surge AI builds a platform that powers the most powerful AI models in partnership with companies like Anthropic, Google, Microsoft, and Meta. They are a profitable, bootstrapped company focused on human intelligence and data quality.