Source Job

$40,000–$80,000/yr
Global

  • Propose and scope a new benchmark or evaluation technique in a domain APEX doesn't yet cover.
  • Design, build, and validate the benchmark with domain experts, including task specifications and scoring.
  • Run frontier models against your benchmark, analyze failures, and publish results as a paper or dataset.

Machine Learning Python Statistics Research Data Analysis

20 jobs similar to Research Fellowship — APEX

Jobs ranked by similarity.

Global

  • Evaluate LLM responses for accuracy, clarity, and completeness.
  • Fact-check technical claims using authoritative references.
  • Validate code and outputs, and annotate model performance.

Prolific builds the largest pool of high-quality human data for AI development, serving over 35,000 AI developers, researchers, and organizations. They connect researchers with a global community to collect ethically sourced behavioral data.

$160,000–$210,000/yr
US

  • Build and maintain evaluation pipelines and environments for AI model benchmarks.
  • Prepare and maintain benchmark datasets, ensuring reproducibility and consistency.
  • Develop analysis tools and leaderboards to track model performance over time.

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. We're a hyper-growth and profitable company, quickly expanding our expert network, launching new offices, and new products.

Global 6w PTO 26w maternity 26w paternity

  • Conduct cutting-edge machine learning research, building and training large language models.
  • Focus on research projects aimed at expanding the frontier of knowledge in language modelling and associated areas.
  • Disseminate your research results through publications, datasets, and code.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products for enterprise AI systems. We are a global team of researchers, engineers, and designers passionate about our craft, with offices in Toronto, San Francisco, London, and more.

Global

  • Set and evolve the research direction for A1’s core intelligence, including context representation, memory, reasoning, planning, and orchestration.
  • Define evaluation frameworks that measure real-world usefulness, robustness, safety, and long-term behavior.
  • Own alignment, safety, and guardrail strategy as first-class product concerns.

A1 builds a proactive smart assistant for everyday users to bring intelligence to conversations, errands, organizing, and workflows. We are a small, high-talent-density team focused on shipping high-quality work and learning at rapid speed.

$150,000–$200,000/yr
Global

  • Lead research on AI capabilities and behaviors using our AI Village platform.
  • Design and execute analyses and experiments to discover important findings.
  • Work independently to drive research efforts and contribute to scaling the Village.

We build interactive AI demos and explainers to help make sense of the future. We're a team of four, focused on long-term open-ended AI agent research.

US

  • Evaluate financial documents and reports to verify accuracy and provide AI training data.
  • Respond to AI prompts using financial expertise to teach models complex fiscal concepts.
  • Validate AI outputs against professional financial standards and provide expert feedback.

Prolific builds the world's largest pool of quality human data for AI training. Over 35,000 AI developers and researchers use Prolific, and the company focuses on ethically sourced, diverse human behavioral data.

UK

  • Compare and rank AI-generated responses for accuracy, logic, and safety.
  • Review CS research papers alongside AI summaries to ensure scientific integrity.
  • Fact-check technical data and code for logical flaws and inaccuracies.

Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers and researchers. They connect researchers with paid participants to gather high-quality, ethically sourced behavioral data for AI development.

US

  • Drive end-to-end ML systems for forecasting products, from scoping and feature engineering to deployment and monitoring.
  • Build AI-assisted development workflows using tools like Claude Code to automate tasks with quality gates.
  • Partner with cross-functional teams to translate business needs into technical roadmaps and measurable impact.

Penguin Random House is the world's leading trade publishing company, with nearly 300 imprints and brands. We are a diverse international community of nearly 300 publishing brands committed to quality and innovation.

Global

  • Own full-cycle recruitment for AI Research, ML, data, evaluations, and ML infra roles.
  • Partner with research leadership and the CEO to define highly specialised candidate profiles.
  • Source talent from AI labs, research organisations, universities, and open-source communities.

White Circle is an AI Safety company building the safety, reliability, and optimization layer for AI systems. We are a small, highly focused team of under 50 people backed by top investors from OpenAI, Anthropic, and DeepMind.

$118,078–$330,661/yr
US

  • Lead cross-line AI/ML research strategy to drive product and pricing innovation across multiple insurance lines.
  • Build and lead a high-leverage team focused on reusable research tooling and workflow modernization.
  • Partner with Product R&D to identify opportunities for improving research speed, depth, and scalability through AI.

Mercury Insurance helps people reduce risk and overcome unexpected events, serving customers for over 60 years. It is a mid-sized employer recognized as one of America's Best Midsize Employers for 2026, fostering a culture of growth, inclusion, and teamwork.

Canada

  • Apply deep subject-matter expertise to AI model evaluation and large language model projects.
  • Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
  • Collaborate with AI research teams to improve training datasets and evaluation methodologies.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.

Global

  • Audit multiple-choice options and correct answers for technical accuracy, eliminating ambiguous distractors.
  • Verify coding question prompts and grading rubrics, and write additional edge test cases.
  • Format and return the final corrected exam in a valid JSON structure.

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

$115,000–$150,000/yr
US

  • Build and deploy AI-driven workflows across content, segmentation, reporting, and lead enrichment.
  • Automate campaign reporting, post-event analysis, and performance insights.
  • Partner cross-functionally with Product, Sales, RevOps, and Marketing to connect data and systems.

Flywire is a global payments enablement and software company that delivers the world's most important and complex payments. With over 1,200 FlyMates across 12 offices and more than 40 nationalities, the company supports over 4,800 clients in education, healthcare, travel, and B2B industries.

Global

  • Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
  • Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
  • Benchmark performance by conducting comparative testing between model outputs based on technical metrics.

Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.

$200,000–$300,000/yr
US Unlimited PTO

  • Design, build, and own agentic systems that produce client-ready financial deliverables from data ingestion to polished output.
  • Build evaluation and quality systems for generated deliverables, including deterministic checks and model-judged review.
  • Turn proprietary data into evidence-backed insight by building pipelines that discover and verify patterns.

Farsight is the agentic AI platform for financial services, helping investment banks and private equity firms automate nuanced workflows. The team comes from leading financial institutions and tech companies, focusing on hiring the best to become the best.

Global

  • Audit multiple-choice and coding questions for technical accuracy in a FastAPI microservices exam.
  • Verify code prompts, evaluate grading rubrics, and add edge test cases.
  • Submit a corrected JSON file with your expert modifications.

Terac builds the world's largest pool of vetted human experts for AI. Researchers and AI labs use Terac to recruit, screen, and pay study participants across many industries and languages.

$120,000–$180,000/yr
Global

  • Apply deep expertise in microbiome biology and metagenomics analysis to define how AI agents reason through complex datasets.
  • Build ground-truth benchmark datasets from published microbiome studies to rigorously test AI agent capability.
  • Work with software engineers and computational biologists to identify confounders and technical artifacts in computational microbiology.

Latch builds intelligent, high-performance agents for biological data analysis, empowering over 5,000 scientists across 150+ R&D labs. They have a waterfront office near Oracle Park and offer perks like 2x free daily meals, unlimited snacks, and a vibrant community with various clubs.

Global 7w PTO 16w maternity 16w paternity

  • Partner closely with researchers to translate research workflows into self-improving agentic flows.
  • Design and build robust backend services using Python and Go to power our agents.
  • Drive adoption of agent-native research workflows.

Poolside is building Artificial General Intelligence to accelerate software development through agentic systems and coding assistants. The team is distributed across Europe and North America with a culture of low ego, collaboration, and kindness.

US

  • Review scientific papers alongside LLM-generated graphical abstracts.
  • Fact-check AI outputs for scientific accuracy and integrity.
  • Verify technical concepts using your neuroscience expertise.

Prolific is building the biggest pool of quality human data in the world, connecting AI developers, researchers, and organizations with paid study participants. With over 35,000 AI developers and researchers using the platform, it enables flexible, ethical data collection for AI training.

Global

  • You will rate and assess the performance of AI models based on their output or behavior.
  • You will label elements of content and assign predefined categories to generate training data.
  • You will create prompts, summaries, and evaluate relevance to improve AI system understanding.

Innodata (Nasdaq: INOD) is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. The company has a 36+ year legacy of delivering high-quality data and outstanding outcomes for customers.