Remote Data Jobs · Research

Job listings

$40,000–$80,000/yr

  • Propose and scope a new benchmark or evaluation technique in a domain APEX doesn't yet cover.
  • Design, build, and validate the benchmark with domain experts, including task specifications and scoring.
  • Run frontier models against your benchmark, analyze failures, and publish results as a paper or dataset.

Mercor organizes human intelligence to power the AI economy by building the layer between human expertise and frontier models. It is a profitable Series C company valued at $10 billion with a culture of in-person collaboration in San Francisco, NYC, or London.

  • Verify and maintain clinic database accuracy through research and outreach.
  • Enter data into platforms like Airtable, Google Sheets, and Monday.com.
  • Collaborate with teams to ensure data accuracy in response to policy changes.

Power to Decide is a national nonprofit, nonpartisan organization that advances sexual and reproductive well-being for all. They are committed to maintaining a diverse staff and an inclusive, multicultural environment.

  • Evaluate and rank AI-generated scientific explanations based on accuracy and logic.
  • Review scientific papers alongside AI-generated abstracts to identify inaccuracies.
  • Verify AI-generated data against source documentation for material properties and formulas.

Prolific builds the largest pool of quality human data for AI training, connecting researchers with a global participant network. We serve over 35,000 AI developers and organizations, focusing on ethical data collection.

  • Apply deep subject-matter expertise to AI model evaluation and large language model projects.
  • Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
  • Collaborate with AI research teams to improve training datasets and evaluation methodologies.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.

  • Review scientific papers alongside LLM-generated graphical abstracts.
  • Fact-check AI outputs for scientific accuracy and integrity.
  • Verify technical concepts using your neuroscience expertise.

Prolific is building the biggest pool of quality human data in the world, connecting AI developers, researchers, and organizations with paid study participants. With over 35,000 AI developers and researchers using the platform, it enables flexible, ethical data collection for AI training.

  • Evaluate model-generated content across multiple modalities including text, images, audio, and video.
  • Apply defined quality rubrics such as factuality, consistency, and aesthetics to assess outputs.
  • Conduct independent research on unfamiliar topics to make well-supported evaluation judgments.

Our client is a global technology company that helps businesses build, train, and manage AI systems. They offer flexible, remote work with variable workload and a focus on high-quality model evaluation.

$84,000–$100,000/yr
US 3w PTO

  • Conduct literature reviews and product evaluations to generate actionable insights for product development.
  • Collaborate with cross-functional teams to identify user needs and support evidence-based product design.
  • Manage research plans, data collection, and analysis to inform K-12 program innovation.

Committee for Children is a social enterprise dedicated to advancing the well-being of children through developing essential human skills, best known for the Second Step family of programs. They are a collaborative, creative team passionate about their work, committed to building a more equitable workplace.