Review domain archives to understand subject matter and extract key facts.
Create accurate question-and-answer pairs covering various complexity levels.
Ensure answers are traceable, unambiguous, and consistent with approved source content.
Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With over 36 years of experience, the company delivers high-quality data and outcomes for Generative AI builders.
Evaluate and assess AI model outputs based on predefined quality, accuracy, relevance, and behavioral guidelines.
Annotate, classify, and label text, images, or audio to support AI model training.
Create prompts and generate high-quality responses to improve language model reasoning capabilities.
Jobgether uses an AI-powered matching process to connect candidates with hiring companies. They focus on efficient, fair recruitment and handle data privacy in compliance with GDPR.
Write prompts in Luxembourgish to test AI models and evaluate responses using a structured rating guide.
Identify areas for improvement in AI responses and provide accurate English translations.
Work remotely on a flexible schedule with competitive pay.
CrowdGen by Appen is a platform that connects freelancers to help improve AI models through data annotation and evaluation. It is part of a large global company with a diverse community of independent contractors working remotely.
Evaluate AI responses using a structured rating guide.
Translate prompts and evaluations into English.
CrowdGen by Appen provides AI training data services to improve machine learning models. They operate a global community of independent contractors and emphasize flexible, remote work.
Review and critically evaluate new AI benchmarks on a regular cadence.
Produce clear, publication-ready research reports for broad audiences.
Analyze benchmark datasets using coding tools while maintaining analytical oversight.
The company focuses on producing rigorous, public-facing evaluations of AI benchmarks. It is a remote-first organization with a collaborative culture that values critical thinking and independence.
Review and assess new AI benchmarks at least every two weeks, evaluating their methodology and implications.
Publish and maintain public-facing reports on benchmarks, updating them as versions and models evolve.
Examine individual tasks within benchmarks in detail, using coding agents while maintaining critical oversight.
Epoch AI is a research institute investigating trends in machine learning and the economic consequences of AI. We aim to build a comprehensive knowledge base on AI, serving policymakers and society, with a small, focused team that values rigor and inclusivity.
Create and curate an evaluation suite of real-world tasks for frontier AI models.
Rigorously evaluate AI systems, analyze results, and communicate findings.
Improve evaluation processes and potentially build out standalone benchmarks.
Epoch AI is a research institute that investigates trends in machine learning and the economic consequences of AI. Our mission is to develop a comprehensive, publicly accessible knowledge base on AI that informs policymakers, industry leaders, and society at large.
Build the evidence and trust architecture to make impact claims audit-ready and donor-ready through rigorous methodologies.
Lead rigorous evaluation and external validation including RCTs, quasi-experimental studies, and partnerships with credible validators.
Build impact intelligence systems and data quality by embedding measurement across the learner journey and strengthening data pipelines.
The African Leadership Group is a family of mission-aligned institutions building a pipeline of ethical, entrepreneurial talent through education and career pathways. With over $1.7 billion raised and 373,000+ graduates across 50+ countries, they aim to develop 3 million leaders by 2035.