Create realistic PowerPoint tasks from your own deal work, including building decks from source documents and filling corporate templates.
Work through each task to define a correct answer and write a yes-or-no checklist precise enough for AI grading.
Apply, attend a 30-minute walkthrough, and submit one original task with input files and rubric for team review.
Aptura is a London-based company studying how AI is changing professional work, with a focus on finance and advisory. Founded by former Lazard, McKinsey and Partners Group professionals, the company values precision and practical insight.
Write realistic PowerPoint tasks based on your own IR or FP&A reporting cycle, including building decks and rolling them forward.
Complete each task yourself to establish a correct answer and create a precise yes/no grading checklist.
Define what a correct deck looks like with enough detail for AI graders to evaluate attempts.
Aptura is a London-based company focused on understanding how AI is transforming professional work, with a close look at finance and advisory roles. Founded by former Lazard, McKinsey, and Partners Group professionals, it operates as a lean, expert-driven team.
Evaluate AI-generated work products in real estate, hospitality, and events using quality rubrics.
Identify factual, aesthetic, and presentation errors and provide actionable feedback.
Apply industry expertise to distinguish realistic, commercially sound work from generic AI content.
The company develops AI systems and evaluates their outputs for quality. They seek experienced industry professionals for flexible remote contract work.
Evaluate AI-generated market research and competitive intelligence artifacts for accuracy, rigor, and quality.
Apply structured rubrics to assess deliverables and identify factual inaccuracies and analytical gaps.
Provide clear, actionable written feedback to support evaluation decisions and improve AI training.
The partner company specializes in evaluating AI-generated market research and competitive intelligence content. They hire independent contractors for flexible remote engagements.
Review and advise on evaluation criteria and scoring rubrics for AI-generated outputs.
Create, edit, and validate high-quality benchmark tasks and reference data for AI training.
Analyze model failures, including hallucinations and flawed reasoning, providing expert explanations.
The company specializes in AI evaluation and training, helping define standards for next-generation AI systems. They operate as a partner company that manages applications and next steps, valuing autonomy and expertise in their consultants.
Designing problems to test state-of-the-art AI models on operations workflows
Working with frontier engineering/research teams to implement and evaluate AI systems
Defining what 'good' looks like precisely enough for someone else to grade against it
Aptura is an applied AI and data company that works with frontier labs to train AI systems on real professional work. The company's size is not specified, but it emphasizes innovation and collaboration.
Evaluate AI-generated documents, spreadsheets, and presentations against domain-specific quality standards.
Assess outputs for factual accuracy, procurement relevance, completeness, clarity, and practical applicability.
Provide structured written feedback identifying issues and opportunities for improvement.
This partner company specializes in evaluating AI-generated work products for public-sector procurement and RFI responses. The company size and culture are not specified in the posting, but the engagement is flexible and fully remote.
Design high-converting landing pages using our AI engine and own the creative process.
Collaborate with marketing managers and designers to deliver polished, timely work.
Give feedback to shape the AI landing page engine.
Uplane builds AI technology for the marketing agency of the future, helping creative teams generate ads and high-converting landing pages. It is a fast-growing, VC-backed startup with an ambitious, fun, and humble culture.
Evaluate developer workflow tasks for technical accuracy, realism, solvability, reproducibility, and alignment with reliable testing and evaluation criteria.
Audit AI-assisted development scenarios to identify technical inconsistencies, logic errors, workflow inefficiencies, or issues affecting task quality.
Provide clear, detailed, and actionable feedback that enables improvements to AI training and evaluation tasks.
Review real user interaction traces with an AI shopping assistant
Identify logical failures, inaccuracies, or poor recommendations in the text
Create structured rubrics and verifiers to judge response quality
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Evaluate AI-generated documents, spreadsheets, and presentation decks for accuracy and professional quality.
Assess visual and aesthetic quality including layout, formatting, and readability.
Provide clear, structured written feedback to identify issues and improve AI outputs.
Our partner is a company focused on improving AI systems through quality evaluation. They offer a flexible, remote work environment for independent contractors.
Dive into messy business processes to find critical insights and build fact bases for large enterprise accounts.
Turn analysis into decision-ready slides, one-pagers, and memos with same-day or next-day turnaround.
Act as surge analytical capacity across multiple accounts and problem types, supporting Transformation Strategists and working-level customer teams.
Duvo builds the Enterprise Execution OS, a shared layer that carries work across existing systems and measures business results. The company is a small, early-stage team with direct access to founders and customers, emphasizing ownership, speed, and AI-native workflows.