Source Job

Global

  • You will rate and assess the performance of AI models based on their output or behavior.
  • You will label elements of content and assign predefined categories to generate training data.
  • You will create prompts, summaries, and evaluate relevance to improve AI system understanding.

English Proficiency Data Annotation AI Evaluation Analytical Thinking

20 jobs similar to Generative AI Associate (PhD)

Jobs ranked by similarity.

Canada

  • Evaluate and assess AI model outputs based on predefined quality, accuracy, relevance, and behavioral guidelines.
  • Annotate, classify, and label text, images, or audio to support AI model training.
  • Create prompts and generate high-quality responses to improve language model reasoning capabilities.

Jobgether uses an AI-powered matching process to connect candidates with hiring companies. They focus on efficient, fair recruitment and handle data privacy in compliance with GDPR.

UK

  • Compare and rank AI-generated responses for accuracy, logic, and safety.
  • Review CS research papers alongside AI summaries to ensure scientific integrity.
  • Fact-check technical data and code for logical flaws and inaccuracies.

Prolific is building the largest pool of quality human data in the world, serving over 35,000 AI developers and researchers. They connect researchers with paid participants to gather high-quality, ethically sourced behavioral data for AI development.

Global

  • Research data collection strategies and design high-impact data slices that uncover model failure modes.
  • Model annotator behavior and design experiments to optimize instruction clarity and reward signal reliability.
  • Develop metrics and frameworks for evaluating dataset quality, diversity, and impact on downstream model alignment.

Surge AI builds a platform that powers the most powerful AI models in partnership with companies like Anthropic, Google, Microsoft, and Meta. They are a profitable, bootstrapped company focused on human intelligence and data quality.

  • Evaluate prompts and AI-generated outputs for accuracy, cultural appropriateness, and brand alignment.
  • Review and correct text, analyze multimedia content, and contribute voice recordings.
  • Apply local cultural insight and consistent evaluation guidelines to ensure high-quality AI training.

Lilt provides multilingual AI and human-verified services to enterprises, governments, and AI developers. They foster a global community of linguists and subject matter experts working on cutting-edge AI and language technology.

United States Latin America

  • Own the design and defense of frontier model evaluations across reasoning, coding, agents, tool use, and multi-modal.
  • Build benchmark packages with expert-verified ground truth, multi-model headroom results, and rigorous QC.
  • Recruit, calibrate, and review a pool of subject-matter experts in coding, agentic/tool-use, and STEM/reasoning.

Anyone AI measures frontier model capability through expert-verified evaluation packages. The company operates as a remote team with a focus on rigorous benchmarking and lab collaboration.

$4–$5/hr
Global

  • Review AI-generated responses against source images and quality guidelines.
  • Identify issues like hallucinations, missing details, or policy violations.
  • Provide structured feedback to improve model performance and output quality.

Jobgether uses AI-powered matching to connect candidates with partner companies. They focus on efficient, objective hiring processes and operate as a platform for remote opportunities.

$22–$22/hr
Canada

  • Evaluate and rank model outputs, stress-test models for failure modes, and create high-quality datasets with detailed rubrics.
  • Annotate and correct multimodal data, maintain consistency through calibration exercises, and adapt to evolving task types.
  • Report on model performance trends and provide clear feedback to cross-functional partners on model successes and failures.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for real-world business problems. It is a global technology company with offices in Toronto, San Francisco, London, New York, Montreal, Seoul, Germany, and Paris, staffed by a team of passionate researchers, engineers, and designers.

$80–$150/hr
UK

  • Review AI-generated responses to clinical scenarios for accuracy and safety.
  • Compare and justify the best responses among multiple model answers.
  • Write improved exemplars and structured feedback to enhance AI model learning.

Prolific is building the largest pool of quality human data in the world, used by over 35,000 AI developers and researchers. It is a platform that connects researchers with a global participant pool for ethically sourced human data.

Global

  • Contribute to AI training through annotation, evaluation, and prompt creation tasks.
  • Work flexibly on remote projects that match your skills and availability.
  • Be part of a global community shaping safer, smarter AI.

Welo Data, part of Welocalize, is a global AI data company with 500,000+ contributors delivering high-quality, ethical data to train the world’s most advanced AI systems. We build smarter AI through the power of human contribution, offering limitless opportunities for our global community to grow, contribute, and work on their terms.

$120,000–$170,000/yr
US

  • Design, build, and maintain automated AI evaluation pipelines for production LLM applications.
  • Develop prompt engineering strategies and evaluate model performance using quantitative methods.
  • Analyze production AI behavior with Python, SQL, and statistical techniques to identify improvement opportunities.

GovWorx provides an AI-powered platform, CommsCoach, that supports 9-1-1 and emergency communications centers by automating quality assurance, training, and real-time call evaluation. The company is a growing technology team focused on public safety, collaborating across AI, engineering, product, and data science.

Europe Middle East

  • Build and evolve the core AI tutoring system with prompt architectures and agentic workflows.
  • Design and implement scalable software integrating AI with platform systems.
  • Collaborate cross-functionally with educators and product teams to translate pedagogical goals into technical solutions.

DataCamp empowers everyone with data and AI skills through practical learning experiences. It serves over 17 million learners and 6,000+ companies, including 80% of Fortune 1000, fostering a culture of data-driven decision-making and transparency.

$4,000–$4,500/mo
Poland

  • Design and iterate complex system prompts and chain-of-thought structures for consumer AI experiences.
  • Partner with engineers and product managers to optimize prompt specifications for latency and cost.
  • Develop evaluation frameworks and playbooks to guardrail LLM outputs against bias and hallucination.

BOLD is a global company that creates digital products to help people build resumes, cover letters, and CVs, empowering job seekers in 180 countries. They are an established organization that values diversity and inclusion, with a culture of growth and professional fulfillment.

Global

  • Evaluate AI-generated text and voice snippets in Punjabi for naturalness and authenticity.
  • Assess audio clips for cultural and tonal accuracy of AI speech.
  • Provide feedback on linguistic nuance and quality of AI outputs.

Prolific builds the biggest pool of quality human data in the world, serving over 35,000 AI developers and researchers. The company connects researchers with paid study participants from diverse backgrounds to gather high-quality, ethically sourced behavioral data.

$115,000–$200,000/yr
Global Unlimited PTO 17w maternity 17w paternity

  • Create and curate an evaluation suite of real-world tasks for frontier AI models.
  • Rigorously evaluate AI systems, analyze results, and communicate findings.
  • Improve evaluation processes and potentially build out standalone benchmarks.

Epoch AI is a research institute that investigates trends in machine learning and the economic consequences of AI. Our mission is to develop a comprehensive, publicly accessible knowledge base on AI that informs policymakers, industry leaders, and society at large.

Global

  • Train and evaluate cutting-edge AI models by completing language tasks in Italian/Spanish/French/German/Dutch.
  • Judge the performance of AI in performing Italian prompts and improve its capabilities.
  • Analyze, edit, and write in target languages with strong attention to detail for up to one hour per task.

Prolific is building the largest pool of quality human data in the world for AI training. Over 35,000 AI developers, researchers, and organizations use the platform to gather data from paid participants with diverse experiences, skills, and knowledge.

Global

  • Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
  • Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
  • Benchmark performance by conducting comparative testing between model outputs based on technical metrics.

Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.

$75–$100/hr
United States

  • Review domain archives to understand subject matter and extract key facts.
  • Create accurate question-and-answer pairs covering various complexity levels.
  • Ensure answers are traceable, unambiguous, and consistent with approved source content.

Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With over 36 years of experience, the company delivers high-quality data and outcomes for Generative AI builders.

Latin America

  • Identify and evaluate high-impact AI use cases across the enterprise, partnering with business stakeholders to prioritize opportunities.
  • Drive adoption of generative AI and agentic technologies through training, coaching, and enablement programs.
  • Configure and support AI-powered workflows and solutions, ensuring governance and compliance with enterprise standards.

South Geeks connects elite, AI-fluent engineers from Latin America with future-shaping companies in Health and Wealth. They offer fully remote, long-term engagements with continuous training and development.

US

  • Evaluate AI quality across the advisor stack, including pre-call briefs, in-call guidance, and post-call outputs.
  • Iterate inside ORA by refining prompts, updating knowledge base entries, and tweaking skills to close the loop on issues.
  • Surface trends and drive continuous improvement by tagging conversations, logging issues, and recommending prioritized improvements.

HighLevel is an AI-powered business operating system that gives agencies, entrepreneurs and SMBs the infrastructure to build, automate and scale. With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership.

$5–$5/hr
Global

  • Perform data collection, evaluation, and annotation for AI training.
  • Conduct pairwise comparisons and counting tasks.
  • Tag and label objects across audio, video, images, or collected data.

RWS provides AI training data services. They are a global company with a focus on diversity, equity, and inclusion, offering flexible remote work opportunities.