Source Job

US

  • Evaluate AI-generated legal research and analysis for accuracy, relevance, and completeness.
  • Verify legal citations, authorities, and reasoning to identify errors and weaknesses.
  • Develop objective evaluation criteria and provide structured feedback to improve AI legal content.

Legal Research Fact-Checking Legal Writing Analytical Reasoning Attention To Detail

20 jobs similar to Legal Domain Expert (SME) – AI Model Evaluation

Jobs ranked by similarity.

Canada

  • Evaluate AI-generated legal and business documents against quality standards and apply professional judgment.
  • Review contracts, diligence materials, redlines for accuracy, consistency, and completeness.
  • Provide clear, structured feedback to improve AI-generated legal content.

US

  • Conduct comprehensive legal analysis on US law topics.
  • Perform in-depth legal research and draft legal content.
  • Participate in remote multidisciplinary collaboration for quality assurance.

The company is a venture-backed AI firm that combines human expertise with machine learning to improve AI models. It has over $40 million in funding and a growing international network of experts.

US

  • Help the team understand real legal workflows and provide domain expertise on commercial agreements.
  • Review AI-generated outputs and develop evaluation criteria for legal reasoning.
  • Collaborate with product and engineering in structured working sessions to shape the AI product.

Brain Co. builds AI-native operating systems for large, regulated institutions. The team consists of elite engineers from top tech companies, with a focus on shipping production AI applications.

US

  • Conduct comprehensive legal analysis across U.S. law topics, applying nuanced reasoning to complex scenarios.
  • Perform detailed legal research, draft and edit legal materials including memoranda and contracts.
  • Collaborate remotely with multidisciplinary teams, evaluate legal content, and support quality assurance.

This company partners with organizations to develop next-generation AI systems by integrating sophisticated legal expertise. It operates with a remote-first culture and engages multidisciplinary teams to deliver high-quality, innovative solutions.

Canada

  • Evaluate AI-generated documents, spreadsheets, and presentation decks against quality rubrics.
  • Identify factual, formatting, visual, and structural issues in professional deliverables.
  • Provide clear, structured feedback to enhance AI output quality and consistency.

This partner company specializes in AI training and evaluation, focusing on improving the quality of AI-generated professional content. Operating as a remote and asynchronous team, they value precision, collaboration, and independent work.

US

  • Combine top-market legal skills with AI-native workflows for drafting, review, and negotiation.
  • Work without billable hours, with modern tooling and a focus on judgment and client issues.
  • Collaborate directly with business and product teams to build trust and deliver clear recommendations.

General Legal is an AI-native law firm delivering high-quality legal work faster and more predictably. We are a startup with a remote-friendly culture and meaningful ownership for our attorneys.

India

  • Evaluate AI-generated documents, spreadsheets, and presentations for privacy and regulatory compliance accuracy.
  • Apply domain expertise to assess outputs against defined quality standards and provide actionable feedback.
  • Contribute to improving AI systems by providing expert judgment and structured feedback.

The company is a partner of Jobgether, offering remote opportunities for privacy and compliance professionals to evaluate AI-generated content. The company values autonomy and flexibility, providing independent contractor arrangements.

Canada

  • Apply deep subject-matter expertise to AI model evaluation and large language model projects.
  • Develop challenging domain-specific problems and assess AI responses for accuracy and reasoning.
  • Collaborate with AI research teams to improve training datasets and evaluation methodologies.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It offers a remote, asynchronous work culture and uses AI tools to support recruitment.

  • Review legal scenarios for accuracy and realism.
  • Construct domain-specific materials based on provided prompts.
  • Provide structured feedback on legal reasoning workflows.

Terac is building the world's largest pool of vetted human experts for AI. They work with researchers, AI labs, and product teams to recruit, screen, and pay study participants across industries.

US

  • Evaluate AI-generated content related to job descriptions, compensation benchmarking, and employee policy interpretation.
  • Create realistic HR workflow scenarios such as performance management cycles, benefits enrollment, and employee investigations.
  • Annotate and label data across offer letter generation, leave administration, and workforce analytics.

Prolific builds the largest pool of quality human data in the world, used by over 35,000 AI developers, researchers, and organizations. The platform connects researchers with paid participants to gather ethically sourced behavioral data, placing itself at the forefront of AI innovation.

Global

  • Perform simulated contract negotiations and redlining exercises.
  • Review and assess AI responses to contract scenarios, providing expert feedback.
  • Collaborate with product and research teams to refine AI-driven contract review solutions.

The company combines world-class human expertise with advanced machine learning workflows to help leading AI organizations build and improve cutting-edge models. Backed by over $40 million in funding, they have a rapidly expanding international network of experts.

US

  • Evaluate AI-generated coding interactions end to end for usefulness, accuracy, and consistency with strong engineering practices.
  • Assess whether coding agents demonstrate sound technical reasoning and practical engineering judgment rather than just producing working-looking code.
  • Provide actionable feedback, distinguishing between adequate and exceptional AI response quality to shape evaluation standards.

Jobgether uses an AI-powered matching process to ensure your application is reviewed quickly and fairly against the role's core requirements. They are a third-party recruitment platform that partners with companies to fill positions, with a streamlined selection process.

$150–$300/hr
Global

  • Develop difficult, novel tasks for models that challenge growing time horizons.
  • Conduct quality assurance to ensure tasks are solvable and appropriately scoped.
  • Baseline and score tasks within your domain of expertise for AI or human performance.

METR is a nonprofit research organization developing scientific methods to assess AI capabilities, risks, and mitigations, focusing on catastrophic AI risk evaluations. It is a mission-driven, tight-knit team with a low-ego, collaborative culture committed to high-quality, trustworthy science.

India

  • Review Indian tax research questions and AI-generated answers for technical accuracy and completeness.
  • Validate interpretations of tax legislation and ensure support by statutory and judicial authorities.
  • Provide structured technical feedback to improve AI model performance and refine evaluation standards.

This partner company focuses on developing advanced AI systems for tax analysis. It offers a remote, asynchronous work environment with a short-term contract, seeking experienced tax professionals.

$20–$20/hr
US

  • Read and analyze published judicial opinions across state and federal jurisdictions.
  • Review AI system outputs and verify accuracy against source legal materials.
  • Collaborate with legal researchers and ML engineers to improve data quality.

Filevine is a Legal AI company that delivers legal operating intelligence through a unified platform. It is a rapidly growing company recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.

  • Dive into a cutting-edge AI benchmarking project focused on highly specific professional domains like software engineering, healthcare, and finance.
  • Design realistic scenarios in English or Korean, adapt evaluation rubrics, and review AI responses for accuracy and cultural appropriateness.
  • Contribute to gold-standard solutions that reflect best practices across your target locale and domain.

Lilt is a company that makes the world's information available to everyone, regardless of language, through multilingual AI and human-verified services for Enterprises, Governments, and AI Developers. The company has a global community of linguists and subject matter experts who thrive on innovation and excellence.

US

  • Review, evaluate, and annotate AI-generated content across text, images, audio, and video.
  • Perform quality checks to ensure accuracy, consistency, and compliance with project guidelines.
  • Identify edge cases and inconsistencies, contribute to high-quality dataset development, and participate in calibration activities.

Welo Data, part of Welocalize, is a global AI data company with over 500,000 contributors that provides high-quality, ethical data for training advanced AI systems. The company supports a diverse, global community across 100+ countries and offers project-based freelance opportunities with flexibility and growth potential.

Global

  • Evaluate model-generated content across multiple modalities including text, images, audio, and video.
  • Apply defined quality rubrics such as factuality, consistency, and aesthetics to assess outputs.
  • Conduct independent research on unfamiliar topics to make well-supported evaluation judgments.

Our client is a global technology company that helps businesses build, train, and manage AI systems. They offer flexible, remote work with variable workload and a focus on high-quality model evaluation.

US

  • Practice labor and employment law at a high level from a fully remote environment.
  • Benefit from flexible billable-hour targets and reduced business-development pressure.
  • Handle employment litigation and counseling for a national client base.

A leading national law firm is expanding its Labor & Employment practice. The firm offers a remote work environment with flexible billable-hour targets and reduced business-development pressure.

$150–$220/hr
India

  • Design and apply evaluation criteria for consulting deliverables such as market analyses and financial models.
  • Assess AI-generated and human work, providing evidence-based scores and justifications.
  • Work independently in a remote, asynchronous environment to improve AI model reasoning.

They are an AI-focused organization improving the quality of AI outputs through expert evaluation. The work is fully remote and asynchronous, with an emphasis on independent problem-solving and collaboration with senior reviewers.