Source Job

Americas

  • Design AI evaluations: define ground truth, metrics, and scoring methods for models and agents.
  • Build repeatable eval loops to track quality over time and catch regressions before release.
  • Translate eval results into actionable recommendations for system improvements.

Python SQL Machine Learning Statistics Model Evaluation

20 jobs similar to Senior Data Scientist, AI Evaluation

Jobs ranked by similarity.

Global

  • Design end-to-end LLM evaluation plans for business scenarios such as dialogue and financial trading.
  • Build evaluation metric systems and rubrics to quantify model performance and drive improvement.
  • Lead construction of evaluation datasets, define dimensions, and ensure high-quality annotation standards.

Binance is a leading global blockchain ecosystem behind the world's largest cryptocurrency exchange by trading volume and registered users. Trusted by 300+ million people in 100+ countries, we offer a diverse range of digital-asset products and services.

Brazil

  • Own model strategy and selection using rigorous benchmarks and statistical analysis.
  • Design and maintain evaluation methodologies for AI systems, including offline sets and LLM-as-judge frameworks.
  • Develop classification, fine-tuned, and agentic AI models to improve accuracy, cost, and latency.

A production agentic AI platform that builds and deploys advanced machine learning systems. The team is collaborative and values innovation, offering a remote work environment with high autonomy.

US

  • Set the technical direction for AI engineering across the team: agent architecture patterns, evaluation methodology, deployment and monitoring strategies
  • Design the AI platform layer, including shared agent frameworks, tool integrations, and evaluation infrastructure
  • Work directly with clients on the most complex engagements, identifying new problem domains and ensuring production quality

Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. With over 1,500 firms in 60 countries managing nearly $10 trillion in assets, Addepar fosters a culture of ownership, collaboration, and innovation.

$145,000–$175,000/yr
US

  • Design and execute technical training programs for AI data and annotation teams, including curricula and certifications.
  • Own quality frameworks, monitor KPIs like accuracy and defect rates, and drive continuous improvement.
  • Provide technical guidance for AI data projects such as LLM evaluation, RLHF, and model benchmarking.

Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, they are committed to delivering the highest quality data and outstanding outcomes for clients.

$160,000–$185,000/yr
United States

  • Conduct independent research in Generative AI, LLMs, NLP, and multimodal AI to design experiments and evaluate models.
  • Develop and implement LLM evaluation frameworks, analyze model performance, and identify data gaps for improvement.
  • Apply strong statistical and data science skills to clean, analyze, and interpret complex datasets for AI/ML research.

Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company is committed to delivering the highest quality data and outstanding outcomes for its customers.

US

  • Review datasets and task outputs to assess accuracy, completeness, and consistency.
  • Apply detailed evaluation rubrics and identify errors, inconsistencies, and quality issues.
  • Provide clear feedback and help improve data quality and AI evaluation processes.

Our partner is a company focused on data analysis and AI evaluation, supporting the development of next-generation AI systems. This remote contractor role offers an opportunity to apply your analytical expertise to emerging AI projects.

$150,000–$165,000/yr
US

  • Find and ship predictive signals by building student-level predictors and analyzing conversational data to improve AI.
  • Own AI data tooling and evaluation, including building frameworks for model assessment and maintaining in-warehouse AI configuration.
  • Make your work reusable by equipping internal teams with data and analysis, and document reasoning and assumptions.

Mainstay, a division of Lemnis, is a public charity dedicated to expanding learning for all through an engagement platform that helps colleges and businesses start and measure conversations. The team is growing, collaborative, and inclusive, with a focus on innovation and mentorship.

$104,000–$170,400/yr
Canada

  • Lead a team producing high-quality training data and evaluations for advanced AI models.
  • Review engineering-domain work, uphold quality standards, and drive operational improvements.
  • Coach and develop AI Tutors while collaborating with engineering and human data teams.

Jobgether is an AI-powered job platform that connects candidates with partner companies using objective matching. This role is with a partner company focused on AI training data, offering a collaborative, remote-first environment emphasizing autonomy and continuous learning.

UK

  • Build tooling for capturing and processing data from agents and humans at significant scale.
  • Solve hard problems around compute, orchestration, scaling, security, and reliability.
  • Help develop approaches for training, benchmarking, and evaluating AI agents.

Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.

US

  • Test and evaluate AI model responses across diverse topics and conversation types.
  • Write and refine system prompts to shape model behavior and personality.
  • Apply detailed rubrics consistently to assess model performance and document findings.

We are a team dedicated to evaluating and improving advanced AI models and products. Our fast-moving environment emphasizes quality, independent judgment, and attention to detail.

$103,000–$117,000/yr
Canada Unlimited PTO

  • Design, build, and deploy LLM-powered product features, including lab summaries and conversational agents.
  • Build backend services integrating LLMs and ML models, primarily using Python with exposure to Elixir.
  • Implement evaluation, monitoring, and CI/CD workflows for AI features, ensuring reliability and clinical relevance.

Fullscript is a health technology platform that helps practitioners deliver better care through clinical insights, lab interpretations, and patient analytics. With over 125,000 practitioners and 10 million patients, the company emphasizes a people-first culture, teamwork, and continuous learning in a remote-first environment.

North America Unlimited PTO

  • Study how engineers and customers use Archie in the field, identify where it succeeds or fails, and turn observations into actionable evaluations.
  • Recreate real-world engineering tasks and failure modes in repeatable environments for development teams.
  • Build and refine evaluation methodologies, including automated judges and human evaluation processes, to align with expert judgment.

P-1 AI is building Archie, an AI engineer agent for the physical world that works alongside human engineering teams. The company recently raised a $50 million Series A led by NEA and is driven by the mission of building superintelligence for engineering.

$30–$30/hr
Global 0w PTO

  • Design and maintain scalable people-data models and architecture within the enterprise data warehouse.
  • Establish automated data-quality monitoring, validation rules, and end-to-end lineage controls.
  • Own the AI-readiness roadmap and drive technical infrastructure investments to enable approved AI use cases.

They are a multinational technology consulting firm that helps businesses leverage technology for success. Founded in 2016, the company has Swiss roots and a development team in Latin America, blending Latin American talent with Swiss organizational culture.

US Unlimited PTO

  • Lead the engineering strategy and execution for evaluations of AI agents, owning the core evaluation platform.
  • Design scalable evaluation infrastructure, APIs, workflows, and production systems across software categories.
  • Mentor and develop a team of engineers as the technical authority on agentic evaluation.

This company focuses on building credible, scalable evaluations of AI agents from software vendors. It operates as a fully remote, inclusive team with a flexible culture and a focus on professional growth.

India

  • Own analytics instrumentation and measurement for AI products including Voice AI, Conversation AI, AI Employee, and Ask AI.
  • Define success metrics and KPIs, and partner on experiments for non-deterministic AI features.
  • Influence product roadmaps with data-driven recommendations and mentor analysts as the team scales.

HighLevel is an AI-powered business operating system that helps agencies, entrepreneurs, and SMBs build, automate, and scale their operations. With over 2,000 team members across 10+ countries, it operates as a global remote-first organization valuing initiative, clarity, and execution.

Canada Unlimited PTO

  • Write and ship production AI code daily as an active contributor.
  • Build the retrieval, graph, and inference layers on top of our data platform.
  • Set patterns other teams build against and turn company goals into shipped systems.

Acquia empowers brands to create digital customer experiences using its Drupal-based Digital Experience Platform. It is a Great Place to Work-Certified company with thousands of global organizations as clients.

AI Strategist

WON
Latin America

  • Drive AI strategy for retail and consumer brands, leading client relationships and program execution.
  • Scope business cases, specifications, and deliver measurable outcomes through AI-enabled systems.
  • Act as the primary point of contact, directing specialists and ensuring adoption of built solutions.

WON redesigns business operations and customer experiences for retail, travel, hospitality, financial services, and private equity clients. The firm is built around senior practitioners who work directly with clients, with an emphasis on AI leverage and accountability for results.

India

  • Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
  • Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
  • Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.

Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.

$239,300–$280,000/yr
US

  • Define and own the technical strategy for Octave's data platform and AI/ML capabilities, partnering with product and business leaders.
  • Lead and mentor a data platform team while remaining hands-on in architecture, prototyping, and complex engineering problems.
  • Balance near-term delivery with foundational investments to scale data and AI initiatives, driving KPIs and technical excellence.

Octave is a modern behavioral health practice offering evidence-based individual, couples, and family therapy in-person and virtually. They focus on quality care and payer partnerships to make therapy affordable, with a growing team and a culture of empathy and collaboration.

$104,000–$170,400/yr
Global

  • Lead a team of AI Tutors to deliver high-quality training data and evaluations for SpaceXAI's models.
  • Own end-to-end quality and delivery for Human Data projects, reviewing work and ensuring consistency.
  • Coach, performance-manage, and develop talent while improving operational processes.

SpaceXAI creates AI systems to understand the universe and aid humanity. The team is small, highly motivated, and focused on engineering excellence with a flat structure.