Source Job

United States

  • Evaluate AI model responses across diverse topics and use cases against established quality standards.
  • Write, test, and refine system prompts to influence model behavior and performance.
  • Identify and document model failures and provide clear, actionable feedback to product teams.

Writing Editing Prompt Engineering Fact Checking

20 jobs similar to Content Specialist III | AI Evaluation & Prompting

Jobs ranked by similarity.

US

  • Test and evaluate AI model responses across diverse topics and conversation types.
  • Write and refine system prompts to shape model behavior and personality.
  • Apply detailed rubrics consistently to assess model performance and document findings.

We are a team dedicated to evaluating and improving advanced AI models and products. Our fast-moving environment emphasizes quality, independent judgment, and attention to detail.

$250,000–$300,000/yr
US Canada Unlimited PTO

  • Own the build out of new agents, skills, and platform capability for teams across TLDR.
  • Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
  • Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.

TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.

US

  • Define end-to-end requirements for AI capabilities, from model behavior to user experience.
  • Translate model capabilities and technical constraints into product decisions.
  • Work closely with ML and engineering teams on system design and iteration.

They develop AI-powered email applications and work at the intersection of user needs and model capability. The company is fast-moving and collaborative with a strong focus on technical ownership and innovation.

Australia

  • Evaluate AI-generated responses for relevance, accuracy, and personalization using personalized prompts and data from connected Google applications.
  • Identify subtle issues such as incorrect assumptions, irrelevant recommendations, inconsistencies, and inappropriate personalization.
  • Provide clear, detailed, and structured feedback to support improvements to AI models and personalization systems.

Our partner company is seeking an AI Response Quality Evaluator to improve AI-generated responses. This is a project-based contract role with a remote, independent working environment and a duration of up to 16 weeks.

India

  • Analyze, evaluate, and review diverse datasets to support AI system training and improvement.
  • Assess AI-generated content for accuracy, relevance, consistency, and quality, providing actionable feedback.
  • Work independently in a remote, digital-first environment, managing multiple tasks and deadlines.

US

  • Partner with product managers, engineers, and operations leads to build AI-powered tools.
  • Design and operate an evaluation framework to ensure AI tools meet quality standards.
  • Own the repository health, governance model, and responsible AI practices.

Maleda Tech is a technology staffing firm that connects professionals with contract opportunities. The company culture is not described in the posting.

US

  • Review datasets and task outputs to assess accuracy, completeness, and consistency.
  • Apply detailed evaluation rubrics and identify errors, inconsistencies, and quality issues.
  • Provide clear feedback and help improve data quality and AI evaluation processes.

Our partner is a company focused on data analysis and AI evaluation, supporting the development of next-generation AI systems. This remote contractor role offers an opportunity to apply your analytical expertise to emerging AI projects.

US

  • Evaluate software engineering tasks for technical accuracy, realism, and reproducibility.
  • Investigate codebases, tests, and integration issues to identify technical weaknesses.
  • Provide clear, actionable feedback that directly improves AI training and evaluation workflows.

Jobgether is an AI-powered job platform that connects candidates to roles through objective, skill-based matching. It focuses on remote and freelance opportunities, with a data-driven recruitment process and a global candidate pool.

US

  • Evaluate AI model performance through real-time, voice-based conversations by roleplaying assigned scenarios with two different models.
  • Compare model responses across five defined dimensions, identify error clusters, and select the stronger performer with a detailed rationale.
  • Maintain consistent conversational turns and voice recording to ensure fair, objective comparisons.

Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, they deliver high-quality data and outcomes for AI builders and adopters.

Global

  • Evaluate AI-generated responses for safety and bias against strict rubrics.
  • Classify harmful content categories like hate speech and self-harm.
  • Verify factual accuracy of model claims using external sources to prevent hallucinations.

TELUS Digital is a leading AI data company that trains models to be safe and accurate. They have a global community of over one million contributors and foster a collaborative, flexible culture.

US

  • Collaborate with client teams to diagnose operational bottlenecks and develop testable hypotheses.
  • Build and test AI-powered prototypes using coding agents like Claude Code or Cursor.
  • Drive adoption by working directly with users and iterating on solutions until they are effective in live workflows.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The platform uses AI to review applications and provide a shortlist to employers, emphasizing efficiency and objectivity.

US

  • Identify, scope, and prioritize AI and automation opportunities based on business impact, feasibility, and adoption potential.
  • Design, build, test, and deploy practical AI-enabled solutions using approved tools and platforms.
  • Develop and deliver AI training programs to raise AI fluency across the organization.

Upstream USA is a national nonprofit dedicated to ensuring equitable, patient-centered contraceptive care is available to everyone. Founded in 2014, the organization operates in 37 states and Washington, D.C., and is supported entirely by venture philanthropy with a culture of aggressive growth and a focus on measurable outcomes.

$15–$15/hr
Global

  • Evaluate and label AI model outputs to improve performance and alignment with project guidelines.
  • Create prompts, rewrite text, and generate training data for large language models.
  • Work on flexible, remote, project-based tasks while helping shape the future of AI.

Innodata is a global data engineering company that enables the responsible advancement of AI by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company delivers high-quality data and outstanding outcomes for customers.

US Canada

  • Curate high-quality code examples and datasets for LLM training and evaluation.
  • Develop and assess AI-generated software across multiple programming languages and the full SDLC.
  • Collaborate with research teams to design verification mechanisms and improve coding benchmarks.

This partner company specializes in evaluating large language models and improving AI systems through rigorous engineering benchmarks. It offers a remote, collaborative culture where engineers and researchers advance AI evaluation workflows together.

Global

  • Review text or media samples based on provided project guidelines
  • Apply accurate labels and categorizations to diverse data sets
  • Evaluate AI-generated responses for clarity, safety, and factual accuracy

Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.

  • Assess the clarity, coherence, and accuracy of written content to ensure it meets project standards.
  • Conduct detailed writing evaluations and provide constructive feedback for continual improvement.
  • Identify and annotate AI-generated content, focusing on detecting low-quality or artificial text elements.

Our client is a rapidly growing, venture-backed AI company building intelligent systems with human expertise and machine learning workflows. Backed by more than $40 million in funding, the company connects a global network of experts to high-impact AI projects.

$104,000–$170,400/yr
Canada

  • Lead a team producing high-quality training data and evaluations for advanced AI models.
  • Review engineering-domain work, uphold quality standards, and drive operational improvements.
  • Coach and develop AI Tutors while collaborating with engineering and human data teams.

Jobgether is an AI-powered job platform that connects candidates with partner companies using objective matching. This role is with a partner company focused on AI training data, offering a collaborative, remote-first environment emphasizing autonomy and continuous learning.

Brazil

  • Evaluate AI-generated responses for logical consistency and real-world business applicability.
  • Analyze business scenarios across operations, management, strategy, and entrepreneurship.
  • Provide structured feedback to improve model prompts, evaluation frameworks, and reasoning quality.

A partner company is hiring a Business and Management Specialist to train AI systems in business reasoning. The freelance role is fully remote and flexible, suitable for both experienced professionals and entry-level candidates.

Europe

  • Annotate and label text, images, audio, or other content for AI training projects.
  • Evaluate AI-generated content for quality, relevance, accuracy, and overall usefulness.
  • Create, review, and test prompts to improve large language model performance.

AI Trainers Network is a global AI contributor network that shapes smarter and more human-centered AI through language and cultural expertise. It connects a diverse international community of independent contributors working on flexible, remote AI training projects.

$100–$150/hr
Canada

  • Evaluate AI-generated slides, spreadsheets, and documents for real-world usability and professional quality.
  • Assess outputs for accuracy, clarity, relevance, and alignment with data science standards.
  • Provide structured written feedback to help improve AI systems and their outputs.

A partner company is seeking a Data Science Expert to evaluate AI-generated work. The company focuses on improving AI systems and operates with a flexible, remote team.