Remote Other Jobs · Critical Thinking

Job listings

  • Evaluate AI model performance through real-time, voice-based conversations by roleplaying assigned scenarios with two different models.
  • Compare model responses across five defined dimensions, identify error clusters, and select the stronger performer with a detailed rationale.
  • Maintain consistent conversational turns and voice recording to ensure fair, objective comparisons.

Innodata is a global data engineering company that enables responsible AI advancement by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, they deliver high-quality data and outcomes for AI builders and adopters.

$7,500–$7,500/yr

  • Conduct research to analyze policies that hinder Illinoisans from thriving.
  • Help craft free-market solutions to bring more liberty and opportunity.
  • Collaborate with a high-performing team and attend KIP programming.

The Illinois Policy Institute is a non-profit organization dedicated to protecting civil and personal liberties, and promoting free-market solutions for Illinois. They are a small, high-performing team focused on impactful research and policy analysis, offering a paid internship through the Koch Internship Program.

$41,200–$66,000/hr

  • Supports special projects and initiatives across various assignments.
  • Gains real-world experience through a 12-week paid internship.
  • Participates in leadership seminars and feedback sessions.

CareSource is a mission-driven health insurance company focused on improving member well-being. With a large employee base, they foster an inclusive, collaborative culture with flexible work and growth opportunities.

  • Evaluate AI-generated work products in real estate, hospitality, and events using quality rubrics.
  • Identify factual, aesthetic, and presentation errors and provide actionable feedback.
  • Apply industry expertise to distinguish realistic, commercially sound work from generic AI content.

The company develops AI systems and evaluates their outputs for quality. They seek experienced industry professionals for flexible remote contract work.

  • Translate a complex workflow into a demanding AI prompt designed to expose model limitations.
  • Test your prompt in ChatGPT, refine it until the AI fails, and write a grading rubric for others to use.
  • Submit your prompt, failure notes, rubric, and a screen recording of your thought process.

Terac builds the world's largest pool of vetted human experts for AI research and evaluation. They are a growing platform used by AI labs and researchers to recruit, screen, and pay study participants globally.