Similar Jobs
See allAI Evaluators: Assessing A Shopping Assistant
Partner Company
US
Data Evaluation
Quality Assurance
Critical Thinking
Domain Advisor Consultant (Subject Matter Expert)
Partner Company
US
Analytical Reasoning
Written Communication
Independent Work
Document / Deck Production QA Evaluator
Jobgether
Canada
Quality Assurance
Microsoft Office
Google Workspace
Designing Challenging AI Prompts
Terac
Prompt Engineering
Workflow Analysis
AI Evaluation
Generative AI Analyst
Welo Data
US
Generative AI
Data Annotation
English
What We're Researching:
- We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant.
- This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic.
- Your analysis will directly feed into improving the underlying model and its response quality.
How It Works:
- You will review real interaction traces between users and the shopping assistant within our custom platform.
- As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations.
- From there, you will create structured rubrics and verifiers to consistently judge future response quality.
Who This Is For:
- This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail.
- We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing.
- You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.
Terac
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.