Similar Jobs
See allSoftware Engineers: Create and Validate AI Coding Tasks
Terac
Global
Software Engineering
GitHub
Docker
Software Engineering Expert (AI Training)
Undisclosed
Global
Python
Java
C++
AI Engineer – Marketing & GTM Systems
Rwazi, Inc.
US
Python
TypeScript
API Integrations
Senior Full-Stack Engineers (React): AI Evaluation Environments
Jobgether
US
React
Full-Stack Development
Automated Testing
Machine Learning Engineer, Evals
Nous Research
Global
Python
LLM Evaluation
Docker
Primary Responsibilities:
- Develop difficult, novel tasks for models, ensuring they remain challenging as time horizons expand.
- Perform quality assurance on existing tasks, verifying solvability and proper information provision.
- Baseline and score tasks to support evaluation methodology and improve processes.
Skills and Experience:
- Several years of software engineering experience with complex projects and codebases.
- Experience building challenging AI evaluations, ideally agent-based, using frameworks like Inspect.
- High attention to detail for spotting misspecifications and ensuring precision.
Work Details:
- Remote worldwide with flexible 20-40 hours per week and at least 1 hour overlap with Pacific Coast Time.
- Contract/freelance employment with compensation of $150-300/hour, top range for exceptional candidates.
METR
METR is a nonprofit research organization developing scientific methods to assess AI capabilities, risks, and mitigations, focusing on catastrophic AI risk evaluations. It is a mission-driven, tight-knit team with a low-ego, collaborative culture committed to high-quality, trustworthy science.