Source Job

$150–$300/hr
Global

  • Develop difficult, novel tasks for models that challenge growing time horizons.
  • Conduct quality assurance to ensure tasks are solvable and appropriately scoped.
  • Baseline and score tasks within your domain of expertise for AI or human performance.

Software Engineering Attention To Detail Python

20 jobs similar to Task Development Engineer

Jobs ranked by similarity.

Global

  • Design self-contained software-engineering problems for AI agents to solve.
  • Build and configure development environments using Docker, and write automated tests.
  • Create accurate reference solutions for complex coding challenges.

Terac is building the world's largest pool of vetted human experts for AI. It is a growing platform that connects AI researchers and labs with skilled professionals across industries.

Global

  • Design and develop software engineering challenges and reference solutions for AI training.
  • Apply strong programming skills across languages including Python, Java, Rust, Go, C++, and TypeScript.
  • Collaborate in a fully remote environment to improve AI reasoning and coding capabilities.

This organization focuses on advancing AI technologies by training next-generation AI systems through expert software engineering contributions. It operates as a remote-first team dedicated to improving AI reasoning and coding capabilities in a flexible, collaborative environment.

US

  • Ship AI-driven GTM automations to replace manual workflows in prospecting, enrichment, campaign ops, and reporting.
  • Build outbound engines, inbound lifecycle flows, content pipelines, and GTM data layers using LLM APIs and modern stacks.
  • Work autonomously and deliver working tools weekly, with measurable impact on hours saved and pipeline moved.

Rwazi is a global consumer intelligence platform that helps businesses understand consumer behavior. The company is a fast-growing startup with a culture of shipping quickly and iterating.

US

  • Design and develop secure AI evaluation environments using modern full-stack development practices.
  • Analyze real-world web application clones to identify bugs and workflow issues for improving evaluation quality.
  • Build automated testing frameworks to accurately assess AI-generated code submissions.

The company focuses on creating secure, reliable testing systems for AI-generated code, directly influencing next-generation AI evaluation. They operate as a remote, collaborative team of highly skilled technical professionals.

Global

  • Run the full eval pipeline end to end, reproducing results and pairing with senior engineers.
  • Build a judge calibration protocol to measure agreement and identify drift zones.
  • Extend benchmarks like GAIA and SWE-bench with new tasks targeting capability gaps.

Nous Research is an AI research lab that develops evaluation infrastructure for LLMs. They are a small, high-growth team valuing ownership and rapid shipping.

$1,860–$4,680/mo
India

  • Design and deploy autonomous AI agents to automate construction and estimation workflows.
  • Build agentic AI workflows using LLM frameworks such as LangChain, LangGraph, and Claude Code.
  • Integrate AI applications with cloud platforms and optimize system performance for scalability and reliability.

The company develops AI-driven solutions for construction and estimation workflows. It is a stable and growing organization that values continuous learning and innovation.

Europe Middle East

  • Build and evolve the core AI tutoring system with prompt architectures and agentic workflows.
  • Design and implement scalable software integrating AI with platform systems.
  • Collaborate cross-functionally with educators and product teams to translate pedagogical goals into technical solutions.

DataCamp empowers everyone with data and AI skills through practical learning experiences. It serves over 17 million learners and 6,000+ companies, including 80% of Fortune 1000, fostering a culture of data-driven decision-making and transparency.

$225,000–$300,000/yr
Unlimited PTO

  • Own software architecture and delivery: design, deploy, and operate systems at scale.
  • Design and evolve AI systems including prompting, retrieval, evaluation infrastructure, and expert feedback loops.
  • Hire, manage, and grow the engineering team from 3 FTEs while staying on the frontier of AI and legal AI.

Inhouse is the #1 AI lawyer for small to midsize businesses, combining AI, their own law firm, and an expert feedback loop to deliver fast, compliant legal work. They grew revenue 1,500% last year and recently raised a $5M seed round from leading VCs.

  • Review learners' AI-assisted code, specs, and governance artifacts, providing constructive feedback.
  • Catch where learners are stuck or falling behind and intervene early with virtual office hours.
  • Evaluate how engineers use AI tools like Claude Code or GitHub Copilot, focusing on judgment and process.

Correlation One is the largest provider of AI and data workforce development programs globally. They have trained over 500,000 professionals across 11 countries and partner with Fortune 500 enterprises and government agencies.

$109,347–$124,460/yr
UK

  • Partner with the incubation Account Executive to lead technical validation for Engineering and product development opportunities.
  • Develop a deep understanding of product development workflows and map them to Command by Asana.
  • Conduct technical discovery sessions, whiteboard conversations, and tailored product demonstrations for Engineering leaders.

Asana is a leading platform for human + AI collaboration, helping teams achieve their most important goals faster. With over 170,000 customers and 13+ offices worldwide, Asana has been recognized for its exceptional workplace culture and innovation.

US

  • Build AI agents that interact with customers and handle accounting autonomously.
  • Construct an accounting ledger from scratch and integrate with financial data providers.
  • Ship new features daily in Typescript/Node/Svelte, making quick decisions and taking risks.

Synthetic synthetically recreates bookkeeping using AI, processing financial data and managing workflows autonomously. Founded in 2025, we are a five-person team backed by Khosla Ventures, building on experience from Bench Accounting and Teal.

$230,000–$380,000/yr
US

  • Designing and advancing the agent harness used across Scout AI's agentic stack.
  • Integrating fullstack software with physical hardware systems and AI agents for robust low-latency communication.
  • Building simulation environments and frameworks for testing agents before physical deployment.

Scout AI develops Fury, the first robotic foundation model for defense, enabling human operators to command fleets of robots through natural language. The company is backed by top investors and fosters a culture of urgency, precision, and relentless work.

Global

  • Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
  • Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
  • Benchmark performance by conducting comparative testing between model outputs based on technical metrics.

Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.

South Africa

  • Evaluate and improve AI model performance on complex infrastructure and platform engineering challenges.
  • Analyze system designs, assess code quality, and provide detailed feedback on architectural soundness and technical accuracy.
  • Create reproducible failure cases and communicate complex technical concepts to enhance AI reasoning capabilities.

Jobgether uses an AI-powered matching process to connect candidates with hiring companies quickly and fairly. As a platform, it facilitates remote freelance opportunities for technical professionals.

Global

  • Teach core engineering design, CAD, and problem-solving skills to clients.
  • Guide clients through certification, project management, and AI automation tools.
  • Provide career coaching including resume review, interview prep, and salary negotiation.

We are building the home for ambition in the age of AI, connecting people with experts, programs, and communities to land jobs, get into schools, or upskill. Since our founding in 2021, we've helped tens of thousands of people and raised $19M from world-class investors, with a collaborative, high-energy culture.

Canada

  • Lead AI validation strategies for LLMs, RAG systems, and AI agent workflows.
  • Build automated validation pipelines and evaluation frameworks for AI output quality.
  • Establish production monitoring, AI guardrails, and quality metrics for AI products.

Jobgether is an AI-powered recruitment platform connecting candidates with hiring companies. It has a collaborative, remote-first culture and uses AI to streamline hiring.

EMEA 5w PTO

  • Own problem spaces end to end: write specs, acceptance criteria, and own the architecture.
  • Build AI into the product, e.g., turning free-text email replies into bookable quotes.
  • Make AI trustworthy with structured outputs, evals, and confidence-gated human review.

Cargo.one operates an AI-native operating system for freight, serving 30,000+ users across 172 countries with customers like Lufthansa Cargo and Kuehne+Nagel, backed by Index and Bessemer. The culture is positive, diverse, hard-working, feedback-heavy, and playful.

Europe Unlimited PTO

  • Own and drive the quality, reliability, and evolution of AI systems in production across multiple products.
  • Design and orchestrate agentic AI workflows, tool-based systems, and multi-step LLM chains with clean tool contracts.
  • Manage production prompt systems, run A/B experiments, and monitor AI system metrics using Langfuse and other observability tools.

Ruby Labs is a leading tech company that creates and operates innovative consumer products across health, education, and entertainment. The company has a fast-growing, ambitious team with a high-performance culture.

$120,000–$170,000/yr
US

  • Design, build, and maintain automated AI evaluation pipelines for production LLM applications.
  • Develop prompt engineering strategies and evaluate model performance using quantitative methods.
  • Analyze production AI behavior with Python, SQL, and statistical techniques to identify improvement opportunities.

GovWorx provides an AI-powered platform, CommsCoach, that supports 9-1-1 and emergency communications centers by automating quality assurance, training, and real-time call evaluation. The company is a growing technology team focused on public safety, collaborating across AI, engineering, product, and data science.

Global

  • Coach students and professionals in software engineering, including architecture, frontend frameworks, and AI-driven development like AI system design, LLM fine-tuning, and prompt engineering.
  • Support career development through resume review, interview prep, salary negotiation, and networking strategies.
  • Set your own hours and rates, host virtual sessions, and partner with the team to deliver a high-quality, supportive experience.

Leland is building the home for ambition in the age of AI, connecting people with the experts, programs, and communities they need to achieve their goals. Since 2021, they've helped tens of thousands of people and raised $19M from world-class investors, fostering a collaborative, high-energy, and AI-native culture.