Source Job

US Canada

  • Curate high-quality code examples and datasets for LLM training and evaluation.
  • Develop and assess AI-generated software across multiple programming languages and the full SDLC.
  • Collaborate with research teams to design verification mechanisms and improve coding benchmarks.

Python JavaScript ReactJS C++ Rust

20 jobs similar to Senior Software Engineer – LLM Evaluation

Jobs ranked by similarity.

Global

  • Evaluate AI-generated coding interactions end to end for correctness and engineering judgment.
  • Assess whether outputs reflect strong engineering taste and provide clear, opinionated feedback.
  • Help define what great looks like for AI coding tools like Codex, Claude Code, and Cursor.

G2i Inc. is a technology staffing company that connects software engineers with remote contract opportunities. The company values engineering excellence and provides flexible, ongoing projects for senior-level developers.

$88,000–$180,000/yr
US Unlimited PTO

  • Direct the agent array on production workstreams by decomposing problems into tasks and integrating agent output into shipped software.
  • Review agent-generated pull requests at volume and depth, identifying correctness, security, and accessibility defects.
  • Author evaluation suites that make quality measurable using eval-driven development and own end-to-end quality within a FedRAMP-authorized environment.

Granicus provides cloud-based solutions for government communications, website design, meeting management, and records management, serving over 5,500 agencies and 300 million citizens. With a globally distributed team and a culture of transparency and inclusion, Granicus has been recognized on the GovTech 100 list for the past 5 years.

US

  • Evaluate software engineering tasks for technical accuracy, realism, and reproducibility.
  • Investigate codebases, tests, and integration issues to identify technical weaknesses.
  • Provide clear, actionable feedback that directly improves AI training and evaluation workflows.

Jobgether is an AI-powered job platform that connects candidates to roles through objective, skill-based matching. It focuses on remote and freelance opportunities, with a data-driven recruitment process and a global candidate pool.

$250,000–$300,000/yr
US Canada Unlimited PTO

  • Own the build out of new agents, skills, and platform capability for teams across TLDR.
  • Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
  • Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.

TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.

US Canada Unlimited PTO

  • Design, build, and ship custom internal AI tooling, agents, and workflows for autonomy and research.
  • Integrate and extend third-party AI tools and evaluate new AI models with structured pilots.
  • Partner with cross-functional teams to prototype solutions and drive adoption of AI tools.

Waabi is a leader in Physical AI, founded by AI visionary Raquel Urtasun. The company is growing quickly with offices in Toronto, San Francisco, Dallas, and Pittsburgh, and seeks diverse, innovative candidates.

US Unlimited PTO

  • Lead the engineering strategy and execution for evaluations of AI agents, owning the core evaluation platform.
  • Design scalable evaluation infrastructure, APIs, workflows, and production systems across software categories.
  • Mentor and develop a team of engineers as the technical authority on agentic evaluation.

This company focuses on building credible, scalable evaluations of AI agents from software vendors. It operates as a fully remote, inclusive team with a flexible culture and a focus on professional growth.

$88,000–$180,000/yr
US

  • Direct autonomous coding agents across production workstreams, decomposing complex problems and reviewing agent output.
  • Design and maintain evaluation suites to establish measurable quality standards and own end-to-end software quality.
  • Operate within security boundaries such as NIST 800-53, WCAG, and SOC 2 while advancing AI-native development practices.

The company is an organization operating within a FedRAMP-authorized environment, focused on AI-native software development and autonomous coding agents. It maintains a remote-first, globally distributed engineering culture with an emphasis on rigorous evaluation, compliance, and engineering quality.

Global

  • Apply your expertise in software engineering to help train next-generation AI systems by creating reinforcement learning environments and solving complex software engineering problems.
  • Contribute expert-level code samples, debugging strategies, and development insights in languages such as Python, Java, Rust, Go, C++, or TypeScript.
  • Refactor and optimize code, review and validate peer contributions, and document technical reasoning to enhance AI training data quality.

The company is a rapidly growing, venture-backed AI company that combines world-class human expertise with advanced machine learning to build and improve cutting-edge AI models. It is backed by over $40 million in funding and has a rapidly expanding international network of experts.

US

  • Implement features that improve MDLinx and drive user impact, with ownership growing over time.
  • Apply software engineering best practices including automated testing, clean code, and peer review.
  • Read AI-generated plans and code critically, and identify work for AI-assisted automation.

M3 is a Japanese global leader in healthcare technology and research solutions, operating physician websites with over 5.8 million members. MDLinx, an M3 company, focuses on transforming pharmaceutical brand promotion through omnichannel engagement.

Global 7w PTO

  • Own product features end-to-end, from problem statement to production.
  • Write specs with AI agents and critically review their code in complex systems.
  • Deliver reliable features handling financial transactions and evolve AI-first development.

Social Discovery Group creates social entertainment platforms that connect people online across cultures, addressing loneliness and disconnection. The company's international remote team works from everywhere, and it earned 'Great Place to Work' recognition in 2024-2025.

EMEA

  • Lead technical evaluations and proofs of concept for an enterprise AI coding platform in real customer environments.
  • Engage directly with CTOs, VPs of Engineering, and senior developers on AI agent architecture and integration.
  • Manage enterprise security and deployment reviews, troubleshoot integration issues, and feed customer insights into product.

The company is an enterprise AI coding platform provider helping engineering organizations accelerate development with AI agents and LLM infrastructure. It operates as a remote-first global team with a fast-paced, innovative culture focused on collaboration and rapid product iteration.

Canada

  • Design and implement the core AI agent runtime in Python, including orchestration loops, tool definitions and execution, memory and context management, and failure handling.
  • Build and maintain evaluation datasets, automated evaluators, LLM-as-judge pipelines, and regression tracking to ensure agent quality and performance.
  • Develop Python and Node.js microservices for agent hosting and LLM interactions, and build reliable Laravel/PHP APIs for integration.

Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities, partnering with companies to manage applications and next steps. They emphasize a fast-growing, collaborative environment with a focus on professional growth and cutting-edge technology.

AI Engineer

Cresteo
Latin America Unlimited PTO

  • Design, develop, and maintain AI-powered software solutions across the full development lifecycle.
  • Work with backend stacks like Python, Node.js, .NET, and Java, plus modern frontend frameworks such as React or Angular.
  • Build and integrate AI agents, RAG pipelines, and LLM-driven products into production environments.

Cresteo is a nearshore tech services company on a mission to be the world leader in people-first and honest software development. Our small, high-trust team values transparency, profit-sharing, and fearless innovation while working with US-based clients and international teams.

Global

  • Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
  • Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
  • Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.

Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.

Global

  • Review coding tasks derived from real GitHub issues and pull requests to assess technical soundness and reproducibility.
  • Evaluate unit tests for correctness, coverage, and robustness, identifying flaky tests or missing dependencies.
  • Provide clear recommendations on whether tasks should be accepted, improved, or excluded.

Anyone AI is a company that focuses on AI-related projects, particularly in evaluating software engineering tasks. They are a smaller organization with a culture that values technical expertise and attention to detail.

Canada

  • Lead architecture and development of production-grade autonomous multi-agent AI systems.
  • Provide technical leadership across AI engineering, cloud, and DevOps practices.
  • Design and implement LLM- and RAG-based solutions using LangGraph, LangChain, and Claude.

The partner company specializes in building production-grade autonomous and multi-agent AI systems for enterprise clients. Team size and culture are not specified in the posting.

Europe

  • Use software engineering expertise and Dutch fluency to evaluate and improve advanced AI systems.
  • Assess model responses for logical accuracy, coding quality, and adherence to best practices across multiple technical domains.
  • Document model failure modes and provide structured feedback to improve AI reasoning and coding performance.

LATAM

  • Apply AI coding assistants and large language models to automate workflows and increase engineering efficiency.
  • Build product features around Model Context Protocol, agentic developer environments, and supporting tooling.
  • Diagnose and resolve coding and AI-related issues while taking technical ownership and communicating solutions.

US

  • Participate in a remote video interview about your experience with AI coding tools.
  • Discuss how tools like GitHub Copilot and Cursor fit into your development workflow.
  • Share feedback on the strengths and limitations of AI-assisted coding.

Our partner company is a research organization focused on AI-assisted software development. It is currently seeking feedback from practicing software engineers through paid interviews.

Turkey

  • Ship real features in production using coding agents like Claude Code, Cursor, and Copilot, setting intent and owning the outcome.
  • Keep the codebase and context lean by maintaining standards, architecture notes, and feeding SAST findings back for real patches.
  • Build team-level tooling, review AI-assisted pull requests, and share workflows that improve how everyone works with agents.

Insider One is a marketing and customer engagement platform that unifies data, personalization, and journey orchestration across channels. The company is powered by 1500+ employees across 30+ offices, and is recognized as a woman-founded, women-led B2B SaaS unicorn.