Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.
Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.
Design high-performance training platform components including orchestration, observability, and performance tuning.
Deliver end-to-end ML pipelines from dataset curation to training, validation, and deployment.
Drive technical direction across ML Platform, Infrastructure, Autonomy, and Safety teams.
Stack is developing revolutionary AI and advanced autonomous systems for safer, more reliable, and efficient operations, focusing on autonomous trucking. The Stack team brings decades of experience in real-world systems and is committed to a culture of inclusion, entrepreneurship, and innovation.
Design and execute technical training programs for AI data and annotation teams, including curricula and certifications.
Own quality frameworks, monitor KPIs like accuracy and defect rates, and drive continuous improvement.
Provide technical guidance for AI data projects such as LLM evaluation, RLHF, and model benchmarking.
Innodata is a global data engineering company that provides data, evaluation frameworks, and human expertise for AI systems. With over 36 years of experience, they are committed to delivering the highest quality data and outstanding outcomes for clients.
Build tooling for capturing and processing data from agents and humans at significant scale.
Solve hard problems around compute, orchestration, scaling, security, and reliability.
Help develop approaches for training, benchmarking, and evaluating AI agents.
Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.
Own the integrity, validation, and reliability of data pipelines powering the product.
Design and execute frameworks for data quality assurance and validation across the organization.
Partner with engineering, quantitative modelers, and solution architects to ensure data is launch-ready and validated.
Jupiter is the global market leader in analytics for resilience planning and enterprise climate risk management. Led by pioneers in data, climate, and technology, the company fosters an inclusive, mission-driven culture and is committed to preparing economies for climate change.
Evaluate software engineering tasks for technical accuracy, realism, and reproducibility.
Investigate codebases, tests, and integration issues to identify technical weaknesses.
Provide clear, actionable feedback that directly improves AI training and evaluation workflows.
Jobgether is an AI-powered job platform that connects candidates to roles through objective, skill-based matching. It focuses on remote and freelance opportunities, with a data-driven recruitment process and a global candidate pool.
Own the quality analysis of annotated data across computer vision and ML programs, making judgment-heavy calls that cannot be outsourced.
Analyze annotation data to detect anomalies and distinguish between annotation error, model error, and genuine field degradation.
Build and maintain performance trackers, root cause analysis documentation, and pattern libraries that the team relies on.
Gather AI builds a vision-powered platform using autonomous drones to digitize warehouse workflows, making operations smarter and safer. The company is a small, fast-moving engineering team with a culture of deep technical accountability and cross-functional collaboration.
Work directly with leading AI labs and enterprises to define research goals and technical requirements.
Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.
Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.
Own and extend the test automation framework across UI, API, and hardware-in-the-loop test surfaces.
Build and maintain CI/CD test infrastructure including containerized environments and pipeline orchestration.
Provide technical leadership and mentorship to software engineers on testing best practices.
Terawatt Infrastructure is a leader in delivering large-scale turnkey charging solutions for autonomous and electric vehicle fleets. The company develops, finances, owns, and operates charging infrastructure with a growing portfolio across the US, guided by values of ownership, candor, and progress.
Analyze repositories for maintainability, dependency risks, and security patterns using AI and static analysis tools.
Use AI coding agents to accelerate code review, vulnerability explanation, and remediation proposal drafting.
Collaborate with security, DevOps, and architecture teams to create evidence packs linking findings to readiness implications.
Deutsche Telekom IT Solutions Slovakia is a leading IT services provider in Kosice, part of the Deutsche Telekom group. With over 3900 employees, they focus on innovation and continuous transformation in ICT services.
Design and maintain reliable, low-latency ML APIs to integrate Safety AI model outputs into cloud applications.
Build scalable data pipelines for continuous model iteration, backtesting, and online evaluation.
Optimize model artifacts for production and monitor rollout health, ensuring predictable failure modes.
Samsara builds a Connected Operations Cloud that helps physical operations use IoT data to improve safety, efficiency, and sustainability. Samsara is a recently public company with an employee-led remote culture and a long-term focus.
Maintain and evolve platform foundations, tools, and processes for agentic development to keep n8n operating as an AI-native engineering organization.
Build and integrate agentic workflows across Linear, GitHub, CI, and developer review processes while ensuring safety and security.
Define quality metrics, benchmarks, and adoption signals to measure agent output and improve engineering productivity.
n8n is an open workflow orchestration platform built for the new era of AI, giving technical teams the freedom of code with the speed of no-code. Since 2019, the company has grown to over 260 employees across Europe and the US, with a community of 650,000+ developers, 190K+ GitHub stars, and a $5.2bn valuation.
Build LLM-powered agents and production ML systems to automate security analysis.
Collaborate with threat-research and data engineering to turn risk identification into scalable solutions.
Own deployment, monitoring, and quality of solutions in CI/CD frameworks.
Zscaler accelerates digital transformation with its Zero Trust Exchange platform, protecting customers from cyberattacks and data loss. The company is the world's largest in-line cloud security platform, distributed across 160+ exchanges, with an AI-native culture focused on ownership, collaboration, and trust.
Design and implement state-of-the-art ML models and training pipelines for robotics.
Develop efficient data/training strategies and evaluation frameworks for rapid experimentation.
Collaborate with engineering to optimize training infrastructure and deployment.
We're revolutionizing real-world automation by making robotic systems accessible to everyone. Our AI-powered platform brings software automation to physical spaces, and we're a small startup team working across the stack to solve customer problems.
Build and operate backend systems serving AI-powered insurance workflows in production.
Design and implement AI orchestration layers connecting models, APIs, workflows, and business logic.
Optimize latency, throughput, and cost across AI services (caching, batching, streaming, routing).
We are a platform that leverages AI for decisioning, automation, and orchestration. We are a global engineering team working across multiple countries to build reliable, scalable AI systems.
Design and maintain automated test frameworks for document-generation workflows.
Validate generated documents against system state, business rules, templates, and expected formats.
Implement CI/CD quality gates and ensure traceability, reproducibility, and compliance.
Nagarro is a digital product engineering company that builds products, services, and experiences. They have 18,000+ experts across 39 countries and a dynamic, non-hierarchical culture.
Lead a team of AI Tutors to deliver high-quality training data and evaluations for SpaceXAI's models.
Own end-to-end quality and delivery for Human Data projects, reviewing work and ensuring consistency.
Coach, performance-manage, and develop talent while improving operational processes.
SpaceXAI creates AI systems to understand the universe and aid humanity. The team is small, highly motivated, and focused on engineering excellence with a flat structure.
Build product-facing tools for browsing environments, inspecting trajectories, and understanding model behavior.
Develop vendor-facing workflows to create, submit, test, and iterate on RL environments and training data.
Create dashboards and observability tools surfacing environment quality, eval results, and pipeline health.
This company builds a reinforcement learning data engine for frontier AI labs, creating and iterating on training data. They have offices in San Francisco and Singapore, offering a culture that values independence and high agency.
Build and operate ingestion pipelines for flight test, simulation, and operational data with quality validation.
Implement dataset curation and versioning with full lineage tracking.
Automate mining and triage for rare events and close the loop on operational signal.
Merlin is a publicly traded aerospace and defense company building a non-human pilot for full-stack aircraft autonomy. Headquartered in Boston, it has flown hundreds of autonomous flights and is expanding its team to accelerate development and commercialization.
Build and ship a project end to end: understand the problem, build a working version, and improve it based on user feedback.
Work with business teams to understand their workflows and translate requirements into actionable solutions.
Use AI-augmented development tools as your default workflow and develop judgment about their effectiveness.
M3 is a global healthcare technology company providing innovative research and technological solutions to the healthcare industry. The M3 Group operates in the US, Asia, and Europe with over 5.8 million physician members and is publicly traded on the Tokyo Stock Exchange, ranked in Forbes' Global 2000 list.