Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.
Work directly with leading AI labs and enterprises to define research goals and technical requirements.
Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.
Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.
Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.
Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.
Own the build out of new agents, skills, and platform capability for teams across TLDR.
Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.
TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.
Design and implement autonomous agents and multi-agent orchestration systems for complex, open-ended tasks.\n- Build and optimize production-grade LLM applications with a focus on reliability, observability, and low-latency performance.\n- Architect advanced RAG pipelines and vector database strategies to provide agents with accurate, real-time context.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective evaluation. The company operates with a distributed, global-first team culture, emphasizing fairness and innovation in recruitment.
Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
Turn research ideas into working prototypes and production-facing capabilities.
This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.
Design, develop, and deploy AI solutions that power business products and operations.
Collaborate with product, data science, and engineering teams to build scalable, production-ready AI systems.
Write clean, maintainable code and integrate AI/ML models into backend and frontend applications.
Zone & Co is the ERP-native AI platform for financial operations, purpose-built for organizations running on Oracle NetSuite. They serve more than 4,500 customers worldwide and operate as a high-velocity, fully remote, global team.
Build LLM evaluation systems and diagnose agent performance issues across prompts and tools.
Netomi is an agentic AI platform for enterprise customer experience, enabling automation for global brands. Backed by Y Combinator and Index Ventures, the company drives efficiency and higher quality experiences.
Design and implement state-of-the-art ML models and training pipelines for robotics.
Develop efficient data/training strategies and evaluation frameworks for rapid experimentation.
Collaborate with engineering to optimize training infrastructure and deployment.
We're revolutionizing real-world automation by making robotic systems accessible to everyone. Our AI-powered platform brings software automation to physical spaces, and we're a small startup team working across the stack to solve customer problems.
Build and ship features of AI/ML and LLM-powered systems with guidance from senior engineers.
Implement and maintain AI/ML and AI agent pipelines from data ingestion through model deployment.
Contribute to LLM-powered features such as prompts, evaluations, retrieval, and tool integrations, while documenting experiments clearly.
Robots & Pencils designs AI systems for a human world, pairing engineering with creativity. Teams average fifteen-plus years of experience and value ownership, craft, direct feedback, and continuous learning.
Build AI agents and automations against specifications, including prompt development and knowledge grounding.
Own quality evaluation, knowledge curation, and monitoring of live AI workloads.
Maintain documentation and support production readiness, with a focus on scaling AI platforms.
Turnitin is a recognized innovator in global education, partnering with educators and institutions for over 25 years. With 16,000+ academic institutions and a remote-first culture, we are a global organization with team members in over 35 countries, unified by a shared desire to make a difference in education.
Design and deliver production-grade agents that investigate, reason, and act on live observability data.
Own agent work from rough prototype to production, including evals.
Extend the agentic workspace with new canvas capabilities, MCP server, and skills.
Honeycomb provides an observability platform that helps engineers understand and debug complex systems. It is a fully distributed company of over 200 people, with a culture that values impact, autonomy, and inclusivity.
Design and improve prompts for classification and extraction tasks, running structured evaluation cycles for accuracy and iteration.
Build datasets for testing and validation, analyze outputs from real usage (logs, SQLite), and write Python scripts for testing and scoring.
Identify failure patterns, propose improvements, validate system behavior in iOS simulations, and document findings clearly.
PiggyBank Ventures is an incubator building AI-first consumer applications, currently developing Nestora, an AI-powered mobile app for intelligent communication management. The team includes repeat founders and operators with experience building and scaling products, from top-tier universities and technical programs.
Build tooling for capturing and processing data from agents and humans at significant scale.
Solve hard problems around compute, orchestration, scaling, security, and reliability.
Help develop approaches for training, benchmarking, and evaluating AI agents.
Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.
Develop, configure, deploy, and optimize AI agents using Cresta's AI platform and tools.
Build AI agent integrations with external systems such as APIs, databases, and CRMs to ensure seamless workflow integration.
Collaborate with customers and internal stakeholders to gather technical requirements and translate business needs into AI agent solutions.
Cresta builds a unified AI platform that combines conversational AI agents, real-time human agent augmentation, and conversation intelligence to drive revenue and efficiency for customer experience channels. Born from Stanford AI Lab, the company has raised over $270 million from top investors and is led by AI industry veterans.
Build production Agentforce agents including topics, actions, grounding, guardrails, and evaluation harnesses.
Implement actions across Flow, Apex, prompt templates, and external APIs, grounding answers in Data 360.
Design guardrails against prompt injection and data leakage, and run evaluation and regression tests before changes.
AspenView Technology Partners builds high-performing nearshore IT teams for North American clients, focusing on innovation and efficiency. They are a people-first, purpose-driven company that believes great culture drives great outcomes.
Build production-grade agentic workflows to automate fan issue triage, seller risk monitoring, and marketplace optimization.
Coach non-technical teammates to use AI tools and create their own automations.
Own the full lifecycle of built solutions, from design to production, and scale successes across operations teams.
Gametime makes it easy for people to discover and access live experiences, with platforms supporting over 60,000 events across the US and Canada. They are a mid-sized company focused on reimagining the event ticket industry, with a culture that emphasizes inclusivity and innovation.
Architect and build scalable Generative AI and agentic AI applications from concept to production.
Design LLM-powered workflows, prompt strategies, and multi-agent solutions using LangChain and LangGraph.
Evaluate, fine-tune, and optimize LLMs, and own end-to-end ML and GenAI pipelines.
We are a partner company specializing in AI-powered recruitment, connecting top talent with innovative companies. We foster a collaborative, fast-paced environment with a focus on technical growth and professional development.
Design agent architectures for planning, reasoning, tool use, and memory integration.
Improve reliability on long-running tasks with failure recovery and evaluation systems.
Build loops for agents to improve with real use, balancing quality, latency, and cost.
Adaption builds AI systems that evolve in real-time, making them flexible and personalized. They are a global-first team focused on talent density and collaboration.
Build and ship a project end to end: understand the problem, build a working version, and improve it based on user feedback.
Work with business teams to understand their workflows and translate requirements into actionable solutions.
Use AI-augmented development tools as your default workflow and develop judgment about their effectiveness.
M3 is a global healthcare technology company providing innovative research and technological solutions to the healthcare industry. The M3 Group operates in the US, Asia, and Europe with over 5.8 million physician members and is publicly traded on the Tokyo Stock Exchange, ranked in Forbes' Global 2000 list.
Set the technical direction for AI engineering across the team: agent architecture patterns, evaluation methodology, deployment and monitoring strategies
Design the AI platform layer, including shared agent frameworks, tool integrations, and evaluation infrastructure
Work directly with clients on the most complex engagements, identifying new problem domains and ensuring production quality
Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. With over 1,500 firms in 60 countries managing nearly $10 trillion in assets, Addepar fosters a culture of ownership, collaboration, and innovation.