Build and improve core components of the Agent Runtime, including intent routing, RAG, and multi-step agent workflows.
Design and optimize Tool Server / Tool Calling capabilities for agent execution and integration.
Develop evaluation datasets and harnesses to measure agent quality and reliability.
Binance is a leading global blockchain ecosystem behind the world's largest cryptocurrency exchange by trading volume and registered users. With over 320 million users in 100+ countries, Binance offers a dynamic, flat-structured environment focused on innovation and growth.
Own the shared AI foundation used by product engineering teams, including model selection, routing, context management, and evaluation.
Design and ship agentic capabilities end to end, from proof-of-concept through production optimization and iterative improvement.
Establish evaluation infrastructure and practices to measure model and agent performance, ensuring quality and accuracy for financial decision-making.
Our partner is building a suite of financial planning and analysis capabilities powered by a shared AI foundation. They operate as a remote-first engineering team with a focus on autonomy, craftsmanship, and high-impact work.
Own and extend the offline evaluation suite for AI products, building datasets and metrics.
Build online quality dashboards and close the production feedback loop by mining failure patterns.
Translate numbers into clear decisions for Product and domain experts.
Finom is a European tech startup developing an all-in-one financial B2B platform integrating banking, accounting, and invoicing for entrepreneurs. With over €115 million in Series C funding and a team dedicated to innovation, they foster a start-up culture that values bold ideas and swift implementation.
Own model strategy and selection using rigorous benchmarks and statistical analysis.
Design and maintain evaluation methodologies for AI systems, including offline sets and LLM-as-judge frameworks.
Develop classification, fine-tuned, and agentic AI models to improve accuracy, cost, and latency.
A production agentic AI platform that builds and deploys advanced machine learning systems. The team is collaborative and values innovation, offering a remote work environment with high autonomy.
Conduct independent research in Generative AI, LLMs, NLP, and multimodal AI to design experiments and evaluate models.
Develop and implement LLM evaluation frameworks, analyze model performance, and identify data gaps for improvement.
Apply strong statistical and data science skills to clean, analyze, and interpret complex datasets for AI/ML research.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company is committed to delivering the highest quality data and outstanding outcomes for its customers.
Design and validate trading factors from multi-source market, fundamental, and on-chain data.
Develop and optimize prediction models using machine learning and deep learning to enhance signal accuracy and stability.
Lead end-to-end strategy design, backtesting, and live deployment, owning P&L and risk performance.
Binance is a leading global blockchain ecosystem and the world’s largest cryptocurrency exchange by trading volume, trusted by over 300 million users in 100+ countries. With a flat structure and world-class talent, the company fosters a fast-paced, results-driven culture focused on innovation and financial inclusion.
Architect and own the technical roadmap for a secure local LLM platform deployed in company-controlled infrastructure.
Build modular inference layers with stable APIs, model routing, and production-grade serving optimizations.
Design and operate retrieval-augmented generation pipelines with permission-aware access and systematic evaluation.
Parallel Wireless is a U.S.-based pioneer in Open RAN innovation, transforming how mobile networks are built and powered. The company is a leader in software-centric, hardware-agnostic network solutions with a focus on reducing complexity and total cost of ownership.
Own the AI foundation product teams build on: model selection, routing, context management, and tool design.
Own eval infrastructure by defining output quality standards with customers and encoding them in evals.
Ship agentic features end to end from v0 through optimization, including prompt engineering, caching, and parallel tool calls.
Aleph is an AI-native platform for Financial Planning & Analysis (FP&A), solving data scattered across systems and spreadsheets. Backed by top VCs, it works with customers like Webflow and Notion, and is founded by former Google and Bain engineers.
Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows.
Develop evaluation methodologies that combine automated metrics with structured human subject matter expert judgment.
Produce defensible evaluation reports comparing candidate capabilities with current mission workflows, including documented limitations and failure modes.
OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by NumPy and SciPy creator Travis Oliphant, the company is built by people with deep roots in the open-source ecosystem.
Design, build, deploy, and monitor LLM-powered document analysis pipelines that serve lawyers daily, meeting defined latency SLAs and accuracy benchmarks.
Architect multi-step, multi-agent systems for complex legal tasks with robust state management, tool calling, and human-in-the-loop checkpoints.
Collaborate with lawyers and product to translate ambiguous legal workflows into structured AI problems and iterate based on user feedback and eval results.
Cleary Gottlieb is a global law firm that provides legal services through 14 offices in major financial centers worldwide. The firm employs approximately 1,100 lawyers from more than 50 countries and operates as a single, integrated global partnership with a culture of intellectual agility and human touch.
Own the build out of new agents, skills, and platform capability for teams across TLDR.
Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.
TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.
Design and implement autonomous agents and multi-agent orchestration systems for complex, open-ended tasks.\n- Build and optimize production-grade LLM applications with a focus on reliability, observability, and low-latency performance.\n- Architect advanced RAG pipelines and vector database strategies to provide agents with accurate, real-time context.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective evaluation. The company operates with a distributed, global-first team culture, emphasizing fairness and innovation in recruitment.
Own the ML strategy for dialogue systems, lead a team of 3 ML engineers, and drive LLM post-training and model adaptation.
Build evaluation layers, cut dialogue failure modes, and keep inference efficient on latency and cost.
Stay hands-on with prototypes, debugging agent traces, and reviewing team work.
Social Discovery Group (SDG) is a group of social discovery companies that solve problems of loneliness and isolation through social entertainment platforms. The company has an international team of digital nomads and has been recognized as a Great Place to Work winner and a top company for remote jobs.
Evaluate AI systems at a scale only possible by combining thousands of vetted experts with model graders.
Innovate at the frontier of QA by shaping industry standards for validating agentic AI and large language models.
Collaborate with global market leaders to architect AI quality blueprints and drive high-impact consultative visibility.
Testlio provides a fully managed crowdsourced testing platform powered by proprietary intelligence technology, LeoCore. They are a female-founded, fully remote company with an inclusive culture, half of their team identifying as women, and are growing profitably.
Work directly with leading AI labs and enterprises to define research goals and technical requirements.
Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.
Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.
Lead a team of AI engineers, setting technical direction and quality standards while staying hands-on in code.
Own the architecture and end-to-end delivery of AI systems using LLMs, RAG, and agentic patterns.
Drive evaluation, monitoring, and continuous improvement loops to ensure production-grade AI capabilities.
Finom is a European tech startup building an all-in-one financial B2B platform for entrepreneurs, integrating banking, accounting, and invoicing. They are a well-funded company with over $346 million in total funding, maintaining a startup culture that values innovation, swift implementation, and employee impact.
Own end-to-end quality of AI product experiences, improving response usefulness and reliability.
Build evaluation frameworks and run experiments to measure and improve AI performance.
Translate AI failure patterns into prioritized fixes for Product and Engineering.
Progress Partners is a fast-growing technology company building and scaling next-generation SaaS and mobile platforms. Our global team of experts collaborates across time zones to deliver category-defining innovations.
Own end-to-end AI solution delivery inside ORA, from design to deployment with measurable success metrics.
Design AI workflows across the call lifecycle, including pre-call briefs, in-call guidance, and post-call grading.
Partner with RevOps and QA to close the loop, raise the technical bar, and stay current with emerging AI tooling.
HighLevel is an AI-powered all-in-one white-label sales & marketing platform that empowers agencies and businesses to drive growth. With over 1,500 team members across 15+ countries, we operate a remote-first environment with a culture rooted in creativity and collaboration.
Architect, build, and optimize high-performance production LLM systems while maintaining a strong personal technical presence on the team.
Spearhead strategic technological changes and champion code refactoring efforts to keep the core codebase cutting-edge and performant.
Lead technical story breakdowns, architectural design, and mentor engineers across the department.
Appian provides AI automation for mission-critical work, automating complex processes in large enterprises and governments. With over 25 years of experience, the company is known for its reliability and scale, and fosters an inclusive culture with employee-led affinity groups.
Design, build, and deploy LLM-powered product features, including lab summaries and conversational agents.
Build backend services integrating LLMs and ML models, primarily using Python with exposure to Elixir.
Implement evaluation, monitoring, and CI/CD workflows for AI features, ensuring reliability and clinical relevance.
Fullscript is a health technology platform that helps practitioners deliver better care through clinical insights, lab interpretations, and patient analytics. With over 125,000 practitioners and 10 million patients, the company emphasizes a people-first culture, teamwork, and continuous learning in a remote-first environment.