Own AI workflows end to end, from problem definition to production and improvement, for finance processes like invoice-to-payment and reconciliation.
Design and build production-grade LLM-based systems with structured outputs, RAG, validation, and monitoring.
Collaborate with product, backend engineering, and customers while setting technical direction and using AI coding tools daily.
Jeeves is a stablecoin-native banking platform providing corporate cards, payments, treasury, and spend management for global enterprises. With over $5B in annualized volume and backing from top investors, the company is scaling rapidly and values relentlessness and tackling hard problems.
Build LLM-based agents on the platform's scaffolding, integrating tool calls, internal APIs, and guardrails.
Take agents to production on AWS with containers, CI/CD, secrets, permissions, and security controls.
Define and run evals, monitor with Langfuse, and document runbooks for independent operation.
Muttdata builds innovative Data Products and Machine Learning solutions to help companies solve complex business challenges. It is a fast-growing, remote-first startup that values collaboration, continuous learning, and a positive, ownership-driven culture.
Build the Learning Gym, a sandboxed environment for training and evaluating AI agents on real Niural workflows with ground-truth reward signals.
Engineer verifiers, run reinforcement learning loops, and establish honest baselines to prove agent improvements generalize.
Write first-author research papers and internal technical reports, and represent the work externally through preprints and talks.
Niural is a global payroll, Employer of Record (EOR), Agent of Record (AOR), and Contractor Management platform. It is a relentlessly ambitious team focused on building foundational internet infrastructure for the AI era.
Work directly with leading AI labs and enterprises to define research goals and technical requirements.
Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.
Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.
Architect AI infrastructure including backend services, data pipelines, and orchestration workflows for LLMs.
Write production-grade, high-performance code for high-throughput AI workflows and data-driven systems.
Integrate foundational models, vector data stores, and cloud microservices into scalable platform components.
Korza builds AI infrastructure and products, specializing in LLM and agent-based systems. The company emphasizes cross-functional collaboration and rapid shipping of cutting-edge AI capabilities.
Build and ship features of AI/ML and LLM-powered systems with guidance from senior engineers.
Implement and maintain AI/ML and AI agent pipelines from data ingestion through model deployment.
Contribute to LLM-powered features such as prompts, evaluations, retrieval, and tool integrations, while documenting experiments clearly.
Robots & Pencils designs AI systems for a human world, pairing engineering with creativity. Teams average fifteen-plus years of experience and value ownership, craft, direct feedback, and continuous learning.
Build new functionality within a fast-moving product, taking ambiguous ideas from design to rigorous experimentation within weeks.
Conduct deep-dive research into state-of-the-art LLM orchestration, retrieval strategies, and data-driven personalization.
Build production-grade ML systems at Twilio scale, leveraging the latest AI development stacks to focus on core differentiation.
Twilio is a cloud communications platform that enables businesses to create personalized customer experiences through APIs and AI-driven solutions. The company is remote-first, serving hundreds of thousands of businesses, and fosters a culture of connection, inclusion, and innovation.
Own agent quality end-to-end for Bliro's chat and voice agents, including prompting, context, and model evaluation.
Build eval sets, LLM judges, and benchmarks to measure quality, latency, and cost with statistical rigor.
Establish a repeatable process for improving agent quality and set the long-term technology roadmap.
Bliro builds an AI assistant that handles desk work for field sales teams, with over one million customer touchpoints documented. Backed by leading investors and trusted by German Mittelstand companies, it's a small, high-performing team that values ownership and continuous improvement.
Set the technical direction for AI engineering across the team: agent architecture patterns, evaluation methodology, deployment and monitoring strategies
Design the AI platform layer, including shared agent frameworks, tool integrations, and evaluation infrastructure
Work directly with clients on the most complex engagements, identifying new problem domains and ensuring production quality
Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. With over 1,500 firms in 60 countries managing nearly $10 trillion in assets, Addepar fosters a culture of ownership, collaboration, and innovation.
Own the build out of new agents, skills, and platform capability for teams across TLDR.
Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.
TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.
Design, build and operate production AI agents, including an autonomous coding agent that works from Jira and Slack, opening PRs and fixing CI failures.
Build the retrieval and context layer behind our AI brain: embeddings, semantic search, knowledge-graph modeling, and integration with Snowflake and internal docs.
Harden agents against prompt injection and over-broad permissions, and drive cost and latency down using caching and structured outputs.
Weedmaps is a global leader in the cannabis industry, providing a platform for consumers and businesses. The company has a collaborative culture focused on innovation and community, serving the U.S. and worldwide.
Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.
Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.
Design, build, and ship custom internal AI tooling, agents, and workflows for autonomy and research.
Integrate and extend third-party AI tools and evaluate new AI models with structured pilots.
Partner with cross-functional teams to prototype solutions and drive adoption of AI tools.
Waabi is a leader in Physical AI, founded by AI visionary Raquel Urtasun. The company is growing quickly with offices in Toronto, San Francisco, Dallas, and Pittsburgh, and seeks diverse, innovative candidates.
Design, develop, and maintain AI-powered software solutions across the full development lifecycle.
Work with backend stacks like Python, Node.js, .NET, and Java, plus modern frontend frameworks such as React or Angular.
Build and integrate AI agents, RAG pipelines, and LLM-driven products into production environments.
Cresteo is a nearshore tech services company on a mission to be the world leader in people-first and honest software development. Our small, high-trust team values transparency, profit-sharing, and fearless innovation while working with US-based clients and international teams.
Design and deploy AI agents that automate sales, operations, and support workflows.
Own technical delivery end-to-end, from architecture to production and post-launch improvements.
Work directly with clients to understand workflows and explain technical tradeoffs.
Stello is an AI transformation firm that partners with mid-market and enterprise companies to automate repetitive work with AI agents. They work across sales, operations, and customer support, deploying solutions inside existing client tools.
Build and deploy AI workflows and skills using AIVA
Integrate AI solutions with APIs, databases and enterprise systems
Rapidly build, test, demonstrate and refine solutions with clients
Williams Lea is the leading global provider of tech-enabled business and marketing services helping clients manage and transform processes. The company serves clients in 20 countries across four continents and has 15,000 employees worldwide.
Take technical ownership of our content and design factory: tooling selection, automation pipelines, quality management, and agent orchestration.
Design and develop complex multistep content and design creation workflows and interfaces using AI model harnesses and tools like AirOps and n8n.
Reduce manual overhead systematically by building automations and agent workflows that replace repetitive processes.
Platform Engineering is home to the world's largest community of platform engineers, connecting thousands of professionals through events like PlatformCon. We are a small, low-ego team with a strong async culture, enabling a community of over 35,000 platform engineers worldwide.
Own product features end-to-end, from problem statement to production.
Write specs with AI agents and critically review their code in complex systems.
Deliver reliable features handling financial transactions and evolve AI-first development.
Social Discovery Group creates social entertainment platforms that connect people online across cultures, addressing loneliness and disconnection. The company's international remote team works from everywhere, and it earned 'Great Place to Work' recognition in 2024-2025.
Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.
Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.
Build the component layer around our layout synthesis engine, including API contracts, services, and evaluation gates.
Turn model retraining into a one-command job with built-in benchmarks and readable results for the whole team.
Own latency and cost budgets for learned capabilities and partner with infrastructure engineers on MLOps and deployment.
Higharc is a VC-backed startup that is changing how new homes are designed and built using spatial AI and generative floor plan technology. The company is fully remote, has raised over $175M, and values flexibility, collaboration, and asynchronous deep work.