Lead research and engineering to develop autonomous AI agents that set and execute complex goals.
Design and implement dynamic planning, memory, tool-use, and evaluation frameworks for safe agent behavior.
Collaborate with multidisciplinary teams to transition research into scalable, production-ready solutions.
Our partner is a large-scale healthcare environment focused on advancing autonomous AI systems. They are looking for a Lead Data Scientist to shape next-generation AI agents in a collaborative, remote U.S. team.
Build sandboxed environments that wrap real Niural workflows, stateful across episodes and seeded for reproducibility.
Design programmatic verifiers from known ground truth, scoring trajectories to prevent reward hacking.
Train agents via RL loops, curriculum schedules, and fine-tuning, then publish negative results internally.
Niural is a global Payroll, Employer of Record (EOR), Agent of Record (AOR), and Contractor Management platform that empowers businesses in the digital economy. We are a team building foundational internet infrastructure with a focus on speed, ownership, and ambition.
Design and build production-grade AI agents that answer complex analytics and marketing questions.
Rapidly prototype and develop proofs of concept for AI-powered data and marketing workflows.
Translate customer needs into intuitive AI-assisted product experiences while balancing experimentation with scalable engineering.
They build AI-powered data and marketing experiences for customers. The environment is highly autonomous, fast-moving, and intellectually curious, with a focus on meaningful impact.
Own the shared AI foundation used by product engineering teams, including model selection, routing, context management, and evaluation.
Design and ship agentic capabilities end to end, from proof-of-concept through production optimization and iterative improvement.
Establish evaluation infrastructure and practices to measure model and agent performance, ensuring quality and accuracy for financial decision-making.
Our partner is building a suite of financial planning and analysis capabilities powered by a shared AI foundation. They operate as a remote-first engineering team with a focus on autonomy, craftsmanship, and high-impact work.
Own research projects end-to-end: identify important questions, formulate hypotheses, design experiments, analyze results, and publish.
Develop rigorous evaluations of misalignment and loss-of-control risks, including evaluation awareness, sandbagging, and dishonesty.
Study behaviors difficult to observe directly, such as long-horizon failure modes and cases where models may conceal relevant behavior.
Neo Research is an independent AI safety research organization based in Singapore that studies frontier risks in increasingly capable AI systems, with a focus on open-weight models. The team is small, offering substantial freedom to pursue important research questions and a collaborative environment with strong engineering support and external relationships.
Lead the development of the AIOS Agent SDK, forming the foundation for world-class agents across the company.
Own agent architecture, model strategy, evals, and production reliability for AI systems handling $100k daily revenue.
Build and scale the applied AI team, setting technical direction and ensuring safe, compliant agent operations.
AIOS is building the world’s first full-stack AI doctor, serving 150k patients monthly via Bolt Pharmacy in the UK. They are a profitable, fast-growing startup with $350M ARR and a founder-led culture of intense execution.
Design and build production-grade agentic AI systems for marketing and customer experience.
Architect agent workflows with reasoning, tool use, retrieval, and performance monitoring.
Evaluate emerging frameworks like LangGraph and LangChain for enterprise applications.
The company specializes in developing production-grade agentic AI systems for enterprise marketing and customer experience. It fosters a remote-first, collaborative culture focused on innovation and technical excellence.
Design and run experiments to understand model behavior and test new ideas.
Create datasets, benchmarks, and evaluation methods for hard problems.
Work directly with AI labs to turn open-ended goals into concrete research projects.
Vetto builds the infrastructure for next-generation AI training data. They partner with the world’s top AI labs and offer a fast, flat, remote-first culture with high ownership.
Architect and build production AI-agent systems like AIDA and NOVA, shaping how Yuno applies LLMs and agentic AI to payment problems.
Design conversational and multilingual agents that hold real customer conversations across channels, optimized for reliability and low latency.
Define evaluation, safety, and guardrails for agentic systems in a compliance-sensitive, multi-market environment.
Yuno is the AI-native operating system of global commerce, powering financial infrastructure for enterprise merchants, banks, and wallets. It connects over 1,000 payment methods across 190+ countries, trusted by global brands like McDonald's and Rappi, with a remote-first culture.
Build the agent execution runtime: Design and ship the layer that compiles a brief into a running team of agents.
Design the tool surface agents work through: Replace "here is a shell, please be careful" with typed, versioned toolpacks.
Make quality measurable, then improve it: Build the evals, benchmarks, and metric sets that score a run.
Ditto is redefining how data moves at the edge by providing a peer-to-peer sync engine for developers to build resilient, real-time applications without internet. With over $145 million in funding and trusted by organizations like Chick-fil-A and Delta Airlines, we are a globally distributed startup committed to building a diverse and inclusive team.
Build and evolve production AI agents on foundation models using AWS Bedrock.
Develop evaluation pipelines and production observability for LLM systems.
Apply context engineering and agent architecture fundamentals to inform orchestration framework decisions.
Smart Working connects exceptional professionals with outstanding global teams through long-term remote opportunities. They are a high-rated workplace on Glassdoor that values integrity, excellence, and professional growth.
Build and ship agentic features for thousands of developers, including orchestration, tool execution, and context management.
Design and implement the foundation for advanced AI agents that handle billions of dollars in purchase volume.
Collaborate with product and engineering teams to scope, prioritize, and drive rollout of new features.
RevenueCat provides infrastructure and tools for app businesses to manage subscriptions and monetization. They are a remote-first team of 150+ people across 25+ countries, guided by values like Customer Obsession, Always Be Shipping, Own It, and Balance.
Build new RL environments targeting different agentic capabilities and industry areas.
Train and evaluate agents in those environments, improving both agents and environments.
Work across modeling and product to identify performance gaps and automate measurement.
Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products. It is a global team of researchers, engineers, and designers headquartered in Toronto with offices worldwide.
Lead rapid prototyping and technical validation of agentic AI solutions and autonomous workflows.
Evaluate foundation models, agent frameworks, and orchestration technologies for enterprise adoption.
Mentor AI engineers and promote strong engineering practices across teams.
This partner company is hiring for a senior technical leadership role focused on shaping next-generation AI agents and autonomous workflows. They are a collaborative enterprise with a culture centered on experimentation, continuous learning, and innovation, offering a remote opportunity in the US.
Design and operate AI agents and automations connecting Salesforce, Gainsight, Slack, and other GTM tools.
Own the architecture of the GTM Systems AI layer, including MCP servers and serverless services on GCP.
Partner with Revenue Operations and GTM leaders to identify automation opportunities and deliver incrementally.
Camunda provides an enterprise platform for agentic orchestration, coordinating AI agents, people, and systems. Trusted by over 700 organizations and recognized as a next unicorn, it is fully remote with an inclusive, innovative culture.
Design, develop, test, and deploy autonomous AI agents using Salesforce Agentforce across various Salesforce clouds.
Build custom Agent Actions (Flows, Apex, Prompt Templates, MuleSoft APIs) and implement agent topics, instructions, and guardrails.
Integrate agents with Data Cloud, external systems, and third-party LLMs, while performing testing and monitoring performance analytics.
We've been forging digital transformation through Salesforce for 20+ years for 3,000+ customers across the globe. We are a Summit status partner with a small but mighty team of ~150 Engagers, driven by values like 'Be Great at What You Do' and 'Be Growth Oriented.'
Conduct independent research in Generative AI, LLMs, NLP, and multimodal AI to design experiments and evaluate models.
Develop and implement LLM evaluation frameworks, analyze model performance, and identify data gaps for improvement.
Apply strong statistical and data science skills to clean, analyze, and interpret complex datasets for AI/ML research.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company is committed to delivering the highest quality data and outstanding outcomes for its customers.
Design and implement novel training optimization techniques for large-scale neural networks.
Investigate approaches to improve training efficiency, stability, and convergence speed.
Collaborate with infrastructure and inference engineering teams to translate research into production performance.
This partner company focuses on AI research and training optimization for large-scale models. They maintain a small, senior-level team that values deep technical thinking and thoughtful execution.
You will design the architecture for specialized subagents operating within live customer conversations, including technical QA, product expertise, and objection handling.
You will build routing and delegation systems that determine when to answer directly, invoke a subagent, or escalate to a human.
You will master the dialogue platform, train AI agents via prompting and fine-tuning, and document workflows to educate the team.
1mind builds autonomous customer experience software that deploys AI-powered 'Superhumans' to engage, demo, onboard, and support customers across the entire buying journey. The company offers a remote-first, fast-moving culture with ownership, autonomy, and impact from day one.
Design and implement model-based RL agents and planning controllers for real industrial systems.
Develop learned world models and training pipelines to make control agents reliable and safe.
Work closely with researchers and engineers to translate cutting-edge RL research into production outcomes.
Phaidra builds AI-powered control systems for industrial facilities, using reinforcement learning to help them learn and improve over time. The company is a fully remote, international team with roots at DeepMind and Google, and values Agency, Velocity, Craft, and Truth.