Define and execute the research strategy for core AI intelligence, including reasoning, planning, memory, and context representation.
Lead technical decisions on model architecture, balancing proprietary solutions with state-of-the-art models.
Establish and oversee AI alignment, safety, and guardrail strategies to ensure responsible system behavior.
This company develops an advanced AI platform to automate complex digital workflows. They are a small, highly experienced team with a fast-moving culture focused on cutting-edge machine learning.
Own AI evaluation methods and operations for Figma's AI-powered experiences, defining quality dimensions and designing measurement frameworks.
Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars using human and automated approaches.
Produce clear readouts and dashboards to enable stakeholders to make confident shipping decisions.
Figma is on a mission to make design accessible to all, empowering teams to bring ideas to life through collaborative design and prototyping tools. With a growing team of passionate creatives and builders, Figma fosters a culture of growth and inclusivity.
Build and ship AI features end-to-end, from model to system to user experience.
Design and iterate on prompts, tools, memory, and agent workflows for real-world reliability.
Debug full-stack issues and optimize for latency, cost, and production performance.
A1 builds a proactive smart assistant for everyday users, bringing intelligence to conversations, errands, organizing, and workflows with minimal prompting. The team is small, world-class, and focuses on rapid iteration and shipping high-quality AI products.
Design and build the agent execution harness, owning the orchestration layer that routes inputs, manages context, and handles multi-step agentic workflows at scale.
Ensure reliability, observability, and fault tolerance in production AI systems while optimizing for latency and cost.
Lead eval engineering and prompt infrastructure to maintain agent quality and drive data-informed decisions across model updates.
ServiceNow is an AI platform company that automates workflows and helps businesses reinvent themselves through its AI control tower. It serves 85% of the Fortune 500 and fosters an AI-native culture focused on innovation and talent.
Identify and design AI-powered agentic architecture to automate repetitive tasks in creative and technical workflows.
Define end-to-end architecture for AI frameworks, integrating with DCC tools and building reusable components.
Drive adoption, create documentation, and report automation impact to studio leadership.
DreamWorks Animation is a leading producer of animated films and series, known for high-quality, award-winning content. They foster a growth-minded community of artists, technologists, and innovators who value transparency, trust, and collaboration.
Lead the delivery of AI projects end-to-end, establishing technical standards and mentoring a high-performing team.
Guide the design and delivery of RAG systems, agentic frameworks, and LLM-powered solutions for production.
Develop evaluation frameworks and quality standards to ensure reliable AI system performance.
Blend is a premier AI services provider that co-creates impactful solutions using data science, AI, and technology. The company fosters a culture of innovation and collaboration, with a focus on aligning human expertise with artificial intelligence.
Collaborate with product and engineering teams to translate product objectives into autonomous agent-based solutions.
Design and build new agent data types, pipelines, and frameworks to coordinate reasoning, function calling, and actions.
Develop and optimize autonomous agents leveraging LLMs, planning algorithms, and multi-step reasoning approaches.
PointClickCare is a leading health tech company that helps providers deliver exceptional care. As a founder-led, privately held company with over 30,000 provider organizations and 400+ integrated partners, they are recognized by Forbes as a top private cloud company and honored as one of Canada's Most Admired Corporate Cultures, offering flexibility and growth opportunities.
Drive the technical vision for AI in 360Learning's product and coach AI Engineers as the team grows.
Set technical direction, own architectural decisions, and stay close to the code.
Partner with Product Manager, Product Designer, and Full-Stack Developers to deliver AI features.
We enable companies to upskill from within by turning their experts into champions for employee, customer, and partner growth. Founded in 2013, we have raised $240 million with 400+ team members across North America and EMEA.
Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
Benchmark performance by conducting comparative testing between model outputs based on technical metrics.
Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.
Build and implement AI features by selecting the right model and approach for each use case.
Evaluate and monitor AI feature performance in production to ensure accuracy and reliability.
Improve and iterate on prompts and implementations based on how features behave in the wild.
Aphex is a construction execution platform that replaces traditional spreadsheets with collaborative tools for delivery teams. They are a remote-first company with a growing engineering team in the Philippines, serving major contractors on multi-billion dollar projects.
Research data collection strategies and design high-impact data slices that uncover model failure modes.
Model annotator behavior and design experiments to optimize instruction clarity and reward signal reliability.
Develop metrics and frameworks for evaluating dataset quality, diversity, and impact on downstream model alignment.
Surge AI builds a platform that powers the most powerful AI models in partnership with companies like Anthropic, Google, Microsoft, and Meta. They are a profitable, bootstrapped company focused on human intelligence and data quality.
Build asynchronous Python/TypeScript services with FastAPI, Pydantic, and asyncio.
Develop agent orchestration, subagent delegation, tool calling, and structured outputs.
Create behavioral evals and regression datasets using Langfuse.
We are a venture-backed defense-tech company building AI-native intelligent software for sovereign institutions and their affiliated organizations. We have offices across the U.S., Europe, and the Middle East, and we are expanding our engineering team with people who demonstrate strong ownership and sound judgment.
You will build AI-assisted tools, workflow automations, agents, prompts, and integrations to reduce manual effort and improve productivity.
You will partner with business stakeholders to understand high-friction workflows and deliver fit-for-purpose AI solutions.
You will implement engineering controls for data handling, access management, prompt safety, and output validation.
Shield AI is a venture-backed defense-tech company that develops intelligent systems, including Hivemind autonomy software and V-BAT and X-BAT aircraft, to protect service members and civilians. With offices across the U.S., Europe, the Middle East, and Asia-Pacific, the company's technology supports operations worldwide.
Own software architecture and delivery: design, deploy, and operate systems at scale.
Design and evolve AI systems including prompting, retrieval, evaluation infrastructure, and expert feedback loops.
Hire, manage, and grow the engineering team from 3 FTEs while staying on the frontier of AI and legal AI.
Inhouse is the #1 AI lawyer for small to midsize businesses, combining AI, their own law firm, and an expert feedback loop to deliver fast, compliant legal work. They grew revenue 1,500% last year and recently raised a $5M seed round from leading VCs.
Serve as the product bridge between safety research and product, translating model evaluations into guardrails.
Own the safety roadmap, prioritizing features based on research findings and customer requirements.
Partner with modeling teams to interpret safety evaluations and understand model behavior across adversarial inputs.
Cohere is a leading security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products for enterprises. They are a global team of researchers, engineers, and designers passionate about their craft, with offices in Toronto, San Francisco, and other major cities.
Own the engineering enablement track for a major iGaming client's AI adoption program, driving continuous diffusion of AI practices across eight engineering organizations.
Package and transfer working practices into reusable artifacts such as playbooks, spec templates, and skills repos, ensuring adoption across multiple companies.
Operate the rhythm of bi-weekly validation calls, monthly cross-company demo meets, and per-company status boards to measure and optimize AI adoption.
Neurons Lab runs a group-wide AI Adoption Program for a major iGaming client, a holding of six game studios plus central business functions with ~800–1,000 employees. The company focuses on enabling engineering organizations to adopt AI practices through continuous diffusion and transfer of working practices.
Work directly with top AI labs to design custom data pipelines and integrations for training and evaluation data.
Build technical infrastructure for advanced quality control workflows, including model-in-the-loop and human-in-the-loop systems.
Ship fast and iterate constantly, translating ambiguous research needs into high-leverage technical systems.
Surge AI is a platform that powers the most advanced AI models in partnership with leading AI labs like Anthropic, Google, Microsoft, and Meta. Founded by engineers and researchers, the company is profitable from day one without venture funding.
Architect and build production systems across a multi-language stack including C#/.NET, Python, and TypeScript.
Design and drive adoption of offline and online evaluation frameworks for AI components, including test suites, golden datasets, and regression monitoring.
Partner with engineering leadership to surface systemic risks in AI features and lead architecture reviews for major initiatives.
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, they offer innovative solutions with integrated intelligence on a single enterprise platform.
Own AI operations and governance across the enterprise, ensuring systems are reliable, compliant, and secure.
Provide hands-on technical leadership in model evaluation, tuning, and troubleshooting alongside program oversight.
Build and lead a team responsible for AI sustainment, including strategic planning and executive communication.
Shield AI builds autonomous systems for defense and aerospace. As a growing company, it emphasizes a culture of innovation and technical rigor, seeking leaders to operationalize AI governance.
Identify and map workflows to find step-change opportunities for AI automation.
Design and build future-state workflows using agents, integrations, and human-in-the-loop checkpoints.
Deploy and run agents in production, tracking KPIs and iterating on performance.
Natera is a global leader in cell-free DNA testing, focusing on oncology, women's health, and organ health. The company employs a diverse team of dedicated professionals from world-class institutions, fostering a collaborative and inclusive culture.