Design, develop, and deploy production-grade AI-powered backend systems integrating LLMs and traditional ML models.
Integrate and optimize vector databases for RAG pipelines, and write clean, well-structured Python code.
Debug complex cross-layer issues and collaborate with product and engineering teams to deliver cohesive solutions.
We are a fast-growing product company integrating cutting-edge AI capabilities into our core offering to deliver exceptional value to customers. Our small, fast-moving team works on practical, real-world AI applications with high autonomy.
Lead technical execution of complex AI initiatives, designing and delivering ML and LLM systems for patient and provider-facing products.
Own problems end-to-end: scope with clinicians, build datasets, iterate on modeling, and ship to production with monitoring and guardrails.
Develop robust evaluation frameworks including offline benchmarks, human-in-the-loop review, and online experiments to ensure model safety and accuracy.
Curai harnesses artificial intelligence and clinical expertise to make healthcare more affordable and accessible for everyone. They are a small, senior team committed to rigorous research, clinical integrity, and improving health outcomes.
Lead the end-to-end lifecycle of language models and AI solutions, from research to production.
Navigate between closed and open-source ecosystems to maximize quality, optimize latency/cost, and ensure data governance.
Conduct applied research, experimentation, and data curation to create new AI capabilities and continuously improve existing ones.
Blip is a technology company that develops conversational AI and customer service platforms. The company fosters a culture of innovation and technical excellence, with a team of engineers and researchers working on cutting-edge AI solutions.
Identify and map workflows to find step-change opportunities for AI automation.
Design and build future-state workflows using agents, integrations, and human-in-the-loop checkpoints.
Deploy and run agents in production, tracking KPIs and iterating on performance.
Natera is a global leader in cell-free DNA testing, focusing on oncology, women's health, and organ health. The company employs a diverse team of dedicated professionals from world-class institutions, fostering a collaborative and inclusive culture.
Build and improve ML components across data, training, evaluation, and inference.
Implement evaluation and testing to understand model behavior.
Debug model issues, performance problems, and production incidents.
This company builds core ML components for large-scale production systems. They emphasize real-world learning, iteration, and collaboration with senior engineers.
Passionate about building AI models from the ground up and deploying them in production
Excited about joining early-stage startups and working directly with founders
Interested in shaping AI strategy and leading the development of AI-powered products
SignalFire partners with top early-stage startups shaping the future of technology. They have a portfolio of over 200 innovative companies across AI, cybersecurity, healthtech, fintech, developer tools, and enterprise SaaS.
Evaluate LLM architecture logic for technical accuracy and audit ML code and notebooks for efficiency.
Refine RLHF frameworks to align models with human intent and analyze model reasoning in complex chain-of-thought prompts.
Benchmark performance by conducting comparative testing between model outputs based on technical metrics.
Prolific connects researchers with a global pool of participants for collecting high-quality human data to train AI models. With over 35,000 users, they focus on ethical data gathering to advance AI capabilities.
Build and implement AI features by selecting the right model and approach for each use case.
Evaluate and monitor AI feature performance in production to ensure accuracy and reliability.
Improve and iterate on prompts and implementations based on how features behave in the wild.
Aphex is a construction execution platform that replaces traditional spreadsheets with collaborative tools for delivery teams. They are a remote-first company with a growing engineering team in the Philippines, serving major contractors on multi-billion dollar projects.
Build the data and model layer for an AI-enabled decision-support system, including ingestion, record matching, scoring, and generation.
Implement probabilistic matching, calibration, and retrieval-augmented generation in a controlled environment.
Ensure traceability, audit trails, and documentation for all data and model outputs.
DEFCON AI leverages AI, mathematical optimization, and data analytics for resilient optimization of complex systems. The team operates in a fully remote, results-based environment with a focus on innovation and collaboration.
Partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten's platform.
Own the journey from initial exploration to production deployment, translating ambiguous goals into reliable services.
Work across product, software development, performance engineering, and customer-facing implementations.
Baseten powers mission-critical inference for dynamic AI companies like Cursor and Notion. They are rapidly growing, recently raised a $1.5B Series F, and foster a collaborative, forward-thinking culture.
Collaborate with product and engineering teams to translate product objectives into autonomous agent-based solutions.
Design and build new agent data types, pipelines, and frameworks to coordinate reasoning, function calling, and actions.
Develop and optimize autonomous agents leveraging LLMs, planning algorithms, and multi-step reasoning approaches.
PointClickCare is a leading health tech company that helps providers deliver exceptional care. As a founder-led, privately held company with over 30,000 provider organizations and 400+ integrated partners, they are recognized by Forbes as a top private cloud company and honored as one of Canada's Most Admired Corporate Cultures, offering flexibility and growth opportunities.
Build and ship AI features end-to-end, from model to system to user experience.
Design and iterate on prompts, tools, memory, and agent workflows for real-world reliability.
Debug full-stack issues and optimize for latency, cost, and production performance.
A1 builds a proactive smart assistant for everyday users, bringing intelligence to conversations, errands, organizing, and workflows with minimal prompting. The team is small, world-class, and focuses on rapid iteration and shipping high-quality AI products.
Set and evolve the research direction for A1’s core intelligence, including context representation, memory, reasoning, planning, and orchestration.
Define evaluation frameworks that measure real-world usefulness, robustness, safety, and long-term behavior.
Own alignment, safety, and guardrail strategy as first-class product concerns.
A1 builds a proactive smart assistant for everyday users to bring intelligence to conversations, errands, organizing, and workflows. We are a small, high-talent-density team focused on shipping high-quality work and learning at rapid speed.
Design, build, and own agentic systems that produce client-ready financial deliverables from data ingestion to polished output.
Build evaluation and quality systems for generated deliverables, including deterministic checks and model-judged review.
Turn proprietary data into evidence-backed insight by building pipelines that discover and verify patterns.
Farsight is the agentic AI platform for financial services, helping investment banks and private equity firms automate nuanced workflows. The team comes from leading financial institutions and tech companies, focusing on hiring the best to become the best.
Build internal tools and coach-facing features on top of LLMs like Claude and OpenAI.
Turn unstructured conversational data into structured, reliable signals for daily prioritisation.
Own the full stack of what you ship: frontend, backend, model integration, and deployment.
Best10 delivers online, highly personalised diet and lifestyle coaching through WhatsApp and its own platform, working with medical insurers to reduce costs from high-risk members. The company fosters a builder culture where teams ship working AI solutions rapidly and take pride in systems that outlast them.
Run the full eval pipeline end to end, reproducing results and pairing with senior engineers.
Build a judge calibration protocol to measure agreement and identify drift zones.
Extend benchmarks like GAIA and SWE-bench with new tasks targeting capability gaps.
Nous Research is an AI research lab that develops evaluation infrastructure for LLMs. They are a small, high-growth team valuing ownership and rapid shipping.
Build AI agents, workflows, and automations that solve real customer and operations problems.
Apply LLMs, retrieval, and tool calling to practical financial workflows.
Create AI-assisted experiences for onboarding, support, KYC, and risk review.
We are the leading insurance platform in Southeast Asia, helping people plan, save, and grow their money. With over 20 nationalities working remotely and from offices, we are expanding to offer spending, saving, investing, and more.
Build internal tools and coach-facing features on top of LLMs to turn unstructured conversational data into structured signals.
Design and maintain prompt chains, RAG pipelines, and agent workflows, rapidly prototyping from idea to demo in days.
Wire AI into existing systems, own full stack, and sit with coaches to build for their reality.
We deliver highly personalised diet and lifestyle coaching at scale through WhatsApp and our own technology platform. We partner with leading health insurers and employers, and our culture is about using AI to remove administrative work and enhance the personal coach-member relationship.
Drive the technical vision for AI in 360Learning's product and coach AI Engineers as the team grows.
Set technical direction, own architectural decisions, and stay close to the code.
Partner with Product Manager, Product Designer, and Full-Stack Developers to deliver AI features.
We enable companies to upskill from within by turning their experts into champions for employee, customer, and partner growth. Founded in 2013, we have raised $240 million with 400+ team members across North America and EMEA.
Architect and build production systems across a multi-language stack including C#/.NET, Python, and TypeScript.
Design and drive adoption of offline and online evaluation frameworks for AI components, including test suites, golden datasets, and regression monitoring.
Partner with engineering leadership to surface systemic risks in AI features and lead architecture reviews for major initiatives.
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. As a leader in SaaS technology for healthcare, they offer innovative solutions with integrated intelligence on a single enterprise platform.