Design and build AI agents and automation to solve real problems across engineering, product, and delivery.
Partner with stakeholders to identify high-leverage opportunities and deliver end-to-end solutions.
Stay current with LLM and agentic frameworks to drive innovation in healthcare technology.
HealtheDGE provides AI-powered operational infrastructure for health insurance companies, helping them modernize operations. The company is experiencing strong market momentum and invests in its people, offering a collaborative culture focused on innovation.
Conduct independent research in Generative AI, LLMs, NLP, and multimodal AI to design experiments and evaluate models.
Develop and implement LLM evaluation frameworks, analyze model performance, and identify data gaps for improvement.
Apply strong statistical and data science skills to clean, analyze, and interpret complex datasets for AI/ML research.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company is committed to delivering the highest quality data and outstanding outcomes for its customers.
Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
Turn research ideas into working prototypes and production-facing capabilities.
This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.
Develop advanced ML models and agentic workflows to accelerate model development.
Use AI-assisted tools like Claude and Cursor to investigate model behavior and automate analysis.
Set technical direction, mentor engineers, and raise the bar for modeling rigor.
Reddit is a community of communities, built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.
Own model strategy and selection using rigorous benchmarks and statistical analysis.
Design and maintain evaluation methodologies for AI systems, including offline sets and LLM-as-judge frameworks.
Develop classification, fine-tuned, and agentic AI models to improve accuracy, cost, and latency.
A production agentic AI platform that builds and deploys advanced machine learning systems. The team is collaborative and values innovation, offering a remote work environment with high autonomy.
Conduct applied research in machine learning, focusing on time-series forecasting and generative AI for supply chain planning.
Design and run experiments using real-world and benchmark datasets, evaluating models against strong baselines.
Collaborate with researchers and engineers to build prototypes and assess product potential for real-world applications.
Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains with an AI-infused platform. We are a global team of over 2,000 employees with a best-in-class HQ in Ottawa, Canada, and we take our culture seriously.
Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.
Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.
Create original graduate- and PhD-level problems within your STEM expertise.
Develop rigorous, step-by-step solutions and review AI-generated responses for accuracy.
Work fully asynchronously through an online platform to contribute to AI reasoning research.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use technology to streamline recruitment and are part of a network of partner companies.
Identify and implement AI/ML opportunities across products, including generative AI and LLM applications.
Build, evaluate, deploy, and maintain machine learning models in production.
Collaborate with product managers, engineers, and experts to translate business challenges into scalable ML solutions.
The company applies machine learning and AI to real-world business and product challenges. It fosters a collaborative, remote-first environment centered on technical learning, innovation, and professional development.
Build the component layer around our layout synthesis engine, including API contracts, services, and evaluation gates.
Turn model retraining into a one-command job with built-in benchmarks and readable results for the whole team.
Own latency and cost budgets for learned capabilities and partner with infrastructure engineers on MLOps and deployment.
Higharc is a VC-backed startup that is changing how new homes are designed and built using spatial AI and generative floor plan technology. The company is fully remote, has raised over $175M, and values flexibility, collaboration, and asynchronous deep work.
Define the technical roadmap for Growth and Engagement ML systems, ensuring scalability and business impact.
Architect and deploy production-grade ML pipelines and real-time decisioning systems for personalization and onboarding.
Mentor senior engineers and collaborate with Product, Data Science, and Marketing leadership to drive core metrics.
Phantom is on a mission to connect the world to the freedom of open markets, providing access to global markets that never close. With around 180 fully remote employees and backed by a $150M Series C investment from a16z, Sequoia Capital, and Paradigm, we foster a culture of innovation and inclusivity.
Design and implement state-of-the-art ML models and training pipelines for robotics.
Develop efficient data/training strategies and evaluation frameworks for rapid experimentation.
Collaborate with engineering to optimize training infrastructure and deployment.
We're revolutionizing real-world automation by making robotic systems accessible to everyone. Our AI-powered platform brings software automation to physical spaces, and we're a small startup team working across the stack to solve customer problems.
Lead the development and optimization of machine learning models to detect and prevent AI security risks like prompt injection and jailbreaks.
Build reproducible training and evaluation pipelines on Reddit's ML platform, partnering with platform engineers to improve performance and reliability.
Set the technical vision and multi-quarter modeling roadmap, mentoring engineers and establishing best practices for responsible ML development.
Reddit is a community of communities, built on shared interests and authentic conversations, with 100,000+ active communities and 130 million daily active visitors. It is one of the internet's largest sources of information, fostering a culture of openness and trust.
Design agent architectures for planning, reasoning, tool use, and memory integration.
Improve reliability on long-running tasks with failure recovery and evaluation systems.
Build loops for agents to improve with real use, balancing quality, latency, and cost.
Adaption builds AI systems that evolve in real-time, making them flexible and personalized. They are a global-first team focused on talent density and collaboration.
Develop and improve ads ranking models, including prediction objectives, feature interactions, user-history modeling, and calibration.
Take end-to-end ownership of machine learning systems from data pipelines to production integration.
Evaluate and apply advances in deep learning and recommendation modeling to improve ads ranking within production constraints.
Quora operates two knowledge-sharing platforms: Quora, a global Q&A platform, and Poe, a platform for interacting with AI language models. They have a remote-first culture with passionate, collaborative, and high-performing global teams focused on transparency and experimentation.
Build RL environments, agentic systems, LLM pipelines, and evaluation frameworks for real-world AI use cases.
Develop benchmarks and evaluation harnesses to assess model quality across accuracy, safety, latency, and cost.
Conduct fine-tuning and model experiments, deploy self-hosted models, and document reproducible methodologies.
The employer is an organization focused on applied AI research, building practical and reusable AI systems for real-world use cases. It values curiosity, accountability, innovation, collaboration, and continuous learning in a remote environment.
Build and refine models of AI takeoff and estimate their key parameters from public data and new experiments.
Work with engineers to execute large-scale experiments on frontier models and design proposals for safely pacing automated AI R&D.
Publish research papers and engage with academic, policy, and industry communities to help prepare for AI's transformative effects.
P-Zero Research is a public benefit corporation working to improve the long term trajectory of artificial intelligence. It aims to forecast and mitigate the risks of automated AI R&D and values alignment with its mission to keep AI safe.
Design, build, and deploy LLM-powered product features, including lab summaries and conversational agents.
Build backend services integrating LLMs and ML models, primarily using Python with exposure to Elixir.
Implement evaluation, monitoring, and CI/CD workflows for AI features, ensuring reliability and clinical relevance.
Fullscript is a health technology platform that helps practitioners deliver better care through clinical insights, lab interpretations, and patient analytics. With over 125,000 practitioners and 10 million patients, the company emphasizes a people-first culture, teamwork, and continuous learning in a remote-first environment.
Work directly with leading AI labs and enterprises to define research goals and technical requirements.
Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.
Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.
Design and implement autonomous agents and multi-agent orchestration systems for complex, open-ended tasks.\n- Build and optimize production-grade LLM applications with a focus on reliability, observability, and low-latency performance.\n- Architect advanced RAG pipelines and vector database strategies to provide agents with accurate, real-time context.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective evaluation. The company operates with a distributed, global-first team culture, emphasizing fairness and innovation in recruitment.