Source Job

Global

  • Develop agentic research systems to automate mechanistic model development for biochemical and physiological processes.
  • Design and train machine learning models for in-context learning and few-shot adaptation on molecular and tabular data.
  • Collaborate with chemists and biologists to drive drug discovery decisions using AI models.

Machine Learning Deep Learning Python PyTorch

20 jobs similar to Machine Learning Researcher - Agentic Science

Jobs ranked by similarity.

$0–$135,000/yr
US

  • Design and build AI agents and automation to solve real problems across engineering, product, and delivery.
  • Partner with stakeholders to identify high-leverage opportunities and deliver end-to-end solutions.
  • Stay current with LLM and agentic frameworks to drive innovation in healthcare technology.

HealtheDGE provides AI-powered operational infrastructure for health insurance companies, helping them modernize operations. The company is experiencing strong market momentum and invests in its people, offering a collaborative culture focused on innovation.

$160,000–$185,000/yr
United States

  • Conduct independent research in Generative AI, LLMs, NLP, and multimodal AI to design experiments and evaluate models.
  • Develop and implement LLM evaluation frameworks, analyze model performance, and identify data gaps for improvement.
  • Apply strong statistical and data science skills to clean, analyze, and interpret complex datasets for AI/ML research.

Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company is committed to delivering the highest quality data and outstanding outcomes for its customers.

$130,000–$200,000/yr
US

  • Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
  • Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
  • Turn research ideas into working prototypes and production-facing capabilities.

This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.

US

  • Develop advanced ML models and agentic workflows to accelerate model development.
  • Use AI-assisted tools like Claude and Cursor to investigate model behavior and automate analysis.
  • Set technical direction, mentor engineers, and raise the bar for modeling rigor.

Reddit is a community of communities, built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.

Brazil

  • Own model strategy and selection using rigorous benchmarks and statistical analysis.
  • Design and maintain evaluation methodologies for AI systems, including offline sets and LLM-as-judge frameworks.
  • Develop classification, fine-tuned, and agentic AI models to improve accuracy, cost, and latency.

A production agentic AI platform that builds and deploys advanced machine learning systems. The team is collaborative and values innovation, offering a remote work environment with high autonomy.

Canada

  • Conduct applied research in machine learning, focusing on time-series forecasting and generative AI for supply chain planning.
  • Design and run experiments using real-world and benchmark datasets, evaluating models against strong baselines.
  • Collaborate with researchers and engineers to build prototypes and assess product potential for real-world applications.

Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains with an AI-infused platform. We are a global team of over 2,000 employees with a best-in-class HQ in Ottawa, Canada, and we take our culture seriously.

India

  • Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
  • Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
  • Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.

Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.

India

  • Create original graduate- and PhD-level problems within your STEM expertise.
  • Develop rigorous, step-by-step solutions and review AI-generated responses for accuracy.
  • Work fully asynchronously through an online platform to contribute to AI reasoning research.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They use technology to streamline recruitment and are part of a network of partner companies.

$85,000–$102,000/yr
US

  • Identify and implement AI/ML opportunities across products, including generative AI and LLM applications.
  • Build, evaluate, deploy, and maintain machine learning models in production.
  • Collaborate with product managers, engineers, and experts to translate business challenges into scalable ML solutions.

The company applies machine learning and AI to real-world business and product challenges. It fosters a collaborative, remote-first environment centered on technical learning, innovation, and professional development.

US Unlimited PTO

  • Build the component layer around our layout synthesis engine, including API contracts, services, and evaluation gates.
  • Turn model retraining into a one-command job with built-in benchmarks and readable results for the whole team.
  • Own latency and cost budgets for learned capabilities and partner with infrastructure engineers on MLOps and deployment.

Higharc is a VC-backed startup that is changing how new homes are designed and built using spatial AI and generative floor plan technology. The company is fully remote, has raised over $175M, and values flexibility, collaboration, and asynchronous deep work.

Global Unlimited PTO

  • Define the technical roadmap for Growth and Engagement ML systems, ensuring scalability and business impact.
  • Architect and deploy production-grade ML pipelines and real-time decisioning systems for personalization and onboarding.
  • Mentor senior engineers and collaborate with Product, Data Science, and Marketing leadership to drive core metrics.

Phantom is on a mission to connect the world to the freedom of open markets, providing access to global markets that never close. With around 180 fully remote employees and backed by a $150M Series C investment from a16z, Sequoia Capital, and Paradigm, we foster a culture of innovation and inclusivity.

$225,000–$250,000/yr
US

  • Design and implement state-of-the-art ML models and training pipelines for robotics.
  • Develop efficient data/training strategies and evaluation frameworks for rapid experimentation.
  • Collaborate with engineering to optimize training infrastructure and deployment.

We're revolutionizing real-world automation by making robotic systems accessible to everyone. Our AI-powered platform brings software automation to physical spaces, and we're a small startup team working across the stack to solve customer problems.

$230,000–$322,000/yr
US

  • Lead the development and optimization of machine learning models to detect and prevent AI security risks like prompt injection and jailbreaks.
  • Build reproducible training and evaluation pipelines on Reddit's ML platform, partnering with platform engineers to improve performance and reliability.
  • Set the technical vision and multi-quarter modeling roadmap, mentoring engineers and establishing best practices for responsible ML development.

Reddit is a community of communities, built on shared interests and authentic conversations, with 100,000+ active communities and 130 million daily active visitors. It is one of the internet's largest sources of information, fostering a culture of openness and trust.

Global

  • Design agent architectures for planning, reasoning, tool use, and memory integration.
  • Improve reliability on long-running tasks with failure recovery and evaluation systems.
  • Build loops for agents to improve with real use, balancing quality, latency, and cost.

Adaption builds AI systems that evolve in real-time, making them flexible and personalized. They are a global-first team focused on talent density and collaboration.

$189,507–$274,604/yr
Global

  • Develop and improve ads ranking models, including prediction objectives, feature interactions, user-history modeling, and calibration.
  • Take end-to-end ownership of machine learning systems from data pipelines to production integration.
  • Evaluate and apply advances in deep learning and recommendation modeling to improve ads ranking within production constraints.

Quora operates two knowledge-sharing platforms: Quora, a global Q&A platform, and Poe, a platform for interacting with AI language models. They have a remote-first culture with passionate, collaborative, and high-performing global teams focused on transparency and experimentation.

India

  • Build RL environments, agentic systems, LLM pipelines, and evaluation frameworks for real-world AI use cases.
  • Develop benchmarks and evaluation harnesses to assess model quality across accuracy, safety, latency, and cost.
  • Conduct fine-tuning and model experiments, deploy self-hosted models, and document reproducible methodologies.

The employer is an organization focused on applied AI research, building practical and reusable AI systems for real-world use cases. It values curiosity, accountability, innovation, collaboration, and continuous learning in a remote environment.

Global

  • Build and refine models of AI takeoff and estimate their key parameters from public data and new experiments.
  • Work with engineers to execute large-scale experiments on frontier models and design proposals for safely pacing automated AI R&D.
  • Publish research papers and engage with academic, policy, and industry communities to help prepare for AI's transformative effects.

P-Zero Research is a public benefit corporation working to improve the long term trajectory of artificial intelligence. It aims to forecast and mitigate the risks of automated AI R&D and values alignment with its mission to keep AI safe.

$103,000–$117,000/yr
Canada Unlimited PTO

  • Design, build, and deploy LLM-powered product features, including lab summaries and conversational agents.
  • Build backend services integrating LLMs and ML models, primarily using Python with exposure to Elixir.
  • Implement evaluation, monitoring, and CI/CD workflows for AI features, ensuring reliability and clinical relevance.

Fullscript is a health technology platform that helps practitioners deliver better care through clinical insights, lab interpretations, and patient analytics. With over 125,000 practitioners and 10 million patients, the company emphasizes a people-first culture, teamwork, and continuous learning in a remote-first environment.

$180,000–$250,000/yr
Global

  • Work directly with leading AI labs and enterprises to define research goals and technical requirements.
  • Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
  • Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.

Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.

Global

  • Design and implement autonomous agents and multi-agent orchestration systems for complex, open-ended tasks.\n- Build and optimize production-grade LLM applications with a focus on reliability, observability, and low-latency performance.\n- Architect advanced RAG pipelines and vector database strategies to provide agents with accurate, real-time context.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective evaluation. The company operates with a distributed, global-first team culture, emphasizing fairness and innovation in recruitment.