Beyond Static Models: Agentic AI and Multi-Agent Systems in Drug Discovery
The model here is the controller. It breaks a goal into steps, writes and runs code, calls docking or a quantum package, reads the result, and revises the plan. One agent proposing, another implementing, and a third critiquing is a multi-agent loop. A static activity, solubility, or toxicity model still answers a single question. This note is about the loop around that question.
The sections below cover what changed, the design patterns, and five GitHub projects that implement them. The last section is where the loop fails.
What actually changed
Three capabilities, layered on top of ordinary LLMs, do most of the work.
1. Tool use. Instead of answering from parametric memory, the agent emits a structured call to an external tool - run a docking job, compute partial charges, query a database - and feeds the result back into its reasoning. The chemistry is still done by the specialised tool; the LLM orchestrates it.
2. The Model Context Protocol (MCP). MCP is an open standard for exposing tools, data sources, and functions to an LLM through a uniform interface. Before MCP, every integration was bespoke. With it, a docking engine, a UniProt query, or an ADMET predictor can be published once as an MCP server and reused across any MCP-aware agent - which is why several of the projects below ship as MCP servers rather than monolithic apps.
3. Retrieval-augmented generation (RAG). Rather than fine-tuning a model on proprietary chemistry (expensive and quickly stale), the agent retrieves relevant facts from external databases at query time and reasons over them. This grounds outputs in real data and makes provenance auditable - important when a prediction has to survive scrutiny.
Together these turn a frozen model into something that can act, look up ground truth, and correct course.
The design patterns
Most agentic drug-discovery systems are variations on a few recurring architectures. Recognising them makes any new project easier to evaluate.
- Planner / executor. One agent operates at a high level of abstraction - laying out a discovery strategy - while a second translates each step into concrete, executable code. Separating “what to do” from “how to do it” keeps the reasoning legible and the code focused.
- Multi-agent critique. A generator proposes a hypothesis or molecule; a separate critic agent challenges it, checks it against constraints, and sends it back for revision. The adversarial framing catches errors a single pass would miss.
- RAG-grounded reasoning. Agents are wired to authoritative sources so claims are anchored in retrieved evidence rather than generated from memory.
- Tool-augmented control. The LLM is a controller over a toolbox - docking, molecular dynamics, QM, property predictors - and contributes orchestration and interpretation, not the underlying physics.
Keep these in mind as a lens for the projects that follow.
Five open-source agents worth examining
The projects below are concrete implementations of the patterns above. They are research-grade and evolve quickly, so verify the repository URL, license, and current status before depending on any of them - links and capabilities drift.
1. CLADD - Genentech
A multi-agent, RAG-driven framework that connects LLMs to biochemical databases without task-specific fine-tuning. Multiple agents collaborate to contextualise a molecule from retrieved data and produce property estimates such as toxicity and binding affinity, with the retrieval step keeping the reasoning grounded in real records. 🔗 github.com/Genentech/CLADD
2. AgentD - hoon-ock
An end-to-end agent for molecular design, virtual screening, and SMILES refinement. It folds drug-likeness filters (Lipinski, Veber) directly into the loop and can generate 3D protein–ligand structures via Boltz. It runs as an MCP server, so its capabilities can be plugged into other MCP-aware clients. 🔗 github.com/hoon-ock/AgentD
3. DrugAgent - FermiQ
A clean planner/executor split. An “LLM Planner” formulates high-level discovery strategies; an “LLM Instructor” turns each into executable Python - for example, assembling and running an ADMET evaluation. Dividing labour between strategy and implementation keeps both halves auditable. 🔗 github.com/FermiQ/drugagent
4. DEDA - Drug Evaluation and Discovery Agent
A local, MCP-based assistant focused on bioinformatics and molecular pathways - effectively a domain-specialised chat agent. It translates plain-English commands into queries against resources like UniProt and AlphaFold and supports real-time binding-pocket mapping. 🔗 github.com/drug-discovery-ai/deda-drug-evaluation-and-discovery-agent
5. Scientific Agent Skills - K-Dense-AI
Less a single agent than a skill library that turns a general LLM (Claude, Gemini, and others) into a scientific operator. It bundles 140+ ready-made skills connecting agents to 70+ scientific databases, spanning tasks from molecular dynamics to genomics to docking. 🔗 github.com/k-dense-ai/scientific-agent-skills
The lineage: this didn’t appear from nowhere
It’s worth situating these tools in the prior work that established the paradigm, because the ideas are a few years old even if the MCP packaging is new.
- ChemCrow (Bran et al., 2024) coupled an LLM to a suite of expert chemistry tools and showed it could plan and execute synthesis and discovery tasks - an early, influential demonstration that tool-augmented LLMs outperform the bare model on real chemical problems.
- Coscientist (Boiko et al., Nature, 2023) used a GPT-4-based agent to autonomously plan and carry out experiments, including driving physical lab hardware - a landmark for closed-loop autonomous experimentation.
The five projects above inherit directly from this lineage; the recent additions are the standardisation (MCP) and the multi-agent collaboration that make such systems easier to assemble and extend.
The caveats the hype skips
Agentic systems are powerful, but they are not a substitute for judgment, and treating them as one is how pipelines produce confident nonsense.
- Hallucination still happens - now with tools. Grounding reduces fabrication but doesn’t eliminate it. An agent can call the right tool with the wrong arguments, or misinterpret a correct result. Outputs need validation, not trust.
- The physics lives in the tools, not the LLM. A docking score is only as good as the docking engine and its setup. The agent’s reasoning cannot rescue a poorly prepared receptor or bad charges (see the partial-charges note).
- Reproducibility is hard. Non-deterministic LLM behaviour makes runs difficult to reproduce exactly. Pin model versions, log every tool call and its inputs, and treat agent transcripts as part of the experimental record.
- Cost and latency add up. Multi-agent loops with critique steps can issue many model calls per task; this matters at virtual-screening scale.
- Research-grade maturity. Most of these repositories are young and fast-moving. Check the license, the test coverage, and how actively the project is maintained before building on it.
The right framing is augmentation: these agents are excellent at orchestrating tedious multi-tool workflows, drafting analysis code, and surfacing candidates - while the scientist sets the question, validates the chemistry, and owns the conclusions.
Takeaway
The shift from static models to agentic, tool-using systems is real and substantive: with MCP for standardised tool access, RAG for grounding, and multi-agent collaboration for planning and self-critique, AI is moving from answering questions to running workflows. The five open-source projects here are a good entry point for seeing how that looks in practice - clone one, read how it wires the LLM to its tools, and you’ll understand the pattern far better than any summary can convey. Just keep the validation discipline of a computational chemist: the agent accelerates the work, but you still answer for the result.