A History of Graph Neural Networks in Drug Discovery
The graph was already the model. ECFP hashed a neighborhood. A GNN learns the neighborhood. The later step, the one that changes a binding model, is putting the protein in the same graph and then refusing to trust a single pose.
Graphs before they were learned
A molecule is a graph: atoms as nodes, bonds as edges. ECFP and Morgan fingerprints, and a lot of QSAR descriptors, hash a local neighborhood into a fixed vector. Graph kernels did the same job with a hand-built similarity, before message passing was trained.
Protein–ligand models already built interaction graphs. A ligand atom or a residue linked to a protein residue if they were in contact, and a summary of that graph went into a classical model. The graph was a feature recipe. It was not trained end to end.
Ligand graphs, mid-2010s
Message passing, graph convolution, and graph attention made the graph the input. The first drug-discovery use was ligand-only: solubility, toxicity, activity, ADMET. The model learned embeddings and passed them along bonds. For several years that was what “a GNN for drug discovery” meant. The protein, if it appeared, arrived as a sequence or a profile, or not at all.
The complex is the graph
The shift is to treat the complex as one object.
- Two graphs, protein and ligand, plus edges between residues and atoms inside a distance cutoff.
- One graph of both atom sets, with edges for contacts, distances, and interaction types.
- Message passing that crosses those edges, so the 3D arrangement is part of the update.
The site is then learnable. Hydrogen bonds, hydrophobic contacts, and metal coordination can be edge labels. Attention decides which of them matter.
Geometry and attention
Three habits stuck.
Distances, radial basis functions, or an SE(3)-equivariant layer keep the coordinates and the symmetry. The model starts to look like a scoring function with trained parameters.
Edges stop meaning only “bonded” or “within some angstrom cutoff.” They carry the interaction class. A cytochrome is the example in this note: iron coordination and a few C–H distances dominate the mechanism, so an edge label is the natural place for them.
GAT-style attention weights the neighbors. Attentive pooling, or multiple-instance learning over poses, pockets, and conformers, chooses which part of a messy input to trust.
Poses as a bag
Docking and MD return an ensemble. The pipeline this note describes:
- One protein–ligand graph per pose, with an edge-aware GAT such as EdgeGATConv.
- A readout to one embedding per pose.
- Attention-based MIL (ABMIL) pools those embeddings and learns which poses matter for binding or for substrate classification.
- The prediction is made from that pooled vector.
The docking score is no longer the only way to pick a pose. The model can learn what a productive pose looks like, and it can also learn to ignore the rest.
Fingerprints made the graph implicit. Ligand GNNs made it trained. Complex graphs added the site. Equivariant layers tied the result to geometry. MIL is the step that stops pretending one frame was the complex.