Cheminformatics

Chemical Space Is Not One Map: PCA, t-SNE, UMAP, and What They Actually Show

PCA, t-SNE, and UMAP can produce very different views of the same molecular dataset. The reason is simple: chemical space depends on how molecules are represented, how similarity is defined, and how that high-dimensional information is projected.

Oct 1, 2026

Sorbent
Sorbent

A deterministic compound triage service for molecular libraries. Ranks, deduplicates, and flags structural liabilities from SMILES inputs using transparent, citable RDKit descriptors and published structural alert catalogs.

Sep 22, 2026

Assayer
Assayer

A curated catalogue of 3,973 tools for medicinal and computational chemistry, paired with 27 actionable protocols guiding step-by-step drug discovery workflows, target validation against UniProt, and real-time PDB structure prioritization.

Sep 13, 2026

Making AI Work in Discovery Chemistry: Precision, Trust, and Practical Value

Key takeaways from the Optibrium panel discussion featuring Nathan Brown, Charlotte Deane, Pat Walters, Chris Swain, and Paul Czodrowski on AI strategy, data leakage, generative risks, and LLMs in drug discovery.

Jul 30, 2026

Protein Descriptors for Machine Learning in Drug Discovery

A practical map of protein and protein-ligand descriptors for machine learning modeling, from sequence-only features to MD-derived structural and interaction features.

Jul 20, 2026

BioLatent
BioLatent

An open-access registry, selection wizard, and interactive benchmark dashboard for chemical and biological vector representations. Indexes foundation models across five modalities: small molecules, proteins, complexes, chemical reactions, and nucleic acids, with a programmatic JSON API.

Jul 16, 2026

Reproducible ≠ Robust: Why One UMAP Seed Isn't Enough to Trust a Split

Pinning random_state=42 makes a UMAP-based train/test split perfectly reproducible - and that is exactly why it can lull you into a false sense of rigor. Reproducibility guarantees you get the same answer every run; it says nothing about whether that answer is typical. Here's the distinction, why it matters for evaluating GNNs on chemical-domain shifts, and how to fix it.

Jun 23, 2026

Choosing the Right Partial Charges for Molecular Docking: AM1-BCC, PM6 and Beyond

Partial charges quietly drive the electrostatics behind every docking score. This guide compares the common semi-empirical options - AM1-BCC, PM6, AM1/PM3, Gasteiger and RESP - and gives a practical recommendation for when to use each.

Jun 23, 2026

Building a 3D Pharmacophore Model from PDB Data: A Free Python Workflow

A step-by-step, fully open-source pipeline that turns raw Protein Data Bank structures into a ligand-based 3D pharmacophore - mining the PDB, aligning binding pockets, clustering ligands, fixing bond orders, and distilling a consensus feature map ready for virtual screening.

Jun 23, 2026

Beyond Static Models: Agentic AI and Multi-Agent Systems in Drug Discovery

AI in drug discovery is moving past one-shot predictions toward autonomous agents that plan experiments, write their own analysis code, call docking and QM tools, and critique their results. Here's what the shift means, the design patterns behind it, and five open-source agents worth examining - with an honest look at the caveats.

Jun 23, 2026