Chemical Space Is Not One Map: PCA, t-SNE, UMAP, and What They Actually Show
When we visualize a chemical dataset, we often talk about plotting its chemical space.
The phrase sounds objective, almost as if every molecular dataset has one underlying map waiting to be discovered.
It does not.
A chemical-space plot is the result of several choices:
- How do we represent each molecule?
- How do we define similarity between those representations?
- How do we compress hundreds or thousands of dimensions into two?
Change any of these choices and the resulting map can change dramatically.
This becomes especially important when comparing three methods that appear constantly in cheminformatics papers: PCA, t-SNE, and UMAP.
So which one should we use?
The answer depends less on which dimensionality-reduction algorithm is fashionable and more on what we actually mean by chemical similarity.
Chemical Space Starts Before PCA or UMAP
Suppose we have a dataset of 5,000 molecules.
Before applying any dimensionality-reduction algorithm, each molecule must first become a numerical representation.
Molecule
↓
Representation
↓
Distance or similarity
↓
Dimensionality reduction
↓
2D visualization
This first step is arguably more important than choosing between PCA, t-SNE, and UMAP.
A molecule represented using molecular weight, logP, TPSA, hydrogen-bond donors, and hydrogen-bond acceptors occupies a very different mathematical space from the same molecule represented using a 2048-bit ECFP fingerprint.
The first representation describes physicochemical property space.
The second describes local structural patterns.
Both are legitimate chemical spaces, but they answer different questions.
This is why saying:
“I visualized the chemical space using UMAP”
is incomplete.
A more informative statement would be:
“I projected ECFP4 fingerprints into two dimensions using UMAP with Jaccard distance.”
Now we know what molecular similarity means in that visualization.
ECFP4: Turning Molecules Into Structural Fingerprints
Extended-connectivity fingerprints, or ECFPs, encode local atomic environments around each atom.
In simplified terms, the algorithm starts from individual atoms and iteratively expands outward through neighboring bonds.
For ECFP4, the maximum diameter of the encoded environments is approximately four bonds, corresponding to a radius of two iterations.
A molecule can then be represented as a binary vector such as:
0 1 0 0 1 0 1 0 0 1 0 ...
Each activated bit indicates the presence of one or more hashed molecular environments.
This representation is particularly useful in medicinal chemistry because molecules sharing scaffolds or local structural motifs tend to activate overlapping fingerprint bits.
It therefore provides a practical definition of structural similarity.
A terminology note is useful here. RDKit’s Morgan fingerprint with radius 2 is commonly referred to as ECFP4 in cheminformatics workflows. Strictly speaking, it is an implementation closely related to the ECFP algorithm rather than a guarantee of bit-for-bit identity with every commercial ECFP implementation.
Jaccard and Tanimoto: What Does “Close” Mean?
Once molecules become binary fingerprints, we need a way to compare them.
For two sets of fingerprint bits, A and B, the Jaccard similarity is:
For binary molecular fingerprints, this corresponds to the familiar Tanimoto coefficient commonly used in cheminformatics.
If two molecules activate many of the same fingerprint bits, their similarity approaches 1.
If they share few structural features, their similarity approaches 0.
The corresponding distance is:
This gives us something chemically meaningful before dimensionality reduction has even started.
We have defined:
similar molecules = molecules sharing fingerprint environments
That is why using a fingerprint-aware metric matters.
Applying ordinary Euclidean geometry blindly to sparse binary fingerprints is not always the most chemically natural choice.
PCA: The Old Method That Is Still Extremely Useful
Principal Component Analysis is sometimes treated as outdated now that nonlinear approaches such as t-SNE and UMAP are available.
That is a mistake.
PCA solves a different problem.
It identifies new orthogonal directions that capture progressively smaller amounts of variance in the original data.
The first principal component captures the largest possible amount of variance.
The second captures the largest remaining amount while remaining orthogonal to the first.
And so on.
The major advantage is interpretability.
If I calculate descriptors such as:
MW
logP
TPSA
HBD
HBA
rotatable bonds
aromatic fraction
ring count
then PCA can tell me which combinations of these descriptors dominate variation across the dataset.
I can also quantify how much variance is represented:
PC1: 37%
PC2: 21%
Total displayed variance = 58%
That information has a clear statistical interpretation.
This is something UMAP and t-SNE do not provide.
Where PCA works particularly well
PCA is excellent for visualizing physicochemical descriptor space.
For example, if one region of the PCA plot corresponds to high molecular weight, high logP, and many aromatic rings, while another corresponds to smaller molecules with higher polarity and fewer rings, the axes can often be investigated through the PCA loadings.
The map is connected directly to the original variables.
Where PCA struggles
The limitation is that PCA is linear.
Chemical structure rarely varies along neat linear directions.
Two compounds can be structurally related through complicated combinations of fragments that are difficult to express as a simple linear projection.
This becomes particularly noticeable for high-dimensional sparse fingerprints.
PCA can still be applied to fingerprints, but the geometry it assumes is often less aligned with the way medicinal chemists normally think about fingerprint similarity.
t-SNE: Excellent at Finding Neighborhoods
t-distributed Stochastic Neighbor Embedding, better known as t-SNE, approaches dimensionality reduction differently.
Instead of asking:
Which directions explain the most variance?
it asks something closer to:
Which points should remain neighbors after projection?
High-dimensional similarities are converted into probability distributions, and the algorithm attempts to produce a low-dimensional map with similar local relationships.
This is why t-SNE often generates remarkably clean clusters.
A chemical series that is difficult to distinguish in PCA may suddenly appear as a compact island.
For exploratory visualization, this can be extremely useful.
But there is an important trap
People naturally interpret scatterplots geometrically.
If two clusters appear far apart, we instinctively assume they are extremely different.
That interpretation is dangerous with t-SNE.
The algorithm prioritizes local neighborhood preservation.
The spacing between distant clusters is much less informative.
Cluster size can also be misleading.
A visually large island does not necessarily represent greater molecular diversity than a smaller island.
This is one reason attractive t-SNE figures should be interpreted cautiously.
UMAP: A Useful Compromise
Uniform Manifold Approximation and Projection, or UMAP, also builds its representation around local neighborhoods.
At a high level, UMAP constructs a graph describing which observations are neighbors in high-dimensional space.
It then searches for a low-dimensional arrangement that reproduces this neighborhood structure as well as possible.
For chemical fingerprints, this creates an attractive workflow:
ECFP4
↓
Jaccard distance
↓
local molecular neighborhood graph
↓
UMAP
↓
2D structural chemical space
This combination is one reason UMAP has become common in modern cheminformatics.
It is usually computationally efficient and often provides clearer structural organization than PCA when working with molecular fingerprints.
Compared with t-SNE, UMAP can also preserve more of the broader organization of the dataset in many situations.
But the word can matters.
UMAP is still a nonlinear projection.
It does not magically reconstruct the true geometry of chemical space.
The Same Dataset Can Produce Three Different Stories
None of these maps is necessarily wrong.
They are preserving different properties of the original high-dimensional dataset.
This is the central idea:
Dimensionality reduction does not simply reveal structure. It decides which structure is worth preserving.
So Why Do People Still Use PCA?
Because PCA gives us things nonlinear embeddings cannot.
It is:
- fast
- deterministic
- mathematically transparent
- interpretable through feature loadings
- accompanied by explained variance
- useful for detecting correlated descriptors
- appropriate for many continuous molecular properties
If I want to understand the physicochemical composition of a library, PCA may be exactly the method I want.
Replacing PCA with UMAP simply because UMAP produces more visually separated clusters would not necessarily improve the analysis.
In fact, it might remove useful interpretability.
And Why Is t-SNE Still Used?
Because local structure matters.
If the goal is exploratory analysis of closely related molecular families, t-SNE can produce informative visualizations.
It also has a long history, is implemented in essentially every machine-learning ecosystem, and is familiar to researchers and reviewers.
The mistake is not using t-SNE.
The mistake is interpreting a t-SNE plot as if it were an ordinary Cartesian map where all global distances are meaningful.
My Preferred Workflow for Structural Chemical Space
For a typical QSAR dataset where I want to inspect structural diversity, I would usually start with:
SMILES
↓
Morgan / ECFP4 fingerprint
radius = 2
2048 bits
↓
Jaccard distance
↓
UMAP
n_neighbors ≈ 15 to 50
min_dist ≈ 0.05 to 0.3
↓
2D chemical-space visualization
A reasonable starting configuration would be:
from rdkit import Chem
from rdkit.Chem import rdFingerprintGenerator
from rdkit import DataStructs
import numpy as np
import umap
generator = rdFingerprintGenerator.GetMorganGenerator(
radius=2,
fpSize=2048
)
fps = []
arrays = []
for smi in smiles:
mol = Chem.MolFromSmiles(smi)
fp = generator.GetFingerprint(mol)
fps.append(fp)
arr = np.zeros((2048,), dtype=np.uint8)
DataStructs.ConvertToNumpyArray(fp, arr)
arrays.append(arr)
X = np.asarray(arrays)
reducer = umap.UMAP(
n_neighbors=30,
min_dist=0.1,
metric="jaccard",
random_state=42
)
embedding = reducer.fit_transform(X)
Then I would reuse the exact same coordinates and color the map according to different variables:
Plot 1: Active vs inactive
Plot 2: Training vs test set
Plot 3: Experimental pIC50
Plot 4: Molecular weight
Plot 5: Scaffold cluster
This is often much more informative than generating a new embedding for every question.
The geometry stays fixed while the interpretation changes.
The Most Useful QSAR Plot Might Be Train vs Test
One of my favorite applications is visualizing whether the external test set occupies similar structural space to the training compounds.
Color the same UMAP embedding by dataset assignment:
● Training
● Test
A heavily intermixed map suggests that many test compounds live within structural neighborhoods represented during training.
Large isolated regions containing only test compounds may indicate extrapolation.
But even here, the UMAP plot should not be treated as proof.
A two-dimensional projection necessarily discards information.
Two points that look close after projection may not actually have exceptionally high fingerprint similarity.
Two points that look distant may still share substantial structural features.
This is why a visual analysis should be paired with quantitative similarity measurements.
UMAP Is the Picture. Tanimoto Is the Measurement.
Suppose I want to know how novel each test compound is relative to the training set.
For every test molecule, I can calculate:
This asks:
What is the most structurally similar molecule this test compound has already seen during training?
For example:
Test molecule A nearest training similarity = 0.87
Test molecule B nearest training similarity = 0.73
Test molecule C nearest training similarity = 0.51
Test molecule D nearest training similarity = 0.29
Molecule A lies close to known chemical territory.
Molecule D represents substantially stronger structural extrapolation.
This quantity is much easier to interpret than measuring distances between points on a UMAP plot.
That leads to a useful distinction:
UMAP is excellent for seeing chemical space. Fingerprint similarity is better for measuring it.
Do Not Forget the Representation
Suppose two molecules occupy neighboring positions in an ECFP4-based UMAP.
What does that mean?
It means they have similar local topological environments according to that fingerprint representation.
It does not automatically mean they have:
- similar conformations
- similar electrostatics
- similar pharmacophores
- similar biological activity
- similar binding modes
- similar metabolism
Chemical similarity is always conditional on the representation.
The same dataset embedded from physicochemical descriptors might reveal a very different organization.
This is not a contradiction.
It means the two maps are answering different questions.
Why I Would Sometimes Show Both PCA and UMAP
Rather than asking which method wins, it can be more useful to combine complementary views.
PCA on physicochemical descriptors
Shows:
size
lipophilicity
polarity
hydrogen bonding
molecular complexity
This describes property space.
UMAP on ECFP4 fingerprints
Shows:
scaffold relationships
local fragments
structural analogues
chemical series
This describes structural space.
Together, the plots provide more information than either one individually.
Two compounds may be structurally unrelated but occupy similar physicochemical space.
Conversely, close structural analogues can sometimes have substantially different properties after relatively small chemical modifications.
That distinction is central to medicinal chemistry.
A Practical Comparison
| Method | What it emphasizes | Strongest advantage | Main caution |
|---|---|---|---|
| PCA | Global linear variance | Interpretability | Misses nonlinear relationships |
| t-SNE | Local neighborhoods | Strong visual separation | Global distances are difficult to interpret |
| UMAP | Local neighborhoods and some broader organization | Flexible and efficient | Geometry remains projection-dependent |
For molecular fingerprints specifically, I would generally prefer:
ECFP4 + Jaccard + UMAP
for visualization.
For physicochemical descriptors, PCA remains extremely valuable.
For investigating strongly local cluster structure, t-SNE remains a legitimate tool.
There Is No Single Chemical Space
Perhaps the most important lesson is that chemical space is not a single object.
Consider the same set of compounds described using:
ECFP fingerprints
MACCS keys
RDKit descriptors
3D shape descriptors
pharmacophore fingerprints
molecular embeddings
protein-ligand interaction fingerprints
Each representation produces a different notion of similarity.
Each therefore creates a different chemical space.
UMAP cannot solve this problem because it appears only at the end of the pipeline.
The most important question comes earlier:
What molecular information do I want proximity in this map to represent?
Once that is clear, choosing the dimensionality-reduction algorithm becomes much easier.
Final Takeaway
There is no universal winner between PCA, t-SNE, and UMAP.
For continuous physicochemical descriptors, PCA provides an interpretable and statistically transparent view of molecular property space.
For discovering local neighborhoods, t-SNE remains powerful, provided that distances between clusters are not overinterpreted.
For high-dimensional structural fingerprints such as ECFP4, UMAP combined with a binary similarity metric such as Jaccard is one of the most useful default approaches for visual exploration.
But even then, the plot should remain what it is:
a visualization.
If the goal is to quantify structural overlap, novelty, or train-test extrapolation, return to the original fingerprints and calculate the similarities directly.
The picture helps us explore the chemical space.
The high-dimensional representation is where that space actually lives.
References
- Rogers, D.; Hahn, M. Extended-Connectivity Fingerprints. Journal of Chemical Information and Modeling 2010, 50, 742-754.
- van der Maaten, L.; Hinton, G. Visualizing Data using t-SNE. Journal of Machine Learning Research 2008, 9, 2579-2605.
- McInnes, L.; Healy, J.; Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv:1802.03426.