<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>T-SNE | Yassir Boulaamane</title>
    <link>https://yboulaamane.github.io/tags/t-sne/</link>
      <atom:link href="https://yboulaamane.github.io/tags/t-sne/index.xml" rel="self" type="application/rss+xml" />
    <description>T-SNE</description>
    <generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://yboulaamane.github.io/media/icon_hu_4d696a8ace2a642b.png</url>
      <title>T-SNE</title>
      <link>https://yboulaamane.github.io/tags/t-sne/</link>
    </image>
    
    <item>
      <title>Chemical Space Is Not One Map: PCA, t-SNE, UMAP, and What They Actually Show</title>
      <link>https://yboulaamane.github.io/blog/chemical-space-is-not-one-map-pca-t-sne-umap-and-what-they-actually-show/</link>
      <pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://yboulaamane.github.io/blog/chemical-space-is-not-one-map-pca-t-sne-umap-and-what-they-actually-show/</guid>
      <description>&lt;p&gt;When we visualize a chemical dataset, we often talk about plotting its &lt;strong&gt;chemical space&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The phrase sounds objective, almost as if every molecular dataset has one underlying map waiting to be discovered.&lt;/p&gt;
&lt;p&gt;It does not.&lt;/p&gt;
&lt;p&gt;A chemical-space plot is the result of several choices:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;How do we represent each molecule?&lt;/li&gt;
&lt;li&gt;How do we define similarity between those representations?&lt;/li&gt;
&lt;li&gt;How do we compress hundreds or thousands of dimensions into two?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Change any of these choices and the resulting map can change dramatically.&lt;/p&gt;
&lt;p&gt;This becomes especially important when comparing three methods that appear constantly in cheminformatics papers: &lt;strong&gt;PCA, t-SNE, and UMAP&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;So which one should we use?&lt;/p&gt;
&lt;p&gt;The answer depends less on which dimensionality-reduction algorithm is fashionable and more on &lt;strong&gt;what we actually mean by chemical similarity&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;chemical-space-starts-before-pca-or-umap&#34;&gt;Chemical Space Starts Before PCA or UMAP&lt;/h2&gt;
&lt;p&gt;Suppose we have a dataset of 5,000 molecules.&lt;/p&gt;
&lt;p&gt;Before applying any dimensionality-reduction algorithm, each molecule must first become a numerical representation.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Molecule
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Representation
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Distance or similarity
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Dimensionality reduction
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;2D visualization
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This first step is arguably more important than choosing between PCA, t-SNE, and UMAP.&lt;/p&gt;
&lt;p&gt;A molecule represented using molecular weight, logP, TPSA, hydrogen-bond donors, and hydrogen-bond acceptors occupies a very different mathematical space from the same molecule represented using a 2048-bit ECFP fingerprint.&lt;/p&gt;
&lt;p&gt;The first representation describes &lt;strong&gt;physicochemical property space&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The second describes &lt;strong&gt;local structural patterns&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Both are legitimate chemical spaces, but they answer different questions.&lt;/p&gt;
&lt;p&gt;This is why saying:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;I visualized the chemical space using UMAP&amp;rdquo;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;is incomplete.&lt;/p&gt;
&lt;p&gt;A more informative statement would be:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;I projected ECFP4 fingerprints into two dimensions using UMAP with Jaccard distance.&amp;rdquo;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Now we know what molecular similarity means in that visualization.&lt;/p&gt;
&lt;div style=&#34;margin: 2rem 0;&#34;&gt;
&lt;svg viewBox=&#34;0 0 1200 300&#34; xmlns=&#34;http://www.w3.org/2000/svg&#34; role=&#34;img&#34; aria-labelledby=&#34;title2 desc2&#34;&gt;
  &lt;title id=&#34;title2&#34;&gt;Workflow for structural chemical-space visualization&lt;/title&gt;
  &lt;desc id=&#34;desc2&#34;&gt;A workflow from molecules to ECFP4 fingerprints, Jaccard or Tanimoto geometry, UMAP, and a two-dimensional plot.&lt;/desc&gt;
  &lt;defs&gt;
    &lt;marker id=&#34;arrow&#34; markerWidth=&#34;10&#34; markerHeight=&#34;10&#34; refX=&#34;8&#34; refY=&#34;3&#34; orient=&#34;auto&#34;&gt;
      &lt;path d=&#34;M0,0 L0,6 L9,3 z&#34; fill=&#34;#666&#34;/&gt;
    &lt;/marker&gt;
  &lt;/defs&gt;
  &lt;style&gt;
    .box{fill:#fafafa;stroke:#cfcfcf;stroke-width:1.5;rx:16}
    .head{font:700 27px system-ui,sans-serif;fill:#222}
    .txt{font:17px system-ui,sans-serif;fill:#333}
    .sub{font:15px system-ui,sans-serif;fill:#666}
    .arr{stroke:#666;stroke-width:2;marker-end:url(#arrow)}
  &lt;/style&gt;
  &lt;text x=&#34;600&#34; y=&#34;34&#34; text-anchor=&#34;middle&#34; class=&#34;head&#34;&gt;A chemically sensible map starts before dimensionality reduction&lt;/text&gt;
  &lt;text x=&#34;600&#34; y=&#34;60&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;Representation defines what “similar” means&lt;/text&gt;
  &lt;rect x=&#34;35&#34; y=&#34;110&#34; width=&#34;190&#34; height=&#34;90&#34; class=&#34;box&#34;/&gt;
  &lt;rect x=&#34;280&#34; y=&#34;110&#34; width=&#34;190&#34; height=&#34;90&#34; class=&#34;box&#34;/&gt;
  &lt;rect x=&#34;525&#34; y=&#34;110&#34; width=&#34;220&#34; height=&#34;90&#34; class=&#34;box&#34;/&gt;
  &lt;rect x=&#34;800&#34; y=&#34;110&#34; width=&#34;160&#34; height=&#34;90&#34; class=&#34;box&#34;/&gt;
  &lt;rect x=&#34;1015&#34; y=&#34;110&#34; width=&#34;150&#34; height=&#34;90&#34; class=&#34;box&#34;/&gt;
&lt;p&gt;&lt;text x=&#34;130&#34; y=&#34;145&#34; text-anchor=&#34;middle&#34; class=&#34;txt&#34;&gt;Molecules&lt;/text&gt;
&lt;text x=&#34;130&#34; y=&#34;172&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;SMILES / SDF&lt;/text&gt;&lt;/p&gt;
&lt;p&gt;&lt;text x=&#34;375&#34; y=&#34;145&#34; text-anchor=&#34;middle&#34; class=&#34;txt&#34;&gt;ECFP4&lt;/text&gt;
&lt;text x=&#34;375&#34; y=&#34;172&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;radius = 2&lt;/text&gt;&lt;/p&gt;
&lt;p&gt;&lt;text x=&#34;635&#34; y=&#34;140&#34; text-anchor=&#34;middle&#34; class=&#34;txt&#34;&gt;Jaccard / Tanimoto&lt;/text&gt;
&lt;text x=&#34;635&#34; y=&#34;168&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;fingerprint geometry&lt;/text&gt;&lt;/p&gt;
&lt;p&gt;&lt;text x=&#34;880&#34; y=&#34;145&#34; text-anchor=&#34;middle&#34; class=&#34;txt&#34;&gt;UMAP&lt;/text&gt;
&lt;text x=&#34;880&#34; y=&#34;172&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;2D embedding&lt;/text&gt;&lt;/p&gt;
&lt;p&gt;&lt;text x=&#34;1090&#34; y=&#34;145&#34; text-anchor=&#34;middle&#34; class=&#34;txt&#34;&gt;Plot&lt;/text&gt;
&lt;text x=&#34;1090&#34; y=&#34;172&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;interpret&lt;/text&gt;&lt;/p&gt;
  &lt;line x1=&#34;225&#34; y1=&#34;155&#34; x2=&#34;270&#34; y2=&#34;155&#34; class=&#34;arr&#34;/&gt;
  &lt;line x1=&#34;470&#34; y1=&#34;155&#34; x2=&#34;515&#34; y2=&#34;155&#34; class=&#34;arr&#34;/&gt;
  &lt;line x1=&#34;745&#34; y1=&#34;155&#34; x2=&#34;790&#34; y2=&#34;155&#34; class=&#34;arr&#34;/&gt;
  &lt;line x1=&#34;960&#34; y1=&#34;155&#34; x2=&#34;1005&#34; y2=&#34;155&#34; class=&#34;arr&#34;/&gt;
&lt;p&gt;&lt;text x=&#34;600&#34; y=&#34;250&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;The projection compresses geometry. It does not define the chemistry.&lt;/text&gt;
&lt;/svg&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id=&#34;ecfp4-turning-molecules-into-structural-fingerprints&#34;&gt;ECFP4: Turning Molecules Into Structural Fingerprints&lt;/h2&gt;
&lt;p&gt;Extended-connectivity fingerprints, or ECFPs, encode local atomic environments around each atom.&lt;/p&gt;
&lt;p&gt;In simplified terms, the algorithm starts from individual atoms and iteratively expands outward through neighboring bonds.&lt;/p&gt;
&lt;p&gt;For ECFP4, the maximum diameter of the encoded environments is approximately four bonds, corresponding to a radius of two iterations.&lt;/p&gt;
&lt;p&gt;A molecule can then be represented as a binary vector such as:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;0 1 0 0 1 0 1 0 0 1 0 ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Each activated bit indicates the presence of one or more hashed molecular environments.&lt;/p&gt;
&lt;p&gt;This representation is particularly useful in medicinal chemistry because molecules sharing scaffolds or local structural motifs tend to activate overlapping fingerprint bits.&lt;/p&gt;
&lt;p&gt;It therefore provides a practical definition of &lt;strong&gt;structural similarity&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;A terminology note is useful here. RDKit&amp;rsquo;s Morgan fingerprint with radius 2 is commonly referred to as ECFP4 in cheminformatics workflows. Strictly speaking, it is an implementation closely related to the ECFP algorithm rather than a guarantee of bit-for-bit identity with every commercial ECFP implementation.&lt;/p&gt;
&lt;h2 id=&#34;jaccard-and-tanimoto-what-does-close-mean&#34;&gt;Jaccard and Tanimoto: What Does &amp;ldquo;Close&amp;rdquo; Mean?&lt;/h2&gt;
&lt;p&gt;Once molecules become binary fingerprints, we need a way to compare them.&lt;/p&gt;
&lt;p&gt;For two sets of fingerprint bits, A and B, the Jaccard similarity is:&lt;/p&gt;
&lt;div style=&#34;text-align:center; margin:1.5rem 0;&#34;&gt;
&lt;strong&gt;J(A,B) = |A ∩ B| / |A ∪ B|&lt;/strong&gt;
&lt;/div&gt;
&lt;p&gt;For binary molecular fingerprints, this corresponds to the familiar &lt;strong&gt;Tanimoto coefficient&lt;/strong&gt; commonly used in cheminformatics.&lt;/p&gt;
&lt;p&gt;If two molecules activate many of the same fingerprint bits, their similarity approaches 1.&lt;/p&gt;
&lt;p&gt;If they share few structural features, their similarity approaches 0.&lt;/p&gt;
&lt;p&gt;The corresponding distance is:&lt;/p&gt;
&lt;div style=&#34;text-align:center; margin:1.5rem 0;&#34;&gt;
&lt;strong&gt;d = 1 - J(A,B)&lt;/strong&gt;
&lt;/div&gt;
&lt;p&gt;This gives us something chemically meaningful before dimensionality reduction has even started.&lt;/p&gt;
&lt;p&gt;We have defined:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;similar molecules = molecules sharing fingerprint environments
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is why using a fingerprint-aware metric matters.&lt;/p&gt;
&lt;p&gt;Applying ordinary Euclidean geometry blindly to sparse binary fingerprints is not always the most chemically natural choice.&lt;/p&gt;
&lt;h2 id=&#34;pca-the-old-method-that-is-still-extremely-useful&#34;&gt;PCA: The Old Method That Is Still Extremely Useful&lt;/h2&gt;
&lt;p&gt;Principal Component Analysis is sometimes treated as outdated now that nonlinear approaches such as t-SNE and UMAP are available.&lt;/p&gt;
&lt;p&gt;That is a mistake.&lt;/p&gt;
&lt;p&gt;PCA solves a different problem.&lt;/p&gt;
&lt;p&gt;It identifies new orthogonal directions that capture progressively smaller amounts of variance in the original data.&lt;/p&gt;
&lt;p&gt;The first principal component captures the largest possible amount of variance.&lt;/p&gt;
&lt;p&gt;The second captures the largest remaining amount while remaining orthogonal to the first.&lt;/p&gt;
&lt;p&gt;And so on.&lt;/p&gt;
&lt;p&gt;The major advantage is interpretability.&lt;/p&gt;
&lt;p&gt;If I calculate descriptors such as:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;MW
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;logP
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;TPSA
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;HBD
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;HBA
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;rotatable bonds
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;aromatic fraction
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;ring count
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;then PCA can tell me which combinations of these descriptors dominate variation across the dataset.&lt;/p&gt;
&lt;p&gt;I can also quantify how much variance is represented:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;PC1: 37%
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;PC2: 21%
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Total displayed variance = 58%
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That information has a clear statistical interpretation.&lt;/p&gt;
&lt;p&gt;This is something UMAP and t-SNE do not provide.&lt;/p&gt;
&lt;h3 id=&#34;where-pca-works-particularly-well&#34;&gt;Where PCA works particularly well&lt;/h3&gt;
&lt;p&gt;PCA is excellent for visualizing &lt;strong&gt;physicochemical descriptor space&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;For example, if one region of the PCA plot corresponds to high molecular weight, high logP, and many aromatic rings, while another corresponds to smaller molecules with higher polarity and fewer rings, the axes can often be investigated through the PCA loadings.&lt;/p&gt;
&lt;p&gt;The map is connected directly to the original variables.&lt;/p&gt;
&lt;h3 id=&#34;where-pca-struggles&#34;&gt;Where PCA struggles&lt;/h3&gt;
&lt;p&gt;The limitation is that PCA is linear.&lt;/p&gt;
&lt;p&gt;Chemical structure rarely varies along neat linear directions.&lt;/p&gt;
&lt;p&gt;Two compounds can be structurally related through complicated combinations of fragments that are difficult to express as a simple linear projection.&lt;/p&gt;
&lt;p&gt;This becomes particularly noticeable for high-dimensional sparse fingerprints.&lt;/p&gt;
&lt;p&gt;PCA can still be applied to fingerprints, but the geometry it assumes is often less aligned with the way medicinal chemists normally think about fingerprint similarity.&lt;/p&gt;
&lt;h2 id=&#34;t-sne-excellent-at-finding-neighborhoods&#34;&gt;t-SNE: Excellent at Finding Neighborhoods&lt;/h2&gt;
&lt;p&gt;t-distributed Stochastic Neighbor Embedding, better known as t-SNE, approaches dimensionality reduction differently.&lt;/p&gt;
&lt;p&gt;Instead of asking:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Which directions explain the most variance?&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;it asks something closer to:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Which points should remain neighbors after projection?&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;High-dimensional similarities are converted into probability distributions, and the algorithm attempts to produce a low-dimensional map with similar local relationships.&lt;/p&gt;
&lt;p&gt;This is why t-SNE often generates remarkably clean clusters.&lt;/p&gt;
&lt;p&gt;A chemical series that is difficult to distinguish in PCA may suddenly appear as a compact island.&lt;/p&gt;
&lt;p&gt;For exploratory visualization, this can be extremely useful.&lt;/p&gt;
&lt;h3 id=&#34;but-there-is-an-important-trap&#34;&gt;But there is an important trap&lt;/h3&gt;
&lt;p&gt;People naturally interpret scatterplots geometrically.&lt;/p&gt;
&lt;p&gt;If two clusters appear far apart, we instinctively assume they are extremely different.&lt;/p&gt;
&lt;p&gt;That interpretation is dangerous with t-SNE.&lt;/p&gt;
&lt;p&gt;The algorithm prioritizes &lt;strong&gt;local neighborhood preservation&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The spacing between distant clusters is much less informative.&lt;/p&gt;
&lt;p&gt;Cluster size can also be misleading.&lt;/p&gt;
&lt;p&gt;A visually large island does not necessarily represent greater molecular diversity than a smaller island.&lt;/p&gt;
&lt;p&gt;This is one reason attractive t-SNE figures should be interpreted cautiously.&lt;/p&gt;
&lt;h2 id=&#34;umap-a-useful-compromise&#34;&gt;UMAP: A Useful Compromise&lt;/h2&gt;
&lt;p&gt;Uniform Manifold Approximation and Projection, or UMAP, also builds its representation around local neighborhoods.&lt;/p&gt;
&lt;p&gt;At a high level, UMAP constructs a graph describing which observations are neighbors in high-dimensional space.&lt;/p&gt;
&lt;p&gt;It then searches for a low-dimensional arrangement that reproduces this neighborhood structure as well as possible.&lt;/p&gt;
&lt;p&gt;For chemical fingerprints, this creates an attractive workflow:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;ECFP4
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Jaccard distance
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;local molecular neighborhood graph
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;UMAP
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;2D structural chemical space
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This combination is one reason UMAP has become common in modern cheminformatics.&lt;/p&gt;
&lt;p&gt;It is usually computationally efficient and often provides clearer structural organization than PCA when working with molecular fingerprints.&lt;/p&gt;
&lt;p&gt;Compared with t-SNE, UMAP can also preserve more of the broader organization of the dataset in many situations.&lt;/p&gt;
&lt;p&gt;But the word &lt;strong&gt;can&lt;/strong&gt; matters.&lt;/p&gt;
&lt;p&gt;UMAP is still a nonlinear projection.&lt;/p&gt;
&lt;p&gt;It does not magically reconstruct the true geometry of chemical space.&lt;/p&gt;
&lt;div style=&#34;margin: 2rem 0;&#34;&gt;
&lt;svg viewBox=&#34;0 0 1200 380&#34; xmlns=&#34;http://www.w3.org/2000/svg&#34; role=&#34;img&#34; aria-labelledby=&#34;title1 desc1&#34;&gt;
  &lt;title id=&#34;title1&#34;&gt;Conceptual comparison of PCA, t-SNE, and UMAP&lt;/title&gt;
  &lt;desc id=&#34;desc1&#34;&gt;A schematic showing overlapping structure for PCA, separated local clusters for t-SNE, and structured neighborhoods for UMAP.&lt;/desc&gt;
  &lt;style&gt;
    .panel { fill:#fafafa; stroke:#d0d0d0; stroke-width:1.5; rx:18; }
    .ttl { font:700 28px system-ui,sans-serif; fill:#222; }
    .sub { font:16px system-ui,sans-serif; fill:#666; }
    .pt1{fill:#4c78a8}.pt2{fill:#f58518}.pt3{fill:#54a24b}.pt4{fill:#b279a2}
  &lt;/style&gt;
  &lt;text x=&#34;600&#34; y=&#34;34&#34; text-anchor=&#34;middle&#34; class=&#34;ttl&#34;&gt;The same chemical dataset can tell different visual stories&lt;/text&gt;
  &lt;text x=&#34;600&#34; y=&#34;61&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;Conceptual illustration only&lt;/text&gt;
  &lt;rect x=&#34;30&#34; y=&#34;85&#34; width=&#34;350&#34; height=&#34;250&#34; class=&#34;panel&#34;/&gt;
  &lt;rect x=&#34;425&#34; y=&#34;85&#34; width=&#34;350&#34; height=&#34;250&#34; class=&#34;panel&#34;/&gt;
  &lt;rect x=&#34;820&#34; y=&#34;85&#34; width=&#34;350&#34; height=&#34;250&#34; class=&#34;panel&#34;/&gt;
&lt;p&gt;&lt;text x=&#34;205&#34; y=&#34;120&#34; text-anchor=&#34;middle&#34; class=&#34;ttl&#34;&gt;PCA&lt;/text&gt;
&lt;text x=&#34;205&#34; y=&#34;146&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;Linear variance&lt;/text&gt;&lt;/p&gt;
&lt;p&gt;&lt;text x=&#34;600&#34; y=&#34;120&#34; text-anchor=&#34;middle&#34; class=&#34;ttl&#34;&gt;t-SNE&lt;/text&gt;
&lt;text x=&#34;600&#34; y=&#34;146&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;Local neighborhoods&lt;/text&gt;&lt;/p&gt;
&lt;p&gt;&lt;text x=&#34;995&#34; y=&#34;120&#34; text-anchor=&#34;middle&#34; class=&#34;ttl&#34;&gt;UMAP&lt;/text&gt;
&lt;text x=&#34;995&#34; y=&#34;146&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;Local structure + broader organization&lt;/text&gt;&lt;/p&gt;
  &lt;!-- PCA overlapping clouds --&gt;
  &lt;g opacity=&#34;0.78&#34;&gt;
    &lt;circle cx=&#34;115&#34; cy=&#34;220&#34; r=&#34;9&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;135&#34; cy=&#34;210&#34; r=&#34;8&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;150&#34; cy=&#34;230&#34; r=&#34;10&#34; class=&#34;pt1&#34;/&gt;
    &lt;circle cx=&#34;168&#34; cy=&#34;218&#34; r=&#34;8&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;180&#34; cy=&#34;242&#34; r=&#34;9&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;197&#34; cy=&#34;223&#34; r=&#34;8&#34; class=&#34;pt1&#34;/&gt;
    &lt;circle cx=&#34;160&#34; cy=&#34;197&#34; r=&#34;7&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;182&#34; cy=&#34;190&#34; r=&#34;9&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;205&#34; cy=&#34;203&#34; r=&#34;8&#34; class=&#34;pt2&#34;/&gt;
    &lt;circle cx=&#34;220&#34; cy=&#34;220&#34; r=&#34;10&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;235&#34; cy=&#34;198&#34; r=&#34;8&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;250&#34; cy=&#34;215&#34; r=&#34;9&#34; class=&#34;pt2&#34;/&gt;
    &lt;circle cx=&#34;215&#34; cy=&#34;245&#34; r=&#34;7&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;235&#34; cy=&#34;240&#34; r=&#34;9&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;258&#34; cy=&#34;250&#34; r=&#34;8&#34; class=&#34;pt3&#34;/&gt;
    &lt;circle cx=&#34;275&#34; cy=&#34;234&#34; r=&#34;10&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;290&#34; cy=&#34;249&#34; r=&#34;8&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;305&#34; cy=&#34;230&#34; r=&#34;9&#34; class=&#34;pt3&#34;/&gt;
    &lt;circle cx=&#34;245&#34; cy=&#34;180&#34; r=&#34;8&#34; class=&#34;pt4&#34;/&gt;&lt;circle cx=&#34;265&#34; cy=&#34;175&#34; r=&#34;10&#34; class=&#34;pt4&#34;/&gt;&lt;circle cx=&#34;285&#34; cy=&#34;188&#34; r=&#34;8&#34; class=&#34;pt4&#34;/&gt;
    &lt;circle cx=&#34;302&#34; cy=&#34;195&#34; r=&#34;9&#34; class=&#34;pt4&#34;/&gt;&lt;circle cx=&#34;318&#34; cy=&#34;180&#34; r=&#34;8&#34; class=&#34;pt4&#34;/&gt;
  &lt;/g&gt;
  &lt;!-- t-SNE islands --&gt;
  &lt;g opacity=&#34;0.82&#34;&gt;
    &lt;circle cx=&#34;505&#34; cy=&#34;205&#34; r=&#34;11&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;525&#34; cy=&#34;192&#34; r=&#34;9&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;540&#34; cy=&#34;215&#34; r=&#34;10&#34; class=&#34;pt1&#34;/&gt;
    &lt;circle cx=&#34;560&#34; cy=&#34;198&#34; r=&#34;8&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;520&#34; cy=&#34;228&#34; r=&#34;8&#34; class=&#34;pt1&#34;/&gt;
    &lt;circle cx=&#34;665&#34; cy=&#34;195&#34; r=&#34;10&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;686&#34; cy=&#34;205&#34; r=&#34;9&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;704&#34; cy=&#34;190&#34; r=&#34;10&#34; class=&#34;pt2&#34;/&gt;
    &lt;circle cx=&#34;718&#34; cy=&#34;212&#34; r=&#34;8&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;680&#34; cy=&#34;225&#34; r=&#34;8&#34; class=&#34;pt2&#34;/&gt;
    &lt;circle cx=&#34;585&#34; cy=&#34;265&#34; r=&#34;10&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;605&#34; cy=&#34;275&#34; r=&#34;9&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;625&#34; cy=&#34;260&#34; r=&#34;10&#34; class=&#34;pt3&#34;/&gt;
    &lt;circle cx=&#34;642&#34; cy=&#34;280&#34; r=&#34;8&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;610&#34; cy=&#34;245&#34; r=&#34;8&#34; class=&#34;pt3&#34;/&gt;
    &lt;circle cx=&#34;575&#34; cy=&#34;165&#34; r=&#34;10&#34; class=&#34;pt4&#34;/&gt;&lt;circle cx=&#34;598&#34; cy=&#34;158&#34; r=&#34;9&#34; class=&#34;pt4&#34;/&gt;&lt;circle cx=&#34;620&#34; cy=&#34;170&#34; r=&#34;10&#34; class=&#34;pt4&#34;/&gt;
    &lt;circle cx=&#34;640&#34; cy=&#34;158&#34; r=&#34;8&#34; class=&#34;pt4&#34;/&gt;
  &lt;/g&gt;
  &lt;!-- UMAP neighborhoods --&gt;
  &lt;g opacity=&#34;0.82&#34;&gt;
    &lt;circle cx=&#34;885&#34; cy=&#34;225&#34; r=&#34;10&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;905&#34; cy=&#34;210&#34; r=&#34;8&#34; class=&#34;pt1&#34;/&gt;&lt;circle cx=&#34;922&#34; cy=&#34;230&#34; r=&#34;10&#34; class=&#34;pt1&#34;/&gt;
    &lt;circle cx=&#34;940&#34; cy=&#34;215&#34; r=&#34;8&#34; class=&#34;pt1&#34;/&gt;
    &lt;circle cx=&#34;955&#34; cy=&#34;255&#34; r=&#34;10&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;978&#34; cy=&#34;245&#34; r=&#34;8&#34; class=&#34;pt3&#34;/&gt;&lt;circle cx=&#34;996&#34; cy=&#34;260&#34; r=&#34;10&#34; class=&#34;pt3&#34;/&gt;
    &lt;circle cx=&#34;1015&#34; cy=&#34;245&#34; r=&#34;8&#34; class=&#34;pt3&#34;/&gt;
    &lt;circle cx=&#34;1005&#34; cy=&#34;190&#34; r=&#34;10&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;1028&#34; cy=&#34;180&#34; r=&#34;8&#34; class=&#34;pt2&#34;/&gt;&lt;circle cx=&#34;1048&#34; cy=&#34;197&#34; r=&#34;10&#34; class=&#34;pt2&#34;/&gt;
    &lt;circle cx=&#34;1070&#34; cy=&#34;185&#34; r=&#34;8&#34; class=&#34;pt2&#34;/&gt;
    &lt;circle cx=&#34;1065&#34; cy=&#34;240&#34; r=&#34;10&#34; class=&#34;pt4&#34;/&gt;&lt;circle cx=&#34;1090&#34; cy=&#34;230&#34; r=&#34;8&#34; class=&#34;pt4&#34;/&gt;&lt;circle cx=&#34;1110&#34; cy=&#34;250&#34; r=&#34;10&#34; class=&#34;pt4&#34;/&gt;
  &lt;/g&gt;
&lt;p&gt;&lt;text x=&#34;600&#34; y=&#34;365&#34; text-anchor=&#34;middle&#34; class=&#34;sub&#34;&gt;Point positions are schematic, not experimental molecular data.&lt;/text&gt;
&lt;/svg&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;h2 id=&#34;the-same-dataset-can-produce-three-different-stories&#34;&gt;The Same Dataset Can Produce Three Different Stories&lt;/h2&gt;
&lt;p&gt;None of these maps is necessarily wrong.&lt;/p&gt;
&lt;p&gt;They are preserving different properties of the original high-dimensional dataset.&lt;/p&gt;
&lt;p&gt;This is the central idea:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Dimensionality reduction does not simply reveal structure. It decides which structure is worth preserving.&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;h2 id=&#34;so-why-do-people-still-use-pca&#34;&gt;So Why Do People Still Use PCA?&lt;/h2&gt;
&lt;p&gt;Because PCA gives us things nonlinear embeddings cannot.&lt;/p&gt;
&lt;p&gt;It is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;fast&lt;/li&gt;
&lt;li&gt;deterministic&lt;/li&gt;
&lt;li&gt;mathematically transparent&lt;/li&gt;
&lt;li&gt;interpretable through feature loadings&lt;/li&gt;
&lt;li&gt;accompanied by explained variance&lt;/li&gt;
&lt;li&gt;useful for detecting correlated descriptors&lt;/li&gt;
&lt;li&gt;appropriate for many continuous molecular properties&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If I want to understand the physicochemical composition of a library, PCA may be exactly the method I want.&lt;/p&gt;
&lt;p&gt;Replacing PCA with UMAP simply because UMAP produces more visually separated clusters would not necessarily improve the analysis.&lt;/p&gt;
&lt;p&gt;In fact, it might remove useful interpretability.&lt;/p&gt;
&lt;h2 id=&#34;and-why-is-t-sne-still-used&#34;&gt;And Why Is t-SNE Still Used?&lt;/h2&gt;
&lt;p&gt;Because local structure matters.&lt;/p&gt;
&lt;p&gt;If the goal is exploratory analysis of closely related molecular families, t-SNE can produce informative visualizations.&lt;/p&gt;
&lt;p&gt;It also has a long history, is implemented in essentially every machine-learning ecosystem, and is familiar to researchers and reviewers.&lt;/p&gt;
&lt;p&gt;The mistake is not using t-SNE.&lt;/p&gt;
&lt;p&gt;The mistake is interpreting a t-SNE plot as if it were an ordinary Cartesian map where all global distances are meaningful.&lt;/p&gt;
&lt;h2 id=&#34;my-preferred-workflow-for-structural-chemical-space&#34;&gt;My Preferred Workflow for Structural Chemical Space&lt;/h2&gt;
&lt;p&gt;For a typical QSAR dataset where I want to inspect &lt;strong&gt;structural diversity&lt;/strong&gt;, I would usually start with:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;SMILES
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Morgan / ECFP4 fingerprint
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;radius = 2
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;2048 bits
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Jaccard distance
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;UMAP
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;n_neighbors ≈ 15 to 50
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;min_dist ≈ 0.05 to 0.3
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;   ↓
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;2D chemical-space visualization
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A reasonable starting configuration would be:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-python&#34; data-lang=&#34;python&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;kn&#34;&gt;from&lt;/span&gt; &lt;span class=&#34;nn&#34;&gt;rdkit&lt;/span&gt; &lt;span class=&#34;kn&#34;&gt;import&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Chem&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;kn&#34;&gt;from&lt;/span&gt; &lt;span class=&#34;nn&#34;&gt;rdkit.Chem&lt;/span&gt; &lt;span class=&#34;kn&#34;&gt;import&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;rdFingerprintGenerator&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;kn&#34;&gt;from&lt;/span&gt; &lt;span class=&#34;nn&#34;&gt;rdkit&lt;/span&gt; &lt;span class=&#34;kn&#34;&gt;import&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;DataStructs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;kn&#34;&gt;import&lt;/span&gt; &lt;span class=&#34;nn&#34;&gt;numpy&lt;/span&gt; &lt;span class=&#34;k&#34;&gt;as&lt;/span&gt; &lt;span class=&#34;nn&#34;&gt;np&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;kn&#34;&gt;import&lt;/span&gt; &lt;span class=&#34;nn&#34;&gt;umap&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;generator&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;rdFingerprintGenerator&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;GetMorganGenerator&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;radius&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;mi&#34;&gt;2&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;fpSize&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;mi&#34;&gt;2048&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;fps&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;[]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;arrays&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;[]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;for&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;smi&lt;/span&gt; &lt;span class=&#34;ow&#34;&gt;in&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;smiles&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;mol&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Chem&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;MolFromSmiles&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;smi&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;fp&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;generator&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;GetFingerprint&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;mol&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;fps&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;append&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;fp&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;arr&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;np&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;zeros&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;((&lt;/span&gt;&lt;span class=&#34;mi&#34;&gt;2048&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;dtype&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;np&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;uint8&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;DataStructs&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;ConvertToNumpyArray&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;fp&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;arr&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;arrays&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;append&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;arr&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;np&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;asarray&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;arrays&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;reducer&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;umap&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;UMAP&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;n_neighbors&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;mi&#34;&gt;30&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;min_dist&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;mf&#34;&gt;0.1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;metric&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;jaccard&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;random_state&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;mi&#34;&gt;42&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;embedding&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;reducer&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;fit_transform&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Then I would reuse the exact same coordinates and color the map according to different variables:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Plot 1: Active vs inactive
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Plot 2: Training vs test set
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Plot 3: Experimental pIC50
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Plot 4: Molecular weight
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Plot 5: Scaffold cluster
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This is often much more informative than generating a new embedding for every question.&lt;/p&gt;
&lt;p&gt;The geometry stays fixed while the interpretation changes.&lt;/p&gt;
&lt;h2 id=&#34;the-most-useful-qsar-plot-might-be-train-vs-test&#34;&gt;The Most Useful QSAR Plot Might Be Train vs Test&lt;/h2&gt;
&lt;p&gt;One of my favorite applications is visualizing whether the external test set occupies similar structural space to the training compounds.&lt;/p&gt;
&lt;p&gt;Color the same UMAP embedding by dataset assignment:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;● Training
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;● Test
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A heavily intermixed map suggests that many test compounds live within structural neighborhoods represented during training.&lt;/p&gt;
&lt;p&gt;Large isolated regions containing only test compounds may indicate extrapolation.&lt;/p&gt;
&lt;p&gt;But even here, the UMAP plot should not be treated as proof.&lt;/p&gt;
&lt;p&gt;A two-dimensional projection necessarily discards information.&lt;/p&gt;
&lt;p&gt;Two points that look close after projection may not actually have exceptionally high fingerprint similarity.&lt;/p&gt;
&lt;p&gt;Two points that look distant may still share substantial structural features.&lt;/p&gt;
&lt;p&gt;This is why a visual analysis should be paired with quantitative similarity measurements.&lt;/p&gt;
&lt;h2 id=&#34;umap-is-the-picture-tanimoto-is-the-measurement&#34;&gt;UMAP Is the Picture. Tanimoto Is the Measurement.&lt;/h2&gt;
&lt;p&gt;Suppose I want to know how novel each test compound is relative to the training set.&lt;/p&gt;
&lt;p&gt;For every test molecule, I can calculate:&lt;/p&gt;
&lt;div style=&#34;text-align:center; margin:1.5rem 0;&#34;&gt;
&lt;strong&gt;max Tanimoto(test, training)&lt;/strong&gt;
&lt;/div&gt;
&lt;p&gt;This asks:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;What is the most structurally similar molecule this test compound has already seen during training?&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Test molecule A    nearest training similarity = 0.87
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Test molecule B    nearest training similarity = 0.73
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Test molecule C    nearest training similarity = 0.51
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Test molecule D    nearest training similarity = 0.29
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Molecule A lies close to known chemical territory.&lt;/p&gt;
&lt;p&gt;Molecule D represents substantially stronger structural extrapolation.&lt;/p&gt;
&lt;p&gt;This quantity is much easier to interpret than measuring distances between points on a UMAP plot.&lt;/p&gt;
&lt;p&gt;That leads to a useful distinction:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;UMAP is excellent for seeing chemical space. Fingerprint similarity is better for measuring it.&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;h2 id=&#34;do-not-forget-the-representation&#34;&gt;Do Not Forget the Representation&lt;/h2&gt;
&lt;p&gt;Suppose two molecules occupy neighboring positions in an ECFP4-based UMAP.&lt;/p&gt;
&lt;p&gt;What does that mean?&lt;/p&gt;
&lt;p&gt;It means they have similar &lt;strong&gt;local topological environments according to that fingerprint representation&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;It does not automatically mean they have:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;similar conformations&lt;/li&gt;
&lt;li&gt;similar electrostatics&lt;/li&gt;
&lt;li&gt;similar pharmacophores&lt;/li&gt;
&lt;li&gt;similar biological activity&lt;/li&gt;
&lt;li&gt;similar binding modes&lt;/li&gt;
&lt;li&gt;similar metabolism&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Chemical similarity is always conditional on the representation.&lt;/p&gt;
&lt;p&gt;The same dataset embedded from physicochemical descriptors might reveal a very different organization.&lt;/p&gt;
&lt;p&gt;This is not a contradiction.&lt;/p&gt;
&lt;p&gt;It means the two maps are answering different questions.&lt;/p&gt;
&lt;h2 id=&#34;why-i-would-sometimes-show-both-pca-and-umap&#34;&gt;Why I Would Sometimes Show Both PCA and UMAP&lt;/h2&gt;
&lt;p&gt;Rather than asking which method wins, it can be more useful to combine complementary views.&lt;/p&gt;
&lt;h3 id=&#34;pca-on-physicochemical-descriptors&#34;&gt;PCA on physicochemical descriptors&lt;/h3&gt;
&lt;p&gt;Shows:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;size
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;lipophilicity
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;polarity
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;hydrogen bonding
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;molecular complexity
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This describes &lt;strong&gt;property space&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&#34;umap-on-ecfp4-fingerprints&#34;&gt;UMAP on ECFP4 fingerprints&lt;/h3&gt;
&lt;p&gt;Shows:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;scaffold relationships
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;local fragments
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;structural analogues
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;chemical series
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This describes &lt;strong&gt;structural space&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Together, the plots provide more information than either one individually.&lt;/p&gt;
&lt;p&gt;Two compounds may be structurally unrelated but occupy similar physicochemical space.&lt;/p&gt;
&lt;p&gt;Conversely, close structural analogues can sometimes have substantially different properties after relatively small chemical modifications.&lt;/p&gt;
&lt;p&gt;That distinction is central to medicinal chemistry.&lt;/p&gt;
&lt;h2 id=&#34;a-practical-comparison&#34;&gt;A Practical Comparison&lt;/h2&gt;
&lt;table&gt;
  &lt;thead&gt;
      &lt;tr&gt;
          &lt;th&gt;Method&lt;/th&gt;
          &lt;th&gt;What it emphasizes&lt;/th&gt;
          &lt;th&gt;Strongest advantage&lt;/th&gt;
          &lt;th&gt;Main caution&lt;/th&gt;
      &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
      &lt;tr&gt;
          &lt;td&gt;PCA&lt;/td&gt;
          &lt;td&gt;Global linear variance&lt;/td&gt;
          &lt;td&gt;Interpretability&lt;/td&gt;
          &lt;td&gt;Misses nonlinear relationships&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;t-SNE&lt;/td&gt;
          &lt;td&gt;Local neighborhoods&lt;/td&gt;
          &lt;td&gt;Strong visual separation&lt;/td&gt;
          &lt;td&gt;Global distances are difficult to interpret&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
          &lt;td&gt;UMAP&lt;/td&gt;
          &lt;td&gt;Local neighborhoods and some broader organization&lt;/td&gt;
          &lt;td&gt;Flexible and efficient&lt;/td&gt;
          &lt;td&gt;Geometry remains projection-dependent&lt;/td&gt;
      &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;For molecular fingerprints specifically, I would generally prefer:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;ECFP4 + Jaccard + UMAP
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;for visualization.&lt;/p&gt;
&lt;p&gt;For physicochemical descriptors, PCA remains extremely valuable.&lt;/p&gt;
&lt;p&gt;For investigating strongly local cluster structure, t-SNE remains a legitimate tool.&lt;/p&gt;
&lt;h2 id=&#34;there-is-no-single-chemical-space&#34;&gt;There Is No Single Chemical Space&lt;/h2&gt;
&lt;p&gt;Perhaps the most important lesson is that chemical space is not a single object.&lt;/p&gt;
&lt;p&gt;Consider the same set of compounds described using:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;ECFP fingerprints
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;MACCS keys
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;RDKit descriptors
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;3D shape descriptors
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;pharmacophore fingerprints
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;molecular embeddings
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;protein-ligand interaction fingerprints
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Each representation produces a different notion of similarity.&lt;/p&gt;
&lt;p&gt;Each therefore creates a different chemical space.&lt;/p&gt;
&lt;p&gt;UMAP cannot solve this problem because it appears only at the end of the pipeline.&lt;/p&gt;
&lt;p&gt;The most important question comes earlier:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What molecular information do I want proximity in this map to represent?&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Once that is clear, choosing the dimensionality-reduction algorithm becomes much easier.&lt;/p&gt;
&lt;h2 id=&#34;final-takeaway&#34;&gt;Final Takeaway&lt;/h2&gt;
&lt;p&gt;There is no universal winner between PCA, t-SNE, and UMAP.&lt;/p&gt;
&lt;p&gt;For continuous physicochemical descriptors, PCA provides an interpretable and statistically transparent view of molecular property space.&lt;/p&gt;
&lt;p&gt;For discovering local neighborhoods, t-SNE remains powerful, provided that distances between clusters are not overinterpreted.&lt;/p&gt;
&lt;p&gt;For high-dimensional structural fingerprints such as ECFP4, UMAP combined with a binary similarity metric such as Jaccard is one of the most useful default approaches for visual exploration.&lt;/p&gt;
&lt;p&gt;But even then, the plot should remain what it is:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;a visualization.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If the goal is to quantify structural overlap, novelty, or train-test extrapolation, return to the original fingerprints and calculate the similarities directly.&lt;/p&gt;
&lt;p&gt;The picture helps us explore the chemical space.&lt;/p&gt;
&lt;p&gt;The high-dimensional representation is where that space actually lives.&lt;/p&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Rogers, D.; Hahn, M. Extended-Connectivity Fingerprints. &lt;em&gt;Journal of Chemical Information and Modeling&lt;/em&gt; &lt;strong&gt;2010&lt;/strong&gt;, 50, 742-754.&lt;/li&gt;
&lt;li&gt;van der Maaten, L.; Hinton, G. Visualizing Data using t-SNE. &lt;em&gt;Journal of Machine Learning Research&lt;/em&gt; &lt;strong&gt;2008&lt;/strong&gt;, 9, 2579-2605.&lt;/li&gt;
&lt;li&gt;McInnes, L.; Healy, J.; Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv:1802.03426.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
  </channel>
</rss>
