Notes on cheminformaticsNo. 02 · July 2026
Embeddings, fingerprints & maps
What does Psiformer think is similar?
On what a representation keeps close
I mapped the same 665,911 ChEMBL molecules with Psiformer and ECFP4. One space follows electronic structure; the other stays closer to molecular scaffolds. Neither wins everywhere.
Go to the resulti.The post
A few weeks ago I saw a LinkedIn post by Woody Sherman that made a simple claim: molecular similarity should look more like chemistry and less like notation.
The post showed four pairs. Benzene and thiophene scored only 0.222 with a Morgan fingerprint and 0.964 with Psiformer. An alcohol and a fluorinated alkane went from 0.250 to 0.962. A ring-opened xanthine analog went from 0.244 to 0.996. Then the last pair reversed the story: a Z/E change kept a fingerprint similarity of 0.538 but fell to 0.026 with Psiformer.

Those are good examples. They are also four examples.
I work on TMAP2, a tool for laying out high-dimensional data as an explorable tree, so I left a comment asking whether the same difference would still be visible over a whole chemical collection. Then I contacted Woody and PsiThera's CTO. They were kind enough to send the Psiformer embeddings for the ChEMBL part of QMugs.
That changed the question. I did not try to reproduce the OpenADMET result in Woody's other post, and this is not a property-prediction benchmark. I asked something narrower:
If I describe the same 665,911 molecules with Psiformer and ECFP4, which molecules does each representation choose as neighbors?
This post is the answer. The short version is that Psiformer and ECFP4 do construct different chemical spaces, and the difference is chemically legible. Psiformer is not simply better. It keeps frontier-electronic properties smoother; ECFP4 keeps scaffolds, ordinary 2D descriptors, and polarizability smoother. Stereochemistry is more complicated still.
ii.Two descriptions
Before the maps, it helps to say what went into them.
ECFP4 is not a text representation. It does not care which valid SMILES string I happened to type. It starts from the molecular graph and records circular atom environments out to a radius of two bonds. I used a 2,048-bit Morgan fingerprint with chirality enabled. Two molecules get a high Tanimoto score when many of those local environments overlap.
This makes ECFP4 a very good parts list. If two molecules share a ring system, a linker, and several substituents, the fingerprint notices. It is fast, interpretable at the fragment level, and has survived a decade and a half of use because it works.
Psiformer begins from a different premise. PsiThera describes it as a 3D, quantum-mechanical embedding trained on millions of quantum-chemistry calculations. The supplied representation gives me 4,776 floating-point values for each molecule. I compare those vectors with cosine similarity: roughly, whether two long arrows point in the same direction.
I cannot inspect a row of 4,776 values and tell you what the model sees. That is the reason for making the maps.
same 665,911 ChEMBL molecules
ECFP4 Psiformer
────────────────────────────── ──────────────────────────────
2,048 bits 4,776 floating-point values
radius-two graph environments learned 3D/QM representation
Tanimoto / Jaccard cosine
chirality enabled geometry encoded by the model
20 nearest neighbors 20 nearest neighbors
This is therefore a comparison of representation plus native metric. Changing both at once is not a perfect causal experiment, but forcing a binary fingerprint and a dense embedding through one artificial metric would answer a less useful question.
iii.Making the maps
TMAP2 takes each molecule, finds its nearest neighbors under the chosen metric, builds a nearest-neighbor graph, and reduces that graph to a minimum spanning tree. The tree has just enough edges to keep every connected part joined. A layout algorithm then spreads those branches over the screen.
That last step is drawing, not measurement. If two branches happen to pass close to one another in the image, that does not make their molecules similar. The edges are the meaningful part.
This is the main difference from UMAP and t-SNE. Those methods optimize the positions of a cloud of points; TMAP2 preserves a sparse set of neighbor relationships as visible edges. It is not a claim that trees are universally better projections. The tree discards many edges and its global screen distances are not a chemical metric. In return, every edge has an audit trail, every two points in a connected component have one path between them, and I can follow a branch without guessing whether two dots only look close because of the projection.
For this study I gave both representations the same corpus, the same neighborhood size of 20, the same seed, and the same mapping and browser viewer. Psiformer stayed as dense continuous data under cosine distance; ECFP4 stayed binary under Jaccard distance. This is a useful change in TMAP2 from the original TMAP, which was designed mainly around sparse fingerprints.
The full maps took about 55 minutes for Psiformer and 5 minutes for ECFP4 on the available machine. Each finished as a self-contained browser file of about 26 MB containing every molecule, not a sample. I can pan, zoom, recolor the tree, and hover a point to inspect its structure. The slower run is the price of comparing 4,776 floating-point values instead of doing fast bit operations.
I also checked the neighbor search rather than trusting the picture. Against exact search on random full-corpus queries, the saved Psiformer graph recovered 99.8% of the true top 20. The tie-aware ECFP4 recall was 98.75%; binary fingerprints have many equally scored neighbors, so there is often no unique twentieth place. Higher-recall graphs changed the property statistics by at most 0.00022 and never changed a winner.
The map is the interface. The original 20-neighbor graphs are the evidence.
iv.Grading the neighborhoods
The first comparison is direct. For every molecule, take its 20 Psiformer neighbors and its 20 ECFP4 neighbors and ask how much the two sets overlap.
The median Jaccard overlap is 0.1765. Because both sets contain 20 molecules, that works out to six shared neighbors. Fourteen of the 20 change when I change the representation.
The next question is what sort of neighbors replace them.
I used Bemis–Murcko scaffolds as a deliberately blunt definition of molecular family. On the Psiformer graph, 18.7% of a molecule's neighbors share its scaffold on average. With ECFP4 the value is 32.3%, or 1.73 times higher. ECFP4 keeps structural relatives together more strongly; Psiformer crosses scaffold boundaries more readily.
That still does not tell me whether either set of crossings is useful. To test that, I colored the same graph by a property that was not used to draw it and asked whether neighboring molecules had similar values.
Imagine a geographical map colored by temperature. A useful map changes color gradually: warm places border other warm places. A bad map looks like salt and pepper. The statistic I used, Geary's C, measures exactly that local roughness. Lower is smoother.
# schematic only
roughness = sum(
(property[i] - property[j]) ** 2
for i, j in nearest_neighbor_edges
)
# Geary's C also normalizes by the global variance and graph size.
# lower C = neighbors are more alike for this property
The properties come from the exact QMugs conformer associated with each embedding. They are computed reference values, not experimental measurements. All 665,911 rows joined back to the source QM calculations. The score is calculated on the full nearest-neighbor graph, not on pixel distances and not only on the tree edges that survived the minimum-spanning-tree step.
v.The result
Here is the full-corpus result.
property Psiformer C ECFP4 C smoother
───────────────────── ─────────── ─────── ─────────
molecular weight 0.1903 0.1423 ECFP4
DFT HOMO 0.2195 0.2751 Psiformer
DFT LUMO 0.1456 0.2079 Psiformer
DFT HOMO–LUMO gap 0.1635 0.2039 Psiformer
DFT total dipole 0.5663 0.6088 Psiformer
GFN2 polarizability 0.2045 0.1413 ECFP4
The cleanest split is between size and frontier electronic structure.
Electrons in a molecule do not have arbitrary energies. In the usual molecular-orbital picture they occupy a ladder of allowed states. The HOMO is the highest occupied state; the LUMO is the lowest empty one. The energy between them—the HOMO–LUMO gap—is a compact description of how difficult it is to promote an electron across that frontier. It is not a universal reactivity number, but it is tied to the electronic response of the molecule in a way that molecular weight or ring count is not.
The HOMO, LUMO, and their gap are mathematically related, so I count them as one family of evidence, not three discoveries. Across that family Psiformer's neighborhoods are 20–30% smoother. Dipole, a measure of how unevenly charge is distributed across the molecule, also favors Psiformer, though by a smaller 7%.


ECFP4 wins molecular weight by 25%. In a separate 50,000-molecule panel it also wins all ten ordinary RDKit descriptors I checked: logP, ring counts, QED, polar surface area, hydrogen-bond counts, fraction sp3, and rotatable bonds among them. That is exactly what I would expect from a representation built out of local graph environments.
The useful result is not that one map looks more chemical. Both do. They preserve different kinds of chemistry.
The electronic result also survives harder controls. I took Psiformer-near, ECFP4-far cross-scaffold pairs and matched them to the reverse kind of pair at the same molecular weight. The Psiformer pairs had HOMO–LUMO gaps 18.9% closer on average. Then I required the same elements, exact counts of each heavy element, and finally the exact molecular formula. Among 4,513 independent exact-formula comparisons, the gap difference was still 23.7% smaller for the Psiformer pairs and the dipole difference was 16.2% smaller.
So this is not just Psiformer grouping molecules of the same size or putting all brominated compounds in one corner.
vi.The counterexample
At this point I had a dangerously neat story: ECFP4 sees structure; Psiformer sees physics.
Polarizability ruined it, which is why it belongs in the post.
Polarizability describes how readily an electron cloud deforms in an electric field. It is plainly an electronic property, and ECFP4 organizes it much more smoothly: C = 0.1413 against 0.2045.


The first explanation is size. In the 50,000-molecule analysis, polarizability correlates 0.971 with molecular weight. Bigger electron clouds are generally easier to deform, and ECFP4 is already very good at organizing molecular size.
That explanation is incomplete. When I held molecular formula exactly fixed, ECFP4 still won. The mean polarizability difference was 0.318 for ECFP4-controlled pairs and 0.426 for Psiformer pairs: Psiformer was 34.1% worse. The bootstrap interval for the paired difference stayed clear of zero.
So I cannot rescue the tidy story by calling polarizability “really just size.” Psiformer is smoother for the frontier-orbital family and dipole, but not for every quantity obtained from a quantum calculation. This comparison was capable of saying no.
vii.Where they disagree
Population statistics need molecular examples, but examples are where it becomes easiest to fool yourself.
I therefore selected the cross-scaffold panel without looking at the QM outcomes. I began with Psiformer-exclusive neighbors that had the exact same molecular formula and different Murcko scaffolds, kept pairs in the high-Psiformer-similarity and low-ECFP4-similarity quartiles, and only then drew the structures and inspected their QM values.

The first pair, CHEMBL3736339 and CHEMBL591470, has Psiformer cosine similarity 0.9970 and ECFP4 Tanimoto 0.1923. Their DFT gaps differ by 0.00666 in the source units. Across the matched population the effect is real, but I would not call these molecules validated bioisosteres. I have no assay pair showing preserved potency or selectivity, and no medicinal chemist annotation saying the replacement serves the same role. They are formula-controlled, cross-scaffold neighbors that the two representations judge very differently.
The arrow also points the other way. I found 37,739 same-connectivity stereochemical pairs: identical canonical non-isomeric SMILES, different canonical isomeric SMILES. A chiral ECFP4 fingerprint retrieved the counterpart within either molecule's top 20 for 99.3% of pairs. Psiformer did so for 83.4%.
In an exact-search sample, the median partner rank was one for ECFP4 and three for Psiformer. For CHEMBL1076442 and CHEMBL509903, ECFP4 puts each molecule first for the other with Tanimoto 0.9726. Psiformer gives cosine 0.4882 and puts each partner beyond rank 1,000.

This is consistent with Psiformer being more sensitive to a particular 3D geometry. Whether that is desirable depends on the job. If geometry controls target recognition, the separation may be useful. If I am trying to retrieve all stereochemical forms of one connectivity, ECFP4 is doing the more useful thing. A representation can notice a difference and still weight it too strongly.
viii.What is missing
There are several reasons not to turn this into “Psiformer beats ECFP4.”
First, Psiformer is not yet described in a paper I can inspect. I know the broad training claim from PsiThera's posts, but not whether these exact QMugs molecules or these quantum targets were present during training. If they were, smooth frontier-orbital neighborhoods would still tell me what the embedding organizes, but they would not be an out-of-sample validation of learned physics.
Second, the analysis uses the supplied embedding for one selected conformer per row. The stereochemical result mixes sensitivity to configuration with sensitivity to the chosen conformation. Re-embedding multiple conformers of the same molecule would show whether the separation is stable or whether one geometry can move a molecule across the map.
Third, neighborhood quality is not predictive performance. I did not test potency, selectivity, ADMET, extrapolation, or the claim that a downstream model needs less data. A smooth property map is encouraging for local learning; it is not a learning curve.
Fourth, ECFP4 is one baseline. Shape overlays, pharmacophore fingerprints, 3D descriptors, and other learned embeddings would make the comparison more informative. The interesting question is not whether 3D beats 2D in one pairing, but which information each method adds at what cost.
My next experiments would be:
- establish which ChEMBL molecules and QM targets, if any, Psiformer saw during training;
- repeat the map over a genuinely held-out corpus and held-out electronic properties;
- embed conformer ensembles and measure within-molecule versus between-stereoisomer variation;
- test a curated bioisostere set and matched molecular pairs with biological measurements;
- run low-data and scaffold-split learning curves for the actual endpoints of interest; and
- compare against 3D shape and pharmacophore baselines, not only ECFP4.
Until then, the safe conclusion is about organization, not drug discovery performance.
ix.Closing
I started because four molecular pairs on LinkedIn made me curious. After 665,911 molecules, the claim became both more convincing and less simple.
Psiformer really does move chemical space away from a scaffold-first organization. Its cross-scaffold neighbors share frontier-orbital energies and dipoles more closely, even after exact-formula controls. ECFP4 does what it was designed to do extremely well: it preserves local graph environments, molecular families, size, and a wide panel of standard descriptors. It also wins polarizability and retrieves same-connectivity stereochemical partners more reliably.
I prefer that result to a leaderboard.
A molecular representation is a choice about which differences to compress and which to amplify. “Similar” is never a property of two molecules by itself. It is shorthand for similar in the variables a representation kept.
TMAP2 did not decide which variables matter. It made the decision visible. It let me use the same tool for a 2,048-bit fingerprint and a 4,776-number embedding, inspect their neighbor graphs at full ChEMBL scale, and find the places where the attractive story stopped working.
The question I would ask of the next molecular embedding is therefore not “is it better?” It is: better at keeping what close?
The Psiformer embeddings were provided by PsiThera. The reference quantum properties and molecular geometries come from QMugs, built from ChEMBL. ECFP is described by Rogers and Hahn. TMAP2 is open source and builds on the original TMAP by Daniel Probst and Jean-Louis Reymond. Thanks to Woody Sherman and the PsiThera team for trusting me with the embeddings and letting me ask an unnecessarily large follow-up question.