REVIEW 2 major objections 5 minor 26 references
Generic pretrained molecular representations do not reproduce human olfactory rating geometry and add no clear predictive value beyond conventional chemical fingerprints.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 03:52 UTC pith:ZGRI3VJZ
load-bearing objection A careful negative-result audit of generic molecular encoders for olfaction; the main claims hold up, but the no-increment result hinges on a single ridge penalty and the mixture claim on a single 20-unit split. the 2 major comments →
Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that, for the evaluated generic encoders (MoLFormer and ChemBERTa), learned embeddings preserve useful chemistry but do not reconstruct the evaluated three-attribute human olfactory geometry, and their apparent predictive advantages disappear once a complementary conventional baseline (217 two-dimensional descriptors plus circular fingerprints) is included. The evidence is organized as an audit: global representational similarity analysis tests structural alignment, repeated cross-validation with paired bootstrap intervals tests incremental value, a 63-molecule shared set tests cross-dataset replication, and a strict component-disjoint split tests mixture transfe
What carries the argument
Representational similarity analysis (RSA): the Spearman correlation between the strict upper triangles of a representation's distance matrix and the human three-attribute rating distance matrix. It is the tool that separates 'does the representation preserve the target's global geometry' from 'does it predict a single outcome.' The paper also uses ΔMAE, the change in mean absolute error when an embedding is added to a chemistry baseline, and a strict component-disjoint split that excludes every mixture touching both training and held-out components.
Load-bearing premise
The mixture-transfer conclusion rests on one prespecified component-disjoint split with only 20 test units; if that split is unrepresentative — and no second qualifying split was found in 5,000 candidate assignments — the outcome- and representation-dependent characterization would not hold.
What would settle it
Run the same protocol with a second qualifying split from the same 72-component graph (11 held-out components, at least 20 strict test units), or find any single-molecule dataset where the combined-baseline-plus-embedding ΔMAE confidence interval excludes zero in the beneficial direction; either result would require revising the corresponding empirical boundary.
If this is right
- If the paper is right, generic pretrained molecular encoders cannot be assumed to encode human perceptual geometry; claims about olfactory representation need separate evidence for structural alignment.
- The disappearance of apparent gains once circular fingerprints are added means evaluations that omit a strong conventional baseline can overstate learned-embedding value.
- Human olfactory rating geometry is replicable within a dataset (RSA around 0.74–0.86) but only partially across protocols (RSA 0.33), so cross-dataset agreement should be reported separately, not assumed.
- Strict component-disjoint transfer for mixtures remains empirically unsupported for these encoders; all evaluated incremental effects are statistically indistinguishable from zero under the one split.
- Evaluation principle: scientific representation quality should be assessed separately for target reliability, structural alignment, incremental information, replication, and out-of-distribution transfer.
Where Pith is reading between the lines
- Editorial inference: the audit design transfers to other perceptual or biological domains — any representation claiming scientific validity should pass the same five-way separation before being adopted.
- Editorial inference: the near-zero Gaussian controls imply that high dimensionality alone does not create the weak positive RSA values, so a direct next step is to test olfaction-specific, receptor-informed, or graph-based encoders under the identical audit.
- Editorial inference: the intensity-versus-pleasantness asymmetry in the mixture split hints that pleasantness may depend more on configural interactions than intensity does; a testable extension is to include concentration and interaction-aware composition models in a larger component-disjoint evaluation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript audits generic pretrained molecular encoders (MoLFormer, ChemBERTa) against conventional RDKit descriptors and Morgan fingerprints for human olfactory perception, using the Keller–Vosshall and Bierling single-molecule datasets and the Ma binary-mixture dataset. It separates four claims: (1) global representational similarity between model embeddings and human three-attribute rating geometry; (2) incremental predictive value of MoLFormer beyond a combined RDKit–Morgan baseline; (3) cross-dataset replication of human geometry across 63 shared molecules; and (4) transfer to mixtures with unseen components under a strict component-disjoint split. The paper reports that human split-half RSA is high (0.743/0.855), model–human RSA is low (0.019–0.158), MoLFormer provides no clear incremental benefit over RDKit+Morgan in either single-molecule dataset, cross-dataset human RSA is 0.331 with a broad interval, and mixture transfer effects are outcome- and representation-dependent with all intervals crossing zero under one prespecified split. The analysis is unusually careful: identity-level joins prevent leakage, molecule-level bootstraps respect pair dependence, hyperparameters and seeds are fixed rather than tuned on the full data, a cohort-definition sensitivity and Gaussian geometry-null control are included, and the paper explicitly discloses an excluded invalid analysis.
Significance. If the conclusions hold, the paper provides a useful template for reliability-aware evaluation of scientific representations and establishes empirical boundaries for generic molecular encoders in human olfaction. The distinction between target reproducibility, structural alignment, incremental information, replication, and out-of-distribution transfer is a valuable contribution, and the negative findings for MoLFormer/ChemBERTa relative to a strong conventional baseline are a credible check on overclaims in representation learning. The paper ships reusable artifacts, fixed seeds, and validated result tables, and the treatment of bootstrap dependence and cohort sensitivity is exemplary. The main risk is that the central 'no clear incremental value' claim rests on a single untuned Ridge penalty, so the headline contribution is conditional on a regularity assumption that is not demonstrated.
major comments (2)
- [§3.4, Eq. (2), Table S8] The headline claim that MoLFormer provides no clear incremental predictive value beyond RDKit+Morgan rests on a single Ridge model with α=10 fixed under an 'existing analysis policy' that is not specified. The three feature blocks are concatenated without weighting and share one penalty: RDKit descriptors and MoLFormer embeddings are standardized, while Morgan bits are unscaled. With 476 and 73 molecules against roughly 3,000 features, the optimal regularization per block may differ substantially; an over-regularized MoLFormer block would bias ΔMAE toward zero or positive even if the embedding carried incremental signal. The Random Forest sensitivity in §4.4 uses fixed hyperparameters and the same unweighted concatenation, so it does not fully rule out this artifact. Because 'no clear increment' is the paper's most load-bearing single-molecule conclusion, I ask for an inner-CV or block-w
- [§3.5, §O, Contribution 4] The mixture-transfer claim is honestly scoped in the main text and limitations, but the introduction lists 'strict component-disjoint transfer to mixtures containing unseen components' as one of the five contributions, and the abstract states that 'incremental effects are outcome- and representation-dependent.' This evidence comes from exactly one prespecified split with 20 test units, 11 held-out components, and no second qualifying partition in a fixed 5,000-candidate search. The test units share components, so the bootstrap intervals may understate dependence; more importantly, the result is a description of one partition, not a general characterization of mixture transfer. I recommend rewording the contribution bullet and any summary language so that the strict-split result is explicitly presented as a single-partition case study. Without this change, readers may cite the paper as es
minor comments (5)
- [Author affiliations] The second affiliation line contains 'Seattle, W A' — presumably 'WA'.
- [§3.1 / Table 1] The entry '4Isoprop (cuminol; CID 325)' is typographically odd; consider a space or hyphen for readability.
- [§3.4] Since the phrase 'existing analysis policy' is used to justify a load-bearing hyperparameter choice, it should be documented or replaced with a reproducible selection rule. This is related to Major Comment 1.
- [Supplement K] The supplement reports RSA as 0.330841323; eight decimal places imply a false precision. Rounding to three decimals is already used in the main text and is preferable.
- [§4.5 / Table S10] For the mixture results, standalone-encoder MAEs are helpful context, but the table caption could state more explicitly that standalone values are not evidence of incremental value, since the paper already makes this point in Section P.
Circularity Check
No significant circularity: all central claims are empirical comparisons against externally supplied frozen representations, baselines, and public datasets.
full rationale
The paper's derivation chain is self-contained against external artifacts. MoLFormer and ChemBERTa are frozen public checkpoints, RDKit descriptors and Morgan fingerprints are deterministic conventional baselines, and the human ratings come from three public datasets. The geometry claims compare representation distance matrices to independently constructed human three-attribute Euclidean geometries; the human matrix is the target being explained, not a fitted input. The incremental ΔMAE analysis trains prediction models on chemistry features with and without a frozen embedding, so the embedding is not derived from the human ratings or from the baseline being compared. Cross-dataset RSA is a direct Spearman correlation between two separately collected rating geometries. The strict unseen-component mixture split is a single prespecified partition with zero component overlap, and the paper explicitly scopes every mixture conclusion to that one split. There are no author self-citations, no imported uniqueness theorem, no ansatz borrowed from prior work by the same authors, and no renaming of a known result. The fixed Ridge alpha and the concatenated-block design are methodological-validity concerns that could affect the negative incremental finding, but they are not circular: the reported numbers are empirical outcomes, not identities forced by definition or by a fitted parameter. The paper also reports negative controls and cohort sensitivities, which further indicate that the conclusions are not constructed from their own inputs.
Axiom & Free-Parameter Ledger
free parameters (4)
- Ridge penalty alpha =
10
- Random Forest hyperparameters =
80 trees, max depth 10, min leaf 2, max_features=sqrt
- Mixture split search thresholds =
exactly 11 held-out components, >=20 test units, >=101 training units
- Fixed seeds and resampling counts =
seeds 20260713, 20260715; 2,000 bootstrap replicates; 1,000 splits
axioms (8)
- domain assumption The human perceptual target is adequately captured by molecule-level mean ratings on intensity, pleasantness, and familiarity, with Euclidean distance in the z-scored 3-attribute space as the geometry.
- standard math Spearman RSA between strict upper triangles is a valid measure of structural alignment between representation and human geometry.
- domain assumption Comparing human geometry (Euclidean) with embedding geometry (cosine) and fingerprint geometry (Tanimoto) is a fair cross-space comparison.
- domain assumption RDKit stereo-aware canonical-SMILES identity resolution is correct across all 476+73+72 molecules and the 63 shared identities.
- ad hoc to paper Mean-pooling component vectors is a valid composition rule for mixture representations.
- domain assumption The combined RDKit (217 descriptors) + Morgan (2048-bit) feature set constitutes a stringent conventional-chemistry baseline.
- domain assumption Frozen public checkpoints (MoLFormer-XL-both-10pct, ChemBERTa-77M-MLM) fairly represent "generic molecular encoders."
- standard math Bootstrap percentile intervals over molecules/mixture units are valid uncertainty estimates despite pair/component dependence.
read the original abstract
Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution. We present a reliability-aware audit of generic molecular representations for human olfaction across four distinct claims: global perceptual geometry, incremental predictive value beyond chemistry, cross-dataset replication, and mixture transfer to unseen components. Using the Keller-Vosshall and Bierling single-molecule rating datasets and the Ma binary-mixture dataset, we compare MoLFormer and ChemBERTa against RDKit descriptors and Morgan fingerprints under identity-controlled and matched evaluations. Human three-attribute rating geometry, based on intensity, pleasantness, and familiarity, is reproducible across participant splits (median RSA 0.743 and 0.855), whereas model-human alignment is substantially weaker (RSA 0.019-0.158). Learned embeddings do not consistently outperform conventional representations in global alignment, and MoLFormer provides no clear incremental predictive value beyond a combined RDKit-Morgan baseline in either single-molecule dataset. Human geometry shows positive but incomplete agreement across 63 shared molecules (RSA 0.331; 95% bootstrap interval [0.204, 0.507]). Under one strict unseen-component mixture split, incremental effects are outcome- and representation-dependent, with all intervals crossing zero. These results establish empirical boundaries for the evaluated generic molecular encoders and motivate a broader evaluation principle: representation quality in scientific domains should be assessed separately for target reliability, structural alignment, incremental information, replication, and out-of-distribution transfer.
Figures
Reference graph
Works this paper leans on
-
[1]
Keller, Andreas and Vosshall, Leslie B. , title =. BMC Neuroscience , year =. doi:10.1186/s12868-016-0287-2 , url =
-
[2]
Bierling, Antonie Louise and Croy, Alexander and Jesgarzewsky, Tim and Rommel, Maria and Cuniberti, Gianaurelio and Hummel, Thomas and Croy, Ilona , title =. Scientific Data , year =. doi:10.1038/s41597-025-04644-2 , url =
-
[3]
2025 , doi =
Bierling, Antonie Louise and Croy, Alexander and Jesgarzewsky, Tim Lukas and Rommel, Maria and Cuniberti, Gianaurelio and Hummel, Thomas and Croy, Ilona , title =. 2025 , doi =
2025
-
[4]
Ma, Yue and Tang, Ke and Xu, Yan and Thomas-Danguin, Thierry , title =. Data in Brief , year =. doi:10.1016/j.dib.2021.107143 , url =
arXiv 2021
-
[5]
Bushdid, C. and Magnasco, M. O. and Vosshall, L. B. and Keller, A. , title =. Science , year =. doi:10.1126/science.1249168 , url =
-
[6]
Nature Machine Intelligence , year =
Ross, Jerret and Belgodere, Brian and Chenthamarakshan, Vijil and Padhi, Inkit and Mroueh, Youssef and Das, Payel , title =. Nature Machine Intelligence , year =. doi:10.1038/s42256-022-00580-7 , url =
-
[7]
Chithrananda, Seyone and Grand, Gabriel and Ramsundar, Bharath , title =. 2020 , eprint =. doi:10.48550/arXiv.2010.09885 , url =
-
[8]
doi:10.5281/zenodo.591637 , url =
RDKit: Open-source cheminformatics , year =. doi:10.5281/zenodo.591637 , url =
-
[9]
Journal of Chemical Information and Modeling , year =
Rogers, David and Hahn, Mathew , title =. Journal of Chemical Information and Modeling , year =. doi:10.1021/ci100050t , url =
-
[10]
Frontiers in Systems Neuroscience , year =
Kriegeskorte, Nikolaus and Mur, Marieke and Bandettini, Peter , title =. Frontiers in Systems Neuroscience , year =. doi:10.3389/neuro.06.004.2008 , url =
-
[11]
Breiman, Leo , title =. Machine Learning , year =. doi:10.1023/A:1010933404324 , url =
-
[12]
Journal of Chemical Information and Computer Sciences , year =
Weininger, David , title =. Journal of Chemical Information and Computer Sciences , year =. doi:10.1021/ci00057a005 , url =
-
[13]
and Guan, Yuanfang and Dhurandhar, Amit and Turu, Gabor and Szalai, Bence and Mainland, Joel D
Keller, Andreas and Gerkin, Richard C. and Guan, Yuanfang and Dhurandhar, Amit and Turu, Gabor and Szalai, Bence and Mainland, Joel D. and Ihara, Yasushi and Yu, Chung Wen and Wolfinger, Russ and Vens, Celine and Schietgat, Leander and De Grave, Kurt and Norel, Raquel and. Predicting human olfactory perception from chemical features of odor molecules , jo...
-
[14]
Khan, Rehan M. and Luk, Chung-Hay and Flinker, Adeen and Aggarwal, Amit and Lapid, Hadas and Haddad, Rafi and Sobel, Noam , title =. Journal of Neuroscience , year =. doi:10.1523/JNEUROSCI.1158-07.2007 , url =
-
[15]
Snitz, Kobi and Yablonka, Adi and Weiss, Tali and Frumin, Idan and Khan, Rehan M. and Sobel, Noam , title =. PLOS Computational Biology , year =. doi:10.1371/journal.pcbi.1003184 , url =
-
[16]
and Ramanathan, Arvind and Chennubhotla, Chakra S
Castro, Jason B. and Ramanathan, Arvind and Chennubhotla, Chakra S. , title =. PLOS ONE , year =. doi:10.1371/journal.pone.0073289 , url =
-
[17]
Koulakov, Alexei A. and Kolterman, Benjamin E. and Enikolopov, Alexei G. and Rinberg, Dmitry , title =. Frontiers in Systems Neuroscience , year =. doi:10.3389/fnsys.2011.00065 , url =
Pith/arXiv arXiv 2011
-
[18]
Lee, Brian K. and Mayhew, Emily J. and Sanchez-Lengeling, Benjamin and Wei, Jennifer N. and Qian, Wesley W. and Little, Kelsie A. and Andres, Matthew and Nguyen, Britney B. and Moloy, Theresa and Yasonik, Jacob and Parker, Jane K. and Gerkin, Richard C. and Mainland, Joel D. and Wiltschko, Alexander B. , title =. Science , year =. doi:10.1126/science.ade4...
-
[19]
Sanchez-Lengeling, Benjamin and Wei, Jennifer N. and Lee, Brian K. and Gerkin, Richard C. and Aspuru-Guzik, Al. Machine Learning for Scent: Learning Generalizable Perceptual Representations of Small Molecules , year =. doi:10.48550/arXiv.1910.10685 , url =. 1910.10685 , archivePrefix =
-
[20]
and Gomes, Joseph and Geniesse, Caleb and Pappu, Aneesh S
Wu, Zhenqin and Ramsundar, Bharath and Feinberg, Evan N. and Gomes, Joseph and Geniesse, Caleb and Pappu, Aneesh S. and Leswing, Karl and Pande, Vijay , title =. Chemical Science , year =. doi:10.1039/C7SC02664A , url =
-
[21]
and Maclaurin, Dougal and Iparraguirre, Jorge and Bombarell, Rafael and Hirzel, Timothy and Aspuru-Guzik, Al
Duvenaud, David K. and Maclaurin, Dougal and Iparraguirre, Jorge and Bombarell, Rafael and Hirzel, Timothy and Aspuru-Guzik, Al. Convolutional Networks on Graphs for Learning Molecular Fingerprints , booktitle =. 2015 , url =
2015
-
[22]
and Riley, Patrick F
Gilmer, Justin and Schoenholz, Samuel S. and Riley, Patrick F. and Vinyals, Oriol and Dahl, George E. , title =. Proceedings of the 34th International Conference on Machine Learning , year =
-
[23]
Journal of Chemical Information and Modeling , year =
Jaeger, Sabrina and Fulle, Simone and Turk, Samo , title =. Journal of Chemical Information and Modeling , year =. doi:10.1021/acs.jcim.7b00616 , url =
-
[24]
and Kaiser,
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N. and Kaiser,. Attention Is All You Need , booktitle =. 2017 , url =
2017
-
[25]
Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , title =. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , year =. doi:10.18653/v1/N19-1423 , url =
-
[26]
Laing, D. G. and Francis, G. W. , title =. Physiology & Behavior , year =. doi:10.1016/0031-9384(89)90041-3 , url =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.