{"id":"7eb7731a-6194-4745-aedd-2c07f5453851","arxiv_id":"2607.24848","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Generic molecular encoders show weak alignment with human olfactory rating geometry and no clear predictive increment over an RDKit-Morgan chemistry baseline, while human rating geometry is reproducible within but only partially across datasets.","lead":"An audit of two AI molecular encoders (MoLFormer, ChemBERTa) against human smell ratings across three datasets finds the models align only weakly with human perceptual geometry and add no clear predictive value beyond standard chemistry descriptors. The paper also reports that human smell-rating geometry is highly reproducible within datasets but only partially consistent across datasets, while mixture-transfer results remain uncertain.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-increment claim rests on a single fixed Ridge α=10 chosen under an unspecified policy; if α is miscalibrated for the concatenated blocks, the headline negative result may be an artifact.","rationale":"I read the paper in good faith and find the audit unusually careful: identity-controlled joins, explicit split-half and molecule-bootstrap controls, fixed seeds, disclosed cohort sensitivity, and negative controls all support the geometry and cross-dataset findings. The reader's weakest assumption, the single mixture split, is genuine but the paper's own abstract and limitations repeatedly scope that claim to one prespecified split and note the absence of a second qualifying partition; a failed robustness replication there would weaken external generalization but would not contradict the text. The more load-bearing issue is the incremental-prediction protocol: the headline negative result that MoLFormer adds nothing beyond RDKit+Morgan depends on a single fixed Ridge α=10 justified by an unexplained 'existing analysis policy,' with heterogeneous feature scaling and one shared penalty. The Random Forest sensitivity is independent in model class but still uses fixed hyperparameters and the same concatenated unweighted design, so it does not eliminate the concern. This is not an accusation of p-hacking; it is a request for a specific sensitivity test that would either strengthen or qualify the abstract's strongest claim. If the inner-CV rerun still crosses zero, the paper's conclusion is robust to reasonable tuning; if not, the central claim is an artifact of the chosen regularization. Because the reader already issued a CONDITIONAL verdict and my concern adds a precise technical condition rather than overturning the paper, I recommend no change to the verdict.","tokens_in":16164,"tokens_out":13004,"duration_ms":154380,"concrete_test":"Re-run the primary Ridge comparison in Table S8 with the same folds, repeats, and master seed, but select α by inner cross-validation on each training fold (e.g., grid 10^-3 to 10^3) separately for the chemistry-only and chemistry+MoLFormer feature sets; also repeat with Morgan bits standardized as a sensitivity. Report ΔMAE and molecule-bootstrap 95% intervals. If either dataset yields a negative ΔMAE interval excluding zero, the no-increment conclusion is not robust; if all intervals still cross zero, the headline claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central negative claim — 'MoLFormer provides no clear incremental predictive value beyond a combined RDKit-Morgan baseline' — is supported by §4.4/Table S8, but the protocol fixes Ridge α=10 under an unspecified 'existing analysis policy' (§3.4) and concatenates RDKit descriptors, unscaled Morgan bits, and standardized MoLFormer embeddings into one feature block with a single shared penalty. With only 476 Keller molecules and 73 Bierling molecules competing against ~3,000 features, the optimal regularization per block could plausibly differ; an over-regularized MoLFormer block would bias ΔMAE toward zero or positive even if the embedding carried incremental signal. The Random Forest sensitivity (§4.4) is reassuring but uses fixed hyperparameters and the same unweighted concatenated design, so it does not fully rule out a calibration artifact. Because 'no clear increment' is a headline boundary claim, the result is load-bearing: if a fair, inner-CV-tuned comparison produced a negative interval excluding zero in either dataset, the paper's strongest single-molecule conclusion would need revision. The single-split mixture limitation identified by the reader is real but explicitly scoped and honestly disclosed; the alpha/preprocessing choice is less visible and cuts directly into the main claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript audits generic pretrained molecular encoders (MoLFormer, ChemBERTa) against conventional RDKit descriptors and Morgan fingerprints for human olfactory perception, using the Keller–Vosshall and Bierling single-molecule datasets and the Ma binary-mixture dataset. It separates four claims: (1) global representational similarity between model embeddings and human three-attribute rating geometry; (2) incremental predictive value of MoLFormer beyond a combined RDKit–Morgan baseline; (3) cross-dataset replication of human geometry across 63 shared molecules; and (4) transfer to mixtures with unseen components under a strict component-disjoint split. The paper reports that human split-half RSA is high (0.743/0.855), model–human RSA is low (0.019–0.158), MoLFormer provides no clear incremental benefit over RDKit+Morgan in either single-molecule dataset, cross-dataset human RSA is 0.331 with a broad interval, and mixture transfer effects are outcome- and representation-dependent with all intervals crossing zero under one prespecified split. The analysis is unusually careful: identity-level joins prevent leakage, molecule-level bootstraps respect pair dependence, hyperparameters and seeds are fixed rather than tuned on the full data, a cohort-definition sensitivity and Gaussian geometry-null control are included, and the paper explicitly discloses an excluded invalid analysis.","tokens_in":16286,"tokens_out":3603,"duration_ms":44055,"significance":"If the conclusions hold, the paper provides a useful template for reliability-aware evaluation of scientific representations and establishes empirical boundaries for generic molecular encoders in human olfaction. The distinction between target reproducibility, structural alignment, incremental information, replication, and out-of-distribution transfer is a valuable contribution, and the negative findings for MoLFormer/ChemBERTa relative to a strong conventional baseline are a credible check on overclaims in representation learning. The paper ships reusable artifacts, fixed seeds, and validated result tables, and the treatment of bootstrap dependence and cohort sensitivity is exemplary. The main risk is that the central 'no clear incremental value' claim rests on a single untuned Ridge penalty, so the headline contribution is conditional on a regularity assumption that is not demonstrated.","major_comments":[{"comment":"The headline claim that MoLFormer provides no clear incremental predictive value beyond RDKit+Morgan rests on a single Ridge model with α=10 fixed under an 'existing analysis policy' that is not specified. The three feature blocks are concatenated without weighting and share one penalty: RDKit descriptors and MoLFormer embeddings are standardized, while Morgan bits are unscaled. With 476 and 73 molecules against roughly 3,000 features, the optimal regularization per block may differ substantially; an over-regularized MoLFormer block would bias ΔMAE toward zero or positive even if the embedding carried incremental signal. The Random Forest sensitivity in §4.4 uses fixed hyperparameters and the same unweighted concatenation, so it does not fully rule out this artifact. Because 'no clear increment' is the paper's most load-bearing single-molecule conclusion, I ask for an inner-CV or block-w","section":"§3.4, Eq. (2), Table S8"},{"comment":"The mixture-transfer claim is honestly scoped in the main text and limitations, but the introduction lists 'strict component-disjoint transfer to mixtures containing unseen components' as one of the five contributions, and the abstract states that 'incremental effects are outcome- and representation-dependent.' This evidence comes from exactly one prespecified split with 20 test units, 11 held-out components, and no second qualifying partition in a fixed 5,000-candidate search. The test units share components, so the bootstrap intervals may understate dependence; more importantly, the result is a description of one partition, not a general characterization of mixture transfer. I recommend rewording the contribution bullet and any summary language so that the strict-split result is explicitly presented as a single-partition case study. Without this change, readers may cite the paper as es","section":"§3.5, §O, Contribution 4"}],"minor_comments":[{"comment":"The second affiliation line contains 'Seattle, W A' — presumably 'WA'.","section":"Author affiliations"},{"comment":"The entry '4Isoprop (cuminol; CID 325)' is typographically odd; consider a space or hyphen for readability.","section":"§3.1 / Table 1"},{"comment":"Since the phrase 'existing analysis policy' is used to justify a load-bearing hyperparameter choice, it should be documented or replaced with a reproducible selection rule. This is related to Major Comment 1.","section":"§3.4"},{"comment":"The supplement reports RSA as 0.330841323; eight decimal places imply a false precision. Rounding to three decimals is already used in the main text and is preferable.","section":"Supplement K"},{"comment":"For the mixture results, standalone-encoder MAEs are helpful context, but the table caption could state more explicitly that standalone values are not evidence of incremental value, since the paper already makes this point in Section P.","section":"§4.5 / Table S10"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically careful and the honest disclosure of limitations is commendable. The main reason for major revision is the untuned Ridge penalty underlying the primary negative result; this is fixable with additional analysis and should strengthen the paper. The mixture single-split issue is a scope/communication problem rather than a fatal flaw. I do not see a circularity or data-integrity concern, and the use of fixed seeds and reusable artifacts is a strong point."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a disciplined negative-result audit, and the negative results are worth taking seriously. It separates four claims—target reproducibility, global model–human alignment, incremental value beyond chemistry, cross-protocol replication, and mixture transfer—and finds that MoLFormer and ChemBERTa don't reproduce the human rating geometry and don't add clear predictive value over an RDKit+Morgan baseline. That is a useful empirical boundary for a field that often reads predictive accuracy as representational validity.\n\nWhat's actually new is the audit design itself: each claim gets its own control (participant split-half for reliability, molecule bootstrap for RSA, identity-level joins for leakage, a strict component-disjoint split for transfer), with fixed seeds, prespecified hyperparameters, a cohort sensitivity, a geometry-null control, and an honest statement that encoder pretraining overlap couldn't be ruled out. The numbers are consistent and the scoping is careful. No p-hacking or circularity—evaluation targets are frozen public checkpoints and external datasets.\n\nThe soft spots are real but manageable. First, the headline no-increment result relies on Ridge α=10 fixed under an unspecified 'existing analysis policy', with one penalty shared across the concatenated RDKit, Morgan, and MoLFormer blocks. A reviewer will want an α sweep or per-block regularization sensitivity. The random-forest sensitivity is reassuring, but it shares the same concatenated design, so it doesn't fully rule out a calibration artifact. Second, the mixture-transfer claim rests on one 20-test-unit split; the authors' own search found no second qualifying partition, so claim 4 is honestly a description of one split, not a robust finding about mixture transfer. Third, the code and derived artifacts are promised but not shipped, which is a real gap for a paper whose contribution is empirical reliability.\n\nThe paper is for anyone evaluating molecular representations or building scientific benchmarks; it is a good template for separating evidence types. It deserves a serious referee. My recommendation: send it out, but ask for the α sensitivity, a clearer statement on the mixture split's status (or a relaxed split), and the release of code/artifacts before acceptance.","headline":"A careful negative-result audit of generic molecular encoders for olfaction; the main claims hold up, but the no-increment result hinges on a single ridge penalty and the mixture claim on a single 20-unit split.","tokens_in":16940,"tokens_out":3327,"would_cite":true,"duration_ms":33754,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generic pretrained molecular representations do not reproduce human olfactory rating geometry and add no clear predictive value beyond conventional chemical fingerprints.","keywords":["olfactory perception","molecular representation learning","representational similarity analysis","model evaluation","mixture transfer","chemical fingerprints","incremental predictive value","reliability"],"falsifier":"Run the same protocol with a second qualifying split from the same 72-component graph (11 held-out components, at least 20 strict test units), or find any single-molecule dataset where the combined-baseline-plus-embedding ΔMAE confidence interval excludes zero in the beneficial direction; either result would require revising the corresponding empirical boundary.","tokens_in":15869,"feed_emoji":"👃","tokens_out":4631,"duration_ms":44962,"temperature":0.7,"pith_summary":"The paper argues that predictive accuracy alone is an insufficient test of whether a learned molecular representation captures real scientific structure. Using three public olfactory datasets, it asks four separate questions: Does a representation preserve the global geometry of human intensity, pleasantness, and familiarity ratings? Does it add information beyond standard two-dimensional chemistry? Does human geometry replicate across independent protocols? Does it transfer to mixtures with unseen components? The answers: human rating geometry is reproducible (split-half agreement around 0.74–0.86), model-to-human agreement is weak (0.02–0.16), MoLFormer provides no clear improvement over a combined descriptor-plus-fingerprint baseline, and cross-dataset human agreement is positive but partial (0.33). Mixture transfer under one strict split is outcome-dependent and statistically indistinguishable from zero.","feed_headline":"Smell AI fails to match human ratings or beat chemistry baselines","feed_subtitle":"A reliability audit shows generic molecular embeddings add no clear predictive value beyond conventional descriptors and fingerprints.","key_machinery":"Representational similarity analysis (RSA): the Spearman correlation between the strict upper triangles of a representation's distance matrix and the human three-attribute rating distance matrix. It is the tool that separates 'does the representation preserve the target's global geometry' from 'does it predict a single outcome.' The paper also uses ΔMAE, the change in mean absolute error when an embedding is added to a chemistry baseline, and a strict component-disjoint split that excludes every mixture touching both training and held-out components.","core_discovery":"The paper's central claim is that, for the evaluated generic encoders (MoLFormer and ChemBERTa), learned embeddings preserve useful chemistry but do not reconstruct the evaluated three-attribute human olfactory geometry, and their apparent predictive advantages disappear once a complementary conventional baseline (217 two-dimensional descriptors plus circular fingerprints) is included. The evidence is organized as an audit: global representational similarity analysis tests structural alignment, repeated cross-validation with paired bootstrap intervals tests incremental value, a 63-molecule shared set tests cross-dataset replication, and a strict component-disjoint split tests mixture transfe","pith_inferences":["Editorial inference: the audit design transfers to other perceptual or biological domains — any representation claiming scientific validity should pass the same five-way separation before being adopted.","Editorial inference: the near-zero Gaussian controls imply that high dimensionality alone does not create the weak positive RSA values, so a direct next step is to test olfaction-specific, receptor-informed, or graph-based encoders under the identical audit.","Editorial inference: the intensity-versus-pleasantness asymmetry in the mixture split hints that pleasantness may depend more on configural interactions than intensity does; a testable extension is to include concentration and interaction-aware composition models in a larger component-disjoint evaluation."],"forward_implications":["If the paper is right, generic pretrained molecular encoders cannot be assumed to encode human perceptual geometry; claims about olfactory representation need separate evidence for structural alignment.","The disappearance of apparent gains once circular fingerprints are added means evaluations that omit a strong conventional baseline can overstate learned-embedding value.","Human olfactory rating geometry is replicable within a dataset (RSA around 0.74–0.86) but only partially across protocols (RSA 0.33), so cross-dataset agreement should be reported separately, not assumed.","Strict component-disjoint transfer for mixtures remains empirically unsupported for these encoders; all evaluated incremental effects are statistically indistinguishable from zero under the one split.","Evaluation principle: scientific representation quality should be assessed separately for target reliability, structural alignment, incremental information, replication, and out-of-distribution transfer."],"fun_headline_variants":["Smell AI fails vs humans and simple chemistry baselines","AI smell detectors don't beat basic chemistry","Reliability audit: smell embeddings lack human alignment","Smell AI: no edge over RDKit, weak human fit"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The mixture-transfer conclusion rests on one prespecified component-disjoint split with only 20 test units; if that split is unrepresentative — and no second qualifying split was found in 5,000 candidate assignments — the outcome- and representation-dependent characterization would not hold.","fun_headline_variants_meta":{"raw":{"variants":["Smell AI fails vs humans and simple chemistry baselines","AI smell detectors don't beat basic chemistry","Reliability audit: smell embeddings lack human alignment","Smell AI: no edge over RDKit, weak human fit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1458,"prompt_tokens":802,"completion_tokens":656,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":601}},"tokens_in":546,"tokens_out":656,"duration_ms":7416,"temperature":1.0,"reasoning_tokens":601,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:52:05.353146+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same protocol with a second qualifying split from the same 72-component graph (11 held-out components, at least 20 strict test units), or find any single-molecule dataset where the combined-baseline-plus-embedding ΔMAE confidence interval excludes zero in the beneficial direction; either result would require revising the corresponding empirical boundary.","supporting_citations":[],"review_version":1}