REVIEW 3 major objections 5 minor 16 references
A graph-transformer platform reconstructs molecular bonding from routine 1H/13C NMR spectra, solving 53% of tested experimental molecules up to 480 Da.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 01:17 UTC pith:K626R5C4
load-bearing objection Concrete multi-stage GNN that recovers real experimental structures for complex molecules up to 480 Da, with honest metrics and clear soft spots on protocol flexibility and shared-model ranking. the 3 major comments →
Inverse-IMPRESSION: A Graph-based Platform for Molecular Structure Elucidation from Experimental NMR Spectroscopic Properties
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The inverse-IMPRESSION platform, built from four task-specific graph transformers that share the IMPRESSION-G2 architecture, reconstructs molecular bonding directly from measurable 1H/13C NMR data. After one-shot bond prediction, stepwise correction of degenerate heteroatoms, and noise-augmented multi-shot ranking by 13C mean absolute error, the system recovers the correct structure for 77.8% of simulated molecules (up to 30 heavy atoms) and for 10 of 19 experimental molecules (molecular weights up to 480 Da).
What carries the argument
The three-stage pipeline (one-shot bond-probability prediction, iterative stepwise re-attachment of NMR-silent heteroatoms, and multi-shot noise-augmented ranking by IMPRESSION-G2 13C MAE) that converts sparse experimental correlations into a ranked list of chemically valid 2D graphs.
Load-bearing premise
That a model trained only on Boltzmann-averaged simulated spectra will transfer, without fine-tuning, to real experimental spectra well enough that ranking candidates by predicted 13C error recovers the true structure.
What would settle it
Run the identical trained platform, without any change of weights or noise schedule, on a fresh set of twenty experimental molecules of comparable size and heteroatom count whose structures are already known by independent methods; if top-1 recovery falls well below 50% the experimental claim fails.
If this is right
- Routine 1H/13C COSY/HSQC/HMBC datasets can be turned into ranked 2D structure candidates without expert fragment assembly or exhaustive CASE enumeration.
- Inclusion of experimentally accessible 15N and 19F data raises top-k accuracy above 80% even for molecules with more than ten heteroatoms.
- The same ranking criterion (13C MAE below ~4.4 ppm) supplies an automatic confidence filter that separates correct from incorrect predictions.
- Molecules whose spectra leave large NMR-silent regions (quaternary clusters, heteroatom-rich substructures) remain systematically harder for both the algorithm and human analysts.
Where Pith is reading between the lines
- Because the platform already ranks by a forward NMR predictor, it can be closed into a self-consistent loop that rejects candidates whose predicted spectra deviate beyond the known error of that predictor.
- The same edge-probability formulation should transfer to other sparse spectroscopic graphs (IR, MS/MS, residual dipolar couplings) once analogous node and edge features are defined.
- Performance cliffs at high heteroatom counts suggest that hybrid human-AI workflows, in which the chemist supplies a few local constraints, could push success rates well above the present 53%.
- The 2000-shot ensembles already generate chemically valid alternatives; these could be used as an automatic “structure-revision” tool when the top-ranked molecule later fails orthogonal tests.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents inverse-IMPRESSION, a multi-stage graph-transformer platform that reconstructs 2D molecular connectivity (with bond orders) from 1H/13C NMR chemical shifts and COSY/HSQC/HMBC-style correlations. A one-shot model predicts bond probabilities on a fully-connected atom graph; a stepwise model then reassigns bonds around NMR-silent heteroatoms (and carbons with large predicted-shift errors) after stripping uncertain edges; noise-augmented multi-shot sampling produces candidate ensembles that are ranked (frequency for simulated data; IMPRESSION-G2 13C MAE for experimental data). On a 2000-molecule simulated test set (NMR Dataset B) the full pipeline reaches 77.8% Top-1 accuracy (Table 1, entry 4) and remains robust with size and nSPS (Fig. 3). On 19 experimental molecules (MW up to ~480 Da for successes) it recovers 10 structures, claimed as the first effective graph-based ML approach for automated elucidation from experimental NMR.
Significance. If the experimental transfer claim holds under a fixed protocol, this is a genuine advance: graph-based edge prediction that recovers complex natural-product-like skeletons (betulin, 24-epibrassinolide) from routine 1H/13C 2D data without reactant priors or exhaustive CASE enumeration. Strengths include an explicit, chemically motivated correction for node degeneracy of N/O/F, clear ablation of NMR information content (Datasets A/B/C), honest reporting of failures on heteroatom-rich and quaternary-dense molecules, public code/data, and a ranking oracle that cleanly separates correct from incorrect candidates when the true structure is generated. The work therefore supplies both a usable platform and a falsifiable baseline for future NMR-to-structure ML.
major comments (3)
- Results “Performance on Experimental Data” and SI §2.4: the reported 10/19 successes are not obtained under a single fixed multi-shot protocol. Compounds 1–7 use the simulated-data defaults (2000 shots, 0.4 Hz coupling noise); compound 8 requires an additional 5000 shots; compounds 9–10 require raising coupling-constant noise to 0.6 Hz. The abstract and conclusions state a 53% experimental success rate without disclosing this variable budget. A load-bearing claim of “first effective … on experimental data” requires either (i) a fixed-budget re-evaluation that reports how many of the 19 are recovered under the original 2000-shot/0.4 Hz settings, or (ii) an explicit, pre-specified adaptive protocol whose computational cost is quantified for every molecule.
- Methods “Multi-Shot Noise-Augmented Molecule Reconstruction” and SI §2.4.1: experimental candidates are ranked by 13C MAE from the identical forward IMPRESSION-G2 model that generated the Boltzmann-averaged training labels. For simulated data the authors deliberately avoided this MAE ranking (using frequency instead) “to avoid bias.” The experimental ranking therefore re-introduces a shared-model oracle that is unavailable in a truly blind CASE setting. The manuscript should either (a) re-rank the experimental ensembles with an independent 13C predictor (or with frequency alone) and report the change in Top-1 recovery, or (b) clearly restate the experimental claim as “recovery under IMPRESSION-G2 MAE ranking” rather than as a general automated elucidation result.
- Table 1 / Fig. 3a and SI Table 4: performance collapses for molecules with >7–10 heteroatoms or multiple quaternary–quaternary C–C bonds, and several experimental failures (baccatin-III, reserpine, ginkgolide A) sit precisely in this regime. The central claim that the platform solves “complex structures … that routinely challenge chemists” is therefore only partially supported. The paper should quantify the fraction of the 19 experimental molecules that lie inside the high-accuracy regime of Fig. 3 and discuss whether the 53% figure is representative of the broader chemical space claimed in the abstract.
minor comments (5)
- Abstract and Introduction: “first effective approach for automated molecular structure elucidation using graph-based machine learning on experimental data” should be tempered by explicit comparison numbers against NMRMind (33% on 12 molecules) and DiffNMR (11% experimental) already cited in the text, so the novelty claim is quantitative rather than absolute.
- Methods / SI §1.2.5: the bond-existence threshold of 0.5 and the structure-correction MAE reopen threshold of 10 ppm are free parameters; a short sensitivity table (or statement that they were fixed a priori) would strengthen reproducibility.
- Fig. 2 caption and main-text discussion of node degeneracy: the example molecule is clear, but the probability heat-maps would be easier to read if atom indices were overlaid on both the predicted and ground-truth skeletons.
- SI §2.4.3 / Supplementary Table 4: the COSY/HMBC-ratio thresholds used to colour-code success/failure are useful; they should be mentioned briefly in the main-text experimental discussion so readers do not have to hunt the SI for the failure-mode analysis.
- Data availability: Zenodo DOI is given; a one-sentence note confirming that the 19 experimental SDF + peak lists are included would help immediate re-use.
Circularity Check
Mild self-reference: training NMR labels and experimental ranking both use the authors' own forward IMPRESSION-G2; bond prediction itself is a learned mapping evaluated on held-out ground-truth connectivity, not forced by construction.
specific steps
-
self citation load bearing
[Methods (Datasets; Multi-Shot Noise-Augmented Molecule Reconstruction); Results (Performance on Experimental Data); SI §2.4.1]
"NMR parameters, including chemical shifts and scalar coupling constants, were subsequently predicted for each conformer using the IMPRESSION-G2 model... For experimental data, candidate structures were ranked using the mean absolute error (MAE) between the experimental 13C chemical shifts and those predicted for each candidate using forward IMPRESSION-G2... the correct structure was ranked first in all cases where it was generated among the noise-augmented candidates. In each case, the corresponding MAE values for the correct structure was below 4.4 ppm, which is in line with the benchmarked p"
Both the simulated training labels and the experimental ranking oracle are produced by the authors' own prior forward IMPRESSION-G2 (self-cited). When a correct structure is generated it is always Top-1 by this MAE (below the model's known error floor), while failures sit higher; the ranking that certifies experimental success is therefore self-consistent with the generative family used for training data rather than an independent external chemical-shift predictor. This is mild load-bearing self-citation for the experimental claim, not a definitional reduction of the inverse bond predictions themselves.
full rationale
The paper is an empirical ML platform paper, not a closed-form derivation. Supervised training of the inverse GTNs (one-shot, stepwise, bond-order, NMR-prediction) uses Boltzmann-averaged chemical shifts and correlations generated by the authors' prior forward IMPRESSION-G2 on MMFF conformers; experimental candidates are ranked by 13C MAE from the same forward model. This is shared tooling and self-citation, not a definitional identity: the inverse models output bond-existence probabilities from spectroscopic node/edge features, structures are checked against known connectivity (simulated Top-1 77.8% via frequency ranking that deliberately avoids the forward MAE; experimental 10/19 via generation + MAE ranking). No equation reduces output bonds to input features by construction, no uniqueness theorem is imported to forbid alternatives, and no fitted parameter is renamed a prediction. The variable multi-shot budget (extra shots/noise for 3 of 10 experimental successes) is a protocol/reproducibility concern outside circularity. Score 2 reflects only the non-load-bearing self-citation of the generative model family for data and ranking; the central claim retains independent content.
Axiom & Free-Parameter Ledger
free parameters (5)
- Bond-existence probability threshold =
0.5
- Structure-correction 13C MAE reopen threshold =
10 ppm
- Noise SDs for multi-shot augmentation =
0.3 ppm / 2.5 ppm / 0.4–0.6 Hz
- Coupling-to-correlation threshold =
2 Hz
- Number of noise-augmented shots =
200 / 2000 / 5000
axioms (5)
- domain assumption Molecules can be represented as fully connected graphs with atom types and NMR features as node/edge attributes, and structure elucidation reduces to edge (bond) prediction.
- domain assumption Boltzmann-weighted IMPRESSION-G2 NMR parameters on MMFF conformers are sufficiently accurate proxies for experimental 1H/13C shifts and COSY/HSQC/HMBC correlations for training and transfer.
- ad hoc to paper NMR-silent heteroatoms (N, O, F under Dataset B) can be recovered by stripping their bonds after one-shot prediction and reassigning them via a stepwise model trained on BFS fragments.
- domain assumption Candidate ranking by 13C MAE (experimental) or prediction frequency (simulated) identifies the true structure when it appears in the ensemble.
- standard math Standard GTN/Transformer message-passing with multi-head attention, gated residuals, BCE/CE/L1 losses, and stated hyperparameters is an adequate inductive bias for bond probability learning.
invented entities (1)
-
Inverse-IMPRESSION multi-stage platform (one-shot + stepwise + bond-order + NMR-prediction models with noise-augmented multi-shot ranking)
independent evidence
read the original abstract
Here, we present a platform built on our inverted Graph Transformer Network, IMPRESSION-G2, which can accurately and rapidly reconstruct molecular bonding directly from experimental nuclear magnetic resonance (NMR) spectroscopic information. It comprises three interconnected stages: a one-shot model that predicts bond connectivity between atoms; a structure-correction stage that corrects the predicted structures by removing uncertain bonds and iteratively reassigning them; noise-augmented multi-shot prediction, generating an ensemble of candidate structures, which are ranked to identify the best-fit structure. By integrating a range of $^{1}$H and $^{13}$C NMR data, including two-dimensional (2D) experiments such as COSY, HSQC, and HMBC, the inverse-IMPRESSION platform correctly identifies the structures of 77.8% of molecules with up to 30 heavy atoms (H, C, N, O and F) using simulated NMR data, and 10 of 19 (53%) molecules using experimental NMR data. The experimental structures solved have molecular weights of up to 480 Da and are representative of the complex structures in synthetic and natural products that routinely challenge chemists. The inverse-IMPRESSION framework thus provides the first effective approach for automated molecular structure elucidation using graph-based machine learning on experimental data.
Reference graph
Works this paper leans on
-
[1]
Yiu, C. et al. IMPRESSION generation 2 - accurate, fast and generalised neural network model for predicting NMR parameters in place of DFT. Chem. Sci. 16, 8377-8382 (2025)
2025
-
[2]
Landrum, G. et al. RDKit: open -source cheminformatics (version 2025.03.6). Zenodo https://doi.org/10.5281/zenodo.16996017 (2025)
-
[3]
Bratholm, L. A. et al. A community -powered search of machine learning strategy space to find NMR property prediction models. PLoS One 16, e0253612 (2021)
2021
-
[4]
He, K., Zhang, X., Ren, S. & Sun, J. Delving deep into rectifiers: Surpassing human -level performance on imagenet classification. Preprint at https://doi.org/10.48550/arXiv.1502.01852 (2015)
-
[5]
Cho, K. et al. Learning phrase representations using RNN encoder -decoder for statistical machine translation. Preprint at https://doi.org/10.48550/arXiv.1406.1078 (2014)
-
[6]
Yun, S., Jeong, M., Kim, R., Kang, J. & Kim, H. J. Graph transformer networks. Preprint at https://doi.org/10.48550/arXiv.1911.06455 (2022)
-
[7]
Vaswani, A. et al. Attention is all you need. Preprint at https://doi.org/10.48550/arXiv.1706.03762 (2023)
-
[8]
You, Y . et al. Large batch optimization for deep learning: Training bert in 76 minutes. Preprint at https://doi.org/10.48550/arXiv.1904.00962 (2020)
-
[9]
Smith, L. N. Cyclical learning rates for training neural networks. Preprint at https://doi.org/10.48550/arXiv.1506.01186 (2017)
-
[10]
Kim, S. et al. PubChem 2025 update. Nucleic Acids Res. 53, D1516-D1525 (2025)
2025
-
[11]
Zdrazil, B. et al. The ChEMBL database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Res. 52, D1180-D1192 (2024)
2023
-
[12]
Halgren, T. A. Merck molecular force field. I. Basis, form, scope, parameterization, and performance of MMFF94. J. Comput. Chem. 17, 490-519 (1996)
1996
-
[13]
pandas -dev/pandas: pandas (version 2.3.2)
The pandas development team. pandas -dev/pandas: pandas (version 2.3.2). Zenodo https://doi.org/10.5281/zenodo.16918803 (2025)
-
[14]
Fey, M. & Lenssen, J. E. Fast graph representation learning with PyTorch Geometric. Preprint at https://doi.org/10.48550/arXiv.1903.02428 (2019)
-
[15]
Wildman, S. A. & Crippen, G. M. Prediction of physicochemical parameters by atomic contributions. J. Chem. Inf. Comput. Sci. 39, 868-873 (1999)
1999
-
[16]
& Waldmann, H
Krzyzanowski, A., Pahl, A., Grigalunas, M. & Waldmann, H. Spacial score - a comprehensive topological indicator for small-molecule complexity. J. Med. Chem. 66, 12739-12750 (2023)
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.