{"id":"963d6889-191e-405b-bff7-fe175fa0b56b","arxiv_id":"2607.09978","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Inverse-IMPRESSION reconstructs 2D molecular bonding from experimental 1H/13C NMR (COSY/HSQC/HMBC) with 77.8% Top-1 accuracy on simulated molecules (≤30 heavy atoms) and 53% (10/19) on experimental structures up to 480 Da.","lead":"A graph neural network platform reconstructs molecular structures from routine 1H/13C NMR data, solving 78% of simulated molecules and 10 of 19 experimental cases up to 480 Da. If reliable, it would automate a bottleneck that still consumes expert chemist time in synthesis and natural-product work.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Experimental 53% success depends on post-hoc increases in shot count and coupling noise for 3 of 10 solved molecules, plus ranking by the same IMPRESSION-G2 family used to generate training labels.","rationale":"The reader already isolates the precise soft spot: transfer of Boltzmann-averaged IMPRESSION-G2 labels (Dataset B) to experimental spectra without fine-tuning, strained by extra shots/noise and shared-model ranking. The paper’s own SI tables confirm the variable budgets and the MAE threshold that cleanly separates successes from failures. No deeper internal inconsistency (e.g., in the GTN architecture or the stepwise degeneracy fix) is required to explain the gap; the experimental claim simply has not been stress-tested under the same fixed protocol used for the 77.8% simulated number. Because the reader already assigns CONDITIONAL precisely for these reasons, and the concrete check above would only quantify rather than invent a new flaw, the verdict needs no adjustment. Code/data release is sufficient to run the proposed test.","tokens_in":23596,"tokens_out":697,"duration_ms":17469,"concrete_test":"Re-run the full inverse-IMPRESSION pipeline on all 19 experimental molecules under a single fixed protocol (exactly 2000 shots, noise σ = 0.3 ppm 1H / 2.5 ppm 13C / 0.4 Hz couplings). Report (i) how many of the original 10 correct structures still appear in any candidate and (ii) how many remain Top-1 when re-ranked by an independent 13C predictor (e.g., DFT or a non-IMPRESSION ML model). If either count falls below 7, the experimental accuracy is protocol- and oracle-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central experimental claim (10/19 structures recovered, MW up to 480 Da) rests on a non-fixed multi-shot protocol. Main text and SI §2.4 state that compounds 1–7 used the simulated-data defaults (2000 shots, 0.3/2.5 ppm, 0.4 Hz), compound 8 required an extra 5000 shots, and compounds 9–10 required raising coupling-constant noise to 0.6 Hz. Ranking of experimental candidates further uses 13C MAE from the identical forward IMPRESSION-G2 model that supplied the Boltzmann-averaged training labels (Methods: Multi-Shot…; SI §2.4.1). When the correct structure is generated it is always Top-1 by this MAE (<4.4 ppm), while failures sit ≥5.3 ppm; this oracle is therefore both generative-family-dependent and unavailable at inference time for a truly blind CASE system. The transfer premise (simulated Dataset B \to real spectra, no fine-tuning) is therefore only demonstrated under variable computational budgets and a shared-model ranking that was deliberately avoided for the simulated Top-k metrics (frequency ranking). If a fixed budget recovers substantially fewer than 10 structures, or if an independent 13C predictor reorders the ensembles, the “first effective graph-based… on experimental data” claim does not hold at the reported rate.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript presents inverse-IMPRESSION, a multi-stage graph-transformer platform that reconstructs 2D molecular connectivity (with bond orders) from 1H/13C NMR chemical shifts and COSY/HSQC/HMBC-style correlations. A one-shot model predicts bond probabilities on a fully-connected atom graph; a stepwise model then reassigns bonds around NMR-silent heteroatoms (and carbons with large predicted-shift errors) after stripping uncertain edges; noise-augmented multi-shot sampling produces candidate ensembles that are ranked (frequency for simulated data; IMPRESSION-G2 13C MAE for experimental data). On a 2000-molecule simulated test set (NMR Dataset B) the full pipeline reaches 77.8% Top-1 accuracy (Table 1, entry 4) and remains robust with size and nSPS (Fig. 3). On 19 experimental molecules (MW up to ~480 Da for successes) it recovers 10 structures, claimed as the first effective graph-based ML approach for automated elucidation from experimental NMR.","tokens_in":23989,"tokens_out":1411,"duration_ms":10716,"significance":"If the experimental transfer claim holds under a fixed protocol, this is a genuine advance: graph-based edge prediction that recovers complex natural-product-like skeletons (betulin, 24-epibrassinolide) from routine 1H/13C 2D data without reactant priors or exhaustive CASE enumeration. Strengths include an explicit, chemically motivated correction for node degeneracy of N/O/F, clear ablation of NMR information content (Datasets A/B/C), honest reporting of failures on heteroatom-rich and quaternary-dense molecules, public code/data, and a ranking oracle that cleanly separates correct from incorrect candidates when the true structure is generated. The work therefore supplies both a usable platform and a falsifiable baseline for future NMR-to-structure ML.","major_comments":[{"comment":"Results “Performance on Experimental Data” and SI §2.4: the reported 10/19 successes are not obtained under a single fixed multi-shot protocol. Compounds 1–7 use the simulated-data defaults (2000 shots, 0.4 Hz coupling noise); compound 8 requires an additional 5000 shots; compounds 9–10 require raising coupling-constant noise to 0.6 Hz. The abstract and conclusions state a 53% experimental success rate without disclosing this variable budget. A load-bearing claim of “first effective … on experimental data” requires either (i) a fixed-budget re-evaluation that reports how many of the 19 are recovered under the original 2000-shot/0.4 Hz settings, or (ii) an explicit, pre-specified adaptive protocol whose computational cost is quantified for every molecule.","section":null},{"comment":"Methods “Multi-Shot Noise-Augmented Molecule Reconstruction” and SI §2.4.1: experimental candidates are ranked by 13C MAE from the identical forward IMPRESSION-G2 model that generated the Boltzmann-averaged training labels. For simulated data the authors deliberately avoided this MAE ranking (using frequency instead) “to avoid bias.” The experimental ranking therefore re-introduces a shared-model oracle that is unavailable in a truly blind CASE setting. The manuscript should either (a) re-rank the experimental ensembles with an independent 13C predictor (or with frequency alone) and report the change in Top-1 recovery, or (b) clearly restate the experimental claim as “recovery under IMPRESSION-G2 MAE ranking” rather than as a general automated elucidation result.","section":null},{"comment":"Table 1 / Fig. 3a and SI Table 4: performance collapses for molecules with >7–10 heteroatoms or multiple quaternary–quaternary C–C bonds, and several experimental failures (baccatin-III, reserpine, ginkgolide A) sit precisely in this regime. The central claim that the platform solves “complex structures … that routinely challenge chemists” is therefore only partially supported. The paper should quantify the fraction of the 19 experimental molecules that lie inside the high-accuracy regime of Fig. 3 and discuss whether the 53% figure is representative of the broader chemical space claimed in the abstract.","section":null}],"minor_comments":[{"comment":"Abstract and Introduction: “first effective approach for automated molecular structure elucidation using graph-based machine learning on experimental data” should be tempered by explicit comparison numbers against NMRMind (33% on 12 molecules) and DiffNMR (11% experimental) already cited in the text, so the novelty claim is quantitative rather than absolute.","section":null},{"comment":"Methods / SI §1.2.5: the bond-existence threshold of 0.5 and the structure-correction MAE reopen threshold of 10 ppm are free parameters; a short sensitivity table (or statement that they were fixed a priori) would strengthen reproducibility.","section":null},{"comment":"Fig. 2 caption and main-text discussion of node degeneracy: the example molecule is clear, but the probability heat-maps would be easier to read if atom indices were overlaid on both the predicted and ground-truth skeletons.","section":null},{"comment":"SI §2.4.3 / Supplementary Table 4: the COSY/HMBC-ratio thresholds used to colour-code success/failure are useful; they should be mentioned briefly in the main-text experimental discussion so readers do not have to hunt the SI for the failure-mode analysis.","section":null},{"comment":"Data availability: Zenodo DOI is given; a one-sentence note confirming that the 19 experimental SDF + peak lists are included would help immediate re-use.","section":null}],"recommendation":"major_revision","confidential_remarks":"The technical core (one-shot + stepwise degeneracy correction + multi-shot) is solid and the simulated results are convincing. The experimental claim is currently overstated relative to the variable protocol and shared-model ranking; once those two points are fixed or clearly caveated the paper becomes a strong candidate for the journal. I do not see evidence of circularity beyond ordinary use of the authors’ own forward model, which is common and disclosed."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is that this is a working graph-transformer pipeline for NMR structure elucidation that actually transfers to experimental data without fine-tuning and solves some hard natural-product-like cases (betulin, 24-epibrassinolide). That is more than most of the recent GNN/LLM CASE papers deliver.\n\nWhat is new is the practical multi-stage fix for node degeneracy: one-shot bond probabilities on the full graph, then strip the silent heteroatoms (N/O/F), re-attach them stepwise with a focused fragment model, reopen high-error carbons via a forward 13C check, and finally multi-shot noise ensembles ranked by frequency (sim) or 13C MAE (exp). Simulated Top-1 hits 77.8% after correction on the 2k-molecule hold-out (Dataset B: only 1H/13C + COSY/HSQC/HMBC correlations), with sensible size/heteroatom trends. Experimental 10/19 is non-trivial; the successes include molecules that challenge chemists, and the MAE ranking cleanly separates correct (<4.4 ppm) from wrong (≥5.3 ppm) when the right graph is generated. Code and data are public. The architecture and training details are reproducible enough.\n\nSoft spots are real but proportionate. The experimental protocol is not fixed: three of the ten successes needed either 5000 extra shots or raised coupling noise (0.6 Hz). Ranking leans on the same IMPRESSION-G2 family that generated the Boltzmann-averaged training labels; they deliberately avoided that oracle for the simulated Top-k numbers, which is honest, but it means the experimental claim is not fully blind. N=19 is small, and performance still drops on heteroatom-rich or quaternary-dense regions—exactly where NMR information is sparse. None of this is fatal; the paper reports the failures and the extra compute clearly.\n\nThis is for people who actually run or automate structure elucidation, and for the ML-for-spectroscopy crowd. It deserves a serious referee. I would engage: read the SI carefully, try the code on a couple of my own spectra, and watch for a larger fixed-budget experimental benchmark. Worth the time.","headline":"Concrete multi-stage GNN that recovers real experimental structures for complex molecules up to 480 Da, with honest metrics and clear soft spots on protocol flexibility and shared-model ranking.","tokens_in":24639,"tokens_out":558,"would_cite":true,"duration_ms":15586,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A graph-transformer platform reconstructs molecular bonding from routine 1H/13C NMR spectra, solving 53% of tested experimental molecules up to 480 Da.","keywords":["molecular structure elucidation","NMR spectroscopy","graph transformer network","inverse IMPRESSION","bond prediction","computer-assisted structure elucidation","COSY HSQC HMBC","experimental NMR"],"falsifier":"Run the identical trained platform, without any change of weights or noise schedule, on a fresh set of twenty experimental molecules of comparable size and heteroatom count whose structures are already known by independent methods; if top-1 recovery falls well below 50% the experimental claim fails.","tokens_in":24440,"feed_emoji":"🔬","tokens_out":1011,"duration_ms":8205,"temperature":0.7,"pith_summary":"Chemists still spend large amounts of time turning NMR spectra into molecular structures by hand. This paper shows that an inverted graph transformer, trained only on simulated 1H and 13C chemical shifts plus COSY/HSQC/HMBC-style correlations, can recover the correct 2D connectivity for most molecules that contain up to thirty heavy atoms. The platform first predicts all bonds at once, then systematically re-attaches NMR-silent heteroatoms (N, O, F) whose nodes lack distinctive features, and finally ranks an ensemble of noise-perturbed candidates by how well their predicted 13C shifts match experiment. On a simulated test set the method reaches 77.8% top-1 accuracy; on nineteen real experimental spectra it correctly identifies ten structures, including polycyclic natural-product-like molecules that routinely challenge expert assignment. The result is the first graph-based machine-learning system that works on experimental NMR data without additional fine-tuning or prior knowledge of reactant fragments.","feed_headline":"Graph AI rebuilds molecules from ordinary NMR spectra","feed_subtitle":"Solves 10 of 19 experimental structures up to 480 Da without fine-tuning or reactant hints","key_machinery":"The three-stage pipeline (one-shot bond-probability prediction, iterative stepwise re-attachment of NMR-silent heteroatoms, and multi-shot noise-augmented ranking by IMPRESSION-G2 13C MAE) that converts sparse experimental correlations into a ranked list of chemically valid 2D graphs.","core_discovery":"The inverse-IMPRESSION platform, built from four task-specific graph transformers that share the IMPRESSION-G2 architecture, reconstructs molecular bonding directly from measurable 1H/13C NMR data. After one-shot bond prediction, stepwise correction of degenerate heteroatoms, and noise-augmented multi-shot ranking by 13C mean absolute error, the system recovers the correct structure for 77.8% of simulated molecules (up to 30 heavy atoms) and for 10 of 19 experimental molecules (molecular weights up to 480 Da).","pith_inferences":["Because the platform already ranks by a forward NMR predictor, it can be closed into a self-consistent loop that rejects candidates whose predicted spectra deviate beyond the known error of that predictor.","The same edge-probability formulation should transfer to other sparse spectroscopic graphs (IR, MS/MS, residual dipolar couplings) once analogous node and edge features are defined.","Performance cliffs at high heteroatom counts suggest that hybrid human-AI workflows, in which the chemist supplies a few local constraints, could push success rates well above the present 53%.","The 2000-shot ensembles already generate chemically valid alternatives; these could be used as an automatic “structure-revision” tool when the top-ranked molecule later fails orthogonal tests."],"forward_implications":["Routine 1H/13C COSY/HSQC/HMBC datasets can be turned into ranked 2D structure candidates without expert fragment assembly or exhaustive CASE enumeration.","Inclusion of experimentally accessible 15N and 19F data raises top-k accuracy above 80% even for molecules with more than ten heteroatoms.","The same ranking criterion (13C MAE below ~4.4 ppm) supplies an automatic confidence filter that separates correct from incorrect predictions.","Molecules whose spectra leave large NMR-silent regions (quaternary clusters, heteroatom-rich substructures) remain systematically harder for both the algorithm and human analysts."],"fun_headline_variants":["Graph AI reconstructs molecules from experimental NMR data","Inverse-IMPRESSION rebuilds bonding via inverted graph transformers","Platform recovers 10 of 19 experimental structures up to 480 Da","One-shot graph model predicts bonds from 1H/13C NMR spectra","Multi-shot ranking solves molecular structures without reactant hints"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That a model trained only on Boltzmann-averaged simulated spectra will transfer, without fine-tuning, to real experimental spectra well enough that ranking candidates by predicted 13C error recovers the true structure.","fun_headline_variants_meta":{"raw":{"variants":["Graph AI reconstructs molecules from experimental NMR data","Inverse-IMPRESSION rebuilds bonding via inverted graph transformers","Platform recovers 10 of 19 experimental structures up to 480 Da","One-shot graph model predicts bonds from 1H/13C NMR spectra","Multi-shot ranking solves molecular structures without reactant hints"]},"model":"grok-4.5","effort":"low","cost_usd":0.004874,"raw_usage":{"total_tokens":1429,"prompt_tokens":829,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":48740000,"prompt_tokens_details":{"text_tokens":829,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":531,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":829,"tokens_out":69,"duration_ms":3945,"temperature":1.0,"reasoning_tokens":531,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T01:17:57.578340+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the identical trained platform, without any change of weights or noise schedule, on a fresh set of twenty experimental molecules of comparable size and heteroatom count whose structures are already known by independent methods; if top-1 recovery falls well below 50% the experimental claim fails.","supporting_citations":[],"review_version":1}