{"id":"f16b4594-b260-481c-aebd-752e417f9182","arxiv_id":"2511.15672","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A hybrid quantum-classical classifier on simulated HH→bbγγ events claims 95% CL limits of 1.9–2.1×SM, but the gain over XGBoost is 21–29%, not the advertised factor of two.","lead":"This paper tests a hybrid quantum-classical neural network on simulated LHC double-Higgs events, reporting better expected search limits than a classical XGBoost baseline. The headline factor-of-two gain only holds against the pure-quantum baseline, and the quoted limits rest on a simplified detector simulation with only one systematic uncertainty.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The expected limits are extracted from the same test set used to choose λ and the optimized binning; no held-out validation or leakage correction is described, so the quoted 1.9/2.1×σSM are optimistically biased.","rationale":"The reader's stated weakest assumption (limited systematics) is reasonable and should be fixed, but I would not make it the primary attack. The most immediate, testable flaw is that the quoted expected limits are produced from the same events that were used to select the model and the binning. The manuscript's own split (Section V.A) sets aside 25% for 'performance evaluation and statistical treatment,' and Section V.D selects λ from ROC curves computed on that test set; Section VI then optimizes the score-region boundaries on the same sample. This is not merely a missing source of uncertainty — it invalidates the point estimate of the claim, since the number is an in-sample maximum rather than an unbiased projection. If the test-set reuse is corrected and the limit moves to ~2.5×σSM or higher, then the headline improvement over XGBoost is no longer established. Fixing systematics alone, while necessary for a physics result, would not address this selection bias. I therefore agree only partially with the reader's weakest_assumption: both threaten the central claim, but the leakage is the more load-bearing and can be checked immediately with a nested validation. Because the reader's verdict is already REJECT and my concern reinforces that rejection, I recommend no change to the verdict.","tokens_in":15626,"tokens_out":4248,"duration_ms":48610,"concrete_test":"Recompute the full pipeline with a three-way split: train on the current 75%, select λ and bin boundaries on a validation subset carved out of that training sample, and run the pyhf limit fit only once on the currently reserved 25% test set. Compare the resulting 95% CL upper limits (and the XGBoost/pure-QML differences) with Table III and Figure 7. If the 1.9×σSM limit degrades by more than ~20% or the claimed 21% improvement over XGBoost disappears, the central claim is an artifact of test-set reuse.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weak point is the reuse of the 25% test sample for both model selection and the quoted limits. Section V.A splits the data 75/25 and reserves the 25% 'exclusively for performance evaluation and statistical treatment.' Section V.D then uses ROC curves computed on this test set to pick the hyperparameter λ=0.1 (Figure 4, caption 'computed using event weights on the test dataset'). Section VI defines 'optimized score regions' whose binning is chosen to maximize the Asimov significance, again with no statement that the binning is fixed before looking at the test events. The same 25% events therefore enter (i) hyperparameter selection, (ii) bin-boundary optimization, and (iii) the final pyhf limit fit. Selecting thresholds and binning on the data that are later used for inference makes the quoted 1.9×σSM and 2.1×σSM expected limits in-sample estimates; they are not unbiased projections of performance on new data. This is distinct from the reader's missing-systematics point: even under the paper's stated 10%/50% normalization-only model, the statistical procedure is not valid after adaptive selection on the test set. The comparison with XGBoost and pure-QML is also affected if their thresholds/binnings were chosen with the same protocol. A proper nested validation or an untouched final evaluation set is required before the central improvement claim can be assessed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a hybrid quantum-classical machine learning (HyQML) framework for the HH→bbγγ search at the LHC, combining a classical meta-network that generates parameterized quantum circuit (PQC) angles with a 4-qubit, 4-layer hardware-efficient ansatz, with amplitude embedding of kinematic features. Using Delphes-simulated events at √s=13.6 TeV and 308 fb⁻¹, the authors train a binary classifier and extract expected discovery significances, 95% CL upper limits on the di-Higgs cross-section, and one- and two-dimensional constraints on κλ and κ2V, under background normalization uncertainties of 10% and 50%. The headline claim is that the hybrid model outperforms both XGBoost and a purely quantum model by a factor of two, achieving an expected 95% CL upper limit of 1.9×σSM (10% systematics) and 2.1×σSM (50% systematics).","tokens_in":16062,"tokens_out":7474,"duration_ms":82614,"significance":"If the quantitative claims were correct, the paper would be a valuable demonstration that hybrid quantum-classical classifiers can improve LHC search sensitivity, with implications for future HH analyses. The architecture—a meta-network conditioning PQC parameters on event kinematics—is a plausible and interesting idea, and the use of a profile-likelihood framework with pyhf is appropriate. However, the manuscript as written contains several load-bearing problems: the abstract's 'factor of two' claim is contradicted by the paper's own Tables and Figures; the κ2V scan refers to an equation that does not contain κ2V; the test set is reused for model selection, binning, and final inference; and only background normalization uncertainties are included. If these issues were corrected, the central finding would be considerably weaker but still potentially publishable as a proof-of-principle study.","major_comments":[{"comment":"The abstract claims the hybrid model 'outperforms both a state-of-the-art XGBoost model and a purely quantum implementation by a factor of two.' This is not supported by the results. Table III gives expected significances of 1.41 (HyQML, 10% sys.) vs 1.09 (XGBoost) — a factor of 1.29 — and vs 0.65 (pure-QML) — a factor of 2.17. Figure 7 gives 95% CL limits of 1.9 (HyQML) vs 2.4 (XGBoost) — a 21% improvement — and vs 3.3 (pure-QML) — a factor of 1.75. The text in Section VI and the conclusion correctly state a 1.75 factor vs pure-QML and 21% vs XGBoost. The abstract must be corrected, and the central claim should be scaled to the actual numbers.","section":"Abstract; Table III; Fig. 7"},{"comment":"The 25% test sample is used for multiple adaptive purposes: (i) the hyperparameter λ is selected from ROC curves 'computed using event weights on the test dataset' (Fig. 4 caption, Sec. V.D); (ii) the 'optimized score regions' in Sec. VI are defined to maximize the Asimov significance without any statement that the binning was fixed before inspecting the test data; and (iii) the final pyhf likelihood fit and the quoted expected limits are computed on the same test events. This constitutes in-sample selection after adaptive optimization. The quoted expected significance of 1.41 and limits of 1.9/2.1×σSM are therefore optimistically biased and are not unbiased projections for new data. A nested validation loop or a separate, untouched evaluation set is required before these numbers can be taken at face value.","section":"Sec. V.A, V.D, VI"},{"comment":"The one-dimensional κ2V scan is introduced by saying that κ2V 'parameterizes deviations in the quartic VVHH interaction entering the VBF Higgs pair production process (Equation 2).' However, Eq. (2) gives σVBF only as a function of κλ and contains no κ2V dependence. Figure 2(b) shows a κ2V variation, but no formula for the κ2V dependence is provided. Consequently, the likelihood scan in Fig. 8(b) and the two-dimensional contour in Fig. 9 are not reproducible from the information in the paper. The authors must supply the full cross-section parameterization used for the κ2V scan.","section":"Sec. VI, Eq. (2)"},{"comment":"The statistical interpretation includes only background normalization uncertainties (10% and 50%) as log-normal nuisance parameters. All other experimental and theoretical systematics—photon identification, b-tagging efficiency, luminosity, PDF, QCD scale, and MC statistical uncertainties—are assumed negligible. The comparison in Section VI with the ATLAS expected limit, and the statement that the HyQML 'achieves a significant improvement' over ATLAS, is not a fair comparison unless the same systematic model is used. At minimum, the paper should frame this as an idealized projection and restrict comparisons to the authors' own XGBoost and pure-QML baselines; better, it should include the dominant systematic uncertainties or a sensitivity study.","section":"Sec. VI, statistical model"}],"minor_comments":[{"comment":"The abstract's 'factor of two' wording should be made consistent with Section VII's more accurate statement ('almost factor-of-two improvement compared to the pure-QML model and a 21% improvement with respect to an XGBoost model').","section":"Abstract / Section VII"},{"comment":"The cross-section parameterizations are cited from Refs. [12] and [19], but Eq. (2) lacks the κ2V dependence mentioned in the text. Please clarify the exact functional form used for VBF and cite the original source.","section":"Eq. (1), Eq. (2)"},{"comment":"The meta-learning objective Lmeta = ⟨log κ(g_ij)⟩ is imported from the authors' own preprint [41]. For reproducibility, the Fubini–Study metric, its condition number, and how the average is estimated should be defined explicitly in the text.","section":"Sec. V.C"},{"comment":"The 'optimized score regions' procedure is under-specified. The authors should state the number of bins, the threshold scan range, the binning granularity, and whether the same binning was used for all models (HyQML, pure-QML, XGBoost) to allow a fair comparison.","section":"Sec. VI"},{"comment":"The γγ+jets cross-section of 48.1 pb is not referenced. Please clarify its provenance and whether it corresponds to the inclusive cross-section after the applied generator-level requirements.","section":"Table I"},{"comment":"References [30] and [34] lack full publication details (journal or arXiv number). Please complete them.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a novel hybrid quantum-classical architecture, but the quantitative conclusions are not reliable in the current form. The most serious issue is the reuse of the test set for hyperparameter selection, binning, and final inference; this alone invalidates the quoted expected limits. The abstract's 'factor of two' claim is contradicted by the paper's own tables. The κ2V scan is not backed by a formula. I would not consider acceptance without a complete re-analysis with a proper validation protocol and corrected claims. The self-citations [12,41] for the baseline and meta-objective are not circular but should be checked for consistency with the new analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a well-written simulation study of a hybrid quantum-classical classifier for HH→bbγγ. The architecture is genuinely interesting: a meta-network maps event features to PQC rotation angles, trained with a Fubini-Study metric-conditioning objective from the authors' own earlier work. That is a real, non-trivial combination, and the paper does a decent job describing the simulation, the event selection, and the likelihood setup. It also compares against XGBoost and a pure-QML variant. So there is substance here.\n\nThe problems are in the claims and the statistics. The abstract says the hybrid model outperforms XGBoost by a factor of two. The paper's own Table III shows significance 1.41 vs 1.09, a 29% gain, and Figure 7 shows limits 1.9 vs 2.4, a 21% gain. The body correctly states 21%, but the abstract is what people will remember. That inconsistency should have been caught.\n\nMore serious: the κ2V scan is said to use Eq. 2, but Eq. 2 contains no κ2V dependence. The VBF cross-section parameterization is quoted as a function of κλ only. So the κ2V constraints are not actually derived from the parameterization the text claims. This is a load-bearing gap in the coupling results.\n\nThe statistical interpretation includes only a background normalization nuisance (10%/50%). All other systematics are ignored. The authors acknowledge this implicitly, but they still compare their 1.9×σSM with ATLAS's 3.7×σSM. That comparison is invalid, as they half-admit, but the number is presented as the main result anyway.\n\nThe stress-test note raises a deeper problem that I think is correct: the same 25% test sample is used to pick λ (Figure 4, ROC on the test dataset), to choose the optimized binning (Section VI, maximizing Asimov significance), and then to fit the limits in pyhf. There is no held-out set or nested validation. That makes the quoted limits in-sample estimates. Even if the systematics were complete, the statistical procedure is not valid after adaptive selection on the same events. This alone would require major revision.\n\nNo code or data are released; 'available upon reasonable request' is not enough for a claim like this.\n\nI want to be clear about what the paper does well: the ML architecture is coherent, the training procedure is described in detail, and the comparison setup is sensible apart from the leakage. The writing is clear. The flaws are not in the core idea; they are in the calibration of the claims and the statistical protocol.\n\nWho should read this? People working on QML for LHC physics will want to know the architecture, but they should not trust the quantitative results as-is. I would send it to peer review rather than desk reject, because the method deserves scrutiny and the errors are fixable. But it needs a major revision, not minor tweaks.","headline":"A genuinely interesting HyQML architecture, but the headline gains are overstated and the limits are in-sample because the same test set drives model selection, binning, and the final fit.","tokens_in":16522,"tokens_out":4811,"would_cite":false,"duration_ms":44709,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid quantum-classical classifier improves expected LHC sensitivity to non-resonant double-Higgs production by roughly a factor of two, reaching an expected 95% CL limit of 1.9 times the Standard Model cross-section.","keywords":["quantum machine learning","Higgs pair production","hybrid quantum-classical model","parameterized quantum circuits","LHC physics","HH→bbγγ final state","Higgs self-coupling","expected limit"],"falsifier":"Re-run the same binned likelihood fit including the usual experimental systematic uncertainties (photon identification, b-tagging, luminosity, PDF, scale) and check whether the expected 95% CL limit remains below about 2×σSM; if it rises substantially, the hybrid model's advantage is an artifact of the omitted systematics. Alternatively, train the hybrid architecture with the PQC replaced by an identity transformation while keeping the meta-network's conditioning; if AUC and limits are unchanged, the quantum component is not load-bearing.","tokens_in":15516,"feed_emoji":"⚛️","tokens_out":8717,"duration_ms":81079,"temperature":0.7,"pith_summary":"This paper proposes a hybrid quantum machine-learning (HyQML) framework for separating Higgs-pair signal from background in the HH→bbγγ final state at the LHC. It claims that a small parameterized quantum circuit, whose rotation angles are generated event-by-event by a classical neural network, outperforms both a conventional gradient-boosted classifier and a purely quantum circuit by roughly a factor of two in expected sensitivity. With 308 inverse femtobarns of simulated Run-2+3 data, the paper quotes an expected 95% CL upper limit of 1.9×σSM (10% background-normalization uncertainty) and 2.1×σSM (50%) on non-resonant double-Higgs production, and correspondingly tighter constraints on the Higgs self-coupling κλ and the quartic coupling κ2V. If correct, this would make hybrid quantum-classical models a practical route to improving rare-process searches at colliders on near-term hardware.","feed_headline":"Hybrid quantum model improves double-Higgs search to 1.9×SM","feed_subtitle":"Expected 95% CL limit falls to 1.9 times the Standard Model rate, tightening Higgs self-coupling constraints.","key_machinery":"The key mechanism is a 4-qubit, 4-layer hardware-efficient parameterized quantum circuit whose per-event rotation angles are not fixed but generated by a 'MetaParamMapNet'—a fully connected neural network with three 64-unit hidden layers and a linear head per circuit parameter. Input features are amplitude-embedded into the quantum state; after entangling CNOT gates, expectation values of Pauli-Z measurements are fed to a linear classification layer. A meta-learning stage minimizes the condition number of the Fubini-Study metric of the PQC, which is intended to keep the quantum landscape trainable and avoid barren plateaus. The interpolation parameter λ controls how strongly the meta-network","core_discovery":"The central claim is that the hybrid architecture—a meta-parameter mapping network that outputs data-dependent PQC rotation angles—learns a quantum feature space in which signal and background separate more cleanly than in the representations learned by a pure quantum circuit (λ=0) or by XGBoost. The paper demonstrates this through ROC/AUC comparisons (AUC 89.75% vs 69.35% pure-QML), through expected discovery significance (1.41 vs 0.65), and through a binned profile-likelihood fit that yields an expected 95% CL upper limit of 1.9×σSM. The authors attribute the gain to the hybrid model's ability to encode high-order correlations in an exponentially large quantum state space while retaining c","pith_inferences":["The reported advantage could partly stem from the Fubini-Study conditioning meta-learning rather than from quantum entanglement itself; replacing the PQC with a classical neural net of comparable capacity under the same meta-conditioning would isolate the quantum contribution.","If the gain survives a full treatment of systematic uncertainties—photon identification, b-tagging, luminosity, and theory-scale uncertainties—the method may push LHC di-Higgs searches toward the SM-rate sensitivity region, but that remains untested.","The same hybrid architecture, with its per-event parameter generation, may be applied to other low-yield, high-dimensional LHC signatures (e.g., HH→4b, ttH, or exotic resonances) where classifier saturation is a known bottleneck.","Since the quantum circuit is simulated exactly, the projection to noisy real hardware is not guaranteed; testing on NISQ devices with error mitigation would reveal whether the 'quantum' part of the advantage survives decoherence."],"forward_implications":["Under 10% background-normalization uncertainty, the expected 95% CL upper limit on non-resonant HH production is 1.9×σSM, improving on the 2.4×σSM quoted for the boosted-tree baseline and 3.3×σSM for the pure quantum circuit.","The expected discovery significance for SM HH production is 1.41 with the hybrid model, versus 1.09 for the boosted-tree baseline and 0.65 for the pure quantum model, under the same 10% uncertainty.","The profile-likelihood scan yields expected 95% CL intervals κλ ∈ [−0.4, 4.9] and κ2V ∈ [−0.6, 2.9], both tighter than the corresponding intervals obtained with the classical and pure-quantum classifiers.","Raising the background-normalization uncertainty from 10% to 50% degrades the expected limit only mildly, from 1.9 to 2.1×σSM, indicating the sensitivity gain is not driven by the assumed normalization constraint."],"fun_headline_variants":["Hybrid quantum ML achieves 1.9×SM double-Higgs limit","Quantum-classical hybrid reaches 1.9×SM in Higgs search","Hybrid quantum model tightens Higgs coupling bounds to 1.9×SM","Quantum neural network hybrid hits 1.9×SM for double-Higgs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The quoted limits treat only the background normalization as uncertain (at 10% or 50%); all other experimental and theoretical systematic uncertainties are assumed negligible, so if those are added the expected limit would worsen.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid quantum ML achieves 1.9×SM double-Higgs limit","Quantum-classical hybrid reaches 1.9×SM in Higgs search","Hybrid quantum model tightens Higgs coupling bounds to 1.9×SM","Quantum neural network hybrid hits 1.9×SM for double-Higgs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000509,"raw_usage":{"total_tokens":2321,"prompt_tokens":757,"completion_tokens":1564,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":1479}},"tokens_in":501,"tokens_out":1564,"duration_ms":13704,"temperature":1.0,"reasoning_tokens":1479,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T21:18:23.069860+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same binned likelihood fit including the usual experimental systematic uncertainties (photon identification, b-tagging, luminosity, PDF, scale) and check whether the expected 95% CL limit remains below about 2×σSM; if it rises substantially, the hybrid model's advantage is an artifact of the omitted systematics. Alternatively, train the hybrid architecture with the PQC replaced by an identity transformation while keeping the meta-network's conditioning; if AUC and limits are unchanged, the quantum component is not load-bearing.","supporting_citations":[],"review_version":1}