{"id":"52b11995-b055-49e3-bbb4-1551924a253e","arxiv_id":"2506.18981","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A neural network trained on simulated m_hh distributions can extract the 2HDM coupling product xi_H^t x lambda_hhH at 10-20% precision in a hypothetical HL-LHC scenario with m_H = 450 GeV.","lead":"This paper simulates HL-LHC data for Higgs pair production in a 2HDM and trains a small neural network to read the product of a heavy Higgs top-Yukawa coupling and a BSM triple Higgs coupling from the shape of the invariant mass distribution. The authors report 10 to 20 percent precision in this simulated setting and claim the network beats maximum likelihood estimation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10-20% precision claim is conditional on a statistical-only uncertainty model; systematics are explicitly deferred, so the external claim is not established.","rationale":"The paper is a coherent simulation-based sensitivity study: it uses HPAIR 2HDM m_hh templates, applies smearing and binning, trains a small NN, and compares against a discrete-grid MLE baseline. The internal demonstration that the NN can recover xi_H^t x lambda_hhH from statistically fluctuated histograms is credible, and the appendix grid-size check plus the 4-fold training splits provide some support. The reader's verdict of CONDITIONAL is appropriate because the abstract and title present an 'experimental determination' with 10-20% precision, while the analysis only includes Poisson statistics and explicitly defers systematics. That is the load-bearing gap: every quantitative statement about final HL-LHC precision depends on the unquantified assumption that systematic uncertainties are subdominant. A concrete systematics-injected rerun would settle whether the claim survives; absent that, the paper's contribution is best read as a proof-of-principle projection. I do not see a separate internal flaw that would overturn the central method, and I agree with the reader that the weakest assumption is the systematic-uncertainty exclusion.","tokens_in":23587,"tokens_out":3033,"duration_ms":36316,"concrete_test":"Repeat the AE95 analysis of Secs. 5.3-5.4 using mock data generated with an added systematic budget: e.g. a 10% log-normal normalization nuisance per m_hh bin (uncorrelated) plus a 5 GeV correlated m_hh scale shift, and retrain/evaluate the NN on augmented smeared samples. If the 1-sigma crossing intervals in Eq. (13) (currently about +/-0.045 for dataset 1) widen by more than a factor of two, the claimed 10-20% precision does not survive plausible systematics and should be rephrased as a statistical-only projection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim, that xi_H^t x lambda_hhH can be determined at 10-20% at the HL-LHC, rests on the statistical-only error model built in Sec. 2.6 and used in Sec. 5.3. There, each bin count is Poisson-fluctuated around the HPAIR prediction using Eq. (7), with fixed total and signal-region efficiencies from Ref. [59] and no theoretical or systematic uncertainties. Section 2.6 explicitly says the latter two are 'harder to estimate' and are assumed 'in an optimistic scenario' to be subdominant; Sec. 6 repeats that systematic uncertainties are beyond scope. The 10-20% number therefore follows only if real HL-LHC systematics, including b-tagging efficiency uncertainties, m_hh scale/calibration uncertainties, background contamination, and theory uncertainties on the m_hh template, are indeed subdominant. This is an external-validity assumption, not an internal regression property, and it is the least secured step between the NN demonstration and the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a machine-learning-based sensitivity study for extracting the product of the heavy-Higgs top-Yukawa coupling modifier and the BSM trilinear coupling, ξ_H^t × λ_hhH, from the invariant mass distribution of gg→hh at the HL-LHC. The study is set in the Type I 2HDM with m_H = 450 GeV, using HPAIR predictions, with 15% Gaussian smearing and 50 GeV binning, and Poisson fluctuations derived from expected b bbar b bbar event counts with ATLAS efficiencies at 3000 fb^-1. A single-hidden-layer neural network is trained on simulated m_hh histograms and compared with maximum-likelihood template estimation and a likelihood-ratio goodness-of-fit test. The authors report that the NN outperforms the classical methods and claim that a 10-20% determination of the coupling product may be achievable by the end of the HL-LHC if efficiencies improve.","tokens_in":23906,"tokens_out":7208,"duration_ms":70790,"significance":"The paper is a useful proof-of-principle: it gives a fully specified benchmark plane with public tools, cross-validated interpolation checks, an explicit comparison metric (AE95), and appendices quantifying grid-size and architecture effects. It is, however, an idealized simulation study rather than a validated experimental projection. The central quantitative claim depends on assumptions—statistical-only uncertainties, specific efficiencies, known m_H, and a single underlying model—that are not tested against independent data. The main value is in demonstrating that a small NN can perform regression on binned m_hh spectra more precisely than a discrete template maximum-likelihood fit, provided the training and test data are generated from the same model.","major_comments":[{"comment":"The headline claim of a 10-20% 'experimental determination' is conditional on an uncertainty model that includes only Poisson statistics. Section 2.6 explicitly defers theoretical and systematic uncertainties ('harder to estimate', 'optimistic scenario'), and Section 6 repeats that systematics are beyond scope. The manuscript does not quantify b-tagging efficiency uncertainties, m_hh scale/calibration uncertainties, background contamination, or theory uncertainties on the m_hh template. The precision in Sec. 5.5 therefore follows only if all of these are subdominant, which is an unvalidated external assumption. The title and abstract should be revised to present this as an idealized sensitivity study with explicit caveats, or the systematics must be incorporated.","section":"Sec. 2.6 and Sec. 6"},{"comment":"Training and evaluation are performed on Poisson-smeared distributions generated from the same 2HDM parameter plane (e.g., the 4-fold split in Sec. 5.1 and the 256 test histograms in Sec. 5.3). This is an in-model interpolation check, not a validation against independent data. The paper acknowledges model dependence in Sec. 6, but the abstract's 'Experimental Determination' claim is much stronger than what the setup can establish. A closure test with out-of-model distributions (e.g., an EFT parameterization or a different scalar sector) or a clear statement that the result is only an internal consistency check is needed before the experimental claim can be sustained.","section":"Sec. 4, Sec. 5.1, Sec. 5.3"},{"comment":"The '10-20% level' precision is never defined quantitatively. The only precision metric, AE95, is an absolute error at 95% CL, and the quoted sensitivity ranges in Eqs. (11)-(13) are absolute thresholds on ξ_H^t × λ_hhH. No relative-error distribution or AE95 divided by |ξ_true| is shown. For the example in Sec. 5.3, the AE95 near ξ ≈ ±0.03-0.05 is comparable to the value itself, which would correspond to 60-100% relative error rather than 10-20%. The improved-efficiency results in Fig. 20 are shown only as scatter/density plots. Please state precisely which interval (e.g., 68% or 95% CL on |pred-θ|/|θ| over a specified θ range) supports the 10-20% claim and give the corresponding numbers.","section":"Sec. 5.4 and Sec. 5.5"},{"comment":"The grid-size check in Table 2 reports deviations Δ of 12%, 8%, and 6% for three points with small ξ, yet the text concludes that the effect is at the 'few percent level'. These deviations are comparable to or larger than the claimed 10-20% precision, and the coarser grid is used throughout. Since the authors themselves interpret the grid coarseness as a source of theoretical uncertainty, this internal check actually weakens the precision claim unless the grid-induced error is included in the reported uncertainty budget.","section":"Appendix A"}],"minor_comments":[{"comment":"The text states that H1 is the full 2HDM prediction including the resonance, but the expression for L(H1) uses the observed counts n_i as the mean (the saturated model). Please reconcile the wording with the formula and justify the number of degrees of freedom used in the χ² approximation.","section":"Sec. 3, Eq. (8)"},{"comment":"The notation σ_i is described as the differential cross section times bin size; please make this explicit in the equation and in the units of N_i.","section":"Sec. 2.6, Eq. (7)"},{"comment":"The claim that MLE 'manages to determine if ξ ≠ 0' while being unable to predict the sign is not evident from the X-shaped scatter plot; please clarify what property of the plot supports the first part of this statement.","section":"Sec. 3, Fig. 8"},{"comment":"The contours and color coding are difficult to compare; consider adding a legend that lists the AE95 thresholds for each contour line.","section":"Sec. 5.4, Fig. 18"},{"comment":"There are minor typos, including 'the a χ2-distributed variable' in Sec. 3, 'T able 2' in the Appendix A header, and inconsistent spacing in reference [29].","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the phenomenological audience, but the title and abstract considerably oversell the result. The authors should either substantially temper the claims or add the missing uncertainty estimates. The in-sample nature of the validation should be made explicit in the abstract. I would not recommend rejection because the core NN-versus-MLE comparison is a sound, reproducible simulation study, but the current framing is not acceptable for publication as an 'experimental determination'."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a quick look if you work on 2HDM phenomenology or ML for Higgs pair production. The genuinely new piece is the first ML sensitivity study aimed at extracting ξ_H^t × λ_hhH from the HL-LHC m_hh spectrum. The paper is internally coherent: HPAIR 2HDM distributions, realistic smearing and binning, Poisson fluctuations, cross-validation on held-out parameter points, and a clean comparison against a maximum-likelihood scan. The NN is almost trivially small, which is a plus—it shows the gain comes from the regression setup, not from a black box. The density plots and the AE95 metric give an honest picture of where the method works and where it scatters. That part deserves credit.\n\nThe soft spots are mostly about the distance between the headline and the evidence. The 10–20% determination claim rests on a statistical-only uncertainty model. Section 2.6 is explicit that systematics are assumed subdominant in an optimistic scenario, and Section 6 repeats that they are beyond scope. That is an external-validity assumption, not a demonstrated result. The title and abstract should say 'projected sensitivity under optimistic systematics' rather than 'experimental determination.' Also, the classical MLE baseline is a discrete grid with bin aggregation; a continuous profile likelihood might do better, so the 'outperforms conventional methods' claim is fair only against that specific baseline. The train/test split comes from the same 2HDM plane, so the reported recovery is an in-model consistency check—appropriate for a sensitivity projection, but not a validation against independent data.\n\nI do not think any of this is fatal. The internal logic holds and the authors are honest about the simplifications. The self-citations to Ref. [29] are appropriate since that is the direct classical-method predecessor.\n\nWho gets value from this: phenomenologists interested in whether NNs are useful for extracting BSM Higgs couplings, and people planning future experimental projections. It deserves a serious refereeing, but the authors should be pushed to recalibrate the abstract and add an explicit systematics caveat, or better, a rough systematics budget. I would accept it for review with that expectation.","headline":"A solid simulation-based sensitivity study showing a small NN beats a grid MLE for extracting a BSM triple-Higgs coupling, but the title's 'experimental determination' overstates what a statistical-only projection can support.","tokens_in":24369,"tokens_out":1278,"would_cite":true,"duration_ms":16862,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a small neural network can read the product of a heavy Higgs boson's top-Yukawa coupling and its trilinear coupling to two light Higgses from the HL-LHC di-Higgs mass spectrum, reaching 10--20% precision with…","keywords":["Higgs pair production","trilinear Higgs coupling","two-Higgs-doublet model","neural network","HL-LHC","invariant mass distribution","maximum likelihood estimation","beyond-Standard-Model Higgs sector"],"falsifier":"Compute the network's $AE_{95}$ on pseudo-experiments generated with a full detector simulation that adds correlated per-bin systematic shifts of a few percent to the $b\\bar b\\, b\\bar b$ event counts at 3000 fb$^{-1}$; if the 95% confidence interval for $\\xi_H^t\\times\\lambda_{hhH}$ widens beyond roughly 0.045 for the $1\\sigma$-trained network, or beyond the reported 10--20% at improved efficiencies, the central projection is contradicted.","tokens_in":23339,"feed_emoji":"⚛️","tokens_out":9082,"duration_ms":81103,"temperature":0.7,"pith_summary":"The paper claims that the product of the top-Yukawa coupling of a heavy CP-even Higgs boson and its trilinear coupling to two 125 GeV Higgs bosons, $\\xi_H^t \\times \\lambda_{hhH}$, can be extracted from the invariant-mass spectrum of Higgs pairs at the HL-LHC using a small neural network. This product controls the resonant $H$ contribution to gluon-fusion di-Higgs production, where it creates a dip-peak or peak-dip interference structure around $m_H$ that ordinary fits struggle to read. Training on 16-bin histograms that include 15% mass smearing, 50 GeV binning, and Poisson fluctuations, the network predicts the coupling product more accurately than maximum-likelihood estimation, including its sign. Assuming the $b\\bar b\\, b\\bar b$ final state, $3\\,\\text{ab}^{-1}$ of data, and improved detector efficiencies, the paper projects a determination at the 10--20% level by the end of the HL-LHC. The result matters because it is the first machine-learning sensitivity study for a beyond-Standard-Model trilinear Higgs coupling, a quantity tied to the shape of the Higgs potential.","feed_headline":"Neural network reads Higgs-pair mass shapes for a BSM coupling","feed_subtitle":"Trained on smeared Higgs-pair mass spectra, it beats maximum-likelihood fits and could reach 10--20 percent precision.","key_machinery":"The machinery is the resonant interference structure in the di-Higgs invariant-mass distribution $m_{hh}$: the $s$-channel exchange of a heavy CP-even scalar $H$ interferes with the continuum box and light-Higgs diagrams, producing a dip-peak or peak-dip feature at $m_{hh}\\approx m_H$ whose orientation encodes the sign of $\\xi_H^t\\times\\lambda_{hhH}$. The paper converts that shape into a regression problem for a small neural network: batch normalization, one hidden layer of 64 ReLU neurons, a linear output, Adam optimization, and mean-squared-error loss, trained on 16-bin histograms (50 GeV bins) after 15% Gaussian smearing in $m_{hh}$ and Poisson resampling of event counts. The 95% confidence interval of the absolute error, $AE_{95}$, is the metric used to compare methods and training sets. This construction is what lets the network learn the coupling product from the spectrum, including the sign, without bin aggregation.","core_discovery":"The central claim is that a deliberately simple one-hidden-layer neural network can infer $\\xi_H^t \\times \\lambda_{hhH}$ from the shape of the $m_{hh}$ distribution in $gg\\to hh$ production, in a Type I 2HDM benchmark with $m_H=450$ GeV that satisfies theoretical and experimental constraints. The network takes as input the 16 bins of a smeared and binned $m_{hh}$ histogram and outputs the coupling product; during training each input is re-sampled with Poisson noise matching the expected event counts in the $b\\bar b\\, b\\bar b$ channel at 3000 fb$^{-1}$. Compared with maximum-likelihood estimation, the NN gives smaller 95% CL errors, resolves the sign of the product, and can serve simultaneously for hypothesis testing and parameter estimation. The paper reports sensitivity ranges of about $0.045$ for the NN versus about $0.076$--$0.080$ for MLE, and finds that training on additional Poisson smearing of about $2\\sigma$ improves robustness. With a hypothetical factor-of-four improvement in experimental efficiencies, the projected precision reaches 10--20%; only statistical uncertainties are included, with systematics left out as beyond the scope.","pith_inferences":["A multi-channel extension the paper leaves implicit would feed NN inputs from $b\\bar b\\gamma\\gamma$ and $b\\bar b\\tau^+\\tau^-$ alongside $b\\bar b\\, b\\bar b$; the per-channel statistical dilution suggests a combined network could reach the low end of the 10--20% band or better.","Because the network resolves the sign of the coupling product, the same regression could be adapted to measure the CP properties of the heavy scalar, since a CP-violating admixture would reshape the interference pattern in $m_{hh}$ in a way a binned likelihood fit would miss.","A testable generalization is to regress jointly on $m_H$ and $\\xi_H^t\\times\\lambda_{hhH}$ rather than fixing $m_H=450$ GeV, letting the position of the dip-peak structure and its distortion be learned together; this would turn the method into a resonance search and parameter measurement in one step."],"forward_implications":["If the projection holds, $\\xi_H^t\\times\\lambda_{hhH}$ becomes a measurable target of HL-LHC Higgs-pair programs, not just a model parameter, with 10--20% precision when efficiencies improve by a factor of four.","The neural network outperforms maximum-likelihood estimation for parameter estimation and matches the classical $p$-value test for hypothesis testing, so a single trained network can replace both steps.","Uncertainties in $m_H$ and $m_{12}^2$ can be absorbed by training: including $m_H$ values spanning twice the expected measurement uncertainty keeps the prediction accurate, and freeing $m_{12}^2$ can even improve performance.","The method transfers to any model with a heavy CP-even scalar resonance in $gg\\to hh$, since the trained quantity and the mass shape are defined model-independently once the resonance is present.","Training on data smeared to about $2\\sigma$ Poisson fluctuations yields better predictions at 95% CL than training at $1\\sigma$, indicating a dedicated noise-augmentation strategy improves the extraction."],"supporting_citations":[{"why":"Establishes the dip-peak/peak-dip structure in the 2HDM $m_{hh}$ distribution and the sensitivity to $\\lambda_{hhH}$ that this paper exploits.","marker":"[29]"},{"why":"Provides the di-Higgs production code used to generate the leading-order $m_{hh}$ distributions.","marker":"[36]"},{"why":"Supplies the QCD-corrected di-Higgs production computation that backs the theoretical predictions.","marker":"[37]"},{"why":"Supplies the experimental efficiencies for the $b\\bar b\\, b\\bar b$ channel that set the expected event counts.","marker":"[59]"},{"why":"Provides the constraint implementation that defines which parameter points are allowed in the benchmark plane.","marker":"[44]"},{"why":"Supplies the interface to experimental Higgs constraints used for compatibility with LHC, Tevatron, and LEP searches.","marker":"[45]"},{"why":"Provides the neural-network implementation used for training the regression model.","marker":"[61]"}],"fun_headline_variants":["Neural net extracts BSM Higgs coupling from mass shapes better than MLE","NN beats maximum-likelihood fits for BSM Higgs self-coupling at HL-LHC","ML extracts BSM triple-Higgs coupling from Higgs-pair mass shapes, outperforming MLE","Simple neural net reads Higgs-pair mass shape to pin down BSM coupling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that systematic uncertainties in the measured invariant-mass spectrum are smaller than the statistical fluctuations; the analysis includes only Poisson statistics and explicitly leaves systematics out because they are harder to estimate, so if real systematics are comparable the 10--20% claim does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Neural net extracts BSM Higgs coupling from mass shapes better than MLE","NN beats maximum-likelihood fits for BSM Higgs self-coupling at HL-LHC","ML extracts BSM triple-Higgs coupling from Higgs-pair mass shapes, outperforming MLE","Simple neural net reads Higgs-pair mass shape to pin down BSM coupling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000771,"raw_usage":{"total_tokens":3487,"prompt_tokens":1093,"completion_tokens":2394,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":709,"completion_tokens_details":{"reasoning_tokens":2303}},"tokens_in":709,"tokens_out":2394,"duration_ms":18252,"temperature":1.0,"reasoning_tokens":2303,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:40:29.082620+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the network's $AE_{95}$ on pseudo-experiments generated with a full detector simulation that adds correlated per-bin systematic shifts of a few percent to the $b\\bar b\\, b\\bar b$ event counts at 3000 fb$^{-1}$; if the 95% confidence interval for $\\xi_H^t\\times\\lambda_{hhH}$ widens beyond roughly 0.045 for the $1\\sigma$-trained network, or beyond the reported 10--20% at improved efficiencies, the central projection is contradicted.","supporting_citations":[{"cited_title":"Paszke, S","cited_arxiv_id":null,"evidence_quote":"Provides the neural-network implementation used for training the regression model."}],"review_version":1}