{"id":"5a41a00a-6c4f-4a6b-83df-b1292366e19c","arxiv_id":"2608.01888","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"In simulated muon-collider data, combining resolved and boosted Higgs pair events with machine learning yields a 68% confidence interval on the Higgs self-coupling modifier of 0.96 to 1.05 at 10 TeV.","lead":"Using simulated collisions, this paper projects that a 3 TeV muon collider could bound the Higgs self-coupling modifier to 0.80-1.29 at 68% confidence, and a 10 TeV machine to 0.96-1.05, both from Higgs pair production in the four b-quark final state. It shows how combining jet pairing networks, topological data analysis, and two classifiers can extract a precision Higgs potential measurement from an event rate submerged in backgrounds.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Beam-induced background and omitted systematics are the load-bearing condition: the quoted few-percent interval is a statistics-only projection.","rationale":"The reader's weakest-assumption identification is exactly the load-bearing point: the paper's headline sensitivity is a statistics-only Asimov projection with no beam-induced background overlay and no systematic uncertainties. My reading of the simulation chain confirms that this is the least secure link between the quoted intervals and a real measurement. The TDA features and jet-level inputs are particularly exposed to BIB, and the quantitative estimate shows that plausible few-percent systematics are the same size as the claimed 10 TeV precision. I do not see an internal inconsistency or a more fundamental flaw in the statistical machinery: the resolved/boosted combination is disjoint, the Asimov procedure is standard, and the comparison with prior work is honest. The reader's CONDITIONAL verdict is therefore appropriate, and no verdict change is needed.","tokens_in":28449,"tokens_out":6541,"duration_ms":84429,"concrete_test":"Rerun the full analysis with beam-induced background overlaid: use the standard muon-collider BIB samples for 3 TeV and 10 TeV, mix them into each Delphes event at the nominal bunch-crossing rate, and repeat the entire chain—VLC clustering, SPANet pairing, TDA descriptors, D_HH/D_kappa3 training, and the two-dimensional likelihood fit. If the 68% intervals widen by more than ~20% or shift by more than the quoted half-width, the central claim fails as stated. A quicker cross-check: add nuisance parameters for b-tagging efficiency (plus/minus 2%) and the leading background normalizations (Hqq nu-nu/ZZ nu-nu at plus/minus 5%) to the profile likelihood; if the 10 TeV interval expands beyond 0.96<kappa3<1.05, systematics dominate the projection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. 4.6 states: 'the beam-induced background is not overlaid on the simulated events, its mitigation being assumed to be achievable by future detector and reconstruction improvements,' and 'Systematic uncertainties ... are not included.' The central claim—0.96<kappa3<1.05 at 10 TeV—depends on this condition. The 10 TeV interval half-width is about 0.045. From Tables 2 and 4, the signal yield changes by roughly 0.6-0.7 relative per unit kappa3 near kappa3=1, so a 3% systematic shift in signal rate, background normalization, or b-tagging efficiency translates into Delta-kappa3 ~ 0.03-0.05, comparable to the quoted precision. BIB is not a perturbative effect for this analysis: muon-collider BIB deposits many low-pT hits, and the TDA descriptors of Sec. 4.1 are built from all EFlow constituents with ET>0, making H0/H1/S0/S1/LB1 directly sensitive to BIB contamination. The resolved-region lepton veto and the boosted-region b-tag and fat-jet mass inputs are likewise BIB-sensitive. The authors are transparent about this, so the paper is not internally inconsistent; but the few-percent claim is only as strong as the BIB-mitigation assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a Monte Carlo sensitivity projection for measuring the trilinear Higgs self-coupling modifier κ3 in VBF HH production at 3 and 10 TeV muon colliders, using the HH→bbbb final state and combining a resolved four-jet region with a boosted two-fat-jet region. The analysis uses SPANet jet pairing, TDA descriptors, and two classifiers (D_HH for signal/background separation and D_κ3 for κ3 shape information) in a two-dimensional binned maximum likelihood fit. With Asimov data at κ3=1, the quoted combined 68% intervals are 0.80<κ3<1.29 at 3 TeV (1 ab^-1) and 0.96<κ3<1.05 at 10 TeV (10 ab^-1), with expected HH observability of 5.4σ and 36σ, respectively. The paper explicitly labels the result as a statistics-only projection: beam-induced background is not overlaid on the simulated events and systematic uncertainties are not included.","tokens_in":28830,"tokens_out":7247,"duration_ms":87177,"significance":"If the projection survives closer scrutiny, the paper demonstrates a qualitatively new level of precision for the Higgs self-coupling at a future multi-TeV muon collider, substantially surpassing the projected HL-LHC sensitivity. The analysis is carefully executed within its stated scope: independent MC samples are used for classifier training, template construction, and Asimov data; all performance figures are based on held-out test partitions; the training code is publicly available; and SHAP-based interpretability is used to validate the two-network design. The central methodological contribution is the combination of SPANet pairing, TDA descriptors, and a two-classifier likelihood, which is shown to improve the extracted κ3 interval by roughly 20–30% relative to using D_HH alone. The main caveat is that the quoted few-percent intervals are conditional on the assumption that beam-induced background can be fully mitigated and that no systematic uncertainties degrade the measurement; the paper is transparent about this, but the headline should be read with that condition.","major_comments":[{"comment":"The headline combined intervals of Sec. 5.2.3 are statistical-only Asimov projections under the explicit assumption that beam-induced background (BIB) can be fully mitigated and that no systematic uncertainties exist. The text states: 'the beam-induced background is not overlaid on the simulated events, its mitigation being assumed to be achievable...' and 'Systematic uncertainties ... are not included.' This assumption is load-bearing: the TDA descriptors in Eq. (4.2) are computed from all EFlow constituents with E_T>0, and the classifiers use low-level particle clouds, b-tags, fat-jet masses, and lepton vetoes, all of which BIB directly degrades. From Tables 2 and 4, d ln N_signal/dκ3 is about 0.6–0.7 near κ3=1, so a few-percent shift in signal rate, b-tag efficiency, or background normalization changes κ3 by O(0.03–0.05), comparable to the 10 TeV interval half-width. I recommend eithe","section":"Sec. 4.6; Eq. (4.3)–(4.4); abstract"},{"comment":"The combined result is defined by adding the -ΔlnL profiles, but the input profiles are not fit over a common range: at 10 TeV the resolved profile uses only κ3∈[0.8,1.2] (Sec. 5.2.1), while the combined figure is said to be fit to the full grid. The quoted combined interval is therefore not tied to a single, reproducible likelihood fit. Please specify explicitly whether the quoted interval is obtained from (a) a single binned fit to the concatenated (D_HH,D_κ3) planes of both regions, (b) the sum of two polynomial fits with different ranges, or (c) something else. A combined fit over a common κ3 range would remove the ambiguity and let the reader verify the 0.96<κ3<1.05 interval.","section":"Sec. 5.2.3; Fig. 16"},{"comment":"At 3 TeV the boosted channel contributes to the combined interval using only a single-bin D_HH>0.95 selection with about six signal events after pre-selection (Table 3) and only of order three expected signal events after the cut. The authors note that the asymptotic conversion is approximate, but with such low yields the Poisson likelihood is discrete and the fourth-order polynomial fit through the coarse κ3 grid may introduce spurious curvature. This channel is not essential to the 10 TeV claim, but it does affect the quoted 3 TeV combined interval. Please check the result with an exact Poisson (or profile-likelihood) treatment, or state how much the 3 TeV interval changes if the boosted channel is removed.","section":"Sec. 5.2.2; Eq. (5.3)"}],"minor_comments":[{"comment":"The fourth-order polynomial fit to -ΔlnL is used without showing residuals or goodness-of-fit. Please include a validation plot or quantify the fit quality, especially in the narrow 10 TeV trough where the quoted interval is read off.","section":"Sec. 4.6"},{"comment":"The comparison with Refs. [45,46] should state explicitly whether the cited intervals are also statistical-only Asimov projections and whether they use the same detector and BIB assumptions. Otherwise the claim of being 'about 20% narrower' is difficult to interpret.","section":"Sec. 5.2.3; Table 5"},{"comment":"The upper pT cuts in the boosted region (800 GeV and 600 GeV) are motivated by MC generation efficiency. Please clarify whether these cuts are optimized for sensitivity or are fixed by the generation choices, and whether the quoted intervals are stable under reasonable variations of these thresholds.","section":"Sec. 3.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know this paper is a careful, honest sensitivity projection. The headline claim—0.96<kappa3<1.05 at 10 TeV at 68% CL—is real, but only under the condition that the muon collider detector performs as assumed. That condition is stated up front in Sec. 4.6, so there's no deception, but it's load-bearing.\n\nWhat's actually new: previous studies at the same energies and luminosities got 0.81-1.44 at 3 TeV and 0.94-1.08 at 10 TeV. This paper improves those to 0.80-1.29 and 0.96-1.05, and the improvement comes from a genuinely new analysis architecture: SPANet jet-to-Higgs pairing, TDA descriptors, a dedicated kappa3-shape classifier, and a two-dimensional likelihood fit. The TDA variables are a nice addition; the SHAP analysis shows they contribute real information beyond the standard kinematics. The paper is also methodologically careful: independent samples for training, templates, and Asimov data, held-out test partitions, and full transparency about what is and isn't included. The training code is public, which is exactly the right level of reproducibility for a projection.\n\nThe soft spot is the one they disclose: beam-induced background is not overlaid, and systematics are not included. I agree with the stress-test note that this is not a minor caveat for the 10 TeV number. The TDA descriptors are built from all EFlow constituents with ET>0, so any BIB contamination directly changes the input features. A 3% systematic shift in signal rate or b-tagging efficiency moves the extracted kappa3 by roughly the same size as the quoted 68% interval. So the few-percent precision is best read as \"if the detector works as designed.\" There's also a lesser caveat at 3 TeV in the boosted region, where fewer than ten signal events force a single-bin cut and approximate asymptotic statistics; the authors acknowledge that too.\n\nWho gets value: anyone evaluating muon collider physics reach, especially for the Higgs self-coupling, and anyone interested in ML methods for collider analyses. The comparison with prior work is honest, and the paper is a clear step forward in technique. It deserves a serious referee, not desk rejection. The result is conditional, but the condition is explicit and the analysis is clean.\n\nMy recommendation: send it to peer review. I'd cite it.\n\nRegards.","headline":"A careful, transparent sensitivity projection that brings the 10 TeV muon collider's kappa3 reach to a few percent, conditional on beam-induced background mitigation working as assumed.","tokens_in":29343,"tokens_out":4318,"would_cite":true,"duration_ms":44122,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 10 TeV muon collider can determine the trilinear Higgs self-coupling to about ±5% (0.96<κ3<1.05 at 68% CL), using Higgs-pair production via vector boson fusion in the four-b-jet final state.","keywords":["Higgs self-coupling","muon collider","Higgs pair production","vector boson fusion","b-quark final state","machine learning classifiers","topological data analysis","kappa3 measurement"],"falsifier":"Overlay the beam-induced background on the Delphes-simulated signal and background at 10 TeV and rerun the two-dimensional likelihood; if the 68% interval widens beyond roughly ±10%, the central projection fails. Alternatively, a 5% shift in b-tagging efficiency or a 10% change in the dominant HZνν background normalization that moves the 10 TeV interval by more than its quoted width would falsify the claim.","tokens_in":28359,"feed_emoji":"⚛️","tokens_out":5802,"duration_ms":59042,"temperature":0.7,"pith_summary":"The paper sets out to show that a multi-TeV muon collider can measure the trilinear Higgs self-coupling, κ3, far more precisely than the HL-LHC, by using Higgs boson pair production through vector-boson fusion with both Higgses decaying to b-quark pairs. At 10 TeV with 10 ab−1, the combined resolved and boosted analysis projects 0.96<κ3<1.05 at 68% confidence, a few-percent determination; at 3 TeV with 1 ab−1 it projects 0.80<κ3<1.29. The analysis separates signal from backgrounds several orders of magnitude larger using a jet-to-Higgs pairing network, topological data analysis of energy flow, and two dedicated classifiers whose outputs feed a two-dimensional likelihood fit. A sympathetic reader would care because κ3 fixes the curvature of the Higgs potential, and a precision measurement would directly probe the shape of the electroweak vacuum and distinguish the Standard Model from extended Higgs sectors.","feed_headline":"10 TeV muon collider pins Higgs self-coupling to ±5%","feed_subtitle":"Vector-boson-fusion Higgs pairs in the four-b-jet final state could beat the HL-LHC's ~50% reach by a factor of ten.","key_machinery":"Two separately trained classifiers per kinematic region, D_HH (signal vs background) and D_κ3 (shape discriminator between κ3=0.4 and 1.6), whose outputs define a two-dimensional binned likelihood. In the resolved region a symmetry-preserving attention network (SPANet) assigns the four jets to two Higgs candidates before classification, and five topological data analysis (TDA) descriptors from persistent homology of the event's (η,ϕ) energy flow are included as global features. The workhorse is the orthogonality of the two scores: D_κ3 carries κ3 shape information that is essentially independent of the signal/background axis, so removing it widens the 68% interval by roughly 30%.","core_discovery":"On its own terms, the paper's central claim is that κ3 can be extracted from the (D_HH, D_κ3) score plane in VBF HH→bbbb at muon colliders, and that the combination of a resolved four-jet region and a boosted two-fat-jet region yields 68% confidence intervals of 0.80<κ3<1.29 at 3 TeV with 1 ab−1 and 0.96<κ3<1.05 at 10 TeV with 10 ab−1, with the SM signal observable at 5.4σ and 36σ respectively. The κ3 information comes from the shape of the m_HH spectrum near threshold, which the trilinear amplitude enhances or suppresses depending on the sign and magnitude of κ3−1; the D_κ3 classifier captures this shape variation, while D_HH handles signal-background separation. The authors emphasize that","pith_inferences":["If the projection holds, a 10 TeV muon collider would make κ3 one of the best-measured Higgs couplings, comparable to the HVV couplings, sharpening model discrimination in extended Higgs sectors.","The same two-classifier plus TDA architecture could be transferred to other lepton colliders or to the quartic HHVV coupling, where VBF also dominates.","Because the paper treats the 10 TeV result as statistics-limited, detector improvements such as higher b-tagging efficiency or stronger beam-background rejection would translate almost linearly into tighter κ3 bounds; conversely, unmitigated beam-induced background is the main threat to the projection.","A dedicated program to calibrate b-tagging and jet-mass scale at the muon collider would be a prerequisite for realizing the quoted precision, since the classifiers lean heavily on those inputs."],"forward_implications":["At 10 TeV, the combined analysis constrains κ3 to 0.96–1.05 at 68% CL and 0.92–1.10 at 95% CL, a direct few-percent measurement of the Higgs self-coupling.","The SM HH signal is observable above background at 5.4σ at 3 TeV and 36σ at 10 TeV, so the κ3 measurement is built on a detected signal.","The 3 TeV projection of 0.80–1.29 already beats the projected HL-LHC reach of about 0.5–1.6 at 68% CL.","Combining resolved and boosted regions narrows the interval beyond either channel alone, and the boosted channel becomes more important as the collision energy rises.","The extracted κ3 interval is asymmetric, with the lower edge tighter than the upper, reflecting the cross-section minimum near κ3≈1.7."],"fun_headline_variants":["Muon collider Higgs self-coupling precision hits ±5%","10 TeV muon collider shrinks κ3 uncertainty to 5%","Higgs pair production at muon collider yields κ3 to 5%","Muon collider: κ3 measured 10× better than HL-LHC"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's own Sec. 4.6 states that beam-induced background is not overlaid on the simulated events, its mitigation being assumed achievable by future detector and reconstruction improvements; if that mitigation falls short, or if unmodeled systematics shift b-tagging or background normalizations, the few-percent 10 TeV interval and even the 5σ observability could degrade substantially.","fun_headline_variants_meta":{"raw":{"variants":["Muon collider Higgs self-coupling precision hits ±5%","10 TeV muon collider shrinks κ3 uncertainty to 5%","Higgs pair production at muon collider yields κ3 to 5%","Muon collider: κ3 measured 10× better than HL-LHC"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1192,"prompt_tokens":891,"completion_tokens":301,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":226}},"tokens_in":635,"tokens_out":301,"duration_ms":4488,"temperature":1.0,"reasoning_tokens":226,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T18:44:28.532344+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Overlay the beam-induced background on the Delphes-simulated signal and background at 10 TeV and rerun the two-dimensional likelihood; if the 68% interval widens beyond roughly ±10%, the central projection fails. Alternatively, a 5% shift in b-tagging efficiency or a 10% change in the dominant HZνν background normalization that moves the 10 TeV interval by more than its quoted width would falsify the claim.","supporting_citations":[],"review_version":1}