{"id":"4630af92-8505-466a-a576-c060df6a329b","arxiv_id":"2512.17874","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A dedicated boosted-jet category in HH→bbγγ is projected to narrow the κ2V constraint to [-0.4, 2.6] at 95% CL and improve heavy-resonance limits by 1–2x at 308 fb⁻¹.","lead":"This paper simulates LHC proton collisions to test whether merging the two bottom-quark jets from a boosted Higgs decay into one fat jet improves searches for new physics in the double-Higgs channel. It finds the boosted category narrows the allowed range of the quartic Higgs-gauge coupling modifier κ2V and roughly doubles the mass reach for heavy scalar resonances compared to the standard resolved selection.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Boosted-category κ2V advantage may be an artifact of asymmetric classifier training: the resolved classifier sees only SM HH, while the boosted classifier is trained on κ2V=0 and 0.5 signal, biasing the comparison in Fig. 5.","rationale":"The reader's weakest assumption focuses on fast-simulation parameterizations (double-b tagging, background normalization, systematics). Those are important for the absolute size of the quoted limits, but even if they shift the numbers they are unlikely to reverse the qualitative conclusion that a boosted category recovers high-mass acceptance, because the resolved acceptance genuinely vanishes once the two b-quarks merge. The classifier-training asymmetry is more directly load-bearing: it biases the actual comparison that supports the κ2V part of the central claim. Section V explicitly states that the resolved classifier trains only on SM signal while the boosted classifier includes κ2V=0 and 0.5. This means the resolved category is handicapped in recognizing the non-SM signals that the κ2V scan is trying to constrain. A retraining test is straightforward and would settle whether the reported κ2V improvement is real or an artifact. The resonant analysis is unaffected by this specific flaw, so the overall verdict remains CONDITIONAL rather than REJECT; the condition should be extended to require the matched retraining and a re-evaluation of Fig. 5. Hence the verdict is unchanged, but the rationale for conditionality is reinforced by a more specific concern than the reader's fast-simulation worry.","tokens_in":14967,"tokens_out":14780,"duration_ms":169655,"concrete_test":"Retrain the resolved XGBoost classifier using the same extended signal definition as the boosted classifier (SM ggF+VBF, plus VBF HH with κ2V=0 and 0.5), with identical event weighting, hyperparameter optimization, and bin-boundary selection. Rerun the profile-likelihood scan of κ2V for the resolved-only and combined categories. Compare the new 95% CL interval to the quoted resolved [-1.4,3.7] and combined [-0.4,2.6]. If the resolved-only upper bound falls below ~2.7 or the combined interval narrows substantially, the boosted advantage in Fig. 5 is a training artifact; if the resolved interval remains near 3.7, the reconstruction itself drives the gain.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the asymmetric training of the two XGBoost classifiers in Section V. The resolved classifier is trained exclusively on SM ggF and VBF HH events, whereas the boosted classifier's signal sample is extended to include VBF HH events with κ2V=0 and κ2V=0.5. The κ2V scan in Fig. 5 then compares a resolved category whose classifier was never shown BSM signal with a boosted category whose classifier was partly optimized to recognize those very signals. If the resolved classifier had instead been trained on the same extended signal definition, its score bins would likely have been re-optimized to retain more κ2V=0 and 0.5 events, narrowing or potentially closing the gap in Fig. 5. This directly threatens the non-resonant half of the central claim ('yielding improved constraints on κ2V'). The resonant limits in Fig. 6 are less exposed because they use rectangular cuts rather than the classifier, but the compound headline claim is weakened if the κ2V improvement is an artifact of training.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a fast-simulation study of the HH→bbγγ final state at √s=13.6 TeV with 308 fb⁻¹, introducing a new boosted category in which the H→bb decay is reconstructed as a single large-radius jet. It defines resolved and boosted categories, trains separate XGBoost classifiers for the non-resonant analysis, and uses rectangular cuts for resonant VBF X→HH production. The central claims are that the boosted category improves sensitivity to κ2V, with a combined 95% CL interval [-0.4, 2.6] versus [-1.4, 3.7] for the resolved category alone, and improves expected 95% CL limits on σ(VBF X→HH) by a factor of 1–2 for mX = 1–5 TeV. The paper concludes that boosted reconstruction should be part of future bbγγ Higgs-pair searches.","tokens_in":15167,"tokens_out":6401,"duration_ms":76019,"significance":"If the projections are reliable, this is a useful phenomenological demonstration that a merged-jet category can recover acceptance at high mHH in the golden bbγγ channel, and it provides a clear template for experimental searches. The use of external higher-order cross-section normalizations and the pyhf statistical framework are strengths. The central qualitative conclusion—that boosted reconstruction helps at high mass—is plausible and likely robust. However, the quantitative comparison and the specific non-resonant claim are weakened by an asymmetric classifier training choice, a rate-only κ2V scan, and unvalidated fast-simulation assumptions for the boosted tagger and continuum background. The code and datasets are not public, which limits reproducibility.","major_comments":[{"comment":"The resolved XGBoost classifier is trained only on SM ggF+VBF HH events, while the boosted classifier's signal definition includes VBF HH with κ2V=0 and 0.5 (Sec. V). The κ2V scan in Sec. VI.A uses the score bins of these classifiers. This asymmetric training is a direct confound: the boosted category's bins are optimized to retain the very BSM samples used to claim superior constraints. Please retrain the resolved classifier on the same extended signal set and repeat the Fig. 5 scan, or quote results with a common classifier-agnostic rectangular selection. This is needed to validate the non-resonant half of the central claim.","section":"Section V, Fig. 5"},{"comment":"The κ2V scan explicitly considers only the overall rate ('without modeling shape modification'), yet the abstract's central claim is that the boosted category is sensitive to effects 'that populate the high-mHH tail.' A rate-only fit cannot distinguish a tail-populating signal from a flat rate increase. The generated κ2V=0,0.5,1.5,-1,-2.5 samples should be used to fit shape-sensitive observables (mHH or classifier output that includes mass information) and the profile-likelihood result shown. If shape information is intentionally not used, the tail-specific claim should be softened.","section":"Section VI.A"},{"comment":"The boosted-category projections rest on a double-b tagging efficiency of 75% with 6% light/gluon mis-tag probability, and a γγ+jets continuum normalized to an LO MadGraph cross-section of 48.1 pb with only a 10% normalization nuisance. Since the boosted category is the paper's new contribution, these assumptions directly set the quoted κ2V intervals and resonant limits. Please add a robustness scan varying the double-b efficiency/mis-tag over a realistic range, increasing the background normalization uncertainty, and comparing with ATLAS merged-jet tagger performance. The qualitative conclusion may survive, but the numerical claims must be shown not to depend on optimistic settings.","section":"Section IV.A, Table I, Sec. VI"}],"minor_comments":[{"comment":"Typo: 'relay' should be 'rely' in 'the resonant analysis will relay only on the rectangular cuts'.","section":"Section VI.B"},{"comment":"The approximate relation ΔR_bb ≈ 2mH/pH_T should define pH_T clearly as the transverse momentum of the H→bb system and state the high-boost regime in which the approximation holds.","section":"Section II, Eq. (1)"},{"comment":"The y-axis label 'Fraction of Events' is appropriate for normalized distributions, but the text should clarify that the limits use absolute yields and how the normalization affects the displayed shapes.","section":"Fig. 3"},{"comment":"The description of the random grid search for XGBoost hyperparameters is brief; reporting the chosen hyperparameters or code repository would improve reproducibility.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The paper is a phenomenological projection rather than an experimental analysis, and its contribution is a timely demonstration of the boosted bbγγ category. My main concern is the fairness of the non-resonant comparison (asymmetric classifier training) and the mismatch between the rate-only κ2V scan and the high-mHH-tail claim. Both are fixable within the paper's scope by retraining and by performing a shape-aware fit. I would be willing to consider a revised version that addresses these points; the resonant limits are less affected and likely robust."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful first projection of a boosted category in HH→bbγγ, and the qualitative conclusion — that merged reconstruction buys real acceptance in the high-mHH tail — is sound. The specific numbers, especially the κ2V intervals, should be treated as indicative rather than definitive.\n\nThe paper's genuine contribution is that it quantifies, for the first time in this final state, what a large-R jet category adds. The two categories are orthogonal; Fig. 3 shows the boosted selection populating the high-mHH region for κ2V≠1 and heavy X, and the resonant limits (Fig. 6) improve by a factor of 1–2 across 1–5 TeV. That part is internally consistent and plausible. The framework uses established ingredients: Powheg/MadGraph samples, Delphes with Run-3 object definitions, and XGBoost from the author's earlier work.\n\nThe soft spots are concentrated where the paper makes its strongest claims. First, the κ2V comparison in Fig. 5 is not apples-to-apples. The resolved category is inclusive by design — no VBF-jet requirement — so it is dominated by ggF, which is nearly blind to κ2V. The boosted category, by contrast, requires two forward jets, enriching VBF and hence the κ2V dependence. That category design, more than the asymmetric classifier training, drives the apparent boosted sensitivity. The stress-test worry about the classifier is real but secondary: retraining the resolved classifier on the same signal mixture would likely narrow the gap, but the resolved category's VBF purity is so low that the qualitative conclusion would probably survive. Second, the κ2V scan uses integrated rates only, explicitly dropping the shape information that is the paper's stated motivation; the 'tail sensitivity' is never actually fitted. Third, the resonant analysis fits mγγ only, not mHH, so the limits are counting limits; fine for a projection, but not a full resonance search.\n\nThe metrology underneath the numbers is the weakest link. The double-b tag efficiency of 75% with ~6% mis-tag looks optimistic for merged topologies; the γγ+jets background is normalized without a k-factor; and a single 10% background nuisance cannot cover the modeling uncertainties. The paper does not report background yields in the fit bins, so the Asimov limits cannot be independently checked. None of this breaks the central kinematic argument, but it does mean the absolute intervals and limits should be taken with a grain of salt.\n\nWho is this for? Phenomenologists who want a quick estimate of the boosted-channel potential in bbγγ, and experimentalists planning Run-3/HL-LHC category strategies. It deserves a serious referee, but the referee should request a fair resolved VBF baseline, a shape-based κ2V treatment, the background yields, and a more realistic double-b tagging model. I'd recommend conditional acceptance after major revisions rather than desk rejection.","headline":"Useful first projection of a boosted bbγγ category, but the κ2V numbers are not apples-to-apples and the fast-simulation assumptions need scrutiny.","tokens_in":15760,"tokens_out":6684,"would_cite":true,"duration_ms":72425,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding a boosted, merged-jet category to the HH→bbγγ search improves the LHC's sensitivity to non-standard quartic gauge-Higgs couplings and to heavy resonances decaying to Higgs pairs.","keywords":["di-Higgs production","boosted jet reconstruction","bbγγ final state","quartic gauge-Higgs coupling (κ2V)","vector-boson fusion","heavy scalar resonance","two-Higgs-doublet model","LHC searches"],"falsifier":"Re-run the analysis with a double-b tagging efficiency of 60% and mis-tag probability of 10%, keeping everything else fixed: if the combined 95% interval on κ2V widens to overlap the resolved-only interval, the claimed boost is an artifact of the tagging assumption. Conversely, compare the predicted boosted-category event yield in a high-mHH window with public LHC data from a resolved bbγγ search; a significant deficit would indicate the background normalization is too low.","tokens_in":14738,"feed_emoji":"⚛️","tokens_out":5058,"duration_ms":49248,"temperature":0.7,"pith_summary":"This paper argues that the standard way of reconstructing double-Higgs events in the bbγγ channel, which resolves the two b-quark jets separately, systematically misses the high-energy events where new physics would show up. It proposes a second, boosted category in which the collimated b-quark pair is caught in a single large-radius jet, and shows, with fast simulation at 13.6 TeV and 308 fb⁻¹, that adding this category sharpens the bound on the quartic gauge-Higgs coupling modifier from [-1.4, 3.7] to [-0.4, 2.6] at 95% CL and improves limits on a heavy scalar decaying to Higgs pairs by a factor of one to two across masses 1–5 TeV. If right, it means a merged-jet category is a cheap and effective extension of existing searches, recovering acceptance that resolved selections lose.","feed_headline":"Boosted jets tighten the LHC's probe of the quartic Higgs coupling","feed_subtitle":"Merged b-jets recover the high-mass tail of HH→bbγγ, giving stronger κ2V limits and better reach for heavy resonances.","key_machinery":"The load-bearing object is the large-radius jet that captures both b-quarks when the Higgs is boosted: the characteristic b-quark separation shrinks as ≈ 2mH/pT^H, so once it falls below the jet radius the two small-radius jets merge into one jet with a two-prong substructure. The analysis defines two mutually exclusive categories—resolved (two small-R b-tagged jets) and boosted (one large-R double-b-tagged jet)—and uses a gradient-boosted classifier in the non-resonant search to define signal-enriched bins; the resonant search uses rectangular cuts because the high-mass signal is sparse but distinctive. The comparison between the two categories is what carries the argument: the same simulat","core_discovery":"The central claim is that the boosted topology—where H→bb is reconstructed as one large-radius jet with a two-prong substructure—carries most of the sensitivity to deviations of κ2V from its Standard-Model value and to resonant X→HH production at high invariant mass. Using two orthogonal categories, a resolved one and this new boosted one, the paper shows that the boosted category alone constrains κ2V to [-0.6, 2.7] at 95% CL, whereas the resolved category alone gives [-1.4, 3.7]; the combination gives [-0.4, 2.6]. For a scalar resonance produced in vector-boson fusion, the boosted category sets 95% CL limits on σ(VBF X→HH) from about 1 fb at 1 TeV to 100 fb at 5 TeV, one to two times strong","pith_inferences":["A natural extension, not explored in the paper, would be to apply the same two-category logic to the HL-LHC full dataset, where the boosted category's statistical deficit at lower mass is reduced; the resonant gain may grow with luminosity.","The 75% double-b tagging efficiency with 6% mis-tag probability assumed here is generous; if real performance is closer to 60%/10%, the boosted limits in this paper degrade, though the qualitative conclusion that boosted helps at high mass likely survives.","The comparison with the resolved category is not fully apples-to-apples: the resolved analysis here is inclusive and does not require VBF forward jets, whereas the boosted category requires two VBF jets; a dedicated resolved-plus-VBF category might recover some of the high-mass sensitivity the paper attributes to boosting.","The paper restricts itself to rate-only κ2V effects and ignores shape modifications; including kinematic shapes in the likelihood, or adding interference effects, could change the derived interval and is a testable extension."],"forward_implications":["Existing bbγγ analyses that run only a resolved selection are leaving a detectable high-mass region unused; adding a boosted category costs little and tightens κ2V constraints by roughly a factor of two in interval width.","The boosted category gives VBF-resonant searches sensitivity to mX up to 5 TeV, where the resolved selection has near-zero acceptance once the b-quarks merge.","The combined category keeps the strong SM sensitivity of the resolved selection (µHH limit of 2.3) while gaining the BSM tail sensitivity, so future Run-3 and HL-LHC analyses should quote both categories.","The method transfers directly to other di-Higgs final states that already use boosted jets, and can be combined with more advanced taggers.","The quoted 95% interval still contains the Standard-Model value 1, but excludes zero, consistent with earlier LHC results that rule out a vanishing quartic gauge-Higgs coupling."],"fun_headline_variants":["Boosted HH→bbγγ tightens κ2V constraints","Merged b-jets recover high-mass HH tail for new physics","Boosted search improves LHC limits on heavy scalars","Boosted HH→bbγγ yields up to 2x stronger resonance limits"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole boost in sensitivity rests on the fast-simulation parameters for the merged jet: 75% double-b tagging efficiency with about 6% mis-tag probability, and a single 10% background normalization uncertainty, with no cross-section correction factor for the dominant γγ+jets background; if any of these is optimistic, the quoted intervals and limits shift.","fun_headline_variants_meta":{"raw":{"variants":["Boosted HH→bbγγ tightens κ2V constraints","Merged b-jets recover high-mass HH tail for new physics","Boosted search improves LHC limits on heavy scalars","Boosted HH→bbγγ yields up to 2x stronger resonance limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001226,"raw_usage":{"total_tokens":4867,"prompt_tokens":726,"completion_tokens":4141,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":4066}},"tokens_in":470,"tokens_out":4141,"duration_ms":29552,"temperature":1.0,"reasoning_tokens":4066,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T15:09:43.905686+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis with a double-b tagging efficiency of 60% and mis-tag probability of 10%, keeping everything else fixed: if the combined 95% interval on κ2V widens to overlap the resolved-only interval, the claimed boost is an artifact of the tagging assumption. Conversely, compare the predicted boosted-category event yield in a high-mHH window with public LHC data from a resolved bbγγ search; a significant deficit would indicate the background normalization is too low.","supporting_citations":[],"review_version":1}