{"id":"59b7f0d1-61bb-4f4c-95d3-b694c6e7d363","arxiv_id":"2506.17022","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A controlled three-generator comparison shows that NLO+PS and LO jet merging differ by up to a factor of two in off-shell gg to H to ZZ to 4l predictions, with MadGraph underproducing sub-leading jets.","lead":"This paper compares three Monte Carlo generators, Powheg, MadGraph, and Sherpa, for off-shell Higgs boson production at the LHC using identical settings. It finds that the NLO-matched Powheg prediction is nearly twice as large as LO jet-merged predictions and has a softer transverse momentum spectrum, which matters for future Higgs width measurements.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Powheg's reweighted massive two-loop gg→VV amplitudes are asserted accurate via Ref. [14] but never checked in this paper; the SBI k-factor conclusion depends on it.","rationale":"The reader identified the reweighting of massive two-loop amplitudes as the weakest assumption, and I agree that this is the most load-bearing point. However, I would refine it: the signal contribution's NLO virtual amplitude is included exactly, so the signal k-factor of ~2.0 is not subject to the reweighting uncertainty. The reweighting only affects the background and SBI contributions. Since experiments ultimately use the SBI combination for width extractions, the accuracy of the reweighted background remains central. The paper cites Ref. [14] as validation but does not display the validation, which is a missing support. A direct comparison against the exact NLO calculation would settle this. The MadGraph sub-leading-jet deficit, while unexplained, does not threaten the main normalization claim because Sherpa independently shows the same k~1.1 factor for the signal and a decrease for background/SBI. The LO cross-section cross-checks in Table 1 provide reasonable confidence that the merged samples are not grossly misconfigured. Overall, the reader's CONDITIONAL verdict is appropriate; I see no new concern that would require changing it.","tokens_in":14377,"tokens_out":9081,"duration_ms":103597,"concrete_test":"Run Powheg/gg4l with the same fiducial setup and replace the reweighted massive two-loop virtual amplitudes with the exact amplitudes of Ref. [14] (or reweight events by the ratio exact/approximate), then recompute the SBI and background inclusive cross-sections and the m4ℓ distribution. If the difference from Table 2 exceeds the Powheg scale uncertainty (~18%), or if the exact-to-approximate ratio deviates from unity by more than ~20% anywhere in 150–500 GeV, the claim that Powheg captures the true NLO correction for the SBI process is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative finding—that NLO+PS yields roughly 1.7–2.0× the LO-merged inclusive rates, and that the difference is due to virtual and unresolved real radiation—adopts, for the background and interference contributions, a reweighting of massless two-loop gg→VV amplitudes to approximate top-quark mass effects (Sec. 2.1). The signal gg→H→VV virtual amplitude is included exactly, but the SBI combination, which drives the experimental width analysis, relies on the reweighted background. The only justification offered is the statement that reweighting 'has been shown to reproduce the exact NLO results very closely [14]'; no quantitative inclusive or differential comparison is provided in the off-shell region 150 < m4ℓ < 500 GeV. If the reweighting error were comparable to the k-factors themselves—for example a 30–50% local error in the interference term—the conclusion that LO jet-merged samples are insufficient for SBI studies would be overstated. The unexplained MadGraph sub-leading-jet deficit (Sec. 3) is a secondary concern: it affects a sub-leading distribution but not the normalization argument, which is supported by both Sherpa and MadGraph.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript compares predictions for off-shell Higgs boson production in gluon fusion, pp -> H* -> ZZ -> 4l, from three event generators: Powheg at NLO matched to PYTHIA8 (the gg4l code), MadGraph5_aMC@NLO at LO with 0+1-jet MLM merging, and Sherpa at LO with 0+1-jet CKKW-L merging. The authors first verify that fixed-order LO cross-sections for signal, background, and signal-background-interference (SBI) agree across all three codes (Table 1). They then show that NLO+PS Powheg predicts inclusive cross-sections roughly 1.7--2.0 times larger than LO, while the LO-merged predictions are close to LO for the signal and slightly smaller for background and SBI (Table 2). Differential distributions for m4l, pT,4l, leading-jet pT, sub-leading-jet pT, and jet multiplicity are compared, with the main qualitative findings being that Powheg produces a softer pT,4l spectrum than the merging codes and that MadGraph underpopulates the high-pT sub-leading-jet distribution. Theoretical uncertainties are estimated by varying hdamp, Qcut, alpsfact, QSF, and CSSKIN. The paper concludes that NLO+PS provides the most robust predictions, especially for the inclusive m4l distribution, and that the bulk of the NLO correction comes from virtual and unresolved real radiation not captured by LO jet merging.","tokens_in":14583,"tokens_out":3748,"duration_ms":42496,"significance":"If the results are reliable, this is a useful and timely comparison for the LHC Higgs-width program, since current experimental analyses normalize LO-merged samples using NLO k-factors. The paper's careful matching of common parameters and the successful LO consistency check in Table 1 are strengths, as is the systematic exploration of generator-specific uncertainty parameters in Section 4. The central claim—that NLO+PS gives substantially larger rates and softer pT spectra than LO jet merging, with implications for experimental modeling—is plausible and supported by the presented tables and figures. However, the significance of the quantitative conclusion depends on the accuracy of the reweighted two-loop amplitudes used in Powheg for the background and interference contributions, which is not validated within this paper. The unexplained MadGraph sub-leading-jet deficit is a secondary concern that also needs attention before the comparison can be considered fully reliable.","major_comments":[{"comment":"The central quantitative claim that Powheg NLO+PS predicts k-factors of about 1.7-2.0 relative to LO, and that this is due to virtual and unresolved real radiation, rests on the accuracy of the reweighted massless two-loop amplitudes used for the background and SBI contributions. The manuscript states that reweighting has been shown to reproduce the exact NLO results very closely, citing Ref. [14], but provides no quantitative check in the off-shell region 150 < m4l < 500 GeV used in this study. Since the exact NLO calculation is now available, the authors should directly compare the reweighted gg4l setup with the exact NLO results, for example by presenting a table of inclusive cross-sections or a ratio plot as a function of m4l for the SBI combination. Without this validation, the magnitude of the claimed NLO correction, and hence the paper's main practical conclusion, cannot be fully assessed.","section":"Sec. 2.1"},{"comment":"The manuscript reports that MadGraph severely underpopulates the high-pT tail of the sub-leading jet distribution, but explicitly states that the authors are unable to explain this observation. Because MadGraph is one of the three generators used to support the central comparison, and because the same deficit appears in the jet-multiplicity distribution of Figure 4, an unresolved technical issue in the MadGraph setup (for example in the MLM merging parameters or the auto_ptj_mjj flag) could affect the interpretation of the MadGraph jet-merged results. The authors should either diagnose the origin of this deficit or, at minimum, quantitatively discuss how it affects the conclusions drawn from the MadGraph cross-sections and distributions. In its current form, the unexplained deficit weakens the claim that the MadGraph results serve as a reliable jet-merged reference.","section":"Sec. 3, Figure 3"},{"comment":"The uncertainty estimates combine all parameter variations in quadrature, but the manuscript does not discuss correlations between the variations or whether the resulting bands are intended to represent a probability interval or an envelope. More importantly, the observation that the parameter-variation uncertainties are smaller than the generator-to-generator differences for pT,4l is presented without further analysis; this suggests that the uncertainty estimate may be inadequate for this exclusive observable. The authors should clarify the interpretation of the uncertainty bands and discuss whether the generator spread should be treated as an additional modeling uncertainty, since the stated aim is to assess the reliability of current MC modeling strategies.","section":"Sec. 4"}],"minor_comments":[{"comment":"The caption states that the table shows 'the large increases from LO to NLO,' but for Sherpa and MadGraph the background and SBI cross-sections decrease relative to the LO values in Table 1. The caption should be reworded to distinguish the Powheg NLO+PS increase from the LO-merged behavior.","section":"Table 2 caption"},{"comment":"The abstract refers to a comparison of 'leading-order and next-to-leading-order plus parton shower' predictions, but the LO generators are actually used with 0+1-jet merging. Clarifying this in the abstract would avoid confusion about what is being compared.","section":"Abstract"},{"comment":"The choice of NNPDF30_NLO_AS_01180 as the PDF set is not motivated; a more modern PDF set or a discussion of PDF uncertainties would strengthen the analysis, especially since the paper aims to provide guidance for future measurements.","section":"Sec. 2.2"},{"comment":"For the pT,4l distribution, the text states that 'the scale uncertainties across the bulk of the pT,4l spectrum are similar for each of the three generators, reflecting the fact that POWHEG only provides LO control on this observable.' This sentence is somewhat confusing because it refers to scale uncertainties that are shown in the lower ratio plots, and the connection to LO control is not immediately obvious; rewording would improve clarity.","section":"Sec. 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful community comparison with a clear experimental motivation, and the authors have taken care to align common settings. The main issue is that the Powheg NLO+PS result, which is the centerpiece of the quantitative conclusion, relies on an approximate treatment of the two-loop amplitudes that is not validated in the paper. Since the exact NLO calculation exists, adding such a validation is both feasible and necessary. The unexplained MadGraph sub-leading-jet deficit is a second concern that should be addressed. I recommend major revision rather than rejection because the overall approach and qualitative findings are defensible, but the current manuscript does not yet establish the reliability of its central quantitative claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time: this is the first controlled three-way comparison of Powheg NLO+PS, MadGraph MLM, and Sherpa CKKW-L for off-shell Higgs production, with signal, background, and SBI all on the same fiducial setup. The LO cross-checks in Table 1 are clean, which gives confidence that the later differences are physics and not parameter artifacts. The central finding—that NLO+PS rates are roughly a factor of two above LO-merged rates and that the pT spectra are softer—is well supported by Tables 1–2 and the differential plots. That is a genuinely useful result for the LHC width-extraction analyses, where the modeling uncertainty will soon dominate over statistics.\n\nThe soft spots are real but not fatal. First, the SBI interpretation rests on the reweighting of massless two-loop amplitudes to approximate top-quark mass effects, validated only by a citation to Ref. [14] and not checked in this paper in the 150–500 GeV region. If that reweighting is locally off by tens of percent in the interference, the claim that the NLO correction is dominated by virtual and unresolved real radiation would be overstated. This does not invalidate the generator comparison, since all three use consistent amplitudes at their respective orders, but it does weaken the quantitative conclusion about SBI. Second, the MadGraph sub-leading-jet deficit is left unexplained; the authors say so themselves, and it is a secondary feature, but it is also the kind of thing that matters for jet-bin cross sections. Third, no run cards, scripts, or data are released, which makes the comparison harder to build on.\n\nThat said, the paper is honest about its limitations, the central qualitative claims hold up, and the comparison is exactly what the LHC Higgs working group needs for systematic studies. The reader's conditional verdict is about right: accept with requests for a reweighting sanity check (or a clear statement that the SBI conclusion is contingent on it) and, ideally, public generator settings.\n\nThis deserves a serious referee and, with those revisions, publication. I'd bring it to our reading group and would cite it in any future off-shell Higgs work.","headline":"Solid three-generator comparison; the two-fold NLO/LO-merged rate gap is real, but the SBI interpretation leans on an unbenchmarked reweighting.","tokens_in":15168,"tokens_out":1467,"would_cite":true,"duration_ms":17662,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes that for off-shell Higgs boson production in gluon fusion, LO jet-merged samples miss most of the NLO QCD correction because that correction is dominated by virtual and unresolved real radiation, not by additional…","keywords":["off-shell Higgs production","Higgs boson width","gluon fusion","NLO+PS matching","jet merging","ZZ four-lepton final state","QCD radiation","Monte Carlo event generators"],"falsifier":"Recompute $gg\\to ZZ$ at NLO with exact massive two-loop amplitudes, the calculation the reweighting approximates, for the same fiducial cuts and overlay the differential $m_{4\\ell}$, $p_{T,4\\ell}$, and jet distributions; if the exact NLO differs from the reweighted Powheg by more than the scale-variation bands, the k-factor comparison collapses. Independently, rerun MadGraph with the MLM parameters varied and with a different shower to see whether the sub-leading-jet deficit near 80 GeV persists.","tokens_in":14143,"feed_emoji":"⚛️","tokens_out":7580,"duration_ms":76271,"temperature":0.7,"pith_summary":"The paper asks whether leading-order jet-merged event samples are good enough for off-shell Higgs boson measurements, and finds that they are not. Comparing three generators with identical fiducial settings, it shows that NLO+PS predictions from Powheg are roughly twice the LO rate for the signal, with k-factors of about 2.0, while LO 0+1-jet merged samples from MadGraph and Sherpa stay within a few percent of LO. The gap implies that most of the NLO correction comes from virtual and unresolved real radiation, which jet merging omits. The paper also finds that MadGraph underpopulates the high-transverse-momentum tail of the sub-leading jet, an effect it cannot fully explain. If correct, off-shell analyses, including indirect Higgs width extractions, should rely on NLO+PS predictions or a reliable NLO normalization rather than LO-merged samples alone.","feed_headline":"Off-shell Higgs rates double at NLO; jet merging misses it","feed_subtitle":"Powheg, MadGraph and Sherpa agree: virtual QCD corrections, not added jets, dominate the NLO shift.","key_machinery":"The instrument is a controlled three-way comparison of event generators with all common parameters fixed: Powheg (NLO matrix elements matched to the Pythia parton shower via the POWHEG formalism), MadGraph (LO matrix elements merged to Pythia with the MLM prescription), and Sherpa (LO matrix elements merged to its own shower with the CKKW-L scheme). To make the NLO benchmark usable, Powheg's two-loop amplitudes for $gg\\to ZZ$ are computed in the massless approximation and reweighted to approximate top-quark mass effects, a procedure validated against the exact NLO result. The argument works by contrasting inclusive observables such as the four-lepton invariant mass, where all generators agree, with radiation-sensitive observables such as the transverse momentum of the four-lepton system and the jet spectra, where they differ, so that the virtual versus real origin of the NLO correction can be separated.","core_discovery":"The paper's central claim is that for $gg\\to H^*\\to ZZ\\to 4\\ell$ in the $150\\,\\text{GeV}<m_{4\\ell}<500\\,\\text{GeV}$ window, the difference between NLO+PS and LO jet-merged predictions is dominated by virtual and unresolved real QCD corrections, not by resolvable jet radiation. This is read off from the fiducial cross sections: Powheg NLO+PS gives signal, background, and signal-background-interference rates about 2.0, 1.7, and 1.7 times its own LO values, whereas the 0+1-jet merged samples in MadGraph and Sherpa stay within a few percent of LO, with signal k-factors of 1.04 and 1.06. Because the fixed-order LO cross sections of all three generators agree within statistical uncertainties, the later differences are attributed to the treatment of higher-order QCD effects. The paper reports that Powheg's transverse-momentum spectrum is softer, Sherpa's is hardest, and MadGraph severely underpopulates sub-leading jets above about 80 GeV and the three-or-more-jet bins, with the sub-leading jet deficit left unresolved. It concludes that Powheg NLO+PS provides the most robust theoretical predictions among the three generators studied.","pith_inferences":["If the NLO correction is genuinely dominated by virtual radiation, then normalizing off-shell samples with on-shell NNLO or NNNLO k-factors may need separate validation in the off-shell window, where the correction is much larger than at 125 GeV.","A direct test of the MadGraph sub-leading-jet deficit would be to vary MLM parameters such as Qcut and alpsfact and to switch the shower; if the deficit persists, it points to the loop-induced MLM merging itself, and MadGraph-based VBF-tagged off-shell analyses should be reweighted or cross-checked.","Extending the same comparison to 0+1+2-jet merging and to NNLO+PS matched predictions would test whether the jet-merged rates still sit near LO while NLO+PS rates remain higher, which would confirm the virtual-dominance interpretation.","Experimental Higgs-width fits could incorporate the Powheg-versus-merged spread directly instead of an additive modeling uncertainty, since the generator spread is smaller than parameter variations for $m_{4\\ell}$ but larger for $p_{T,4\\ell}$."],"forward_implications":["Off-shell Higgs analyses that rely on LO jet-merged samples must either add an NLO reweighting or normalization, since the missing NLO correction is roughly a factor of two and is not reproduced by adding resolved jets.","Radiation-sensitive observables such as $p_{T,4\\ell}$ and jet multiplicity cannot be trusted from LO merged samples in this process; the NLO+PS prediction is softer and carries the virtual corrections that jet merging omits.","The unexplained MadGraph deficit in the sub-leading jet at high $p_T$ and in three-or-more-jet bins will propagate into measurements that use jet-pair tags, such as vector-boson-fusion-style selections in off-shell Higgs studies.","Scale-variation uncertainties on the four-lepton invariant mass are conservative relative to the generator spread, while on $p_{T,4\\ell}$ they may underestimate the spread, so generator comparison should be part of the uncertainty budget.","Sherpa's large uncertainty bands are driven by its shower-parameter variations (QSF and CSSKIN), making parameter tuning a priority if Sherpa is used in off-shell analyses."],"supporting_citations":[{"why":"Validates the massless-two-loop reweighting that gives Powheg its NLO virtual corrections in the off-shell region.","marker":"[14]"},{"why":"Computes the two-loop interference correction used in the Powheg gg4l implementation.","marker":"[16]"},{"why":"Provides the NLO QCD corrections including interference that define the fixed-order benchmark and the experimental normalization scheme.","marker":"[17]"},{"why":"Implements $gg\\to ZZ$ at NLO matched to parton showers in the POWHEG-BOX-RES framework; this is the Powheg prediction compared throughout.","marker":"[24]"},{"why":"Supplies the MadGraph generator and its automated one-loop amplitudes used for the LO 0+1-jet merged samples.","marker":"[32]"},{"why":"Defines the Sherpa setup for QCD corrections to four-lepton plus jets that underlies the Sherpa merged samples.","marker":"[33]"},{"why":"Documents the Sherpa 2.2 event-generation infrastructure and the CKKW-L merging used here.","marker":"[34]"},{"why":"Defines the MLM matrix-element-parton-shower merging prescription used by MadGraph.","marker":"[49]"},{"why":"Introduces the CKKW-L merging scheme used by Sherpa.","marker":"[53]"}],"fun_headline_variants":["NLO doubles off-shell Higgs rates; jet merging lags","Off-shell Higgs NLO: 2x rate, jet merging fails","NLO QCD doubles off-shell Higgs; merging misses it","Jet merging undercounts off-shell Higgs NLO QCD effects","Off-shell Higgs rates double at NLO, but not with jet merging"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparative conclusion assumes that Powheg's reweighted two-loop amplitudes faithfully reproduce the true top-quark mass dependence in the 150-500 GeV off-shell region; if that reweighting fails there, the claimed NLO k-factors and the verdict that jet merging is insufficient would be wrong.","fun_headline_variants_meta":{"raw":{"variants":["NLO doubles off-shell Higgs rates; jet merging lags","Off-shell Higgs NLO: 2x rate, jet merging fails","NLO QCD doubles off-shell Higgs; merging misses it","Jet merging undercounts off-shell Higgs NLO QCD effects","Off-shell Higgs rates double at NLO, but not with jet merging"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000876,"raw_usage":{"total_tokens":3830,"prompt_tokens":1027,"completion_tokens":2803,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":2712}},"tokens_in":643,"tokens_out":2803,"duration_ms":18987,"temperature":1.0,"reasoning_tokens":2712,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:13:41.344748+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute $gg\\to ZZ$ at NLO with exact massive two-loop amplitudes, the calculation the reweighting approximates, for the same fiducial cuts and overlay the differential $m_{4\\ell}$, $p_{T,4\\ell}$, and jet distributions; if the exact NLO differs from the reweighted Powheg by more than the scale-variation bands, the k-factor comparison collapses. Independently, rerun MadGraph with the MLM parameters varied and with a different shower to see whether the sub-leading-jet deficit near 80 GeV persists.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the MLM matrix-element-parton-shower merging prescription used by MadGraph."},{"cited_title":"Catani, F","cited_arxiv_id":null,"evidence_quote":"Introduces the CKKW-L merging scheme used by Sherpa."}],"review_version":2}