{"id":"2b4cf222-3fc4-4cdd-a75b-0e75af5fd5e8","arxiv_id":"2509.05409","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Simulation-based inference produces stronger combined LHC constraints on SMEFT Wilson coefficients than histogram-based inference for four di-boson processes.","lead":"This paper combines four di-boson production processes at the LHC to compare machine-learning-based, unbinned inference with traditional histogram-based analyses for constraining Standard Model Effective Field Theory parameters. Across every channel and in a combined global analysis, the unbinned approach gives tighter expected limits, up to a factor of two for some operators.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed SBI advantage may be inflated by comparing against deliberately coarse one-dimensional experimental binnings; an optimized histogram baseline could erase much of the factor-of-two gain.","rationale":"I read the paper as a phenomenological demonstration that unbinned SBI improves expected SMEFT limits over binned histograms in a four-channel combination. The methodology is coherent: derivative learning, fractional smearing, and coverage checks provide real supporting evidence. The main weakness is that the comparison baseline is fixed to current experimental binning choices (Sec. 3.2), which the authors themselves acknowledge are improvable. The abstract and conclusions, however, generalize to 'SBI clearly outperforms histogram-based methods', and the quantitative factors in Sec. 4.2 are derived only against the chosen non-optimized bins. Real histograms can be optimized, and neglecting systematics further weakens the robustness of the stated factors. These are external limitations rather than internal contradictions, so a conditional verdict is appropriate. My proposed check targets the stronger of the two caveats: redoing the comparison with an optimized histogram baseline. If the factor persists, the central claim is robust; if not, only the qualitative direction remains supported.","tokens_in":19102,"tokens_out":4514,"duration_ms":54202,"concrete_test":"Recompute the combined profiled limits in Sec. 4.2 using an optimized histogram baseline: for each process, build a 2D histogram from (pT of the vector boson or leading lepton, invariant mass of the VV/Higgs candidate) with adaptive quantile binning (e.g., 10x10 bins), or equivalently bin a per-coefficient gradient-boosted score trained on signal-vs-SM likelihood ratio, and use the same Poisson likelihood as the existing histogram limits. If the SBI/histogram limit ratio drops below ~1.2 for cPhiD, cPhiW, cPhiWB, and cPhiB, the quantitative claim is partly an artifact of the fixed binnings; if it remains near 2, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (Sec. 4.2, Fig. 11: SBI limits a factor ~2 stronger than histogram limits, equivalent to ~4x luminosity) is measured against a deliberately minimal histogram baseline. Sec. 3.2 states: 'The choice of observables and binnings can be improved to provide better sensitivity to our SMEFT operators... we deliberately adopt the standard STXS binning.' The histogram for each process is a single one-dimensional distribution (pT^l1 for WW, mT^WZ for WZ, pT^V for WH/ZH), while real LHC diboson analyses use multiple observables, control regions, and optimized discriminants. If a moderately optimized 2D or BDT-binned histogram retrieves even half the information SBI extracts from the same phase space, the factor-of-two and '~4x luminosity' claims would shrink substantially. The qualitative direction may survive, but the headline 'SBI clearly outperforms histogram-based methods' in global analyses is not established against a competitive histogram baseline. Systematics are also neglected, with the argument that they should affect histograms more, but no demonstration is given. This is a baseline-choice risk, not an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares simulation-based inference (SBI), implemented via the derivative-learning likelihood-ratio method, with traditional histogram-based inference for SMEFT constraints in a combined analysis of WW, WZ, WH, and ZH production at the LHC. Six dimension-six Wilson coefficients are considered. Expected limits at 300 fb^-1 are derived using parton-level reference limits, total-rate limits, one-dimensional histograms based on standard experimental/STXS binnings, and unbinned SBI. The central result is that SBI yields significantly stronger expected limits than the histogram approach, with factor-of-two improvements in several profiled directions (Fig. 11), corresponding to a claimed factor of roughly four in effective luminosity. The paper includes coverage checks of the learned likelihoods, uses fractional smearing to control high-weight events, and includes backgrounds for the Higgs-associated channels.","tokens_in":19371,"tokens_out":7648,"duration_ms":84287,"significance":"If the quantitative claim holds, the paper is a useful step toward global SMEFT interpretations with unbinned inference, and it demonstrates an interesting multi-process combination rather than a single-channel proof of principle. The strength is that the SBI likelihood ratios are derived from theory matrix elements and checked against parton-level ratios and coverage curves, so the comparison is not circular. The method is technically credible. However, the headline comparison is made against deliberately coarse histogram baselines and neglects systematic uncertainties and several backgrounds. These are acknowledged limitations, but they directly affect the magnitude of the claimed SBI advantage, so the significance as stated is somewhat narrower than the abstract/conclusion suggests.","major_comments":[{"comment":"The central quantitative claim—SBI limits a factor ~2 stronger, equivalent to ~4x more luminosity—is measured against a deliberately unoptimized histogram baseline. The authors state that 'the choice of observables and binnings can be improved... we deliberately adopt the standard STXS binning.' The histograms are one-dimensional (pT^l1, mT^WZ, pT^V), while real diboson analyses use multiple observables, control regions, and optimized discriminants. Since the abstract and conclusion claim that 'SBI clearly outperforms histogram-based methods,' the factor-of-two statement is not established for a competitive histogram baseline. The paper should either benchmark against an optimized histogram approach (e.g., 2D binnings or a binned classifier discriminant) or qualify the claim to comparison with existing standard experimental binnings.","section":"Sec. 3.2 (Histogram observables) and Sec. 4.2 (Fig. 11)"},{"comment":"Systematic uncertainties are neglected, with the expectation that 'systematics to degrade the histogram limits more than the SBI limits,' but no demonstration is provided. All expected limits in Figs. 1–11 are pure-statistical. In realistic LHC diboson analyses, normalization and shape systematics are often comparable to statistical power, and they can reduce the advantage of unbinned methods, especially for operators where sensitivity is dominated by rate information. A stylized systematic model, or an explicit restriction of the conclusions to statistical-only limits, is needed to support the quantitative improvement claim.","section":"Sec. 3.2 (Event generation)"},{"comment":"For WW and WZ, the dominant backgrounds (di-top for WW; Z+jets, Zgamma, ttbar for WZ, ~20–35% of the signal region) are omitted, with the justification that the two methods should be affected similarly. This is a plausibility argument rather than a demonstrated equivalence. Backgrounds enter both the Poisson rate term and the signal–background classifier, so omitting them could change the relative performance of SBI and histograms, particularly for rate-sensitive operators. Including these backgrounds or showing numerically that the ranking is unchanged would remove this uncertainty.","section":"Sec. 3.2 (Pre-selection cuts and backgrounds)"}],"minor_comments":[{"comment":"The term 'global LHC analyses' overstates the scope: the analysis combines four diboson processes and a subset of six operators, with no systematic uncertainties. Consider 'multi-process' or 'a global analysis of the electroweak/Higgs diboson sector' for precision.","section":"Title/Abstract"},{"comment":"The caption states that parton-level limits are compared in the combined result, but Sec. 4.1 explicitly says parton-level bounds are not shown for WH (and none are shown for ZH). Clarify how the combined parton-level curve is obtained, or remove it from Fig. 10 to avoid an apparent inconsistency.","section":"Fig. 10 caption and Sec. 4.1"},{"comment":"The phrase 'small central ellipsis' is confusing; 'ellipsis' should be 'ellipse.' Also clarify that the expected SM point is included in the 1-sigma region and what exactly the small ellipse represents.","section":"Fig. 7 caption"},{"comment":"Typos: 'emperical' should be 'empirical' and 'slighlty' should be 'slightly'.","section":"Appendix B"},{"comment":"For WW and WZ the histogram limits are shown without backgrounds, consistent with the SBI setup, but this should be stated in the table/caption so that the comparison is not misinterpreted as including full experimental backgrounds in either method.","section":"Sec. 3.2 (Histogram observables)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and properly builds on the authors' previous derivative-learning work (Ref. [14]). The main issue is not the technical soundness of the SBI method, but the breadth of the headline claim relative to the deliberately minimal histogram baseline and the absence of systematics. With a benchmark against an optimized histogram baseline or a suitably qualified conclusion, the paper would be a useful contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The first thing to know: this is the first systematic application of simulation-based inference to a global combination of diboson processes in SMEFT, and it's a competent one. The method (derivative learning, fractional smearing, classifier-based backgrounds) comes mostly from the authors' own published work, so the novelty is the combination and the quantitative claim: for several Wilson coefficients, SBI gives expected limits about 2x tighter than histograms in a combined WW/WZ/WH/ZH fit, and the gap persists even for cWWW and cPhiq^(3) where the chosen histogram observables are reasonably sensitive (~30% better). That global-combination result is new.\n\nWhat I like: the formalism is correct, the ratios come from theory matrix elements rather than from the limits being quoted, coverage checks are included (App. B), they sanity-check against parton-level limits, and they are unusually honest about simplifications. The ZH piece is the most interesting physics—SBI picks up Z polarization, which a pT histogram cannot, and that is what breaks the cPhiW/cPhiB/cPhiD/cPhiWB degeneracy in the profiled limits. The paper explicitly says it deliberately adopts the standard STXS/experimental binnings, and cites Ref. [16] showing better binnings close the gap for WH. Right kind of self-awareness.\n\nThe soft spots are the baseline choice and systematics. The factor-of-two is measured against a deliberately coarse single 1D observable per process, and the stress-test note is right that a moderately optimized 2D or BDT-binned histogram would likely reclaim a good chunk of the gap. The paper's framing 'in comparison to existing experimental analyses' is defensible, but the abstract and conclusions state it more broadly ('SBI clearly outperforms histogram-based methods'), and that overreach is where a referee should push. One caveat to the stress-test: the ZH gain is not just a binning artifact, since no pT-only histogram captures the polarization information SBI uses. Systematics are asserted to hurt histograms more but not demonstrated, and the omitted WW/WZ backgrounds get the same 'affected similarly' treatment—plausible, but asserted rather than shown. Neither issue is fatal: the qualitative result, that SBI recovers information current binned analyses throw away even in a multi-process combination, is believable and consistent with prior single-process studies. Minor: no code release attached, which makes the coverage checks harder to build on.\n\nWho gets value: people building global SMEFT fits and anyone weighing whether unbinned methods are worth the engineering. It deserves a serious referee: clean setup, falsifiable claim, and the baseline question is exactly what review should adjudicate. I'd send it out.","headline":"Competent first global SBI-vs-histogram SMEFT comparison; the factor-of-two headline is conditional on a deliberately coarse baseline, so the direction likely survives but the magnitude is the open question.","tokens_in":19821,"tokens_out":8753,"would_cite":true,"duration_ms":78731,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Unbinned machine-learning inference beats binned histograms when four LHC diboson channels are combined to constrain six SMEFT Wilson coefficients.","keywords":["simulation-based inference","SMEFT","diboson production","Wilson coefficients","likelihood ratio estimation","unbinned inference","LHC","machine learning"],"falsifier":"Repeat the same four-channel, six-operator combination with histogram observables optimized per operator, such as polarization-sensitive angular variables, and with correlated systematic uncertainties folded in. If the SBI limits are no longer roughly twice as strong as the histogram limits, the paper's central quantitative claim fails. A quicker check is whether the SBI advantage persists when the background classifier is replaced by a different background-modeling treatment.","tokens_in":19026,"feed_emoji":"⚛️","tokens_out":5622,"duration_ms":57218,"temperature":0.7,"pith_summary":"This paper asks whether simulation-based inference (SBI)—using machine learning to estimate the full unbinned likelihood ratio from event data—keeps its known advantage over traditional histogram-based inference when several LHC processes are combined into one global analysis. The authors combine WW, WZ, WH, and ZH production to constrain six dimension-six Wilson coefficients of the Standard Model Effective Field Theory. In every single channel and in the combined fit, the unbinned SBI limits are tighter than histogram limits; for coefficients most entangled with Z-boson polarization, the combined SBI limits are roughly twice as strong, equivalent to about four times more luminosity in the histogram approach. The paper concludes that global SMEFT analyses, not just single high-profile processes, stand to gain from unbinned inference.","feed_headline":"Unbinned ML doubles sensitivity of diboson SMEFT limits","feed_subtitle":"Full-event inference tightens limits on six SMEFT coefficients about 2x over binned analyses.","key_machinery":"The central mechanism is 'derivative learning': because the SMEFT squared matrix element is a quadratic polynomial in the Wilson coefficients, the likelihood ratio can be expanded as a second-order polynomial whose coefficient functions R_i and R_ij depend only on event kinematics. Neural networks are trained to regress these coefficient functions from parton-level ratios, yielding an unbinned estimate of the reconstruction-level likelihood ratio. Backgrounds are incorporated through a signal–background classifier, and the per-process log-likelihood-ratio test statistics are then summed to form a combined test statistic for the global fit.","core_discovery":"Working at 13.6 TeV with a simplified detector simulation, the paper derives expected exclusion limits on six Wilson coefficients from four diboson processes. It finds that SBI consistently outperforms histograms built from standard experimental observables and binnings: in profiled one-dimensional combined limits, the SBI constraints are roughly a factor of two stronger for coefficients to which the chosen histograms are only indirectly sensitive, and about 30% stronger even for coefficients that the histogram observables directly probe. The advantage is traced to SBI's use of the full phase space, in particular its sensitivity to Z-boson polarization, which breaks degeneracies that the tra","pith_inferences":["Because the paper deliberately used standard, unoptimized binnings, a fair next test would compare SBI against histograms built from polarization-sensitive observables or per-operator optimized binnings; such a test would reveal how much of the factor-of-two gap is intrinsic to unbinned inference versus an artifact of binning choice.","If realistic correlated systematic uncertainties were included, the ranking could shift: rate systematics affect both methods equally, but shape systematics and limited Monte Carlo statistics may degrade histograms more, potentially enlarging the SBI advantage—or, conversely, shrinking it if the learned likelihoods inherit simulation inaccuracies.","The same derivative-learning combination strategy should extend to more processes, such as vector-boson fusion or gluon-fusion Higgs production, and to the full SMEFT operator set, where degeneracies are even more severe; the paper's architecture appears ready for that scaling.","A concrete cross-check would be to reproduce the combined SBI limits with an independent unbinned method, such as the matrix-element method or a different likelihood-ratio estimator, to confirm that the reported limits are not sensitive to the specific neural-network training choices."],"forward_implications":["A global SMEFT fit built on SBI would set stronger limits on electroweak and Higgs–gauge operators than the same fit built on rate or binned differential measurements.","Operators that are degenerate in low-dimensional histograms—especially those distinguished by gauge-boson polarization—become individually constrainable once full event information is used.","The expected gain corresponds to roughly a factor of four in integrated luminosity for the most degenerate coefficients, so unbinned inference can substitute for additional data in a histogram-based analysis.","Combining channels does not erase the SBI advantage; the improvement persists after profiling over all other Wilson coefficients.","The gain is not limited to one high-profile process: the method transfers to a multi-process, multi-parameter global analysis.","The reported quantitative gains assume no systematic uncertainties and use unoptimized binnings, so they represent the kinematic-information gain under idealized conditions."],"supporting_citations":[{"why":"Supplies the derivative-learning approach, fractional smearing, and the WZ reference limits that this paper extends with an additional operator and a global combination.","marker":"[14]"},{"why":"Provides the comparison of histogram binning versus ML sensitivity in WH production and motivates questioning whether standard template binnings capture the available kinematic information.","marker":"[16]"},{"why":"Establishes the methodological path from single-process SBI to unbinned multivariate observables in global SMEFT analyses, which the present paper tests quantitatively.","marker":"[19]"},{"why":"Supplies the signal–background classifier likelihood-ratio formalism used to include backgrounds in the unbinned likelihood.","marker":"[21]"},{"why":"Demonstrates that SBI works in a real LHC measurement, providing the experimental precedent that makes the phenomenological comparison meaningful.","marker":"[24]"},{"why":"Establishes the four diboson processes as key constraints in global electroweak and Higgs-sector SMEFT fits, motivating the chosen process combination.","marker":"[27]"},{"why":"Source of the WZ histogram observable and binning adopted for the histogram comparison.","marker":"[40]"},{"why":"Source of the WW histogram observable and binning adopted for the histogram comparison.","marker":"[41]"},{"why":"Provides the simplified template cross-section stage-1.2 binnings used for the WH and ZH histogram comparisons.","marker":"[46]"}],"fun_headline_variants":["Global diboson fits: SBI doubles sensitivity","Full-event inference beats histograms for SMEFT","Machine learning sharpens diboson EFT bounds","Unbinned analysis squeezes diboson limits twofold","SBI outperforms binned in global diboson EFT"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The comparison assumes that the binnings taken from existing experiments and the neglect of systematic uncertainties are representative of real LHC analyses; if better binnings or realistic systematics are used, the reported factor-of-two gain may shrink.","fun_headline_variants_meta":{"raw":{"variants":["Global diboson fits: SBI doubles sensitivity","Full-event inference beats histograms for SMEFT","Machine learning sharpens diboson EFT bounds","Unbinned analysis squeezes diboson limits twofold","SBI outperforms binned in global diboson EFT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000487,"raw_usage":{"total_tokens":2144,"prompt_tokens":558,"completion_tokens":1586,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":302,"completion_tokens_details":{"reasoning_tokens":1521}},"tokens_in":302,"tokens_out":1586,"duration_ms":12304,"temperature":1.0,"reasoning_tokens":1521,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:23:51.870268+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the same four-channel, six-operator combination with histogram observables optimized per operator, such as polarization-sensitive angular variables, and with correlated systematic uncertainties folded in. If the SBI limits are no longer roughly twice as strong as the histogram limits, the paper's central quantitative claim fails. A quicker check is whether the SBI advantage persists when the background classifier is replaced by a different background-modeling treatment.","supporting_citations":[{"cited_title":"Benchmarking simplified template cross sections in $WH$ production","cited_arxiv_id":"1908.06980","evidence_quote":"Provides the comparison of histogram binning versus ML sensitivity in WH production and motivates questioning whether standard template binnings capture the available kinematic information."}],"review_version":1}