{"id":"a16835cb-9f8a-4bc2-a732-f450a6b0cd27","arxiv_id":"2506.14700","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Measurements of WbWb differential cross sections in the dilepton channel at 13 TeV with 140 fb^-1 of ATLAS data provide new constraints on ttbar and tW interference modeling.","lead":"ATLAS reports new measurements of WbWb production in proton-proton collisions at 13 TeV, using 140 fb^-1 of data, with differential cross sections unfolded to particle level for 11 kinematic variables. The results test how well next-to-leading-order Monte Carlo models describe the interference between top-quark pair and single-top production.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; the unfolding-convention risk is covered by the closure-based signal-model systematics.","rationale":"The central claim is a differential m_bl^minimax measurement that favours the DR interference scheme over DS and bb4l. The weakest link, as the reader notes, is the model dependence of the unfolding corrections in the tail. I examined whether this is load-bearing. Section 5.3.1 performs a closure test: alternative signal samples (including DS) are unfolded with the nominal DR corrections and compared with their own particle-level spectra, and the difference is symmetrized into the systematic uncertainties. These uncertainties are part of the covariance used for the chi2 values in Table 2. Therefore, if the data had truly been produced by DS dynamics, the DR-based unfolding would shift the measured points toward DR by an amount comparable to the quoted tW-modelling uncertainty; the chi2 for the DS prediction would then be reduced by that covariance. The observed p<0.01 for DS despite this inflation indicates the incompatibility is larger than the expected correction shift, so the discrimination is not merely a convention artifact. The one thing not shown is the result of unfolding directly with DS (or bb4l) corrections. Such a cross-check would be definitive and is inexpensive, but its absence does not invalidate the claim given the closure systematic. The bb4l comparison is appropriately caveated in the paper because that generator lacks a complete set of systematic uncertainties, so it is not a hidden flaw. I therefore see no reason to change the ACCEPT verdict; I would note the alternative-unfolding cross-check as a strengthened-validation request.","tokens_in":68414,"tokens_out":10257,"duration_ms":112691,"concrete_test":"Recompute the m_bl^minimax particle-level cross section and its covariance using the DS PWG+PY8 sample (and, if MC statistics allow, the bb4l sample) for f_acc, epsilon, and the migration matrix in Eq. (2), and recompute the chi2/p-values for the DR and DS predictions. If the ranking flips or DS becomes compatible at p>0.01, the discrimination is convention-dependent; if DR still wins, the central claim is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The only substantive risk to the central claim is that the m_bl^minimax result is unfolded with the nominal DR PWG+PY8 corrections (Eq. (2)), and the acceptance/efficiency corrections drop sharply in the interference-sensitive tail (Sec. 5.2). If the true tW/tbar composition differed from DR, the central values would be shifted. However, Sec. 5.3.1 directly evaluates this: alternative signal samples, including the DS sample, are unfolded with the nominal DR corrections and compared with their own particle-level spectra, and the resulting difference is symmetrized into the signal-modelling systematic uncertainty. This uncertainty is part of the covariance matrix used for the chi2 values in Table 2. Consequently, the DS p<0.01 means the DS prediction disagrees with the DR-unfolded data by more than the expected DR-DS correction shift, so the discrimination is not merely an artifact of the unfolding convention. The remaining soft spot is that an alternative-unfolded version of the data is not shown, so the reader must trust the closure systematic; this is a standard limitation rather than a flaw. The bb4l comparison is also appropriately caveated by the paper's statement that the generator lacks a complete uncertainty set.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents measurements of particle-level differential cross-sections for WbWb production in the dilepton (eμ) channel using 140 fb^-1 of pp collisions at 13 TeV recorded by ATLAS during Run 2. Events with one electron, one muon, and at least two b-tagged jets define the 2b-inclusive fiducial region; a region with exactly two b-jets defines the 2b-exclusive fiducial region used for the interference-sensitive observable m_bl^minimax. Detector effects are corrected using acceptance, efficiency, and iterative Bayesian unfolding (Eq. (2)) derived from the nominal Powheg+Pythia8 sample with the DR scheme for t-tbar/tW overlap removal, and eleven observables are unfolded. The headline result is the normalised m_bl^minimax cross-section in the 2b-exclusive region: PWG+PY8 with DR gives the best agreement (chi2/NDF = 12.8/9, p = 0.17), while PWG+PY8 with DS (39.1/9, p < 0.01) and the Powheg bb4l full-final-state generator (28.6/9, p < 0.01) are disfavored. In the 2b-inclusive region, none of the NLO+PS predictions describes all observables simultaneously, with the NNLO-reweighted PWG+PY8 sample performing best overall. Integrated fiducial cross-sections of 5.77 pb (2b-exclusive) and 5.97 pb (2b-inclusive) are reported. Uncertainties follow standard ATLAS practice, with closure-based signal-modelling systematics and 10k pseudo-experiments for the statistical component.","tokens_in":68591,"tokens_out":21989,"duration_ms":201997,"significance":"If the results hold, this is the most precise measurement of WbWb production in the dilepton channel at 13 TeV, improving on the earlier 36 fb^-1 measurement by roughly a factor of two in uncertainty and providing the first particle-level m_bl^minimax measurement with full Run-2 statistics. The central discrimination (DR-scheme PWG+PY8 in agreement with p = 0.17, while DS and the bb4l generator are disfavored at p < 0.01) is a genuinely useful constraint on t-tbar/tW interference modelling. I examined the principal correctness risk, namely that the unfolding corrections in Eq. (2) are computed from the nominal DR sample in the tail where acceptance and efficiency drop sharply (Sec. 5.2, Fig. 5); the closure-based signal-modelling systematic of Sec. 5.3.1 directly propagates alternative-generator corrections into the covariance used for Table 2, so the DS exclusion is conservative with respect to this effect, and the residual limitation (no alternative-unfolded data set shown) is a standard, non-blocking one.","major_comments":[],"minor_comments":[{"comment":"The quoted systematic uncertainties for the integrated fiducial cross-sections are inconsistent across the paper: Section 6.4 gives 5.77 +0.27/-0.29 pb and 5.97 +0.27/-0.30 pb, Section 7 gives 5.77 +0.27/-0.30 pb and 5.97 +0.28/-0.31 pb, and neither set rounds exactly to the percentages in Table 6 (+4.6/-5.2% and +4.6/-5.1%); please harmonize these numbers.","section":"Sec. 6.4, Sec. 7, Table 6"},{"comment":"There are typographical errors in Section 4.2: 'satify' should be 'satisfy', 'identificationMedium' is missing a space before the word 'Medium', and the sentence 'The large number of events collected ... allow' should use the singular verb 'allows'.","section":"Sec. 4.2"},{"comment":"The definition of the 2b-exclusive region would benefit from an explicit statement of how the two b-tagging working points combine: the selection is N_b-tags >= 2 at the 70% WP together with N_b-tags = 2 at the 85% WP, which (because 70%-WP jets are a subset of 85%-WP jets) is equivalent to exactly two 70%-WP-tagged jets with no additional 85%-WP-tagged jet; the current phrasing mixes the two working points without spelling out the logical relation.","section":"Sec. 4.2, Sec. 4.3"},{"comment":"The p-value quoted for PWG+PY8 (bb4l) in Table 2 and the statement in Section 6.2 that the bb4l sample 'does not fulfil this expectation' should carry the caveat directly at the point of the comparison, namely that the bb4l generator is still under internal review with an incomplete systematic set and that its own theoretical uncertainties are not included in the chi2; the caveat currently appears only in Sections 3.1.2 and 7, and I do not consider this blocking because the central DR-versus-DS discrimination does not rely on the bb4l comparison.","section":"Sec. 3.1.2, Table 2, Sec. 6.2"},{"comment":"Reference [63] lists the year as '2013' for an ATL-PHYS-PUB-2023-029 note; the year should be 2023.","section":"References"},{"comment":"Section 5.1 states that the first and last bins do not contain under- or overflow events except for N_jets, whereas the captions of Figures 2-4 state that events beyond the horizontal axis are included in the last bin for presentation purposes; please clarify that the overflow handling in the detector-level control distributions differs from the binning used for the cross-section extraction.","section":"Sec. 5.1, figure captions"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a competent and transparent ATLAS measurement that fits JHEP well, and the changes I request are straightforward. The only editorial consideration is that the comparison of the data with the bb4l generator rests on a sample that the collaboration itself states is still under internal review; the paper is transparent about this and the conclusion is appropriately hedged, but editors may wish to confirm that the collaboration is comfortable quoting a public p-value for an under-review generator. There are also minor internal inconsistencies in the quoted integrated cross-section uncertainties between Sections 6.4, 7, and Table 6 that should be fixed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this about the paper: it is the full Run-2 update of the ATLAS WbWb differential measurement, and the main result is that the m_bl^minimax observable in the two-b-jet-exclusive region now has about half the uncertainty of the 36 fb^-1 version and clearly separates the interference schemes. The nominal DR-scheme Powheg+Pythia8 agrees with the data (chi2/NDF = 12.8/9, p = 0.17), while the DS scheme and the bb4l final-state generator are both disfavored at p < 0.01. That discrimination is the paper's real value.\n\nWhat's genuinely new: 140 fb^-1, eleven differential observables, several (pT of bbll, bb4l, mT of bb4l, etc.) measured for the first time for the combined ttbar+tW signal. The integrated fiducial cross-sections are also given in both the inclusive and exclusive regions. The analysis follows the usual ATLAS playbook, and it does it carefully: the unfolding uses the nominal DR sample, but the signal-modeling systematic is evaluated by unfolding alternative generators (DS, hdamp, pT,hard, Herwig7) with the nominal corrections and symmetrizing the difference. That closure test is what makes the DS disfavoring meaningful. The paper is also upfront that the bb4l generator is still under internal review and lacks a complete uncertainty set, so the bb4l comparison is provisional.\n\nSoft spots, in proportion: they are minor. First, the quoted uncertainties on the integrated fiducial cross-sections differ slightly between Section 6.4 and the Conclusions (e.g., +0.27/-0.29 vs +0.27/-0.30 for the exclusive region); that is a typo-level inconsistency that should be fixed. Second, the chi2 values are computed with experimental covariance only, without theoretical uncertainties. That is defensible, but it means the p-values are not Bayesian evidence against DS or bb4l; they are statements about the measured particle-level cross-section under the DR unfolding convention. A reader could easily over-interpret the p < 0.01. Third, the model dependence of the unfolding, while covered, is not eliminated; the data points themselves would shift under a different signal model. That is standard for this kind of measurement, and the closure test is the right way to handle it, but it is worth remembering.\n\nConclusion: this belongs in JHEP and deserves a serious referee. It is not a paradigm shift, but it is exactly the kind of precise differential data that generator tuning and interference modeling need. I would cite it. Bring it to the reading group if you care about top physics or MC modeling.","headline":"Full Run-2 WbWb differential measurement with about half the uncertainty of the 2018 result; the m_bl^minimax variable clearly disfavors the DS and bb4l interference schemes, making this a key new input for generator tuning.","tokens_in":69118,"tokens_out":3243,"would_cite":true,"duration_ms":33888,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Measurements of $WbWb$ production in the dilepton channel favor the diagram-removal scheme for $\\bar{t}t$/$tW$ interference, disfavoring diagram subtraction and the $bb4l$ generator.","keywords":["WbWb production","top-quark pair production","single top tW","ttbar/tW interference","differential cross-section","unfolding","LHC","particle-level fiducial cross-section"],"falsifier":"Recompute the unfolded $m_{bl}^{\\text{minimax}}$ cross-section using acceptance and efficiency corrections from the DS or $bb4l$ generator in place of the nominal DR sample; if the high-mass tail shifts by more than the quoted systematic uncertainties, or if the $\\chi^2$ ordering of DR and DS reverses, then the reported preference is an artifact of the unfolding model. As a second check, evaluate the published covariance matrix against a future $bb4l$ prediction that includes a full set of systematic uncertainties; a $p$-value above 0.01 would overturn the paper's disfavoring of that model.","tokens_in":68197,"feed_emoji":"⚛️","tokens_out":8211,"duration_ms":75855,"temperature":0.7,"pith_summary":"This paper attempts to determine, with data, which scheme for handling the overlap between top-quark pair production and single-top $tW$ production best describes the $WbWb$ final state. Using 140 fb$^{-1}$ of 13 TeV proton--proton collisions, the analysis measures particle-level differential cross-sections in the dilepton channel and compares them with next-to-leading-order predictions. The comparison shows that the diagram-removal (DR) scheme, in which doubly resonant diagrams are removed from the $tW$ sample, agrees with the interference-sensitive $m_{bl}^{\\text{minimax}}$ spectrum better than diagram subtraction or a generator of the full $bb4l$ final state, both of which are excluded at $p<0.01$. The measurement also reveals that no single generator describes every observable in the inclusive region, and it achieves a factor-of-two improvement in precision over the previous ATLAS result. If the ranking is correct, generator developers now have a concrete target: reproduce the data's interference tail and the simultaneous pattern of the eleven measured distributions.","feed_headline":"Data favor one top-interference scheme over rivals","feed_subtitle":"DR-scheme Powheg+Pythia8 matches the measured WbWb interference tail best; DS and bb4l fail at p<0.01.","key_machinery":"The argument is carried by the variable $m_{bl}^{\\text{minimax}} \\equiv \\min\\{ \\max(m_{b_1 l_1}, m_{b_2 l_2}), \\max(m_{b_1 l_2}, m_{b_2 l_1}) \\}$, which stays below the top-quark mass for doubly resonant $\\bar{t}t$ events and therefore becomes sensitive to singly resonant $tW$ interference at high values. The measurement chain corrects the detector-level distribution for acceptance, efficiency, and resolution via iterative Bayesian unfolding, with all corrections derived from the nominal DR-scheme Powheg+Pythia8 signal sample, and then compares the unfolded spectra with model predictions. The DR and DS schemes are the two standard prescriptions for removing the double counting between the $\\bar{t}t$ and $tW$ samples; the difference between them is the dominant modeling uncertainty in the interference-sensitive bins.","core_discovery":"The central result is the unfolded particle-level cross-section for $WbWb$ production as a function of $m_{bl}^{\\text{minimax}}$, measured in a fiducial region requiring exactly two $b$-jets, one electron and one muon of opposite charge. In the interference-sensitive tail above about 180 GeV, where doubly resonant $\\bar{t}t$ contributions are kinematically suppressed, the data agree best with the DR-scheme Powheg+Pythia8 prediction ($\\chi^2/9 = 12.8$, $p = 0.17$). The DS scheme gives $\\chi^2/9 = 39.1$ and the $bb4l$ final-state generator gives $\\chi^2/9 = 28.6$, each with $p<0.01$, implying that neither describes the interference as implemented. Even the favored DR prediction does not fully capture the extreme tail, and the paper notes this measurement should inform future versions of the $bb4l$ generator.","pith_inferences":["The $bb4l$ generator's failure in the tail despite being theoretically the most complete treatment suggests its interference prescription or its matching scheme needs adjustment; the published spectra give a precise target for that work.","Since the unfolding corrections depend on the relative normalisation of $\\bar{t}t$ and $tW$, the data could be reinterpreted to constrain the $tW$/ $\\bar{t}t$ ratio, effectively turning the analysis into a measurement of the single-top fraction in $WbWb$ events.","A complementary measurement requiring three $b$-tagged jets would probe interference effects in a phase-space region where combinatorics, not the doubly-resonant suppression, governs $m_{bl}^{\\text{minimax}}$, providing a cross-check on the DR preference.","The measured failure of all NLO+PS predictions for the transverse momentum of the $bbll$ and $bb4l$ systems points to missing higher-order QCD corrections for additional radiation, which a future full $bb4l$ NNLO+PS calculation could address."],"forward_implications":["The DR-scheme Powheg+Pythia8 prediction is the best available description of the interference-sensitive $m_{bl}^{\\text{minimax}}$ spectrum, while the DS scheme and the $bb4l$ final-state generator are excluded at $p<0.01$.","In the two-$b$-jet-inclusive region, none of the tested NLO-plus-parton-shower generators describes all eleven observables simultaneously; reweighting the $\\bar{t}t$ component to NNLO makes PWG+PY8 (NNLO rew.) the only prediction that describes every variable.","The measured integrated fiducial cross-sections, $5.77^{+0.27}_{-0.29}$ pb (exclusive) and $5.97^{+0.27}_{-0.30}$ pb (inclusive), are consistent with all predictions within uncertainties and benchmark how generators model the fiducial acceptance.","Uncertainties in the interference-sensitive bins are dominated by the choice of $\\bar{t}t$/$tW$ scheme and by parton-shower modeling, with per-bin total uncertainties near 10% or below, roughly half the previous ATLAS measurement's level."],"supporting_citations":[{"why":"Previous ATLAS measurement that first explored $m_{bl}^{\\text{minimax}}$ sensitivity to $\\bar{t}t$/$tW$ interference and defined the variable; the present analysis extends it with the full Run 2 dataset.","marker":"[26]"},{"why":"Introduces the diagram-removal (DR) and diagram-subtraction (DS) schemes used to treat the double counting between $\\bar{t}t$ and $tW$ samples.","marker":"[16]"},{"why":"Provides the diagram-subtraction (DS) alternative scheme used as one of the comparison models in the interference study.","marker":"[17]"},{"why":"Describes the NLO+PS generator of the full $pp \\to l^+\\nu l^- \\bar{\\nu} b \\bar{b}$ process including interference, off-shell and non-resonant effects; the $bb4l$ sample is a central comparison model.","marker":"[25]"},{"why":"Supplies the Powheg Box Res framework used to generate the $bb4l$ final-state sample.","marker":"[65]"},{"why":"Establishes the resonance-aware NLO+PS matching used by the $bb4l$ generator.","marker":"[66]"},{"why":"Supplies the NLO Powheg generation of $tW$ associated production used for the nominal $tW$ sample.","marker":"[59]"}],"fun_headline_variants":["WbWb data pick DR over DS interference model","ATLAS WbWb tail rules out DS and bb4l schemes","WbWb interference favors DR Powheg prediction","Dilepton WbWb cross-sections discriminate top models","Top interference: ATLAS data vindicate DR scheme"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The unfolding corrections that shape the measured cross-section come from the nominal DR-scheme Powheg+Pythia8 signal sample; if that sample mis-models the relative fraction of $\\bar{t}t$ versus $tW$ events, or their kinematics in the high-$m_{bl}^{\\text{minimax}}$ region, the data-model comparison and the preference for DR could be biased.","fun_headline_variants_meta":{"raw":{"variants":["WbWb data pick DR over DS interference model","ATLAS WbWb tail rules out DS and bb4l schemes","WbWb interference favors DR Powheg prediction","Dilepton WbWb cross-sections discriminate top models","Top interference: ATLAS data vindicate DR scheme"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1537,"prompt_tokens":1001,"completion_tokens":536,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":450}},"tokens_in":617,"tokens_out":536,"duration_ms":5010,"temperature":1.0,"reasoning_tokens":450,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:47:56.208058+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the unfolded $m_{bl}^{\\text{minimax}}$ cross-section using acceptance and efficiency corrections from the DS or $bb4l$ generator in place of the nominal DR sample; if the high-mass tail shifts by more than the quoted systematic uncertainties, or if the $\\chi^2$ ordering of DR and DS reverses, then the reported preference is an artifact of the unfolding model. As a second check, evaluate the published covariance matrix against a future $bb4l$ prediction that includes a full set of systematic uncertainties; a $p$-value above 0.01 would overturn the paper's disfavoring of that model.","supporting_citations":[],"review_version":2}