{"id":"1e585d38-332e-4568-b992-29409c8ed7f7","arxiv_id":"2510.23755","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"ATLAS measures the ttH cross-section at 0.63+0.20−0.19 of the Standard Model prediction (3.3σ observed, 5.3σ expected) and excludes top-Higgs CP-mixing angles |α|>62° at 68% confidence.","lead":"The ATLAS experiment measured how often the Higgs boson is produced together with top quarks in 140 fb−1 of proton–proton collisions at 13 TeV, finding the rate to be 63% of the Standard Model expectation with roughly 20% uncertainty. The data are consistent with the Standard Model, and the analysis adds new constraints on the CP properties of the top quark's Higgs coupling.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Inclusive ttH signal strength may be biased low by the 2ℓSS1τhad channel, whose best-fit μ is −0.72; the paper's 'downward fluctuation' interpretation is not quantitatively demonstrated.","rationale":"The paper is a thorough, internally consistent measurement with well-defined control regions and a full systematic accounting. The reader's CONDITIONAL verdict is appropriate. Our concern is a specific channel-dependence that the paper itself raises but does not resolve: it labels the 2ℓSS1τhad deficit a 'downward fluctuation' without a p-value or an alternative-model test. This is not an accusation of error; it is a standard robustness check for a cross-section measurement with an outlier channel. The reader's weakest_assumption (MC shape modelling) is broader; our concern is the specific manifestation in the channel with the largest deviation, so agreement is partial. We recommend keeping the verdict CONDITIONAL (UNCHANGED) pending this check.","tokens_in":72240,"tokens_out":10704,"duration_ms":105665,"concrete_test":"Rerun the global profile-likelihood fit with the 2ℓSS1τhad channel removed from the likelihood (or with its background normalisation factors independently floated in the SR), and compare the resulting σ_ttH/σ_SM and observed significance to the nominal 0.63+0.20/−0.19 and 3.3σ. If the central value shifts by more than about half of its quoted uncertainty (i.e., to ≳0.75) or the observed significance changes by more than 1σ, the nominal result is not robust to this channel; if the shift is smaller, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The combined result σ_ttH/σ_SM = 0.63+0.20/−0.19 (Sec. 8.1) is sensitive to the 2ℓSS1τhad channel, which returns a best-fit signal strength of −0.72+0.56/−0.59 (Fig. 12). The paper attributes this to a downward fluctuation in data at high BDT score (Fig. 11a), but no quantitative test of that hypothesis is given. If the deficit instead reflects an unmodelled shape of the misidentified-τhad or non-prompt-lepton backgrounds in this channel—whose CRs use a relaxed 2–3 jet selection while the SR requires ≥4 jets (Table 5)—the global fit would absorb it as negative signal and bias the inclusive μ downward. This is load-bearing because the observed significance (3.3σ) is far below the expected (5.3σ), and the deficit in this single channel is the largest contributor to that gap; the reported central value and its interpretation as 'compatible with SM' rest on this channel behaving as a statistical fluctuation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a measurement of t-tbar-Higgs (ttH) production in multilepton final states using 140 fb^-1 of pp collisions at 13 TeV recorded by ATLAS. Six orthogonal channels (2ℓSS0τhad, 3ℓ0τhad, 4ℓ, 2ℓSS1τhad, 1ℓ2τhad, 2ℓ2τhad) are combined in a profile-likelihood fit. The inclusive measured cross-section ratio is σ_ttH/σ_SM = 0.63+0.20/−0.19, corresponding to σ_ttH = 321+102/−99 fb against an SM prediction of 507+35/−50 fb, with observed (expected) significance of 3.3σ (5.3σ). A STXS fit measures the cross-section in three bins of Higgs pT (0–120, 120–200, >200 GeV); a simultaneous fit gives μ_tHqb = 7.2+4.6/−4.0 with μ_ttH = 0.59+0.22/−0.20; and a CP-mixing scan excludes |α| > 62° at 68% CL (expected 43°). All results are reported as compatible with the SM within about 2σ.","tokens_in":72517,"tokens_out":22857,"duration_ms":220243,"significance":"If correct, this is the most precise ttH cross-section measurement in the multilepton final state with the full Run-2 dataset, and the STXS/CP results are the first in this channel at this luminosity. The manuscript's strengths are substantial: the six-channel likelihood structure with pre-defined control regions (Tables 4–5) is standard and internally consistent (321/507 = 0.63); the background normalisation factors are cross-checked against dedicated external measurements (λ_ttW = 1.18±0.07 vs Ref. [31]); the systematic decomposition (Table 6) is detailed; and validation regions are shown (Fig. 8). The paper is also careful to quote the SM-theory component of σ_SM separately (footnote 10) rather than folding it into the measurement uncertainty. The observed central value is 1.8σ below the SM (compatibility 7.2%), and the observed significance is well below expectation; the interpretation of the 2ℓSS1τhad-channel deficit is therefore the main point to be resolved.","major_comments":[{"comment":"The 2ℓSS1τhad channel returns μ = −0.72+0.56/−0.59, the largest per-channel deviation and a natural suspect for the shortfall between the observed (3.3σ) and expected (5.3σ) significances. The text attributes this to 'a downward fluctuation in data at high value of the 2ℓSS1τhad BDT', but no quantitative test supports that interpretation. The CRs that anchor the misidentified-τhad and non-prompt normalisations use a relaxed 2–3 jet selection while the SR requires ≥4 jets (Table 5), and the extrapolation uncertainty between them is described but not sized (Sec. 7). Please add: (a) a local goodness-of-fit or spurious-signal p-value for the SR BDT deficit; (b) the combined-fit result with this channel removed, to show how much the inclusive μ and the observed significance move; (c) the numerical value of the jet-multiplicity extrapolation uncertainty. This is load-bearing: the inclusive cen","section":"§8.1, Figs. 11(a) and 12, Table 5"},{"comment":"The CP interpretation highlighted in the abstract (|α|>62° excluded at 68% CL) is derived from a fit whose best-fit point is stated in §8.3 to be 'sensitive to numerical instabilities' because the likelihood is flat around the minimum. This is a self-declared limitation of a headline claim. Please quantify the effect of the instability on the 68%/95% contours (grid density, minimiser variations, alternative profiling) and state whether it affects only the best-fit coordinates or the excluded-region boundary. Without this, the CP exclusion cannot be evaluated.","section":"§8.3, Fig. 17"},{"comment":"The STXS measurement and the CP analysis rely on the GNN/BDT estimate of the Higgs pT, but the text only asserts that 'no significant discrepancies between data and simulation were observed' for the MVA input variables; the supporting comparison is not shown. The migration matrices in Fig. 3 show diagonal fractions as low as ~61% in the 2ℓ2τhad channel, so the unfolding is sensitive to the response model. Please provide the data/MC validation of the pT,H-estimator inputs or a closure test of the migration, and state how a GNN response bias would be covered by the STXS migration systematics. The STXS uncertainties are large, so the conclusions are unlikely to change, but the current support is an unshown assertion.","section":"§5 (STXS/GNN paragraph), Fig. 3, Table 7"}],"minor_comments":[{"comment":"The colour-axis label reads 'Number of events' while the caption states that colours denote the fraction of events; relabel the axis (e.g., 'Fraction of events') to avoid confusion.","section":"Fig. 3"},{"comment":"Please specify how the 7.2% SM-compatibility value is computed (one- or two-sided p-value; which nuisance parameters are profiled) and whether the ±35/−50 fb uncertainty on the SM prediction enters the test or only the quoted ratio.","section":"§8.1"},{"comment":"The per-channel best-fit values and their stat./total decompositions are dense and partly illegible in the current rendering; a companion table listing each channel's μ with statistical and total uncertainties separately would improve usability.","section":"Fig. 12"},{"comment":"The 'Non-prompt normalisation' row contributes only ±0.02 to σ_ttH/σ_SM even though some individual factors (e.g., λ_eHF = 1.20±0.41) carry 20–40% uncertainties. One sentence explaining why the CR constraints are so strong would pre-empt a natural reader question.","section":"§7 / Table 6"}],"recommendation":"major_revision","confidential_remarks":"The low inclusive central value (0.63) relative to the earlier 36 fb^-1 result (μ = 1.6) and the shortfall between observed (3.3σ) and expected (5.3σ) significance will attract scrutiny, especially since the 2ℓSS1τhad channel drives the deficit. The requested leave-one-out fit and local p-value would substantially strengthen the paper against that scrutiny. The CP-fit numerical-instability disclaimer in §8.3 should be resolved before publication because the abstract highlights the CP exclusion. These are quantitative additions rather than a re-analysis; the measurement itself appears sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the most precise ATLAS multilepton ttH measurement to date, and the central result holds up; the flagged 2ℓSS1τhad concern is speculative and does not undermine it.\n\nWhat's genuinely new: full Run-2 140 fb⁻¹ dataset, a GNN-based pT,H estimator that enables the first ATLAS multilepton differential STXS measurement, rebuilt ttW/ttZ modeling with in-situ normalizations, and a simultaneous tHqb extraction. The internal consistency is good — 321/507 = 0.63 — and the systematic breakdown is detailed and honest, including the admitted flat-likelihood instability in the CP fit.\n\nSoft spots, in proportion: the CP exclusion (|α|>62°) is fragile, as the paper itself states; treat it as fragile. The 2ℓSS1τhad channel returns μ = −0.72 ± 0.59, and the paper calls it a downward fluctuation. The stress-test speculation that an unmodelled background shape biases the inclusive μ is not demonstrated: the validation regions and BDT distributions show no glaring shape mismatch, and the six channels are jointly compatible with a single μ at 12.4%. A quantitative p-value for the deficit would have been welcome, but the current evidence does not make the fluctuation interpretation implausible. The observed-vs-expected significance gap (3.3σ vs 5.3σ) is real but compatible with the SM at 7.2%. No public data/code accompanies the preprint; that is standard for ATLAS and limits independent audit, but does not affect the internal logic.\n\nThis is a reference measurement for ttH, top-Yukawa, and Higgs CP studies. I would cite it. It deserves a serious referee.","headline":"Solid, standard-setting ATLAS multilepton ttH measurement; the flagged 2ℓSS1τhad concern is speculative and does not undermine the central result.","tokens_in":73112,"tokens_out":2904,"would_cite":true,"duration_ms":28898,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A simultaneous fit to six multilepton final states measures the ttH cross-section at 0.63 of the Standard Model prediction (3.3σ observed) and excludes |α|>62° for the top-Higgs CP-mixing angle at 68% confidence.","keywords":["Higgs boson","top-quark Yukawa coupling","ttH associated production","multilepton final states","simplified template cross sections","CP-mixing angle","profile likelihood fit","LHC proton-proton collisions"],"falsifier":"Replacing the default Monte Carlo model for the dominant ttW background with an independent generator in the same fit and checking whether the fitted signal strength shifts by more than the quoted systematic uncertainty would directly test the assumption; alternatively, a future measurement at roughly three times the luminosity would show whether the 0.63 ratio moves toward 1 (SM) or stays below.","tokens_in":72034,"feed_emoji":"⚛️","tokens_out":10874,"duration_ms":107834,"temperature":0.7,"pith_summary":"This paper reports a measurement of the associated production of a Higgs boson with a top-quark pair (ttH) in final states with multiple charged leptons, based on 140 inverse femtobarns of 13 TeV proton–proton collisions recorded by a multipurpose detector at the Large Hadron Collider. Six orthogonal channel definitions—two same-sign leptons, three leptons, four leptons, and three channels with hadronically decaying taus—are combined in a single profile-likelihood fit, with normalisations of the dominant backgrounds determined from control regions in the same fit. The central result is a measured ttH cross-section of 321 fb, i.e. 0.63 times the Standard Model prediction of 507 fb, with an observed (expected) significance of 3.3σ (5.3σ); this is compatible with the SM at about the 1.8σ level. The same events are used to measure the rate in three bins of Higgs transverse momentum, to fit the single-top-plus-Higgs process tHqb (signal strength 7.2+4.6/−4.0), and to constrain the CP-mixing angle of the top-Higgs coupling, excluding |α|>62° at 68% confidence. The paper itself notes two caveats: the quoted ratio excludes the theory uncertainty on the SM prediction, and the CP-fit minimum is sensitive to numerical instabilities.","feed_headline":"Measured Higgs-plus-top rate: 63% of Standard Model","feed_subtitle":"Full 13 TeV dataset, six lepton channels, 3.3-sigma excess: still compatible with SM, CP angles above 62 degrees excluded","key_machinery":"The analysis rests on a simultaneous profile-likelihood fit over six mutually exclusive final states, defined by light-lepton and tau multiplicities, with BDT/DNN discriminants separating ttH from ttW, ttZ, diboson, and misidentified-object backgrounds; normalisation factors for the main backgrounds are determined in control regions within the same fit. Because the Higgs cannot be unambiguously reconstructed, a graph neural network (or a boosted decision tree in the two-tau channels) estimates the Higgs transverse momentum, providing the migration matrices for the three measured pT,H bins. The CP interpretation interpolates between simulated samples with different mixing angles α and couplin","core_discovery":"The central claim is a measured ttH signal strength of σ_ttH/σ_SM = 0.63+0.20/−0.19 (σ_ttH = 321+102/−99 fb vs 507+35/−50 fb SM), with observed (expected) significance 3.3σ (5.3σ); the result is compatible with the Standard Model at about 1.8σ. The differential STXS fit gives ratios of 0.78, 0.08, and 1.19 in the Higgs-pT bins 0–120, 120–200, and >200 GeV, the low middle value tracing underfluctuations in the two-same-sign-lepton and one-lepton-plus-two-tau channels. A simultaneous fit returns μ_tHqb = 7.2+4.6/−4.0, and the CP interpretation excludes |α|>62° at 68% confidence level (expected 43°), with the pure CP-odd hypothesis excluded at 1.8σ observed.","pith_inferences":["If the low ratio persists with Run-3 statistics, the most economical interpretation would be a modified top-Higgs coupling, since the same events point to a tHqb rate above the Standard Model.","The 120–200 GeV STXS bin (10+81/−76 fb vs 127 fb SM) is where an altered pT spectrum would first appear; a dedicated unfolded measurement with finer bins would sharpen this.","The analysis's dependence on simulated shapes for ttW and ttZ means the central value could move if those shapes are wrong; an independent re-fit using a different Monte Carlo generator for those backgrounds would quantify this.","Combining this multilepton result with the diphoton and bottom-quark channels cited in the paper would reduce the total uncertainty and test whether the deficit is multilepton-specific."],"forward_implications":["At face value, the top-quark–Higgs coupling is within about 1.8σ of the Standard Model; no new physics is required to describe the inclusive rate.","The differential measurement shows no significant shape deviation in Higgs pT, but the 120–200 GeV bin is compatible with zero, so an enhanced or depleted coupling at intermediate pT is not excluded.","The tHqb signal strength of 7.2+4.6/−4.0 is above the Standard Model expectation and, if real, points to new physics in single-top-plus-Higgs production.","The CP constraint disfavours large CP-odd admixtures: |α|>62° is excluded at 68% confidence, consistent with a mostly CP-even Higgs-top interaction.","Since the observed significance (3.3σ) is below the SM expectation (5.3σ), repeating this measurement with more data is the direct route to deciding whether the low central value is a fluctuation."],"fun_headline_variants":["ttH rate measured at 63% of SM, 3.3σ excess","ATLAS measures ttH signal strength 0.63, CP angle <62°","Higgs+top quarks: rate 0.63 SM, CP exclusion extended","ttH production: 3.3σ evidence, consistent with SM","New ttH cross-section: 0.63 SM, CP mixing limited"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the simulated generators and the machine-learned Higgs-pT estimator reproduce the shapes of signal and backgrounds in the multilepton phase space—the estimator's calibration is checked only via migration matrices computed in the same simulation (Fig. 3)—since the fit floats only overall normalisations; the paper explicitly notes the quoted ratio excludes the theory uncertainty on the SM prediction and that the CP fit's minimum is numerically u","fun_headline_variants_meta":{"raw":{"variants":["ttH rate measured at 63% of SM, 3.3σ excess","ATLAS measures ttH signal strength 0.63, CP angle <62°","Higgs+top quarks: rate 0.63 SM, CP exclusion extended","ttH production: 3.3σ evidence, consistent with SM","New ttH cross-section: 0.63 SM, CP mixing limited"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000655,"raw_usage":{"total_tokens":2953,"prompt_tokens":973,"completion_tokens":1980,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":717,"completion_tokens_details":{"reasoning_tokens":1884}},"tokens_in":717,"tokens_out":1980,"duration_ms":14090,"temperature":1.0,"reasoning_tokens":1884,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T07:50:38.527960+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replacing the default Monte Carlo model for the dominant ttW background with an independent generator in the same fit and checking whether the fitted signal strength shifts by more than the quoted systematic uncertainty would directly test the assumption; alternatively, a future measurement at roughly three times the luminosity would show whether the 0.63 ratio moves toward 1 (SM) or stays below.","supporting_citations":[],"review_version":1}