{"id":"f85b08e7-37d0-4950-bb32-3c2da8858972","arxiv_id":"1908.06463","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"CMS measures the four-top-quark production cross section to be 12.6 +5.8 -5.2 fb with an observed significance of 2.6 standard deviations, consistent with the standard model prediction of 12.0 fb.","lead":"The CMS experiment measured the production rate of four top quarks at the LHC using 137 inverse femtobarns of proton-proton collision data, finding a value consistent with the standard model. The result tightens constraints on the top quark's coupling to the Higgs boson and on several new physics models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ttbb/ttjj correction to ttW/ttZ/ttH is the largest systematic and assumes process-independence; a direct test is missing.","rationale":"The reader's weakest-assumption identified the nonprompt-lepton background, but Table 2 shows its impact on sigma(tttt) is only 3%, and the paper documents closure tests and a 30-60% uncertainty. The larger systematic is the ttbb/ttjj correction applied to ttW/ttZ/ttH; it has the largest single impact (11%) and rests on a process-independence assumption that is not directly validated in the paper. The text explicitly states the correction is based on inclusive tt measurements and is applied to associated-production processes; this is an extrapolation. Because the BDT analysis has no dedicated ttW control region, the correction is only indirectly constrained by the fit. Still, the measurement is statistically dominated (observed significance 2.6 vs expected 2.7), and an 11% shift would not change the central conclusion of consistency with the SM. Thus the verdict remains ACCEPT, but a direct check of the ttbb/ttjj ratio in ttW/ttZ/ttH would materially increase confidence in the quoted central value and uncertainty.","tokens_in":40833,"tokens_out":13618,"duration_ms":144828,"concrete_test":"Using the same MG5_aMC@NLO setup as the analysis, generate ttW+2 jets with the analysis phase-space cuts and compute the ratio of events with an additional b-quark pair to events with only light additional jets; compare with 1.7 from inclusive tt. If the ttW ratio differs by more than 0.6, repeat the BDT fit with a process-specific scale factor and check whether the measured cross section shifts by more than the quoted 11% impact. Alternatively, add CRW with b-tag multiplicity to the BDT likelihood and float the scale factor; a pull beyond 1 standard deviation would confirm under-coverage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing uncertainty is the flavor correction applied to the dominant prompt backgrounds (Section 5): ttW, ttZ, and ttH events with an additional b-quark pair are scaled by the ratio sigma(ttbb)/sigma(ttjj) = 1.7 +/- 0.6 measured in inclusive tt production (Ref. [60]). The 35% uncertainty on this ratio is the largest single systematic in Table 2, with an 11% impact on sigma(tttt). The assumption that this ratio is process-independent is not self-evident: in ttW production, an additional b jet can arise from W radiation off a b quark, a topology absent from inclusive tt, so the b-quark/light-jet ratio may differ. The primary BDT analysis has no ttW control region (only CRZ), so this shape correction is constrained only by the signal regions themselves; if the correction is wrong, the high-Nb SRs, which are also signal-rich, would be biased, shifting the measured cross section and significance. The nonprompt-lepton method, by contrast, has a small 3% impact and is supported by closure tests; it is not the weakest link.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a search for four-top-quark production in final states with same-sign dileptons or at least three leptons, using 137 fb^-1 of proton-proton collisions at sqrt(s)=13 TeV recorded by CMS. Two analysis strategies are developed: a cut-based categorization with 14 signal regions and dedicated ttW and ttZ control regions, and a BDT-based analysis with 17 signal regions and a ttZ control region. Signal and prompt backgrounds are modeled with Monte Carlo simulation, with corrections for ISR/FSR jet multiplicity and for the flavor of additional jets based on the measured sigma(ttbb)/sigma(ttjj) ratio. Nonprompt leptons are estimated with the tight-to-loose ratio method and charge-misidentified leptons from simulation with data-derived correction factors. A profile maximum-likelihood fit yields sigma(pp to tttt) = 12.6 +5.8 -5.2 fb in the BDT analysis with an observed (expected) significance of 2.6 (2.7) standard deviations, consistent with the standard model prediction of 12.0 +2.2 -2.5 fb. The cut-based analysis gives 9.4 +6.2 -5.6 fb and is found to be statistically compatible. The results are interpreted as constraints on the top-quark Yukawa coupling (|yt/yt^SM| < 1.7), the H-hat oblique parameter (H-hat < 0.12), and on heavy scalar and pseudoscalar production in Type-II 2HDM and simplified dark matter models.","tokens_in":40993,"tokens_out":13941,"duration_ms":134522,"significance":"If correct, this is the most precise measurement of the four-top-quark production cross section at 13 TeV to date and represents the first CMS result with the full Run 2 dataset in this final state. The analysis is thorough and well documented: it uses a profile-likelihood fit with dedicated control regions for ttZ in both analyses and for ttW in the cut-based analysis, data-driven nonprompt-lepton estimates with simulation closure tests, and a detailed uncertainty treatment summarized in Table 2. The dual cut-based and BDT strategies provide an important internal cross-check, and the paper extends the physics reach with several BSM interpretations. The main caveat is the reliance on the inclusive ttbb/ttjj ratio to correct the flavor of additional jets in ttW, ttZ, and ttH backgrounds; this is the largest single systematic and is not directly validated in the BDT analysis, which is the primary result.","major_comments":[{"comment":"The largest single systematic in the measurement is the correction of the ttW, ttZ, and ttH backgrounds based on the inclusive ratio sigma(ttbb)/sigma(ttjj) = 1.7 +/- 0.6 from Ref. [60], which has an 11% impact on sigma(tttt). The BDT analysis, which is the primary result, has no dedicated ttW control region (only CRZ), so this shape correction is constrained only by the signal regions themselves. The assumption that the inclusive ttbb/ttjj ratio applies to ttW, ttZ, and ttH is not self-evident, because the additional b-quark production mechanisms differ (e.g., W radiation from a b quark in ttW). I request a direct validation: for example, include the CRW in the BDT fit and check the change in the measured cross section, or compare the predicted and observed Nb distribution in the CRW under the BDT selection. If such a test is not feasible, please provide a quantitative argument that the 35% uncertainty on the ratio covers the expected process-dependent variation.","section":"Section 5, Section 6, Table 2"}],"minor_comments":[{"comment":"The sentence about the fitted nuisance parameters states that the ttW and ttZ normalizations are both scaled by 1.3 +/- 0.2 by the fit, but it is not clear whether this refers to the BDT analysis, the cut-based analysis, or both; please specify, and for the BDT analysis, explain how the ttW normalization is constrained without a dedicated control region.","section":"Section 7"},{"comment":"The BDT input list includes the pT of the sixth, seventh, and eighth jets; please state what value is used for these variables when fewer jets are present in an event.","section":"Section 4"},{"comment":"The phrase 'The tttt cross section and the 68% CL interval is measured to be' is grammatically awkward; please rephrase, for example as 'The tttt cross section is measured to be ... with a 68% CL interval of ...'.","section":"Section 7"},{"comment":"The sentence 'These limits exclude couplings larger than 1.2 for m_phi in the 25-340 GeV range and larger than 0.1 (0.9) for m_Z' = 25 (300) GeV' is hard to parse; please rephrase for clarity.","section":"Section 8"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong experimental result and the central measurement is well supported. The only substantive concern is the missing validation of the largest systematic (the ttbb/ttjj-based flavor correction) in the BDT analysis, which I believe can be addressed with an additional cross-check involving the ttW control region. The authors should also clarify the fit details for the ttW normalization in the BDT analysis. I recommend major revision to address these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the full Run 2 CMS four-top search, and it is a real measurement. The new pieces are the 137 fb^-1 dataset, the BDT-based signal categorization, and the BSM interpretations (top Yukawa, H-hat, 2HDM, simplified dark matter) that go beyond the previous 36 fb^-1 result. The measured cross section, 12.6 +5.8 -5.2 fb with 2.6 sigma observed (2.7 expected), agrees with the SM prediction, and the analysis looks methodologically sound. This is the kind of paper that should be refereed, not desk-rejected.\n\nThe analysis does several things well. The profile likelihood fit is standard but carefully built, with a ttZ control region in both analyses and an additional ttW control region in the cut-based one. The nonprompt-lepton estimate uses the tight-to-loose ratio method with closure tests and a 30-60% uncertainty; the stress-test note is right that this is not the weakest link. The charge-misidentification background gets a data-derived correction factor. The systematic table is thorough, and the post-fit yields in the BDT bins look reasonable.\n\nThe soft spot, and it is a real one, is the flavor correction applied to the dominant prompt backgrounds. The ttW, ttZ, and ttH samples are reweighted by sigma(ttbb)/sigma(ttjj) = 1.7 +/- 0.6, measured in inclusive tt production. That 35% uncertainty is the largest single systematic, with an 11% impact on sigma(tttt). The assumption that this ratio is process-independent is not self-evident: in ttW, an extra b jet can come from W radiation off a b quark, a topology that does not exist in inclusive tt pair events. The BDT analysis has no ttW control region (only CRZ), so this shape correction is constrained mainly by the signal regions themselves. If the correction is off, the high-Nb bins, which are also signal-rich, could shift both the cross section and significance. I would have liked to see a direct test of the process-dependence, or at least an explicit discussion of why the inclusive ratio should transfer. That said, the authors assign a sizable uncertainty to it, and the effect is included in the reported error bars. It is a legitimate caveat, not a fatal flaw or a circular argument.\n\nOn the citation and framing side, the paper is honest. It builds on Ref. [27] and the ATLAS searches, and the BSM limits are compared to external predictions. Nothing looks overclaimed; the significance is modest and stated as such.\n\nWho should read this: anyone working on top-quark physics, top-Higgs coupling constraints, or BSM searches in multilepton final states. It is also a useful reference for the ttbb/ttjj transfer uncertainty in rare top processes. I would bring it to a reading group and would cite it. Send it to peer review; the flavor-correction point should be raised in the report, but it is a discussion point, not a rejection.","headline":"A solid, incremental four-top search on the full Run 2 dataset; the largest systematic is the ttbb/ttjj flavor correction, not the nonprompt-lepton method, and the paper deserves serious peer review.","tokens_in":41580,"tokens_out":1730,"would_cite":true,"duration_ms":22518,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper measures the four-top-quark production cross section in 13 TeV proton-proton collisions and finds it consistent with the standard model, with an observed significance of 2.6 standard deviations.","keywords":["four top quark production","same-sign dilepton final state","multilepton final state","top quark Yukawa coupling","boosted decision tree","two-Higgs-doublet model","effective field theory","dark matter simplified models"],"falsifier":"Compare the tight-to-loose prediction with data in a same-sign dilepton sideband with exactly two jets and at most one b-tagged jet, where the four-top signal is negligible; a disagreement beyond the quoted 30-60% uncertainty in the nonprompt estimate would shift the measured cross section directly.","tokens_in":40562,"feed_emoji":"⚛️","tokens_out":10188,"duration_ms":98336,"temperature":0.7,"pith_summary":"The paper seeks to establish that the standard-model process producing four top quarks in proton-proton collisions is present in LHC data and to measure its rate. Using $137\\,\\mathrm{fb}^{-1}$ of $\\sqrt{s}=13$ TeV collisions and final states with two same-sign leptons or at least three leptons, the authors report a cross section of $12.6^{+5.8}_{-5.2}$ fb from their boosted decision tree analysis, with an observed $2.6$ standard deviation excess over background alone and an expected excess of $2.7$. This agrees with the standard model prediction of $12.0^{+2.2}_{-2.5}$ fb, so the measurement would confirm the predicted rate and sharpen the search for new physics that could enhance four-top production. A correct measurement matters because four-top production is a rare process with relatively clean final states, making it a sensitive probe of the top quark's Higgs coupling and of particles that couple strongly to top quarks.","feed_headline":"Four-top rate measured at 12.6 fb, matching the standard model","feed_subtitle":"A 2.6-sigma excess in same-sign and multilepton events puts the rare process near its predicted 12 fb rate.","key_machinery":"The analysis is carried by a selection of same-sign dilepton or multilepton events with high jet and b-jet multiplicity, divided into signal regions and control regions and fitted with a profile likelihood. For the primary result, a boosted decision tree (BDT), a multivariate classifier trained on 19 kinematic variables, separates the $\\mathrm{t\\bar{t}t\\bar{t}}$ signal from backgrounds, and its output is discretized into 17 signal regions plus a $\\mathrm{t\\bar{t}Z}$ control region. The dominant fake-lepton background is estimated with the tight-to-loose ratio method, which measures the probability for a loosely identified nonprompt lepton to also pass the tight selection; the paper redefines the lepton $p_{\\mathrm{T}}$ to include isolation-cone energy so that a single efficiency can be applied across different parent-parton momenta. The profile likelihood fit then extracts the signal cross section while constraining the $\\mathrm{t\\bar{t}W}$ and $\\mathrm{t\\bar{t}Z}$ normalizations.","core_discovery":"The central discovery claim is a measured $\\sigma(pp\\to \\mathrm{t\\bar{t}t\\bar{t}}) = 12.6^{+5.8}_{-5.2}$ fb in the boosted decision tree analysis, with an observed (expected) significance of $2.6$ ($2.7$) standard deviations relative to the background-only hypothesis and a 95% CL upper limit of 22.5 fb. The cut-based analysis gives a compatible value of $9.4^{+6.2}_{-5.6}$ fb with an observed significance of $1.7$ standard deviations. The paper treats the BDT result as primary because it provides higher expected precision, and uses it to derive a 95% CL limit $|y_{\\mathrm{t}}/y_{\\mathrm{t}}^{\\mathrm{SM}}|<1.7$, an effective-field-theory bound $\\hat{H}<0.12$, and mass exclusions up to 470 (550) GeV for a heavy scalar (pseudoscalar) in Type-II two-Higgs-doublet and simplified dark matter models.","pith_inferences":["If the standard-model rate is correct, the expected significance should scale roughly with the square root of integrated luminosity, so a combined Run 2 plus Run 3 data set should push this search past the 5-sigma discovery threshold.","The paper notes that dedicated top-tagging algorithms did not improve sensitivity because only a few events reconstruct all top-quark decay products; with more data those algorithms could become useful for the boosted heavy-scalar interpretation, where the current analysis relies on BDT binning.","The tight-to-loose method's single-efficiency assumption could be stress-tested by measuring the efficiency separately for muon- and electron-like nonprompt leptons in control samples; a failure there would directly change the measured cross section.","The limits on light scalar and vector particles coupling to top quarks suggest that four-top production is a uniquely sensitive probe of top-philic new physics below the on-shell top-pair threshold, a region that other LHC searches constrain only weakly."],"forward_implications":["If the central measurement is right, the standard model's next-to-leading-order prediction for four-top production is confirmed at the level of precision reached by this data set.","The observed 2.6-sigma excess over background alone would grow as more data are analyzed, making a definitive observation of this rare process plausible with the full LHC data set.","The limit $|y_{\\mathrm{t}}/y_{\\mathrm{t}}^{\\mathrm{SM}}|<1.7$ constrains the top quark's Yukawa coupling without assumptions about the Higgs boson width, complementary to constraints from Higgs-rate measurements.","The exclusions of heavy scalar and pseudoscalar bosons up to 470 and 550 GeV translate into limits on two-Higgs-doublet parameter space and on simplified dark matter mediators that couple to top quarks.","The BDT analysis's better expected precision compared with the cut-based approach establishes a reusable pattern for future searches in this final state."],"supporting_citations":[{"why":"Supplies the standard model cross section of 12.0 fb used as the reference prediction for expected significance and theory comparison.","marker":"[1]"},{"why":"Defines the previous iteration of this search, whose selection, lepton identification, and control-region strategy this analysis inherits and improves.","marker":"[27]"},{"why":"Provides the most sensitive previous measurement by another experiment, which this result is compared against.","marker":"[25]"},{"why":"Introduces the tight-to-loose ratio method used to estimate the nonprompt lepton background.","marker":"[59]"},{"why":"Describes the b-tagging algorithm whose efficiency and misidentification rates drive the signal-region categorization.","marker":"[57]"},{"why":"Gives the leading-order prediction for the dependence of the four-top cross section on the top-Higgs Yukawa coupling, used for the coupling limit.","marker":"[2]"},{"why":"Defines the H-hat oblique parameter and provides the reference constraint used for the effective-field-theory interpretation.","marker":"[11]"},{"why":"Supplies the measured ratio of ttbb to ttjj cross sections used to correct the flavor composition of additional jets.","marker":"[60]"}],"fun_headline_variants":["Four-top production measured at 12.6 fb, close to SM","CMS sees 2.6-sigma hint of four-top quark events","Four-top cross section constrains top-Higgs coupling","Rare four-top events: CMS measures 12.6 fb, limits BSM"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The nonprompt-lepton background estimate rests on the tight-to-loose ratio method, which assumes that a single efficiency for loose leptons to pass the tight selection, parameterized by flavor, $p_{\\mathrm{T}}$, and $|\\eta|$, applies to all sources of nonprompt leptons after a specific momentum redefinition, and that the prompt-lepton contamination subtracted from the control sample is correctly described by simulation.","fun_headline_variants_meta":{"raw":{"variants":["Four-top production measured at 12.6 fb, close to SM","CMS sees 2.6-sigma hint of four-top quark events","Four-top cross section constrains top-Higgs coupling","Rare four-top events: CMS measures 12.6 fb, limits BSM"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000346,"raw_usage":{"total_tokens":1993,"prompt_tokens":1136,"completion_tokens":857,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":752,"completion_tokens_details":{"reasoning_tokens":778}},"tokens_in":752,"tokens_out":857,"duration_ms":8955,"temperature":1.0,"reasoning_tokens":778,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:44:00.810983+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the tight-to-loose prediction with data in a same-sign dilepton sideband with exactly two jets and at most one b-tagged jet, where the four-top signal is negligible; a disagreement beyond the quoted 30-60% uncertainty in the nonprompt estimate would shift the measured cross section directly.","supporting_citations":[{"cited_title":"Measurements of tt-bar cross sections in association with b jets and inclusive jets and their ratio using dilepton final states in pp collisions at sqrt(s) = 13 TeV","cited_arxiv_id":"1705.10141","evidence_quote":"Supplies the measured ratio of ttbb to ttjj cross sections used to correct the flavor composition of additional jets."}],"review_version":1}