{"id":"f132c03e-4145-48c8-a249-030d9e2c102b","arxiv_id":"2603.08315","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"No excess found in 137 fb^-1 of 13 TeV pp data; new 95% CL exclusions reach chargino masses up to 880 GeV and tau-sleptons up to 320 GeV at ~1 ns lifetimes.","lead":"An ATLAS search for long-lived charginos and tau-sleptons using 'disappearing track' signatures found no excess over the Standard Model, setting new 95% CL mass limits that extend sensitivity to shorter lifetimes via 3-hit tracks and a machine-learned soft-pion tagger.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fake-tracklet pT template for the 4-layer SRs is taken from 3-hit tracklets; the absence of a layer-count transfer factor biases SR4 background if 3- and 4-hit fake shapes differ.","rationale":"The central claim is a null result and associated 95% CL limits; both depend on the accuracy of the dominant fake-tracklet background. The reader correctly identifies the fake-tracklet estimate as the weakest assumption, focusing on the low-to-high EmissT extrapolation and MC-derived transfer factors. My concern sharpens this to a specific unstated assumption in the same chain: the fake pT template is built from 3-hit tracklets even for SRs requiring 4-hit tracklets, with no layer-count correction in f^fake_SR. This is load-bearing because SR4High and SR4Mid drive the strongest exclusions, and a shape difference at pT > 60 GeV would propagate directly into the SR yields and the binned likelihood limits. The existing validation regions do not cover this kinematic corner with a fake-enriched, 4-hit, high-pT sample: VR4Mid has pT < 60 GeV and VR4MidS is hadron-dominated. The concern is not a demonstrated error, but an untested assumption whose failure would alter the quoted limits. It does not overturn the null-result claim outright, since even a factor-of-two background change would not create a 5σ excess, but it could materially change the mass limits, so the appropriate verdict is CONDITIONAL acceptance pending a 3-hit/4-hit shape closure test.","tokens_in":63753,"tokens_out":12970,"duration_ms":119005,"concrete_test":"From V+jets MC (or from the low-EmissT TR in data), extract fake-tracklet pT templates separately for exactly 3-hit and exactly 4-hit tracklets, using the same |d0|/sigma and |z0 sinθ| selections. Apply TF^fake_hybrid and SF^fake_EmissT and compare the two shapes for pT > 60 GeV. If the 4-hit/3-hit ratio deviates from unity by more than the sum in quadrature of statistical and systematic uncertainties, recompute SR4High/SR4Mid fake yields and the wino/higgsino/stau 95% CL contours using the 4-hit template. A quick closure: fit CR4Low and CR4High with a 4-hit template and compare the predicted pT distribution to data in bins; significant pulls would confirm the bias.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 6 defines f^fake_SR(pT) = N_TR^fake(pT) × TF^fake_d0 × TF^fake_hybrid(pT) × SF^fake_z0 × SF^fake_EmissT. The text states: 'All three TRs select tracklets reconstructed from hits in three layers to ensure a high-statistics pT template.' Yet SR4High and SR4Mid require four-layer tracklets (Table 1). There is no term correcting the template for layer multiplicity. Thus the dominant background in the 4-layer SRs, which provide the main sensitivity and drive the strongest limits, assumes the pT spectrum of 3-hit fake tracklets equals that of 4-hit fake tracklets at pT > 60 GeV and high EmissT. The validation shown does not test this: VR4Mid has tracklet pT < 60 GeV, and VR4MidS is hadron-dominated (Tracklet E_clus > 5 GeV). A harder 4-hit fake spectrum would mean the SR4High background (0.68 ± 0.14) is underestimated, making the 1.9σ excess less significant and the observed limits (880 GeV wino, 720 GeV higgsino) weaker than claimed; a softer 4-hit spectrum would have the opposite effect. The MC-derived TF^fake_hybrid and SF^fake_EmissT also have no dedicated methodology uncertainty (Section 7), so the quoted uncertainties may undercover this extrapolation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a search for long-lived charginos and tau-sleptons using the disappearing-track signature in 137 fb^-1 of 13 TeV pp collisions recorded by ATLAS. Four signal regions are defined: two requiring four-pixel-layer tracklets and two requiring three-pixel-layer tracklets, with the latter further split by a BDT-based low-energy pion tag. The background is estimated with a data-driven strategy using template regions, transfer factors, and control-region normalizations; the dominant fake-tracklet component is modeled from a low-EmissT, high-|z0 sin theta| template with MC-derived transfer factors. No significant excess is found (largest local significance 1.9 sigma in SR4High), and 95% CL exclusion limits are set on wino and higgsino charginos and on tau-sleptons in CMSSM- and GMSB-inspired scenarios. The observed (expected) limits reach 880 GeV (1020 GeV) for wino production and 720 GeV (840 GeV) for higgsino production at ~1 ns lifetime, and 320 GeV (390 GeV) for CMSSM staus.","tokens_in":64103,"tokens_out":8562,"duration_ms":90709,"significance":"If the background estimate is unbiased, the paper presents a solid experimental result with improved sensitivity over the previous ATLAS disappearing-track search, particularly for short lifetimes due to the use of three-pixel-layer tracklets and the dedicated pion tag. The paper is unusually transparent: detailed selection tables, control and validation regions, post-fit distributions, and an explicit discussion of the local excess are provided. The systematic treatment is thorough for the electron, muon, and hadron backgrounds, with data-driven tag-and-probe methods where possible. However, the central exclusion limits rely sensitively on the fake-tracklet background in the four-layer regions, and the manuscript itself states that no dedicated uncertainty is associated with the overall background estimation methodology. Because the validation regions do not cover the high-pT, high-EmissT, calo-veto phase space of SR4High, this omission is load-bearing for the central claim.","major_comments":[{"comment":"The fake-tracklet pT template N_TR^fake(pT) is explicitly taken from tracklets with hits in three layers ('All three TRs select tracklets reconstructed from hits in three layers to ensure a high-statistics pT template'), while SR4High and SR4Mid require four-layer tracklets. No transfer factor or shape correction is applied for the layer multiplicity. The four-layer SRs dominate the sensitivity and drive the strongest limits. The validation regions do not test the extrapolation: VR4Mid has pT<60 GeV, and VR4MidS requires E_clus>5 GeV (hadron-dominated), so neither probes the high-pT, calo-veto, high-EmissT region relevant to SR4High. If the 3-hit and 4-hit fake tracklet pT spectra differ at pT>60 GeV, the SR4High background of 0.68 +/- 0.14 and the resulting mass limits would be biased. Please either introduce a layer-count transfer factor, validate the 4-layer fake shape in a dedicated","section":"Section 6, Eq. (1) and Table 1"},{"comment":"The manuscript states: 'No dedicated uncertainty is associated with the overall background estimation methodology as, within other sources of uncertainties including the statistics of the data, the predicted post-fit tracklet pT distributions in the VRs are consistent with the data.' This is not sufficient because the MC-derived terms in the fake background, TF^fake_hybrid(pT) and SF^fake_EmissT, are based on V+jets simulation and a fitted exponential-plus-constant function, and the validation regions (VR4Mid, VR4MidS, VR3Mid1pi, VR3High0pi) do not cover the SR4High phase space. The quoted background uncertainties in SR4High and the derived limits therefore may undercover the extrapolation uncertainty. A quantitative methodology uncertainty, or an additional validation specifically in the high-pT, low-E_clus, high-EmissT region, should be provided.","section":"Section 7"},{"comment":"The scale factor SF^fake_EmissT is derived from the ratio of events in CR4Low (EmissT<150 GeV) to CR4High (EmissT>300 GeV). Both regions require the EmissT trigger to pass. The text justifies the low-Emiss TR by saying the trigger efficiency 'does not need to be well understood' because the TR is not used for the overall yield, but the CR ratio is used for the normalization and is therefore directly affected by any trigger inefficiency in CR4Low. Since CR4Low is below the 230 GeV online trigger threshold quoted in Section 3, the trigger efficiency is not on the plateau. Please demonstrate that the trigger efficiency cancels in the ratio or apply a correction; otherwise the SF is biased.","section":"Section 6, CR4Low/CR4High (Table 2)"}],"minor_comments":[{"comment":"Typo: 'in both the CSSM and GMSB models' should be 'CMSSM'.","section":"Section 3"},{"comment":"Typo: 'The data obervations' should be 'observations'.","section":"Section 9"},{"comment":"The text says 'the previous analyses used five-layer tracking', but Section 2 describes four pixel layers and Ref. [43] is described as requiring at least four pixel hits. Please clarify what 'five-layer tracking' refers to (perhaps it includes SCT hits).","section":"Section 8"},{"comment":"The reproduction of the text contains numerous formatting artifacts (missing spaces, stray characters such as 'Tτ-sleptons' in the contents). These should be corrected in the final published version.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-executed ATLAS search and the central null result is plausible, but the layer-count extrapolation in the fake-tracklet estimate is a genuine concern because the four-layer SRs are the main sensitivity drivers and the validation regions do not cover this phase space. The absence of a methodology uncertainty, explicitly acknowledged in Section 7, strengthens the concern. I would like to see either a dedicated 4-layer validation or a systematic uncertainty, along with clarification of the trigger-efficiency issue in the low-Emiss CR ratio, before recommending acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent, incremental but genuinely useful ATLAS search. The new bits are the three-hit tracklet reconstruction and the BDT soft-pion tag; they close part of the short-lifetime higgsino gap, raising the excluded mass from 152 to 225 GeV in the loop-driven splitting scenario. The wino and stau limits are also improved. The paper is transparent about the small excess in SR4High (1.9 sigma), and the analysis is documented to the level where a referee can follow the logic.\n\nWhat's done well: the background estimation is mostly data-driven with clear definitions of TRs, CRs, and VRs. The electron and muon backgrounds use tag-and-probe transfer factors from Z decays. The fake background is split into pure and hybrid fakes with a pT-dependent transfer factor from V+jets MC. The systematics are broken down and the dominant uncertainties identified.\n\nThe soft spot, and it's real but not fatal: the fake-tracklet pT template for the four-layer signal regions is extracted from three-layer tracklets. There is no layer-count transfer factor. The stated reason is statistics, but if the pT spectrum of fake tracklets differs between 3- and 4-hit configurations at high pT, the SR4High background (0.68±0.14) could be off, which directly affects the observed limits and the local excess. The validation regions don't cover this: VR4Mid has tracklet pT<60 GeV and VR4MidS is dominated by hadrons. The paper has no dedicated methodology uncertainty for this extrapolation. This should be a referee question: either add a systematic or show MC closure for the 3-to-4 hit shape at high pT.\n\nAlso worth noting: the muon background in SR4High is comparable to the fake background (0.34 vs 0.33), so the fake shape issue is not the whole story in that SR.\n\nBottom line: the central null result and the limits are probably fine, but the SR4High numbers carry a bit more uncertainty than the quoted error bars suggest. I'd send this to peer review. The techniques will be reused, and the short-lifetime higgsino limit is a real step forward.","headline":"Solid, incremental ATLAS search with genuinely useful new techniques and one extrapolation concern worth asking about.","tokens_in":64589,"tokens_out":2217,"would_cite":true,"duration_ms":27217,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A search for long-lived charginos and tau-sleptons in 137 fb^-1 of 13 TeV proton-proton collisions finds no significant excess and sets new 95% CL mass limits, excluding higgsino-like charginos up to 225 GeV at lifetimes below 0.03 ns.","keywords":["disappearing track","tracklet","chargino","tau-slepton","long-lived particles","supersymmetry","missing transverse momentum","LHC"],"falsifier":"Take the three observed SR4High events (tracklet pT near 139, 142, and 152 GeV, expected background 0.68 ± 0.14) and re-analyze the next ~140 fb^-1 of Run-3 data with the same selection: if the yield grows to several events while the scaled background stays near one, the null result is contradicted; if the events disappear or match the scaled background, the central limit claim survives.","tokens_in":63640,"feed_emoji":"🎯","tokens_out":6637,"duration_ms":67241,"temperature":0.7,"pith_summary":"This paper is trying to establish whether supersymmetric partner particles — charginos and tau-sleptons — that travel a few centimetres before decaying show up as short 'disappearing tracks' in the ATLAS detector. Using the full Run 2 dataset of 137 fb^-1, it finds no significant excess over Standard Model backgrounds and converts that null result into 95% confidence level lower bounds on particle masses in several benchmark models. The key advance is reconstructing tracks with only three pixel hits and using a machine-learned tagger for the very soft pion from chargino decay, which extends sensitivity to the short lifetimes expected for higgsino-like dark-matter candidates. If correct, the result tightens the exclusion of light higgsino and wino supersymmetric scenarios and of long-lived tau-sleptons in CMSSM- and GMSB-inspired models.","feed_headline":"No new excess in LHC disappearing-track search","feed_subtitle":"The limits press on higgsino and wino dark-matter models and long-lived tau-sleptons.","key_machinery":"The carrying mechanism is the disappearing-track signature: a short charged track left by a heavy charged particle, such as a chargino (the charged supersymmetric partner of the electroweak and Higgs states) or a tau-slepton (the partner of the tau lepton), that decays after crossing three or four of ATLAS's innermost pixel layers, leaving no hits in the outer silicon tracker. A dedicated tracklet reconstruction allows tracks as short as three pixel hits, and for three-hit tracklets a boosted decision tree identifies the low-energy charged pion from chargino decay. Backgrounds are estimated data-drivenly from template regions, transfer factors, and control regions; the fake-tracklet componen","core_discovery":"On the paper's own terms, the central result is the absence of an excess and the setting of 95% CL exclusion limits in the 0.01–10 ns lifetime window. Observed (expected) limits reach 225 GeV (250 GeV) for pure-higgsino charginos at lifetimes below 0.03 ns, 720 GeV (840 GeV) for the same particles at around 1 ns, 880 GeV (1020 GeV) for wino-like charginos near 1 ns, and 320/300 GeV (390/380 GeV) for tau-sleptons in CMSSM/GMSB-inspired scenarios. The largest local excess, in the high missing-energy four-hit region, has a significance of 1.9σ, which the paper treats as consistent with background.","pith_inferences":["One thing the paper leaves open is the fate of the 1.9σ excess in SR4High: if the three events around 140–150 GeV tracklet pT persist in more data, they could become a real signal or expose an underestimated fake background.","The same three-hit tracklet plus pion-tagging technique could be applied to Run-3 data at 13.6 TeV, where the larger dataset should push the higgsino limit to higher masses.","Because the signal regions are defined model-independently around tracklet kinematics, other new physics with a decaying charged track and missing energy could be reinterpreted with the same results.","The reliance on V+jets simulation for the pure-to-hybrid fake ratio suggests that a dedicated high missing-energy fake-enriched control sample would directly test the extrapolation that carries the main background uncertainty."],"forward_implications":["Wino-like charginos with masses up to 880 GeV and lifetimes around 1 ns are excluded at 95% CL.","Higgsino-like charginos below 225 GeV are excluded for lifetimes below 0.03 ns, covering the loop-induced mass-splitting region.","Long-lived tau-sleptons with lifetimes around 1 ns are excluded up to about 320 GeV (CMSSM) and 300 GeV (GMSB).","No signal region shows more than a 1.9σ local excess, so the Standard Model background prediction is consistent with data in these final states.","The improved tracklet and pion reconstruction extends the expected mass reach by about 100 GeV compared with the earlier Run-2 analysis."],"fun_headline_variants":["ATLAS excludes wino charginos up to 880 GeV in disappearing-track search","No excess in ATLAS disappearing-track search; chargino limits reach 880 GeV","LHC disappearing-track search: no excess, but chargino limits up to 880 GeV","ATLAS tightens chargino and slepton limits with disappearing-track search","No signal, new limits: ATLAS probes long-lived charginos and sleptons"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The estimate of the dominant fake-tracklet background assumes that the tracklet transverse-momentum shape and the ratio of 'pure' to 'hybrid' fakes measured in low missing-transverse-momentum regions and simulated V+jets events correctly describe the high missing-transverse-momentum signal regions.","fun_headline_variants_meta":{"raw":{"variants":["ATLAS excludes wino charginos up to 880 GeV in disappearing-track search","No excess in ATLAS disappearing-track search; chargino limits reach 880 GeV","LHC disappearing-track search: no excess, but chargino limits up to 880 GeV","ATLAS tightens chargino and slepton limits with disappearing-track search","No signal, new limits: ATLAS probes long-lived charginos and sleptons"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001183,"raw_usage":{"total_tokens":4813,"prompt_tokens":923,"completion_tokens":3890,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":667,"completion_tokens_details":{"reasoning_tokens":3779}},"tokens_in":667,"tokens_out":3890,"duration_ms":25491,"temperature":1.0,"reasoning_tokens":3779,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T02:31:34.085859+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the three observed SR4High events (tracklet pT near 139, 142, and 152 GeV, expected background 0.68 ± 0.14) and re-analyze the next ~140 fb^-1 of Run-3 data with the same selection: if the yield grows to several events while the scaled background stays near one, the null result is contradicted; if the events disappear or match the scaled background, the central limit claim survives.","supporting_citations":[],"review_version":1}