{"id":"6136081c-3bc0-4515-a016-3806a6f5fd53","arxiv_id":"2501.11420","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"ATLAS's Run 3 b-jet trigger, based on the DL1d and GN1 neural-network taggers plus a FastDIPS preselection, achieves about 50% higher efficiency for HH to bbbb and HH to bbtautau than the Run 2 trigger strategy.","lead":"ATLAS describes how it selected jets containing b-quarks in real time during the 2022 and 2023 LHC runs, using new neural-network taggers and a fast preselection step. The new trigger menu records about 50% more Higgs-boson-pair events with b-quarks, sharpening the search for the Higgs self-interaction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 50% HH efficiency gain relies on MC extrapolation to low-pT b-jets; the only data/MC validation is above the trigger turn-on, leaving the soft-jet regime where the gain is largest untested.","rationale":"The reader's weakest_assumption correctly identifies the same load-bearing concern: the headline efficiency gain for HH final states is evaluated entirely in Monte Carlo, with the only direct data validation performed on a single b-jet trigger chain in ttbar-enriched events. My stress-test sharpens this by noting that the validation deliberately sits above the trigger turn-on plateau, so it does not exercise the low-pT jet acceptance region where the Run 3 menu gains the most efficiency. This matters because the largest gain occurs at low mHH, where the Higgs boson pair is produced close to threshold and the b-jets are soft; the online tracking and fast b-tagging preselection are most fragile precisely in that regime. This is not an internal inconsistency or a circular argument: the taggers are trained on ttbar MC and evaluated on independent HH MC, and the rate estimates use enhanced-bias data. The paper also provides independent support through ROC improvements, rate measurements, and a data/MC comparison for one chain. The concern is therefore about the quantitative robustness of the central number, not about its validity, and it is addressable by a targeted low-pT data/MC check. The reader's CONDITIONAL verdict already captures this caveat, so no verdict change is needed.","tokens_in":56896,"tokens_out":4270,"duration_ms":50970,"concrete_test":"Re-run the Figure 6 efficiency calculation replacing the nominal MC b-tagging efficiencies for jets with pT < 60 GeV with per-jet scale factors derived from the ttbar-enriched 2b2jasym data/MC comparison of Figure 5, restricted to the same low-pT range. Alternatively, measure the 2b2jasym event-level efficiency in a ttbar control region with third/fourth jets in the 20-60 GeV range without the 'above plateau' requirement, and apply the resulting data/MC event-level correction to HH->bbbb. If the inclusive gain drops by more than about 15% (e.g., from 44% to below roughly 30%), the headline should be rephrased as a simulation-based projection with a low-pT systematic caveat.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 6.4, Figures 6-7; abstract) is a simulated roughly 50% efficiency gain for HH->bbbb and HH->bbtautau. The load-bearing assumption is that the online trigger response in HH events, especially at low mHH where the gain reaches about 75%, is correctly modeled by simulation. The paper's data validation (Section 6.3, Figure 5) only covers the 2b2jasym chain in ttbar-enriched events and deliberately requires jets 'above the plateau of the jet trigger turn-on' (Section 6.3.2), so it tests b-tagging at moderate/high pT but not the low-pT jet acceptance whose loosening is the main source of the gain. HH b-jets near threshold are softer (pT roughly 20-40 GeV), and the online FTF/precision tracking and FastDIPS preselection are known to degrade at low pT; a few-percent data/MC inefficiency there would directly reduce the headline gain. The abstract calls the gain 'observed' although it is a simulation-based projection; this is not an internal inconsistency, but it elevates an untested extrapolation to a result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper documents the configuration, commissioning, and performance of the ATLAS b-jet triggers in the first two years of LHC Run 3 (2022-2023). It describes the online inputs (FTF and precision tracking, primary-vertex reconstruction, EMTopo and particle-flow jets), the low-level taggers (JetFitter, SSVF, DIPS) and high-level taggers (DL1d in 2022, GN1 in 2023), the new two-step FastDIPS preselection strategy, and the full deployed trigger menu (Tables 1-4). Trigger rates are predicted from enhanced-bias data and compared with measured 2023 rates as a function of instantaneous luminosity (Figures 3-4). The event-level efficiency of the 2b2jasym chain is compared between data and ttbar Monte Carlo in events above the jet trigger turn-on plateau (Figure 5). The paper concludes that the new menu achieves about a 50% inclusive efficiency gain for HH->bbbb relative to the Run 2 strategy, an efficiency above 50% for HH->bbtautau relative to the fiducial selection, and up to a factor 1.7 improvement over the Run 2 tau-based strategy (Figures 6-7), and it reports the online and offline monitoring infrastructure used for data-quality assessment.","tokens_in":57078,"tokens_out":14514,"duration_ms":141493,"significance":"The reported gain is of direct physics interest: if it holds in data, the Run 3 menu records roughly 40-70% more SM HH events than the Run 2 strategy at comparable rate, directly improving the reach of the HH self-coupling program. The paper has clear strengths: the rate predictions are data-driven (enhanced-bias data) and validated against measured 2023 trigger rates, the deployed menu and operating points are documented transparently, the tagger comparison (Figure 2) quantifies the algorithmic improvement, and the ttbar-enriched data/MC comparison (Figure 5) is an honest, explicit anchor for the online tagger performance. I found no internal inconsistency and no circularity in the central claim: the taggers are trained on ttbar MC and evaluated on independent HH MC, while operating points and rates are anchored in data. The load-bearing caveat is that the headline HH efficiency gain is a pure-MC projection whose dominant regime of improvement, the low-pT acceptance near 20-40 GeV, is only partially covered by the data validation, and the headline number is quoted without a total uncertainty.","major_comments":[{"comment":"The headline claim of the paper, an inclusive efficiency gain of about 50% for HH->bbbb and an efficiency above 50% for HH->bbtautau, rests entirely on a Monte Carlo projection: the efficiencies in Figures 6 and 7 are evaluated on Powheg+Pythia HH samples processed through the full Geant4 detector simulation (Section 3), and they are quoted without any statistical or systematic uncertainty, so the reader cannot tell whether the gain is about 40% or about 60%. The data validation in Section 6.3 covers only the 2b2jasym chain in ttbar-enriched events, and the abstract's wording that the improvement 'is observed' is therefore stronger than the evidence supports; notably, Section 6.4 itself qualifies the bbtautau figure as 'predicted, in simulation.' Additionally, the abstract conflates two different statements: for HH->bbbb the claim is a relative improvement of about 45% over Run 2 (59% vs 41% in Figure 6), whereas for HH->bbtautau the claim is an absolute efficiency above 50% relative to the fiducial selection with up to a factor 1.7 gain. I request: (i) a total uncertainty on the efficiencies in Figures 6 and 7, at minimum the MC statistical precision with the dominant systematics identified (for example tagger training, pile-up modeling, and jet/track response at low pT), and an explicit statement of whether the quoted values are unrounded point estimates; (ii) rewording of the abstract so that the gain is described as an expected improvement from simulation, validated in data only in a restricted phase space; and (iii) confirmation that the Run 2 baseline values (41% for bbbb, from the 13 TeV analysis of Ref. [8]) and the Run 3 values are computed with the same fiducial definitions, since the Figure 6 caption indicates the comparison mixes 13 TeV and 13.6 TeV simulated samples.","section":"Section 6.4, Figs 6-7; Abstract"},{"comment":"The data/MC validation excludes the kinematic regime in which the claimed gain is largest. The ttbar control selection requires at least four offline jets with pT above 120, 70, and 30 GeV, and the 2b2jasym efficiency in Figure 5 is measured only for events 'above the plateau of the jet trigger turn-on,' where the jet-pT requirements are chosen so that at least 95% of the selected events satisfy the L1 and HLT jet thresholds. This design isolates the performance of the b-tagging discriminants, which is a genuine strength, but it removes all sensitivity to the jet reconstruction and tracking acceptance at pT between 20 and 30 GeV. That is precisely the loosened acceptance responsible for the gain claimed in Section 6.4: the 2b2jasym chain in Table 2 accepts b-tagged jets with pT > 20 GeV, and the paper attributes the gain to loosened jet-pT thresholds enabled by the improved taggers. The largest improvement (about 75%) is quoted at low mHH, where the signal b-jets are softest and the FTF/precision tracking and FastDIPS preselection are known to degrade. I request either a data/MC comparison of the full chain efficiency, including the FastDIPS preselection and jet acceptance, in a phase space extending down to pT ~ 20 GeV (accepting the turn-on as an additional systematic), or a quantitative sensitivity study of how the quoted 50% gain changes under plausible data/MC differences in the low-pT acceptance, together with an explicit statement of this limitation at the point where the number is introduced.","section":"Section 6.3.2, Fig 5, Table 2"}],"minor_comments":[{"comment":"The sentence comparing GN1 with DL1d refers to Figure 2(a), but that panel shows the DL1d/DIPS versus DL1r comparison; the GN1 comparison appears in panel (b).","section":"Section 5.6, Fig 2"},{"comment":"There are typos: 'degrees of freeedom' should be 'degrees of freedom,' and 'Trigger and Data Acquistion' should be 'Trigger and Data Acquisition.'","section":"Section 7"},{"comment":"Please state the pile-up profile assumed in the HH simulation used for Figures 6 and 7 and whether the efficiencies are averaged or reweighted over the Run 3 luminosity profile, given that the 2022-2023 data cover an average number of interactions per crossing of about 30 to 70 (Section 1).","section":"Section 6.4"},{"comment":"The data/MC agreement in Figure 5 is characterized only as 'overall good'; adding a quantitative measure per panel (for example a chi-square value or per-bin uncertainties including correlated components) would make the validation reproducible for the reader.","section":"Section 6.3.2, Fig 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a well-executed ATLAS performance paper, squarely within JINST scope, and I see no circularity or fabrication concerns. The abstract overstates an MC-only projection as 'observed,' which is the main editorial risk; the requested uncertainty statement and wording fixes should be enforced before acceptance. The citation pattern is typical for an ATLAS collaboration paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. This is a solid, useful ATLAS trigger paper, and the headline claim is genuinely important: the Run 3 b-jet menu gives roughly a 50% efficiency gain for HH->bbbb and up to a factor 1.7 for HH->bbtautau relative to Run 2. But read the fine print: that gain is a Monte Carlo projection, not a data measurement. The abstract's 'observed' overstates it.\n\nWhat's new and good: deploying GN1, a graph neural network tagger, in the HLT is a first, and it gives up to a factor two better light-jet rejection than DL1d. The FastDIPS two-step preselection is a sensible engineering solution to the CPU problem, and the trigger menu tables are detailed enough to reproduce the logic. The rate predictions use enhanced-bias data, and Figure 4 shows the deployed rates are stable in 2023 data. The data/MC comparison in ttbar-enriched events (Figure 5) shows good agreement for the 2b2jasym chain, which is the main chain for HH->bbbb.\n\nThe soft spots are proportionate to the paper's scope. The headline efficiency gain is computed on simulated HH samples, and no total uncertainty is quoted. The data validation deliberately requires jets above the trigger turn-on plateau, so the low-pT b-jet region where the gain is largest (about 75% at low mHH) is not directly tested in data. That's a real limitation, especially since online tracking efficiency degrades at low pT. But this is standard practice for ATLAS trigger papers; the full detector simulation and the good plateau agreement make the projection believable. It should be labeled as a projection, not 'observed'.\n\nThe taggers are trained on ttbar jets and evaluated on HH jets. That's not circular—there are independent anchors in data rates and ttbar validation—but it does mean the HH-specific gain is one extrapolation step away from data.\n\nBottom line: this paper deserves a serious referee. It's an enabling result for Run 3 HH and SUSY analyses, and the technical content is sound. I'd ask for two revisions: change the abstract wording to 'predicted' or 'expected' rather than 'observed', and add a sentence (or a small table) giving the dominant systematic uncertainties on the efficiency gain, or at least an explicit statement that the gain is MC-only. Both are minor. Send it to peer review.","headline":"Solid Run 3 b-jet trigger paper with a real but MC-only headline gain; deserves review after small wording and uncertainty fixes.","tokens_in":57686,"tokens_out":2397,"would_cite":true,"duration_ms":24578,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that ATLAS rebuilt its real-time b-jet identification for LHC Run 3, replacing the Run 2 tagger with two neural-network classifiers, DL1d and GN1, and adding a fast calorimeter-based preselection step before precision…","keywords":["b-jet triggers","high-level trigger","flavour tagging","neural network taggers","Higgs boson pair production","LHC Run 3","DL1d","GN1"],"falsifier":"Measure the per-jet online b-tagging efficiency for b-jets in a data sample enriched in HH-like low-mHH topology, for example HH→bbττ events selected by the tau triggers, and propagate the measured scale factors to the HH→bbbb and HH→bbττ trigger efficiencies; if the corrected efficiency gain drops below about 30% at mHH below 450 GeV, the central claim fails.","tokens_in":56617,"feed_emoji":"⚛️","tokens_out":4331,"duration_ms":42284,"temperature":0.7,"pith_summary":"The paper reports how ATLAS rebuilt its real-time b-jet identification for LHC Run 3, replacing the Run 2 tagger with two neural-network classifiers, DL1d and GN1, and adding a fast calorimeter-based preselection step before precision tracking. The central claim is that this menu raises the trigger efficiency for Standard Model Higgs-boson-pair production by roughly 50% in the HH→bbbb channel and by up to a factor 1.7 in the HH→bbττ channel, relative to the Run 2 trigger strategy, at comparable trigger rates. If true, the same recorded luminosity now delivers about half again as many fully hadronic di-Higgs events, which are the key channel for measuring the Higgs self-coupling. The efficiency gain is largest at low di-Higgs invariant mass near the 2mH threshold, precisely the region most sensitive to the self-coupling.","feed_headline":"New b-jet triggers record 50% more Higgs-pair events","feed_subtitle":"Run 3 trigger upgrades keep rates flat while catching about 1.5 times as many HH→bbbb and HH→bbττ decays.","key_machinery":"The load-bearing objects are the neural flavour taggers DL1d, an eight-layer perceptron taking low-level vertex and DIPS inputs, and GN1, a graph neural network consuming about twenty track quantities per jet, together with a two-step HLT structure: a loose FastDIPS preselection on EMTopo jets that runs before precision tracking, and a final selection on particle-flow jets using DL1d or GN1 operating points. The b-jet discriminant is the log-odds D_b = log(p_b/(f_c p_c + (1-f_c) p_u)) with f_c = 0.018, and operating points are defined by b-jet efficiency in simulated ttbar events. The mechanism that yields the gain is that stronger light-jet rejection allows lower jet pT thresholds and looser b-tag working points without raising the trigger rate.","core_discovery":"On the paper's own terms, the central discovery is that a trigger strategy built from the DL1d and GN1 b-tagging discriminants, a FastDIPS preselection running on calorimeter jets, and asymmetric multi-b-jet chains recovers about 50% more SM HH→bbbb events and more than 50% of HH→bbττ events passing the fiducial selection, compared with the Run 2 strategy. The 2b2j_asym chain alone gives 55% efficiency on HH→bbbb and is the single most efficient chain; combining it with 3b1j_asym, 2b1j and 2b2j_cent reaches 59%. The gain reaches about 75% near mHH ~ 2mH, and the bbττ efficiency exceeds 50% with up to a 1.7 factor gain over the Run 2 tau-based strategy.","pith_inferences":["The largest efficiency gain appearing in the low-mHH region means the practical boost to Higgs self-coupling measurements may be larger than the inclusive 50%, since that region carries the most coupling information.","A direct data-driven closure test inside the HH→bbbb fiducial region, using for example tag-and-probe b-jets reweighted to the HH kinematics, would test whether the MC-only gain holds; this is an extension the paper does not perform.","The same two-step preselection recipe could be applied to c-tagging or to future high-pileup runs, where the fast preselection's rejection factor would need to be re-optimized.","The reported efficiency is relative to inclusive or fiducial HH events, not to the offline analysis selection, so the analysis-level gain will depend on how the trigger overlaps with the offline b-tagging working points."],"forward_implications":["The di-Higgs analyses can collect roughly 50% more fully hadronic signal events at the same integrated luminosity, directly increasing sensitivity to the Higgs self-coupling.","The looser jet thresholds push trigger turn-on curves lower, so HH→bbbb events with softer b-jets, which dominate near the 2mH threshold, are recorded rather than lost.","Switching the 2023 baseline from DL1d to GN1 adds up to a factor two in light-flavour jet rejection at fixed b-efficiency, which will also benefit other all-hadronic searches such as supersymmetry and resonances decaying to b-quarks.","The FastDIPS preselection keeps CPU usage manageable, meaning the improved menu scales to higher pile-up without requiring more trigger farm capacity.","Data-MC agreement in ttbar-enriched events shows the online taggers behave as simulated for b-jets from top decay, supporting extrapolation to other b-rich signatures."],"supporting_citations":[{"why":"Defines the Run 2 ATLAS b-jet trigger strategy that serves as the baseline for the efficiency-gain comparison.","marker":"[16]"},{"why":"The Run 2 HH→bbbb analysis whose trigger selection is used to compute the Run 3 gain in Figure 6.","marker":"[8]"},{"why":"The Run 2 HH→bbττ analysis whose tau-based trigger strategy is the baseline for the 1.7 factor gain.","marker":"[69]"},{"why":"Documents the DL1d and DIPS flavour-tagging algorithms and their training prescription, which the online taggers inherit.","marker":"[5]"},{"why":"Describes the GN1 graph neural network tagger that provides the additional light-jet rejection used in 2023.","marker":"[57]"},{"why":"Describes the FastDIPS fast b-tagging preselection that enables the two-step HLT strategy.","marker":"[55]"},{"why":"Provides the Run 3 ATLAS trigger system description, including the HLT and tracking steps the b-jet chains use.","marker":"[4]"},{"why":"Defines the offline b-tagging efficiency measurement used to validate data-MC agreement in ttbar-enriched events.","marker":"[27]"}],"fun_headline_variants":["ATLAS b-jet triggers boost Higgs-pair yield by 50%","New ATLAS triggers catch 1.5x more Higgs-pair events","Run 3 b-jet triggers net 50% gain in Higgs-pair selection","b-jet trigger redesign improves HH→bbbb efficiency by 50%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 50% gain is computed in full Monte Carlo simulation of HH events, and the paper's data validation covers only the 2b2j_asym chain in ttbar-enriched events, so the claim assumes the online taggers and tracking reproduce their simulated response in the HH signal phase space, especially at low mHH.","fun_headline_variants_meta":{"raw":{"variants":["ATLAS b-jet triggers boost Higgs-pair yield by 50%","New ATLAS triggers catch 1.5x more Higgs-pair events","Run 3 b-jet triggers net 50% gain in Higgs-pair selection","b-jet trigger redesign improves HH→bbbb efficiency by 50%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000631,"raw_usage":{"total_tokens":2929,"prompt_tokens":974,"completion_tokens":1955,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":1870}},"tokens_in":590,"tokens_out":1955,"duration_ms":13180,"temperature":1.0,"reasoning_tokens":1870,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:17:27.925649+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the per-jet online b-tagging efficiency for b-jets in a data sample enriched in HH-like low-mHH topology, for example HH→bbττ events selected by the tau triggers, and propagate the measured scale factors to the HH→bbbb and HH→bbττ trigger efficiencies; if the corrected efficiency gain drops below about 30% at mHH below 450 GeV, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the GN1 graph neural network tagger that provides the additional light-jet rejection used in 2023."}],"review_version":1}