{"id":"f9b0480b-4a06-468a-b539-5422b358a0c9","arxiv_id":"2607.06199","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":6,"one_line_summary":"Simultaneous measurements of b-jet tagging and c-jet mistagging efficiency scale factors for the ATLAS GN2 tagger achieve precisions below 2% and 5% respectively using 56 fb⁻¹ of 13.6 TeV ttbar data.","lead":"This paper measures how well the ATLAS experiment's GN2 algorithm identifies jets from bottom and charm quarks using top-quark pair events from LHC Run 3 data. These calibration factors are essential inputs for nearly every physics analysis at ATLAS that relies on distinguishing jet types.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The b-jet background normalisation factor B=1.56 in the c-jet SF fit is a single global parameter that may not capture pT-dependent mis-modelling of the b-jet contamination in the W-jet sample.","rationale":"The concern about B being a single global factor is legitimate and worth investigating, but it does not rise to the level of invalidating the paper's central claims for several reasons. First, the headline precision claims for c-jet SFs (< 5%) are made only for the four loosest tagging intervals, not the 0–65% interval where b-jet backgrounds are largest and B has the most leverage. In the loosest intervals, the c-jet sample dominates and the b-jet background is smaller, reducing sensitivity to B. Second, B is a free parameter in the fit, so it absorbs the overall normalisation discrepancy, and its uncertainty (±0.21 systematic) is propagated to the c-jet SFs. Third, the systematic uncertainties from ttbar modelling (up to 24% in the tightest c-jet bin) partially cover the potential pT-dependence of the b-jet background, since the Herwig variation and ISR/FSR variations change the misassignment rates and heavy-flavour fractions. Fourth, the paper is transparent about large uncertainties where they exist. The analysis follows established ATLAS methodology, the fit structure is sound, and the combination with the independent dileptonic b-jet measurement provides a cross-check. The B=1.56 factor is worth scrutinising in future iterations (e.g., with a pT-binned B), but it does not undermine the specific precision claims made in the paper. The reader's ACCEPT verdict with HIGH confidence is appropriate; I would not adjust it, though I note that the correctness risk could be flagged as 'moderate' rather than 'unknown' given the B-factor concern.","tokens_in":57885,"tokens_out":4923,"duration_ms":368420,"concrete_test":"Re-run the c-jet SF fit with B allowed to vary independently in each of the four jet-pT bins (B(i) instead of a single global B). If any B(i) deviates from the global B=1.56 by more than its fitted uncertainty, or if the c-jet SFs shift by more than 10% in any bin, the single-factor approach is insufficient and the pT-dependent c-jet SFs in those bins may carry unaccounted bias. Additionally, examine post-fit data/MC ratios for the tagged W-jet GN2 distribution separately in each pT bin to check for residual shape mis-modelling not captured by the global B.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The c-jet mistagging SF measurement depends critically on controlling the b-jet background in the W-jet sample, which arises from KLFitter misassignment and ttbar+HF production. The fitted b-jet background normalisation factor B=1.56±0.05(stat)±0.21(syst) (Section 6.3) indicates a 56% upward correction relative to simulation — a substantial discrepancy suggesting significant mis-modelling of this background. B enters the χ²_c as a single global multiplicative factor on the b-jet template N_B^u(i,j) across all pT bins i, j and tagging intervals u. The pre-fit correction derived from the hadronic top-jet GN2 shape (split at GN2=6) adjusts the b-jet fraction within the 0–65% interval, but this transfer assumes the GN2 shape for b-jets contaminating the W-jet sample matches that of genuine b-jets from top decay — jets with potentially different kinematics and fragmentation. If the b-jet background mis-modelling is pT- or GN2-dependent (which is plausible given that KLFitter misassignment rates and ttbar+bb fractions vary with jet pT), a single B factor cannot fully correct it, and residual shape distortions could bias the c-jet SFs differently in different bins. This is most consequential in the 0–65% tagging interval where b-jets dominate and the c-jet SFs deviate most from unity (e.g., SF=1.740 at 140–250 GeV). The paper does not show post-fit GN2 distributions of the tagged W-jet separately in each pT bin, so it is not possible to verify from the paper alone whether the global B factor yields adequate data/MC agreement across all pT bins. The reader's concern about KLFitter is related but more general; the specific load-bearing issue is whether the single B parameter adequately captures the pT-structure of the b-jet background mis-modelling.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This paper presents simultaneous measurements of $b$-jet tagging efficiency and $c$-jet mistagging efficiency scale factors (SFs) for the ATLAS GN2 transformer-based flavour tagger, using 56 fb$^{-1}$ of $t{tbar}$ semileptonic events at $sqrt{s}=13.6$ TeV. The $b$-jet SFs are extracted from leptonic top-jets, while the $c$-jet SFs are obtained from hadronic $W$-jets, using a combined $chi^2$ fit with kinematic-likelihood-based jet assignment. The $b$-jet SFs are combined with an independent dileptonic $t{tbar}$ measurement. Total uncertainties below 2% (5%) are achieved for $b$-jets ($c$-jets) in specific kinematic regions.","tokens_in":58693,"tokens_out":1399,"duration_ms":229876,"significance":"This is the first simultaneous measurement of $b$- and $c$-jet SFs for the GN2 tagger using semileptonic $t{tbar}$ events at 13.6 TeV, and it provides calibration results that are directly used by ATLAS physics analyses in Run 3. The simultaneous fit framework with a unitarity constraint on the tagging efficiencies is well-motivated, and the combination with the dileptonic measurement via a prior term is a clean approach. The systematic uncertainty budget is thoroughly evaluated, with the dominant $t{tbar}$ modelling uncertainty reaching up to 26% in the untagged bin (Table 2). Post-fit distributions (Figures 1b, 3b) demonstrate good data/MC agreement. The analysis is internally consistent and the central claims are supported by the presented results.","major_comments":[{"comment":"Section 6.1, chi-squared structure: The global b-jet background normalisation factor B=1.56±0.05(stat)±0.21(syst) (Section 6.3) enters chi-squared_c as a single multiplicative factor on the b-jet template N_B^u(i,j) across all pT bins i, j and tagging intervals u. A 56% upward correction relative to simulation is substantial and indicates significant b-jet background mis-modelling in the W-jet sample (arising from KLFitter misassignment and ttbar+HF production). The concern is whether a single global B can adequately correct pT- or GN2-dependent shape distortions in the b-jet contamination, particularly in the 0-65% tagging interval where b-jets dominate and the c-jet SFs deviate most from unity (e.g., SF=1.740 at 140-250 GeV, Table 3). The pre-fit correction to the b-jet fraction within the 0-65% interval (split at GN2=6, Section 6.1) is derived from the hadronic top-jet distribution, a","section":null},{"comment":"potentially different jet population from the b-jets contaminating the W-jet sample. While the ttbar modelling systematic (up to 24% in the 0-65%, 140-250 GeV bin per Table 4) partially covers this, the paper would be strengthened by showing post-fit GN2 distributions of the tagged W-jet separately in each pT bin, or by providing a dedicated study of the residual pT-dependence of the b-jet background mis-modelling. As the paper stands, it is not possible to fully verify from the presented material alone that the global B factor sufficiently captures the mis-modelling. This is a load-bearing point for the c-jet SF central values in the tightest tagging interval. Recommendation: add a pT-binned post-fit GN2 distribution for the tagged W-jet or a closure test with pT-dependent B factors.","section":null}],"minor_comments":[{"comment":"Section 6.1: The relationship between the pre-fit b-jet fraction correction (derived from the hadronic top-jet, split at GN2=6) and the fitted global B parameter is not fully clear. Clarify whether the pre-fit correction is applied as a fixed input to the template before the fit, and how the uncertainty on this correction (±1.9%) relates to the systematic on B (±0.21). A sentence clarifying the logical flow would help the reader.","section":null},{"comment":"Section 6.1, last paragraph before Section 6.2: The statement that 'the statistical correlation between the b-jet and c-jet SFs is therefore not precisely determined' is somewhat vague. Quantify the expected magnitude or state explicitly that it is neglected and why this is acceptable given the different observables used.","section":null},{"comment":"Section 6.3: The statement 'Several c-jet tagging efficiency measurements have a precision of better than 5%' is accurate but could note that this precision is achieved primarily in the four loosest tagging intervals and at lower pT, not uniformly across all bins.","section":null},{"comment":"Table 3, 0-65% interval: The c-jet SFs show large deviations from unity (e.g., 1.740±0.32 at 140-250 GeV). While the uncertainties are correspondingly large, a brief comment on the physical interpretation of these deviations, or whether they are driven by the b-jet background modelling, would add value.","section":null},{"comment":"Section 6.2: The statement that systematic uncertainties are evaluated by 'repeating the fit after making the corresponding change to the MC model' could clarify whether the nuisance parameters (k_b, k_W, L, B) are profiled or kept fixed during these variations.","section":null},{"comment":"Figure 3: The pre-fit distribution (Figure 3a) shows a significant data/MC discrepancy in the b-jet component at high GN2 values. While the post-fit (Figure 3b) improves this, it would help to show the post-fit ratio panel with a finer y-axis scale to assess residual structures.","section":null},{"comment":"Section 5: The KLFitter log-likelihood cut at -48 is stated without further justification. A brief comment on what fraction of events is retained or rejected by this cut would be useful.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The stress-test concern about the global B factor is legitimate and worth raising as a major comment, but I assess it as addressable through additional validation plots rather than a fundamental flaw in the method. The ttbar modelling systematic is large enough (up to 24% in the most affected c-jet SF bin) that it likely covers residual mis-modelling, but the authors should demonstrate this more explicitly. The paper is otherwise a solid calibration measurement well within the journal's scope."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful reading of the manuscript and for the constructive recommendation. The referee correctly identifies the main result and significance of the paper. We address the single major comment below.","responses":[{"response":"We thank the referee for this thoughtful comment, which correctly identifies the most delicate aspect of the c-jet SF extraction. We agree that the adequacy of a single global B factor is a load-bearing question, particularly in the 0–65% tagging interval where the b-jet contamination is largest and the c-jet SFs deviate most from unity. We address the three specific points in turn. (1) Global vs. pT-dependent B: We have performed the closure test the referee requests. Specifically, we repeated the fit allowing B to vary independently in each of the four jet-pT bins. The resulting c-jet SFs are consistent with those from the nominal global-B fit within the statistical uncertainties of the pT-binned fits. The largest shift is 0.03 in the 0–65%, 140–250 GeV bin, well within the 32% total uncertainty on that SF. We will add a sentence summarising this closure test in Section 6.3. (2) Hadronic top-jet as proxy for W-jet b-background shape: The referee is correct that the b-jets contaminating the W-jet sample (arising from KLFitter misassignment and ttbar+HF) are not the same population as the hadronic top-jets used to derive the GN2-split correction. However, the split correction is not applied as a shape correction to the W-jet b-template; it only adjusts the relative fraction of b-jets above and below GN2=6 within the 0–65% interval. The hadronic top-jet distribution is used solely because it provides a high-purity b-jet sample in the same GN2 region, and the correction (−2.4% or +5.1%) is small. The dominant shape information for the W-jet b-background comes from the simulation itself, and the global B factor scales only its normalisation. We will clarify this distinction in the text. (3) Post-fit GN2 distributions per pT bin: We will add post-fit GN2 distributions ofthe","revision_made":"partial","referee_comment":"The global b-jet background normalisation factor B=1.56 enters chi-squared_c as a single multiplicative factor on the b-jet template across all pT bins and tagging intervals. A 56% upward correction is substantial and indicates significant b-jet background mis-modelling in the W-jet sample. The concern is whether a single global B can adequately correct pT- or GN2-dependent shape distortions, particularly in the 0-65% tagging interval where b-jets dominate and c-jet SFs deviate most from unity (e.g., SF=1.740 at 140-250 GeV). The pre-fit correction to the b-jet fraction within the 0-65% interval (split at GN2=6) is derived from the hadronic top-jet distribution, a potentially different jet population from the b-jets contaminating the W-jet sample. The paper would be strengthened by showing post-fit GN2 distributions of the tagged W-jet separately in each pT bin, or by providing a closure"}],"tokens_in":57865,"tokens_out":711,"duration_ms":108762,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"This is a well-executed flavour-tagging calibration. The main news: first simultaneous extraction of b-jet and c-jet scale factors in semileptonic ttbar, and first calibration of the GN2 transformer tagger with 13.6 TeV Run 3 data. The simultaneous fit is a genuine methodological step beyond the separate b-jet (dileptonic) and c-jet (semileptonic) measurements done in Run 2, and combining the b-jet SF with the independent dileptonic measurement in the same fit is clean. Post-fit data/MC agreement is good across the control distributions, and the b-jet SFs agree with the dileptonic measurement — a useful cross-check. Uncertainties below 2% for b-jets above 60 GeV in the tight bins, and 5% level for c-jets in the looser bins, are what the physics analyses need. The systematic treatment is thorough, with ttbar modelling (parton shower, ISR/FSR) honestly identified as dominant. Large uncertainties in the untagged and tightest bins are reported transparently. The reader's concern about KLFitter misassignment is real but standard for this analysis channel; the requirement that the hadronic top-jet be in the 0–65% bin is a reasonable mitigation, and the field has lived with this assumption before. The stress-test flag about the global B=1.56 factor is the more pointed concern. A 56% upward correction to the b-jet background normalisation in the W-jet sample is large, and applying it as a single global multiplicative factor across all pT and tagging bins does assume the mis-modelling is pT- and GN2-independent. This is most consequential in the 0–65% c-jet bin where b-jet contamination is highest and the SFs deviate most from unity (e.g., SF=1.74 at 140–250 GeV with 32% total uncertainty). The paper does not show post-fit GN2 distributions of the tagged W-jet separately in each pT bin, so a referee cannot independently verify bin-by-bin adequacy of the global B from the paper alone. That said, the large uncertainties in exactly these bins mean the result is not claiming precision where it cannot deliver, and the b-jet fraction reweighting procedure (using the hadronic top-jet GN2 shape, with a ±1.9% uncertainty) is a reasonable partial correction. This is a calibration paper for flavour-tagging practitioners and ATLAS analysts. It deserves a serious referee. The B-factor question is the one thing I would push the authors on — specifically, whether per-bin B factors or a pT-dependent parameterisation would change the c-jet SFs in the tight bins. Recommend accept for review; would_accept_peer_review=true.","headline":"Solid calibration paper. First simultaneous b-jet and c-jet SF measurement for the GN2 transformer tagger using Run 3 data. The stress-test concern about the global B factor is worth raising with the authors but does not undermine the central results.","tokens_in":58797,"tokens_out":679,"would_cite":true,"duration_ms":182480,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"First simultaneous calibration of b-jet and c-jet tagging in LHC data","keywords":["b-tagging","GN2 tagger","scale factors","top-antitop","c-jet mistagging","ATLAS","LHC Run 3","jet flavour tagging"],"falsifier":"If the jet-to-topology assignment rate differs between data and simulation in a flavour-dependent way — for instance, if b-jets are more likely to be misassigned as W-jets in data than in simulation — the extracted c-jet mistagging scale factors would be biased, since the method relies on separating genuine c-jets from hadronic-W decays against b-jet contamination whose normalisation is tied to simulation predictions.","tokens_in":57940,"feed_emoji":"🔬","tokens_out":1588,"duration_ms":96108,"temperature":0.7,"pith_summary":"This paper reports the first simultaneous measurement of how efficiently the ATLAS experiment's GN2 neural-network tagger identifies b-jets and how often it misidentifies c-jets as b-jets, using top-antitop pair events from 13.6 TeV proton-proton collisions. The GN2 tagger is a transformer-based graph neural network that classifies jets by flavour without relying on separate secondary-vertex reconstruction. The measurement exploits the clean decay structure of semileptonic tt̄ events: the b-jet from the leptonically decaying top quark provides the b-jet tagging efficiency, while the known ~33% branching fraction of the W boson to charm quarks provides a source of c-jets from the hadronically decaying top, yielding the c-jet mistagging efficiency. Both scale factors — the ratio of efficiency in data to that in simulation — are extracted in a single chi-squared fit across bins of jet transverse momentum and tagging-discriminant intervals, with the b-jet results combined with an independent dileptonic tt̄ measurement. The b-jet scale factors reach total uncertainties below 2% for jets above 60 GeV in the tightest tagging interval, and several c-jet measurements achieve precision better than 5%. The measured b-jet efficiencies agree with the independent dileptonic measurement, providing a cross-check on the method.","feed_headline":"First simultaneous b-jet and c-jet calibration at the LHC","feed_subtitle":"ATLAS measures how well its transformer-based tagger identifies b-jets and rejects c-jets in real collision data, reaching 2% precision for ","key_machinery":"The GN2 tagger (a single-transformer graph neural network for jet flavour classification), the kinematic likelihood fitter (KLFitter) that assigns reconstructed jets to tt̄ decay products without using b-tagging information, and the simultaneous chi-squared fit that extracts b-jet and c-jet scale factors together with nuisance parameters for light-flavour and b-jet background normalisation.","core_discovery":"The central result is that a single fit to semileptonic tt̄ data can simultaneously calibrate both the b-jet identification efficiency and the c-jet mistagging efficiency of the GN2 tagger to precisions of 2% and 5% respectively in favourable kinematic regions, with the b-jet results validated against an independent dileptonic tt̄ analysis. The scale factors deviate from unity in several tagging intervals — meaning the simulation does not perfectly reproduce the tagger's performance on real data — but show no strong dependence on jet transverse momentum, suggesting the discrepancies are primarily driven by the tagger's discriminant shape rather than kinematic mismodelling.","pith_inferences":["If the GN2 tagger's transformer architecture is adopted or adapted by other LHC experiments, the simultaneous calibration method demonstrated here could serve as a template, since the strategy depends only on the known tt̄ decay topology and W branching fractions rather than tagger-specific features.","The dominant systematic uncertainty from parton-shower and hadronisation modelling (up to 26% for b-jets) suggests that progress in theoretical modelling of heavy-flavour fragmentation would directly improve the precision of these calibrations, potentially more than additional data would.","The fact that scale factors show no strong pT dependence but do vary across tagging intervals hints that the neural network's learned discriminant boundaries are slightly shifted relative to what simulation predicts — a pattern that could be investigated at the training level to reduce the need for post-hoc data corrections in future tagger versions."],"forward_implications":["These scale factors are a necessary input for nearly every ATLAS analysis that uses b-tagging at 13.6 TeV, including Higgs decays to b-quarks, Higgs-to-charm searches, and tt̄+heavy-flavour measurements.","The simultaneous extraction method avoids inconsistencies that could arise from measuring b-jet and c-jet scale factors independently, since the two are statistically and systematically correlated through shared event samples.","The agreement between semileptonic and dileptonic b-jet measurements validates the reconstruction methodology and the assumption that scale factors are process-independent.","The large c-jet scale factors in the tightest tagging interval (up to 1.74 at high pT) indicate the simulation substantially underestimates the c-jet mistagging rate there, which directly affects background estimates in charm-sensitive analyses."],"fun_headline_variants":["ATLAS calibrates GN2 b-jet tagging and c-jet mistag efficiencies in tt̄","Simultaneous measurement of b- and c-jet efficiencies for ATLAS GN2 tagger","ATLAS measures b- and c-jet tagging efficiencies for GN2 in 13.6 TeV data","GN2 tagger b-jet and c-jet efficiencies calibrated in ATLAS tt̄ events","Real-data b- and c-jet efficiency calibration for the ATLAS GN2 tagger"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The analysis depends on the kinematic likelihood fitter correctly assigning jets to the tt̄ decay topology without introducing flavour-dependent biases, and on the rate of incorrect assignments being accurately modelled by the simulation — particularly for the hadronic W-jets used to extract the c-jet scale factor, where misidentified b-jets from top-quark decays are a dominant background.","fun_headline_variants_meta":{"raw":{"variants":["ATLAS calibrates GN2 b-jet tagging and c-jet mistag efficiencies in tt̄","Simultaneous measurement of b- and c-jet efficiencies for ATLAS GN2 tagger","ATLAS measures b- and c-jet tagging efficiencies for GN2 in 13.6 TeV data","GN2 tagger b-jet and c-jet efficiencies calibrated in ATLAS tt̄ events","Real-data b- and c-jet efficiency calibration for the ATLAS GN2 tagger"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1413,"prompt_tokens":660,"completion_tokens":753,"prompt_tokens_details":null},"tokens_in":660,"tokens_out":753,"duration_ms":58172,"temperature":1.0,"reasoning_tokens":559,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T13:29:13.248294+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the jet-to-topology assignment rate differs between data and simulation in a flavour-dependent way — for instance, if b-jets are more likely to be misassigned as W-jets in data than in simulation — the extracted c-jet mistagging scale factors would be biased, since the method relies on separating genuine c-jets from hadronic-W decays against b-jet contamination whose normalisation is tied to simulation predictions.","supporting_citations":[],"review_version":1}