{"id":"68a56b21-062a-4c54-bb52-eb0dc2759b1a","arxiv_id":"2607.06353","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"The average muon detection efficiency of 137 CMS GE1/1 triple-GEM chambers was measured at 93.3% using 2023 collision data, with defect-free chambers reaching ~96%, confirming readiness for HL-LHC operation.","lead":"The CMS collaboration measured the detection efficiency of 137 newly installed GEM muon detectors using 2023 LHC collision data, finding an average efficiency of 93.3%, rising to ~96% for chambers without hardware defects. This validates the detector technology for the upcoming High-Luminosity LHC era.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"No uncertainty quantification on central efficiency values, and the abstract misstates which chamber subset achieves ~96% efficiency (108 vs 78 per Table 2).","rationale":"The reader correctly identified the back-propagation/matching-window concern and the absence of a systematic uncertainty budget as the primary weakness. I broaden this slightly: the most load-bearing issue is the complete absence of any uncertainty quantification — not just the extrapolation bias, but also statistical errors, matching-window dependence, and background contamination. The reader also caught the abstract inconsistency regarding 108 vs 78 chambers, which is a factual error that should be corrected. The CONDITIONAL verdict is appropriate: the methodology is sound (standard tag-and-probe-style back-propagation, well-established in CMS), the results are internally consistent across the body of the paper, and the fine-grained VFAT-level analysis (§3.5) and short-circuit impact study (§3.7) add real value. But without uncertainties, these are approximate characterizations rather than precision measurements. The paper is adequate as a commissioning report — which is essentially what it is, being the first full-system efficiency evaluation — but the efficiency values should not be cited as precise numbers until uncertainties are added. No adjustment to the verdict is needed.","tokens_in":12613,"tokens_out":2482,"duration_ms":136148,"concrete_test":"Recompute efficiency for the 78 optimal chambers using matching windows of ±2, ±3, ±4, and ±6 cm in RΔφ. If efficiency varies by more than 1% across these windows, the matching window choice is a dominant systematic. Simultaneously, compute per-chamber binomial statistical uncertainties (ε ± 1.96·√(ε(1-ε)/N)) and add them to Figs. 9 and 14. This would establish whether the ~96% value is precise enough to claim consistency with the 97% design target and whether the pile-up independence is statistically supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper presents ~93.3% and ~96% as point estimates with no error bars on any figure (Figs. 9, 10, 14), no binomial statistical uncertainties, and no systematic uncertainty budget. The key unquantified sources are: (1) the ±4 cm RΔφ matching window in §3.3 — the paper never studies efficiency vs. window size, so we cannot tell if the large window introduces background contamination or if a smaller window would give a different result; (2) track extrapolation bias from multiple scattering, alignment uncertainties, and magnetic field non-uniformity in the endcap — the back-propagation from ME1/1 to GE1/1 traverses material and field regions where systematic offsets could shift the residual distribution; (3) background contamination from delta rays or punch-through within the ±4 cm window. Without these uncertainties, one cannot assess whether ~96% is statistically consistent with the 97% design specification cited in §2.3, or whether the pile-up independence shown in Figs. 13/15 is statistically meaningful rather than just visually flat. Additionally, the abstract states '108 detectors that had no shorts were operated at the nominal HV working point with average efficiency of ~96%,' but Table 2 in §3.8 clearly shows 108 chambers at nominal HV of which 30 had shorts, leaving 78 good chambers — the 96% figure applies to the 78, not 108. This is a factual error in the abstract that misrepresents the result.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This paper reports the first full-system efficiency measurement of the CMS GE1/1 triple-GEM chambers using 2023 pp collision data at sqrt(s) = 13.6 TeV. The measurement uses a back-propagation technique: global muon tracks are re-fitted without GE1/1 hits and extrapolated to the GE1/1 surface, where a matched hit within a +/-4 cm R*Delta*phi window defines a successful detection. The average efficiency across 137 operational chambers is ~93.3%, and for a subset of 78 chambers without electrical shorts operated at nominal high voltage, the average efficiency is ~96%, meeting the design specification. Efficiency is found to be independent of pile-up. The methodology is sound and the results represent an important milestone for the CMS muon upgrade program.","tokens_in":13396,"tokens_out":1273,"duration_ms":252164,"significance":"This is the first comprehensive efficiency characterization of the full GE1/1 system installed in CMS, and as such it is a significant contribution to the detector physics literature. The measurement validates the GE1/1 design specification (~97% efficiency) for the subset of optimally operating chambers, which is a key result for the HL-LHC muon program. The VFAT-level efficiency maps and the characterization of short-circuit impacts provide valuable operational diagnostics. The back-propagation methodology using an external tracking benchmark (CSC+tracker tracks without GEM hits) is a standard and appropriate tag-and-probe approach, and the result is free of fitted parameters or circular logic. The paper would be substantially strengthened by the addition of a quantitative uncertainty budget, which is currently absent.","major_comments":[{"comment":"Abstract: The abstract states 'A subset of 108 detectors that had no shorts were operated at the nominal HV working point with average efficiency of ~96%.' However, Table 2 in Section 3.8 shows that 108 chambers were operated at nominal HV, of which 30 had shorts, leaving 78 good chambers. The ~96% efficiency figure applies to the 78 good chambers, not 108. This is a factual error in the abstract that misrepresents the central result and must be corrected.","section":null},{"comment":"Sections 3.4, 3.8, and all efficiency figures (Figs. 9, 10, 13, 14, 15): No uncertainties are reported on any efficiency value. The paper presents ~93.3% and ~96% as point estimates with no statistical or systematic uncertainties. At minimum, binomial statistical uncertainties should be reported for the average efficiencies and on the per-chamber values in Fig. 9. Without any uncertainty, one cannot assess whether the ~96% result is statistically consistent with the 97% design specification cited in Section 2.3, or whether the pile-up independence shown in Figs. 13/15 is statistically meaningful rather than visually flat. This is load-bearing for the central claims.","section":null},{"comment":"Section 3.3: The +/-4 cm R*Delta*phi matching window is the sole parameter controlling hit-track association, yet no systematic study of the efficiency dependence on window size is presented. A scan of efficiency versus matching window (or at least a comparison of results for a smaller window, e.g. +/-2 cm) would demonstrate that the result is not sensitive to background contamination from delta rays or punch-through within the window. Additionally, no systematic uncertainty is assigned for potential track extrapolation biases from multiple scattering, alignment uncertainties, or magnetic field non-uniformities in the endcap region. A quantitative assessment of these sources, or at minimum an estimate of their magnitude, is needed to validate the efficiency value itself.","section":null}],"minor_comments":[{"comment":"Section 3.3, Fig. 8: The residual R*Delta*phi distributions are described as being shown for 'entire rings of chambers' but the caption mentions 'four individual chambers.' Please clarify whether these are per-chamber or per-ring distributions.","section":null},{"comment":"Section 3.7: The efficiencies for chambers with 1, 2, 3, and 4 shorts (~89%, ~79%, ~69%, ~54%) are presented without specifying whether these are averages and without uncertainty estimates. Please state the number of chambers in each category and add uncertainties.","section":null},{"comment":"Fig. 9: The shaded chambers with shorts are difficult to distinguish from unshaded ones in the current plot. Consider using a more distinct marker or color for better clarity.","section":null},{"comment":"Section 2.3: The design specification is stated as '97% efficiency for detecting minimum ionizing particles.' Please clarify in Section 3.8 whether the ~96% measured efficiency is directly comparable to this specification (same definition, same particle spectrum) or whether differences in methodology should be noted when the comparison is made.","section":null},{"comment":"Section 3.6, Fig. 13: The pile-up independence claim would be strengthened by reporting the slope of a linear fit or the chi-square of a flat-line hypothesis, rather than relying on visual inspection alone.","section":null},{"comment":"The paper would benefit from a brief statement of the total number of propagated tracks and matched hits used in the efficiency calculation, and the pT threshold of the trigger (24 GeV) versus the offline selection (pT > 10 GeV) should be clarified to explain the track sample composition.","section":null}],"recommendation":"major_revision","confidential_remarks":"The factual error in the abstract (108 vs 78 chambers) is a clear oversight that must be fixed, but the more substantive issue is the complete absence of an uncertainty budget. For a JINST detector performance paper, reporting efficiency point estimates without any statistical or systematic uncertainties is below the standard of the field. The methodology itself is sound and the results are valuable, so I expect the authors can address these issues within a revision cycle. The paper is appropriate in scope for JINST."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful reading of the manuscript and for the constructive and well-targeted comments. All three major points are well taken. Below we address each in turn.","responses":[{"response":"The referee is correct. This is a factual error in the abstract. The number 108 refers to the total number of chambers operated at nominal HV (column A in Table 2), which includes 30 chambers with shorts. The ~96% average efficiency applies to the 78 chambers that were operated at nominal HV and had no shorts (column A−B). The abstract will be corrected to read: 'A subset of 78 detectors that were operated at the nominal HV working point and had no electrical shorts showed an average efficiency of ~96%.' We will also review the full text for any other instances where the 108 and 78 figures might be conflated.","revision_made":"yes","referee_comment":"Abstract: The abstract states 'A subset of 108 detectors that had no shorts were operated at the nominal HV working point with average efficiency of ~96%.' However, Table 2 shows that 108 chambers were operated at nominal HV, of which 30 had shorts, leaving 78 good chambers. The ~96% efficiency applies to the 78, not 108. This is a factual error."},{"response":"We agree that the absence of uncertainties is a significant omission that weakens the paper's central claims. We will add a quantitative uncertainty budget in the revised manuscript. Specifically: (1) For the average efficiencies (93.3% and 96%), we will report binomial statistical uncertainties computed as sqrt(ε(1−ε)/N), where N is the number of probe tracks. Given the large sample sizes from 17.8 fb⁻¹ of Z-enriched data, these statistical uncertainties are expected to be at the level of a few ×10⁻³ or smaller. (2) For the per-chamber efficiencies in Fig. 9, we will add binomial error bars. (3) For the pile-up dependence plots (Figs. 13 and 15), we will add statistical error bars to each bin so that the flatness of the efficiency versus pile-up can be assessed quantitatively. (4) We will also add a systematic uncertainty component, as discussed in our response to the third comment below. The total uncertainty will be quoted with each central value, and the consistency with the 97% design specification will be explicitly discussed.","revision_made":"yes","referee_comment":"No uncertainties are reported on any efficiency value, including the ~93.3% and ~96% averages and per-chamber values in Fig. 9. Binomial statistical uncertainties should be reported at minimum."},{"response":"We agree that a systematic assessment of the matching window dependence and extrapolation-related biases is needed. In the revised manuscript we will add the following: (1) A scan of average efficiency versus matching window size, comparing results at ±2 cm, ±3 cm, ±4 cm, and ±5 cm. Based on the residual distributions shown in Fig. 8, which are well-contained within ±4 cm with prominent central peaks, we expect the efficiency to be stable for windows of ±3 cm and larger, with a possible slight decrease at ±2 cm if the window begins to exclude genuine matches near the tails. This scan will be included as a new figure and discussed in Section 3.3. (2) The difference in efficiency between the ±4 cm and ±2 cm windows will be taken as a systematic uncertainty associated with the matching criterion. (3) For extrapolation biases, we will estimate the systematic uncertainty from multiple scattering and alignment by examining the residual distribution widths and their centrality (the residuals in Fig. 8 are well-centered at zero, indicating no significant bias), and by noting that the back-propagation uses the full Kalman-filter track fit including material effects. We will provide a summary table of systematic uncertainty sources and their estimated magnitudes. We note that a full alignment-based systematic study using dedicated alignment constants is beyond the scope of this paper, but the residual distributions themselves provide a direct empirical constraint on the magnitude of any extrapolation bias, and we will state this explicitly.","revision_made":"partial","referee_comment":"No systematic study of efficiency dependence on the ±4 cm RΔφ matching window is presented, and no systematic uncertainty is assigned for track extrapolation biases from multiple scattering, alignment, or magnetic field non-uniformities."}],"tokens_in":12411,"tokens_out":1385,"duration_ms":156538,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The bottom line: this is the first full-system efficiency measurement of the installed CMS GE1/1 chambers using 2023 collision data, and it shows the system is working. The 78 chambers without shorts at nominal HV average ~96% efficiency, meeting the TDR design spec. The pile-up independence is a useful confirmation. This is a solid commissioning result and the measurement technique — back-propagation of tracker+CSC tracks without GEM hits, tag-and-probe with Z muons, ±4 cm RΔφ matching — is well-established and correctly applied here. The VFAT-level efficiency maps and the systematic study of how electrical shorts degrade performance (one short → ~89%, two → ~79%, etc.) are genuinely useful for the collaboration. Credit is earned for the granularity of the analysis and the honest reporting of operational issues like opto-hybrid failures and HV instabilities. The paper does not oversell what it has; the framing as a commissioning validation is appropriate. The central numbers are credible as approximate characterizations of detector performance. Now the soft spots. The most important is the complete absence of a systematic uncertainty budget. There are no error bars on any efficiency figure, no binomial statistical uncertainties quoted, no study of efficiency versus matching window size, and no quantification of extrapolation bias from multiple scattering, alignment, or magnetic field non-uniformity. The ±4 cm window is never validated against a scan. Without these, one cannot assess whether ~96% is statistically consistent with the 97% design target, or whether the pile-up independence is real or just visually flat. This is the main thing a referee should push on. Second, the abstract contains a factual error. It says '108 detectors that had no shorts were operated at the nominal HV working point with average efficiency of ~96%.' Table 2 in §3.8 shows 108 chambers at nominal HV, of which 30 had shorts, leaving 78 good chambers. The ~96% applies to the 78, not 108. This is a straightforward misstatement that should be corrected. The stress-test note flagged this correctly. I checked the text: §3.7 says 'Out of 108 chambers in both endcaps operated at nominal HV, 30 had shorts,' and §3.8 confirms 78 good chambers. The abstract is wrong. These issues are fixable. The measurement itself is sound, the methodology is standard, and the result — that GE1/1 meets its design efficiency — is important for the CMS upgrade program. This paper is for detector physicists and collaboration members who need to know the GE1/1 system is ready for HL-LHC. It deserves a serious referee who should require the uncertainty budget and the abstract correction before acceptance.","headline":"First full-system GE1/1 efficiency measurement with Run-3 data; missing uncertainty budget and abstract misstatement need fixing.","tokens_in":14406,"tokens_out":628,"would_cite":false,"duration_ms":139356,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"New CMS muon detectors hit 96% efficiency in first collision test","keywords":["CMS","GEM detector","GE1/1","muon detection efficiency","triple-GEM","LHC Run 3","pile-up","back-propagation"],"falsifier":"If future measurements with improved track reconstruction or independent calibration sources (e.g., cosmic ray muons, or tracks reconstructed with GE1/1 included and then cross-checked) yield efficiencies that differ systematically from the back-propagation results by more than a few percent, the method's assumption of unbiased extrapolation would be challenged.","tokens_in":12685,"feed_emoji":"","tokens_out":1648,"duration_ms":213454,"temperature":0.7,"pith_summary":"This paper reports the first full-system measurement of how efficiently the newly installed GE1/1 triple-GEM muon detectors in the CMS experiment register muon tracks, using 17.8 inverse femtobarns of 2023 proton-proton collision data at 13.6 TeV. The authors use a back-propagation technique: they reconstruct muon tracks from other detectors (silicon tracker, cathode strip chambers, resistive plate chambers) without any GE1/1 hits, then project those tracks onto the GE1/1 surface and check whether a GE1/1 hit appears within 4 cm of the predicted position. Across 137 operational chambers, the average detection efficiency is 93.3%. For the 78 chambers free of electrical shorts and operated at their designed high voltage, the average efficiency rises to 96%, meeting the 97% design target closely enough to validate the system for the High Luminosity LHC era. Efficiency shows no dependence on pile-up—the number of simultaneous proton-proton collisions per bunch crossing—up to the levels observed in 2023 data. The dominant causes of suboptimal efficiency are electrical shorts on GEM foils (each short drops a chamber's efficiency by roughly 7-10 percentage points) and optical readout failures from adhesive outgassing on opto-hybrid boards. The paper argues that with hardware refurbishment planned for the 2026-2030 long shutdown, the system will meet its performance goals.","feed_headline":"New CMS muon detectors hit 96% efficiency in first collision test","feed_subtitle":"First full-system check of 137 GE1/1 triple-GEM chambers shows design-spec performance when hardware faults are excluded, with no efficiency","key_machinery":"The triple-GEM chamber: a gaseous ionization detector using three stacked polyimide foils perforated with microscopic holes, each held at successively lower amplification voltage, immersed in an Ar/CO2 gas mixture. Ionization electrons drift through the foils, undergo avalanche multiplication at each stage, and induce a readout signal on strip electrodes. The back-propagation method is the measurement tool: muon tracks reconstructed from other subdetectors are extrapolated to the GE1/1 surface, and a hit within a 4 cm window in RΔφ counts as a successful detection.","core_discovery":"The central finding is that the GE1/1 triple-GEM chamber system, newly installed in the CMS forward muon region, achieves an average per-chamber muon detection efficiency of approximately 96% when operating under nominal conditions (no shorts, full high voltage), measured for the first time with real LHC collision data. This is within reach of the 97% design specification. The efficiency is stable against pile-up, meaning the detectors do not lose performance as collision density increases. The gap between the 93.3% all-chamber average and the 96% best-case average is attributable to identifiable hardware problems—foil shorts and optical connection failures—rather than any fundamental limit,","pith_inferences":["If the 32 chambers with shorts and the 7 turned-off chambers are refurbished during Long Shutdown 3, the system-wide average efficiency could rise from 93.3% to approximately 96%, assuming no new faults develop. This would bring the full system within 1 percentage point of the 97% design goal.","The pile-up independence was measured at 2023 pile-up levels (roughly 50-70 interactions per crossing). Whether it holds at HL-LHC levels (140-200) depends on whether the dominant inefficiency sources—readout bandwidth, cluster overlap, and electronic deadtime—scale with hit rate rather than with the number of interactions. The paper's data cannot confirm this extrapolation directly.","The ~1% gap between the 96% measured efficiency and the 97% design target may reflect the combined effect of the fiducial region definition, the 4 cm matching window, and residual track extrapolation uncertainties. A tighter matching window or improved alignment could narrow or close this gap, but the paper does not decompose the residual inefficiency into these components.","The fact that optical connection failures from adhesive outgassing caused entire VFAT sectors to be excluded suggests that the readout chain, not the GEM gas amplification itself, is the primary operational risk for the system. This implies that reliability engineering on the opto-hybrid boards may matter more for long-term performance than further optimization of chamber high-voltage settings."],"forward_implications":["The 96% efficiency for fault-free chambers at nominal voltage validates the GE1/1 design for the High Luminosity LHC, where muon fluxes in the forward region will be roughly an order of magnitude higher than in Run 3.","The pile-up independence observed in 2023 data, if it persists at HL-LHC pile-up levels (140-200 interactions per crossing), would mean GE1/1 can serve as a reliable trigger and reconstruction station without efficiency corrections for collision density.","The clear correlation between foil shorts and efficiency loss (roughly 7-10% per short) provides a quantitative basis for quality control: chambers with shorts can be identified and refurbished before HL-LHC operation, potentially bringing the system-wide average close to the 96-97% target.","The VFAT-level efficiency map identifies specific readout sectors with reduced performance, enabling targeted repairs during Long Shutdown 3 rather than full chamber replacement.","The back-propagation method demonstrated here can be reused for future GE1/1 efficiency monitoring and extended to the GE2/1 and ME0 GEM stations currently under construction."],"fun_headline_variants":["CMS GE1/1 chambers achieve 96% muon efficiency in 2023 collisions","New CMS GEM detectors reach 96% efficiency, unaffected by pile-up","First collision test of CMS GE1/1 muon chambers shows 96% efficiency","CMS GE1/1 detectors reach 96% muon efficiency under nominal conditions"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The back-propagation method assumes that muon tracks reconstructed without GE1/1 hits provide an unbiased and accurate prediction of where each muon actually crossed the GE1/1 surface, and that a hit within 4 cm of that prediction is a genuine GEM detection rather than background. If track extrapolation has systematic biases from multiple scattering, detector alignment errors, or magnetic field non-uniformities in the endcap region, the measured efficiency could be shifted in","fun_headline_variants_meta":{"raw":{"variants":["CMS GE1/1 chambers achieve 96% muon efficiency in 2023 collisions","New CMS GEM detectors reach 96% efficiency, unaffected by pile-up","First collision test of CMS GE1/1 muon chambers shows 96% efficiency","CMS GE1/1 detectors reach 96% muon efficiency under nominal conditions"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1164,"prompt_tokens":562,"completion_tokens":602,"prompt_tokens_details":null},"tokens_in":562,"tokens_out":602,"duration_ms":61581,"temperature":1.0,"reasoning_tokens":470,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T08:37:21.768528+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If future measurements with improved track reconstruction or independent calibration sources (e.g., cosmic ray muons, or tracks reconstructed with GE1/1 included and then cross-checked) yield efficiencies that differ systematically from the back-propagation results by more than a few percent, the method's assumption of unbiased extrapolation would be challenged.","supporting_citations":[],"review_version":1}