{"id":"db59448f-e1c9-4bbc-b202-c76d22f917fa","arxiv_id":"1908.04589","paper_version":5,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"At a surface RPC stack in Madurai, the rates of cosmic ray events with two, three, or four parallel muons exceed CORSIKA/QGSJET predictions by factors of roughly 3, 5, and 6 to 8.","lead":"A 12-layer stack of glass resistive plate chambers at IICHEP-Madurai recorded around 250 million cosmic ray events and counted how often two, three, or four parallel muons arrive together. The measured rates are several times higher than CORSIKA air-shower simulations predict, adding a new data point to the unresolved muon multiplicity problem in cosmic ray physics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CORSIKA comparison is truncated at 10^6 GeV; if high-multiplicity events are driven by primaries above this cutoff, the claimed discrepancy may be an artifact.","rationale":"The reader's weakest-assumption point about detector simulation fidelity is reasonable, but the more load-bearing and more easily testable issue is the hard truncation of the CORSIKA primary-energy range at 10^6 GeV (Section 3.1). High-multiplicity muon events in a small surface detector are expected to be biased toward high primary energies; without knowing the primary-energy spectrum of the accepted multi-track events, the simulation prediction is incomplete. The proposed check, extending the energy range and re-evaluating the fractions, would settle whether the excess is physical or an artifact of the cutoff. The verdict remains CONDITIONAL because the paper can be revised with this sensitivity study; the conclusion should be tempered until the check is performed.","tokens_in":11213,"tokens_out":6669,"duration_ms":73712,"concrete_test":"Rerun the CORSIKA stage with the same settings but extend the maximum primary energy from 10^6 GeV to 10^8 GeV, keeping the spectral index and composition weighting fixed, and pass the resulting particles through the identical GEANT4 and reconstruction chain. Compute the composition-weighted fractions for 2, 3, and 4 tracks and compare them with the current Table 2 values. As a diagnostic, also histogram the primary energy of the simulated 2-, 3-, and 4-track events that survive all cuts and report the fraction lying within a factor of 3 of the 10^6 GeV upper edge. If the high-energy extension raises the 4-track fraction by a factor of 6 or more, the claimed discrepancy is not robust evidence for new hadronic physics.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central comparison rests on CORSIKA showers generated with a fixed power-law spectrum E^-2.7 over a primary-energy range of 10-10^6 GeV, as stated in Section 3.1. The paper never reports the primary-energy distribution of the simulated events that actually pass the 2-, 3-, and 4-track selection cuts. This matters because the rate of high muon multiplicity in a 2 m x 2 m detector grows steeply with primary energy, and the observed 4-track fraction (1.94e-8) is roughly six times the truncated-simulation prediction (3.21e-9). If most accepted 3- and 4-track events in the simulation lie within about an order of magnitude of the 10^6 GeV upper edge, then extending the generated energy range to 10^8 GeV, with the same spectral index and composition, could raise the predicted fractions by factors comparable to the reported excess. In that case the statement that 'there is a discrepancy between the observed data and predictions from the cosmic ray particle spectrum, the CORSIKA and finally the GEANT4 simulation' would not be established; the discrepancy would instead reflect an energy-range truncation. The paper's own remark that a significant fraction of high-multiplicity events come from primaries beyond current collider energies reinforces the need to check the energy coverage explicitly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using a 12-layer stack of 2 m x 2 m glass RPCs at IICHEP-Madurai, the authors analyze roughly 250 million cosmic-ray triggered events (206 million with at least one reconstructed track). They reconstruct tracks with a Hough transform and select events with 2, 3, or 4 parallel tracks using a 2.5-degree skewed-angle cut. The measured fractions relative to single-track events are (6.35 +/- 0.05) x 10^-5, (5.82 +/- 0.53) x 10^-7, and (1.94 +/- 0.97) x 10^-8. These are compared with CORSIKA v7.6300 simulations using an E^-2.7 spectrum over 10-10^6 GeV, six primary species, and two QGSJET hadronic models, followed by a GEANT4 detector simulation that includes measured detector parameters. The paper finds that the data exceed the simulations by factors of roughly 2.7, 5, and 6-8 for 2, 3, and 4 tracks, respectively, and concludes that the EAS simulation cannot reproduce the observed multi-muon rates.","tokens_in":11443,"tokens_out":6395,"duration_ms":63677,"significance":"The measurement is potentially valuable: the data set is large, the detector response is calibrated from data, and the reconstructed angular distributions match the simulation, giving confidence in the basic reconstruction pipeline. If the discrepancy survives the additional checks described below, it would provide a useful ground-level constraint on high-energy hadronic interaction models and on the muon content of air showers. However, the paper does not currently establish the discrepancy quantitatively because of an unexamined energy cutoff in the CORSIKA sample and the absence of a systematic uncertainty budget for the simulation prediction.","major_comments":[{"comment":"The CORSIKA primary-energy range is stated as 10-10^6 GeV, but the paper does not report the primary-energy distribution of simulated events that survive the 2-, 3-, and 4-track selections. The rate of high-multiplicity events rises steeply with primary energy, and the manuscript itself notes in Section 5 that a significant fraction of the interactions responsible for high multiplicities are beyond current collider energies. If the selected simulated events cluster near the 10^6 GeV boundary, extending the range to 10^7-10^8 GeV with the same spectrum and composition could raise the predicted fractions by factors comparable to the reported excess. The paper must either show that the selected events are safely below the cutoff or extend the simulation range; otherwise the claimed discrepancy may be an artifact of the truncation.","section":"Section 3.1, Tables 1-2"},{"comment":"The statement that 'Systematic error due to uncertainties of roof thickness, material in the detector setup, strip multiplicity, noise, efficiencies and the physics models used in GEANT4 are much smaller than the observed discrepancy' is not supported by any quantitative estimate. In particular, the primary-composition weights, the spectral index gamma, the low-energy hadronic model, and the parallel-track angle cut are not varied. Since the Introduction identifies composition and spectral index as dominant factors for multiplicity, a scan over these inputs is required before the discrepancy can be considered model-independent.","section":"Section 5, paragraph after Table 2"},{"comment":"For the 4-track fraction, the data value (1.94 +/- 0.97) x 10^-8 and the QGSJET-II-04 prediction (3.21 +/- 0.87) x 10^-9 differ by only about 1.7 standard deviations when the quoted errors are combined; the QGSJET01d prediction gives a similar conclusion. The paper therefore overstates the evidence for a 4-track discrepancy. This limited significance should be stated explicitly and the conclusions adjusted accordingly.","section":"Table 2"}],"minor_comments":[{"comment":"The header 'QGSJET-II-042' appears to be a typo for 'QGSJET-II-04'; please correct it.","section":"Table 1"},{"comment":"The phrase 'within one order of magnitude less' is imprecise; Table 2 shows factors of about 2.7, 5.2, and 6-8 for 2, 3, and 4 tracks, respectively.","section":"Section 5"},{"comment":"The statement that 'the maximum number of tracks reconstructed in an event is 4' should clarify whether this is a hard algorithmic cap and, if so, how it affects the comparison of data and simulation.","section":"Section 4"},{"comment":"Reference [15] is from 1969; more recent cosmic-ray composition measurements would be more appropriate for weighting the simulated primary species.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The main risk is that the energy cutoff in the CORSIKA sample explains part or all of the discrepancy. I recommend major revision rather than rejection because the missing energy-range check and systematic-error estimate are obtainable within the scope of a revision. The authors should also be asked to report the statistical significance of each track-multiplicity comparison separately, since the 4-track excess is not statistically significant as quoted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [name],\n\nQuick take on arXiv:1908.04589. The paper reports measured fractions of 2-, 3-, and 4-track cosmic ray events in a 2x2 m^2 RPC stack at IICHEP-Madurai, using 206M reconstructed events. The 2- and 3-track fractions are (6.35±0.05)e-5 and (5.82±0.53)e-7, while a composition-weighted CORSIKA/QGSJET prediction gives 2.35e-5 and 1.12e-7. That is a factor 2.7 and 5.2 excess, consistent with the known muon puzzle. The 4-track fraction is 1.94e-8 vs 3.21e-9, but that is based on only ~4 observed events, so it is a weak hint.\n\nWhat the paper does well: the detector description is thorough, the efficiency maps, noise, and strip multiplicity are measured from data and put into the GEANT4 simulation. The Hough transform tracking and the parallel-track cut (2.5 degrees) are well motivated, and they make a good effort to reject random coincidences by time separation. The energy and angular distributions of reconstructed tracks match the simulation in shape, which gives confidence in the detector model. Citation pattern to previous multi-muon experiments is fair and complete.\n\nSoft spots. The MC comparison is not as clean as the data analysis. The primary composition, spectral index, and hadronic model are not varied; only QGSJET-II-04 and QGSJET01d are used, and those two look similar. No systematic uncertainty is attached to the composition weights or the spectral index. The paper's claim that the simulated high-multiplicity events come from primaries beyond collider reach is internally odd: the CORSIKA simulation is truncated at 10^6 GeV, which is actually below LHC-equivalent energies. If the true contributing energies extend beyond 10^6 GeV, the prediction could rise. The paper should report the primary energy distribution of accepted 2-, 3-, and 4-track events, or at least run a simulation with the upper bound pushed to 10^8 GeV. That is the most important missing check.\n\nThe 4-track excess is not convincing on its own, but the 2- and 3-track excesses look solid statistically. The lack of sensitivity studies means the central discrepancy is established only at the level of a particular simulation setup, not as a robust model-independent excess.\n\nWho it is for: cosmic ray physicists working on the muon puzzle and hadronic interaction models. A serious referee could ask for the sensitivity study and the energy distribution, but the paper is worth refereeing. My recommendation: send it to review, expect major revision focused on the MC systematics.","headline":"A careful surface muon-multiplicity measurement that reports a real-looking excess over CORSIKA, but the MC comparison is under-specified at the high-energy end and the 4-track bin is statistically thin.","tokens_in":12025,"tokens_out":7347,"would_cite":true,"duration_ms":70691,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 12-layer muon detector stack measured 2-, 3-, and 4-parallel-track cosmic-ray event rates that exceed air-shower simulation predictions by factors of about 3, 5, and 6.","keywords":["cosmic ray experiments","cosmic ray detectors","hadronic interaction models","resistive plate chambers","muon multiplicity","extensive air showers","CORSIKA simulation","GEANT4 simulation"],"falsifier":"Remove the 22 cm concrete roof from the GEANT4 geometry and rerun the full analysis on the same CORSIKA input: if the measured 2-track fraction drops from its observed $6.35\\times10^{-5}$ toward the simulated $2.35\\times10^{-5}$ while the single-track rate barely changes, the extra parallel tracks are produced by roof interactions that the simulation mishandles; if the factor-of-about-3 excess survives without the roof, the discrepancy is genuinely atmospheric.","tokens_in":11030,"feed_emoji":"🌌","tokens_out":14363,"duration_ms":131616,"temperature":0.7,"pith_summary":"Using roughly 206 million reconstructed cosmic-ray events from a 12-layer, 2 m by 2 m glass resistive plate chamber stack at ground level, this paper measures the normalized fractions of events with 2, 3, and 4 parallel tracks: $(6.35\\pm0.05)\\times10^{-5}$, $(5.82\\pm0.53)\\times10^{-7}$, and $(1.94\\pm0.97)\\times10^{-8}$. Feeding CORSIKA/QGSJET air showers into a GEANT4 model of the same detector gives composition-weighted fractions smaller by factors of about 3, 5, and 6. The authors argue that detector systematics are too small to explain the gap, so the discrepancy belongs to the chain of primary cosmic-ray composition, high-energy hadronic interaction models, and air-shower simulation. This matters because multi-muon event rates are a background and calibration ingredient for underground neutrino detectors and a rare ground-level handle on hadronic physics beyond collider energies.","feed_headline":"Muon data show 3–6x more multi-track events than simulations","feed_subtitle":"A ground-based 2 m x 2 m RPC stack measured 2-, 3-, and 4-muon event rates that the standard air-shower simulation cannot reproduce.","key_machinery":"The load-bearing element is the parallel-track selection applied to events reconstructed by the Hough transform. In each event, strip hits are grouped into straight-line candidates in the X-Z and Y-Z projections, and a candidate is a track if at least four layers fit with $\\chi^2/\\text{ndf}<10$. A pair of tracks is accepted as coming from the same cosmic-ray shower only when the skew angle between the pair is below $2.5^\\circ$, about $3\\sigma_0$ of the triple-Gaussian resolution function measured from simulated pairs; this cut removes random coincidences between independent showers and secondaries produced in the roof. On the simulation side, the same reconstruction runs on CORSIKA showers for hydrogen, helium, carbon, oxygen, silicon, and iron primaries that have been propagated through a GEANT4 detector model whose digitization uses measured efficiency maps, noise, and strip multiplicities, and the per-primary fractions are combined with standard abundance weights.","core_discovery":"The central claim is that the normalized fractions of cosmic-ray events containing 2, 3, and 4 parallel reconstructed tracks are $(6.35\\pm0.05)\\times10^{-5}$, $(5.82\\pm0.53)\\times10^{-7}$, and $(1.94\\pm0.97)\\times10^{-8}$, several times larger than the corresponding composition-weighted CORSIKA/QGSJET-II-04 predictions of $(2.35\\pm0.13)\\times10^{-5}$, $(1.12\\pm0.13)\\times10^{-7}$, and $(3.21\\pm0.87)\\times10^{-9}$. The same deficit appears with QGSJET01d, so it is not specific to one hadronic model. Because the reconstructed directions show no anisotropy and the quoted systematic uncertainties from roof thickness, material budget, strip multiplicity, noise, efficiency, and GEANT4 physics are much smaller than the gap, the paper concludes that the simulation chain underestimates the rate of parallel muons from cosmic-ray showers and that the missing ingredient most plausibly lies in the extrapolation of hadronic interactions to energies beyond current collider coverage, or in the assumed primary spectrum and composition.","pith_inferences":["Beyond the paper: a direct extension would be to repeat the analysis with other hadronic models such as EPOS-LHC or Sibyll and with the primary composition weights left as free parameters; if no physically reasonable composition reaches the 4-track rate, the problem would be isolated to forward hadron production in the simulation.","Beyond the paper: because the stack spans only 2 m by 2 m, the observed fractions are sensitive to muon lateral density at small separation; a wider array or a longer-baseline pair of stacks could test whether the simulation underestimates muon density in small cells or simply produces too few high-energy muon bundles.","Beyond the paper: if the excess is physical, correlated-muon backgrounds for rare-event searches in underground detectors would be underestimated by the same simulations; the paper does not quantify this, but the direction of the discrepancy makes it a relevant check.","Beyond the paper: the same data set could be used to derive a lower bound on the rate of high-energy primaries that produce multi-muon final states, by unfolding detector response and comparing with collider-constrained models; this would give a quantitative handle on the forward region that colliders cannot see."],"forward_implications":["Standard CORSIKA/QGSJET simulations underproduce ground-level muon bundles: by factors of roughly 3, 5, and 6 for 2, 3, and 4 parallel muons, so underground experiments that use these simulations for atmospheric-muon backgrounds will undercount correlated multi-muon events.","Because the measured fractions are normalized to single-track events and the same reconstruction is applied to data and simulation, a uniform inefficiency would cancel; the discrepancy therefore lives specifically in how often the simulation produces additional parallel muons.","The disagreement appears for both QGSJET-II-04 and QGSJET01d, so it is common to those high-energy hadronic models rather than a quirk of one; the paper places the likely cause in the forward-region extrapolation beyond collider energies or in the assumed primary composition and spectral index.","Multiplicity ratios of this kind offer a rare surface-level constraint on the high-energy tail of the cosmic-ray spectrum and on hadronic models, which is why the authors propose using the result to tune hadronic parameters or composition.","The paper notes that earlier underground multi-muon measurements and KASCADE-Grande's shorter simulated muon attenuation length point in the same direction, so the new ground-level result adds to a pattern of simulations underestimating atmospheric muons."],"supporting_citations":[{"why":"Supplies the air-shower simulation package and the QGSJET hadronic-interaction models whose track-multiplicity predictions are compared with data.","marker":"[6]"},{"why":"Provides the GEANT4 detector simulation that propagates CORSIKA particles through the RPC stack and building material.","marker":"[7]"},{"why":"Earlier study of the same stack; supplies the measured detector parameters (efficiency, noise, strip multiplicity) folded into the GEANT4 digitization.","marker":"[4]"},{"why":"Provides the primary cosmic-ray composition and spectrum and the vertical cutoff estimate used to set simulation energy ranges.","marker":"[5]"},{"why":"Supplies the relative cosmic-ray abundances used to weight per-primary simulated fractions into the combined prediction.","marker":"[15]"},{"why":"Hough transformation algorithm used for hit-to-track association in event reconstruction; the measured multiplicity depends on its output.","marker":"[11]"},{"why":"Independent ground-level result that simulated muon attenuation length in the atmosphere is smaller than observed, cited in support of the same muon discrepancy.","marker":"[21]"}],"fun_headline_variants":["RPC stack sees 3–6x more multi-muon events than simulations","Cosmic-ray muon multiplicities exceed CORSIKA predictions","Ground RPC array finds muon counts 3–6x above simulations","Multi-muon excess: RPC stack data defy CORSIKA model","Muon multiplicity discrepancy: 3–6x above simulation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison stands on the assumption that the detector simulation and the reconstruction cuts, including the 2.5-degree parallel-track definition, translate a given number of true muons into a recorded number of parallel tracks in data and simulation with equal fidelity; if the simulation mis-models multi-track response or the reconstruction fabricates extra tracks in data, the inferred excess would be an artifact, not a property of cosmic-ray showers.","fun_headline_variants_meta":{"raw":{"variants":["RPC stack sees 3–6x more multi-muon events than simulations","Cosmic-ray muon multiplicities exceed CORSIKA predictions","Ground RPC array finds muon counts 3–6x above simulations","Multi-muon excess: RPC stack data defy CORSIKA model","Muon multiplicity discrepancy: 3–6x above simulation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001639,"raw_usage":{"total_tokens":6511,"prompt_tokens":937,"completion_tokens":5574,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":5477}},"tokens_in":553,"tokens_out":5574,"duration_ms":39645,"temperature":1.0,"reasoning_tokens":5477,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:38:15.673828+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Remove the 22 cm concrete roof from the GEANT4 geometry and rerun the full analysis on the same CORSIKA input: if the measured 2-track fraction drops from its observed $6.35\\times10^{-5}$ toward the simulated $2.35\\times10^{-5}$ while the single-track rate barely changes, the extra parallel tracks are produced by roof interactions that the simulation mishandles; if the factor-of-about-3 excess survives without the roof, the discrepancy is genuinely atmospheric.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the air-shower simulation package and the QGSJET hadronic-interaction models whose track-multiplicity predictions are compared with data."},{"cited_title":"Agostinelli et al., GEANT4: A Simulation toolkit , Nucl","cited_arxiv_id":null,"evidence_quote":"Provides the GEANT4 detector simulation that propagates CORSIKA particles through the RPC stack and building material."},{"cited_title":"Pethuraj et","cited_arxiv_id":null,"evidence_quote":"Earlier study of the same stack; supplies the measured detector parameters (efficiency, noise, strip multiplicity) folded into the GEANT4 digitization."},{"cited_title":"Tanabashi et al","cited_arxiv_id":null,"evidence_quote":"Provides the primary cosmic-ray composition and spectrum and the vertical cutoff estimate used to set simulation energy ranges."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the relative cosmic-ray abundances used to weight per-primary simulated fractions into the combined prediction."},{"cited_title":"Duda, Peter E","cited_arxiv_id":null,"evidence_quote":"Hough transformation algorithm used for hit-to-track association in event reconstruction; the measured multiplicity depends on its output."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Independent ground-level result that simulated muon attenuation length in the atmosphere is smaller than observed, cited in support of the same muon discrepancy."}],"review_version":1}