{"id":"1e19c442-7844-4e5d-b705-d51f0e28105c","arxiv_id":"2607.22323","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using a universal relation found across hadronic interaction models, Auger's deeper-than-predicted shower maxima imply increased elasticity and hadronic energy fraction in proton–air collisions, needing 2.8–4.6× amplification to match muon counts.","lead":"This paper connects measurements of cosmic-ray air showers to properties of the very first proton–air collision, showing that all major simulation models need more energy flowing into hadrons to explain both shower depth and muon counts. It gives a concrete range (2.8–4.6×) for amplifying the hadronic energy fraction in the first interaction to resolve the 'muon puzzle'.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2.8–4.6 muon amplification is not predicted; it is assumed via an unmeasured f_g profile in Eq. (3). Different f_g choices yield very different g_mod, so the capstone claim is untested.","rationale":"The reader's verdict (CONDITIONAL) identifies the calibration curves and the assumption that later hadron–air interactions do not contribute sizeably to the Xmax shift as the weakest premise. That is a valid concern about the mapping from Xmax to first-interaction variables. However, the single most load-bearing element of the paper's central claim is the numerical capstone: the amplification factor 2.8–4.6 that connects the Xmax-derived increase in alpha_had to the muon puzzle. That factor comes from Eq. (3), which depends entirely on f_g, the per-generation modification factor. The paper never specifies or derives f_g; it merely asserts that f_1=1 and that it decreases with generation. As a result, g_mod is not a prediction of the model—it is the ratio of the muon rescaling to the alpha_had shift, essentially a fit. Different plausible f_g profiles (e.g., a modification confined to the first interaction, or a modification that persists through all generations) would change g_mod by factors of order 1 to 26, completely altering the conclusion. The paper's own statement in Sec. IV that the interpretation assumes modifications 'evolve continuously with projectile energy' is exactly the untested f_g assumption. This is more load-bearing than the number of calibration points, because even a perfect universal relation and a correct Xmax shift would not yield the quoted muon-puzzle explanation without a physically grounded f_g. The reader mentioned the 'unmeasured f_g profile' in the rationale but did not center it as the weakest assumption, so I partially agree. A direct simulation test could settle whether the assumed amplification is physical, and the verdict remains CONDITIONAL pending that check.","tokens_in":22609,"tokens_out":5765,"duration_ms":51294,"concrete_test":"Perform a dedicated CONEX simulation in which the first proton–air interaction is reweighted/replaced so that <alpha_had> increases by the tabulated shifts (3.6%, 6.0%, 3.1% for EposLHC, QGSjet-II.04, Sibyll2.3d) while all later hadronic interactions remain at default model settings; compute the resulting ratio δln<N_mu>/δln<alpha_had>. If this ratio is not within the quoted g_mod ranges (2.8–4.6, including uncertainties), Eq. (3) is not a valid predictor and the muon-puzzle link fails. Alternatively, derive f_g by comparing N_mu and alpha_had between available model pairs (e.g., EposLHC vs EposLHC-R, QGSjet-II.04 vs QGSjet-III.01) at the same energy/zenith; the implied g_mod should match if the cascade-propagation assumption is physical.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that the Xmax-preferred increase in <alpha_had> must be amplified by 2.8–4.6 to account for the muon puzzle—rests on Eq. (3): δln<N_mu> = g_mod δln<alpha_had>, with g_mod = Σ f_g. The text specifies only f_1=1 and 'decreasing progressively with g'; no functional form, no derivation from the hadronic models, and no simulation is provided to constrain f_g. g_mod is then inferred backward from the Auger muon rescaling R_mu and the alpha_had shift, so the statement that the alpha_had increase 'must be amplified' is a restatement of the assumed propagation, not an independent consequence. If f_g=0 for g>1 (modification confined to first interaction), g_mod=1 and the same Xmax shift cannot explain the muon excess; if f_g=1 for all g, g_mod≈c≈26 and the required alpha_had increase would be tiny. The paper's own caveat—that modifications must 'evolve continuously with projectile energy'—is precisely the untested f_g assumption. Without an external constraint on f_g, the 2.8–4.6 factor is a free parameter dressed as a prediction.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses hybrid Pierre Auger data at sqrt(s) ≈ 100 TeV to infer properties of the first proton-air interaction. Using six LHC-tuned hadronic interaction models (EPOS LHC-R, QGSJet-III.01, Sibyll2.3e and their predecessors), it shows that predicted <X_max> lies on approximately universal linear relations with mean values of first-interaction production variables: the elasticity, the hadronic energy fraction α_had, and the entropy-like variables ζ_had and ζ_EM. Assuming these relations hold in Nature, the Auger <X_max> shifts of Ref. [2] are mapped into changes of these variables, yielding increased elasticity and α_had and decreased ζ_had and ζ_EM for all models. Using Eq. (3), the paper claims that a 2.8–4.6 amplification of the α_had increase suffices to explain the muon puzzle. The claim rests on assumed generation-dependent modification factors f_g, with only f_1=1 specified.","tokens_in":22917,"tokens_out":12270,"duration_ms":86337,"significance":"If the universal relations and their extrapolation to Nature hold, the paper offers a new way to convert X_max measurements into constraints on forward particle production that are testable at LHCf, FASER, and in proton-oxygen collisions. The paper is commendably explicit about its assumptions, quantifies several systematics, validates the relations with pre-LHC models, and tabulates all inferred values. However, the quantitative capstone—the 2.8–4.6 amplification—is not an independent prediction: it is the ratio of the inferred muon rescaling to the inferred α_had shift, divided by an assumed, unmeasured f_g profile. This limits the strength of the muon-puzzle connection.","major_comments":[{"comment":"The amplification factor g_mod = Σ f_g is the load-bearing element of the muon-puzzle claim, but it is unconstrained: only f_1=1 is specified ('decreasing progressively with g'), with no functional form, derivation, or simulation. The quoted values g_mod=2.8–4.6 are obtained by dividing the Auger muon rescaling R_μ by the inferred δln⟨α_had⟩ (Fig. 2), so the claim that the α_had shift 'must be amplified' restates the assumed propagation. If modifications were confined to the first generation, g_mod=1 and the same X_max shift would not explain the muon excess; if f_g=1 for all g, g_mod≈c≈26 and the required α_had change would be tiny. Provide explicit CONEX simulations with chosen f_g profiles, or reframe the 2.8–4.6 factor as a scenario rather than a prediction.","section":"Sec. IV, Eq. (3)"},{"comment":"The 'universal relations' for ζ_had and ζ_EM are partly by construction: these variables were designed in Ref. [8] to maximize correlation with X_max, and SM Eq. (4) is a linear estimator of ΔX_max in terms of α_had, ζ_had, ζ_EM. The small residual scatter in the ζ panels of Fig. 1 is therefore not an independent test of a universal law. The independent content resides in α_had and κ_el, which were not constructed for this purpose, and in the model-to-model scatter about the common lines. Please state this explicitly and restrict the 'discovered universality' claim accordingly.","section":"Sec. III / SM Eq. (4)"},{"comment":"The calibration uses six points from three model families at a single energy/zenith (E0=10^18.7 eV, θ=55°). The claim that Nature lies on the model locus is assumed, not tested. The pre-LHC validation in the Supplemental Material (E0=10^19 eV, θ=60°) is helpful but uses a different energy/zenith and shows increased scatter. Please add cross-checks against model-independent observables—e.g., X_max fluctuations, the elongation rate, or the measured proton-air cross section at √s=57 TeV—or validate the curves at a second energy/zenith with the same model set.","section":"Sec. III, Fig. 1"},{"comment":"The pion-air cross-section freedom, estimated at ~10 g/cm^2, is of the same order as the systematic uncertainties and as a substantial fraction of the X_max shift being interpreted. The estimate rests on a single model-pair comparison (EPOS LHC-R vs. Sibyll2.3e) that confounds σ_π–air with other model differences. A systematic two-dimensional scan (production variables vs. σ_π–air) is needed to quantify how much of the inferred first-interaction changes could be absorbed by pion-air modifications. The paper's stated caveat is clear, but the current estimate understates the degeneracy.","section":"Sec. IV / SM §3"}],"minor_comments":[{"comment":"The range '2.8–4.6' refers to central values across the three models; the asymmetric uncertainties (e.g., 4.0^{+7.5}_{-2.5}) are much larger. Please quote the range as central values across models and give the uncertainties.","section":"Abstract / Sec. V"},{"comment":"The meaning of the open 'Data' markers is hard to parse; the caption should state explicitly that they are the Auger-shifted <X_max> projected onto the x-axis along the fit curves.","section":"Fig. 1"},{"comment":"The sentence linking the inferred changes to Ref. [37] conflates a generation-averaged α_had with the first-interaction α_had used in Fig. 2; clarify the distinction.","section":"Sec. IV"},{"comment":"The naming of the EPOS model is inconsistent ('Epos LHC', 'EPOS LHC', 'EposLHC'); please standardize.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The central mapping (X_max → α_had, κ_el) is plausible and the paper is unusually candid about its caveats. The blocker is Eq. (3): g_mod is unconstrained, so the 2.8–4.6 amplification is not a prediction. This is fixable (explicit CONEX runs with f_g profiles, or a softened claim), hence major revision rather than reject. The ζ-based universality is partly circular; the α_had-based inference is not."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper's real contribution is a clean demonstration: across six LHC-tuned hadronic models — and checked against legacy models in the supplement — <Xmax> falls on tight linear universality bands against several first-interaction production variables, not just the ζ variables but also α_had, ln(κ_el/m_total), etc. That is a useful, testable statement about the hadronic-model landscape. Second, the headline muon amplification factor, 2.8–4.6, is not a prediction. It is the slope of a line drawn through the default model point and the shifted data point after assuming a propagation kernel f_g that is only specified as 'decreasing progressively'. The stress-test note is right: different f_g choices change g_mod by an order of magnitude, so the capstone number is a restatement of the assumed propagation, not an independent consequence.\n\nWhat the paper does well: the mapping of Auger's <Xmax> shift into first-interaction variables is internally consistent and physically coherent — more hadronic energy, higher elasticity, more forward energy flow. The authors also state the key failure modes plainly, particularly that pion–air interactions could change <Xmax> without touching the first interaction. The residual dispersion about the universal curves (1.7–5.2 g/cm²) is smaller than the Xmax systematic uncertainty, so the universality claim is credible. The validation with pre-LHC models at a different energy and zenith helps, and the energy-scale systematics are addressed, though briefly.\n\nThe soft spots, in proportion: the main calibration is six model points at one energy and zenith (10^18.7 eV, θ=55°), though the supplement's energy-evolution and legacy checks broaden that somewhat. The inferred shifts inherit the assumption that the whole Xmax deficit lives in the first proton–air interaction; the paper says so in Section IV, but it means the quoted values are conditional. The worst issue is indeed Eq. (3). f_g is not derived, not fitted to the models, not constrained by any simulation. g_mod is just inferred from R_mu and the α_had shift, so 'must be amplified by a factor of 2.8 to 4.6' is a restatement of the assumed propagation. The paper should either constrain f_g using the generation-by-generation hadronic-energy evolution in the models, or present g_mod as an assumption-dependent parameter with a sensitivity study. As written, the abstract and conclusions overstate the quantitative claim.\n\nVerdict: this deserves a serious referee. The universal relation is a genuine contribution to cosmic-ray and forward-physics discussions, and the qualitative link to the muon puzzle is plausible. A referee should ask for an explicit f_g treatment and softer language on the amplification factor. With that fix, the paper is solid. I'd take the referee assignment; for a reading group, I'd bring it with the caveat about the muon number.","headline":"The universal-relation part is real and worth reading; the 2.8–4.6 muon amplification number is assumed, not derived, and the abstract oversells it.","tokens_in":25588,"tokens_out":3280,"would_cite":true,"duration_ms":30034,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The deeper-than-predicted maximum of air showers and the excess of ground muons are traced to one modification of the first proton–air collision: a small increase in hadronic-energy fraction, amplified three- to fivefold down the cascade.","keywords":["extensive air showers","ultra-high-energy cosmic rays","hadronic interaction models","depth of shower maximum","muon puzzle","proton-air interactions","forward physics","Pierre Auger Observatory"],"falsifier":"A concrete falsifier would be a hadronic interaction model, or a dedicated accelerator measurement, that reproduces the same mean first-interaction production variables at sqrt(s) about 100 TeV but predicts an <X_max> that deviates from the universal calibration curve by more than the observed residual scatter of about 5 g/cm^2; alternatively, an air-shower measurement in which the zenith-angle dependence of the muon excess and the X_max shift require different amplification factors g_mod would break the proposed link between the two observables.","tokens_in":22418,"feed_emoji":"☄️","tokens_out":6187,"duration_ms":55995,"temperature":0.7,"pith_summary":"This paper tries to show that two long-standing anomalies of ultra-high-energy cosmic-ray air showers—the measured depth of shower maximum being deeper than every LHC-tuned hadronic model predicts, and the muon count at ground being higher than predicted—can be traced to a single modification of the very first proton–air interaction. The paper demonstrates that, across six versions of leading hadronic interaction models, the predicted average shower-depth maximum is nearly linearly controlled by the mean values of a few variables characterizing the energy spectrum of secondary particles in that first collision: the hadronic-energy fraction, the elasticity of the leading secondary, and entropy-like partition variables. Assuming these universal relations hold in Nature, the Auger-favored deeper shower maxima map into an increased mean elasticity and hadronic-energy fraction, together with a more asymmetric partition of the primary energy. A small increase in the hadronic-energy fraction, amplified over roughly ten shower generations, is then sufficient to account for the reported muon excess, quantitatively connecting the two observables to one physical effect at sqrt(s) roughly 100 TeV.","feed_headline":"Amplified hadronic boost links shower depth to the muon excess","feed_subtitle":"Shower-depth shifts require a 3–6% hadronic-energy increase, amplified 2.8–4.6 times to match the muon count.","key_machinery":"The load-bearing object is the set of first-interaction production variables of a proton–air collision: alpha_had (the sum of energy fractions carried by hadronically interacting secondaries), kappa_el (the energy fraction of the leading secondary), and the entropy-like variables zeta_had and zeta_EM, each defined as minus the energy-weighted sum of x ln x over the hadronic or electromagnetic secondaries. The paper establishes nearly model-independent linear calibration curves of <X_max> versus the mean of each variable, with residual scatter of 1.7–5.2 g/cm^2, far smaller than the ~34 g/cm^2 spread of model predictions. These curves do the work: they allow the single measured quantity <X_ma","core_discovery":"The central claim is that the predicted average depth of the shower maximum, <X_max>, is controlled to within a few g/cm^2 by the mean values of first-interaction production variables, with only a mild residual dependence on the rest of the shower development. Denoting by alpha_had the fraction of the primary energy carried by hadronically interacting secondaries, by kappa_el the elasticity of the leading secondary, and by zeta_had and zeta_EM entropy-like measures of energy partition, the paper fits common linear calibration curves to six model versions (EPOS-LHC and EPOS-LHC-R, QGSjet-II.04 and QGSjet-III.01, Sibyll2.3d and Sibyll2.3e) at E0 = 10^18.7 eV and theta = 55 degrees. The Auger d","pith_inferences":["If the universal calibration curves extend beyond the fitted six model versions, the same mapping could be applied at other primary energies and zenith angles, turning the full energy and angular dependence of Auger <X_max> into a continuous measurement of first-interaction variables; the paper only demonstrates the curves at a single energy and zenith angle.","The preference for low zeta values implies more diffraction-like, rapidity-gap-like, or strongly asymmetric non-diffractive configurations; a testable extension is to compare the predicted shape of the muon production depth distribution and the fluctuations of X_max, not just the mean, since the paper notes these distributions could distinguish cross-section changes from production-variable change","The amplification-factor derivation assumes that a small modification recurs, with decreasing weight, at every shower generation; an alternative reading is that a single very large change confined to the first interaction would need to be much bigger and would likely alter fluctuation patterns in ways that the current mean-based analysis does not test.","The inferred production variables are, in principle, measurable in collider phase space, so a direct comparison of energy flow at high pseudorapidity in proton–oxygen collisions at the LHC with the values implied here would test whether the trend toward far-forward energy flow extrapolates from accelerator energies to 100 TeV."],"forward_implications":["All three model families require the same qualitative modification of the first proton–air interaction: increased mean elasticity and hadronic-energy fraction, reduced zeta values, and an energy flow shifted toward high pseudorapidities with a very asymmetric partition of the primary energy.","The required increase in the hadronic-energy fraction is small, 3.1–6.0%, but is amplified by a factor of 2.8–4.6 over the cascade generations to reproduce the 14–17% muon excess, directly linking the shower-depth shift and the muon excess to a single cause.","The updated model EPOS-LHC-R, with more diffractive proton–air interactions and a smaller pion–air cross section, satisfies the X_max shift inferred for its predecessor at this zenith angle, indicating the required modification is plausible within standard-physics variations, whereas QGSjet-III.01 only partially matches and Sibyll2.3e does not.","The inferred modification of forward, far-forward energy flow can be probed, at lower energies, by forward and far-forward measurements in proton–oxygen collisions at the LHC, which are sensitive to energy flow at high pseudorapidity.","The framework turns air-shower measurements into constraints on hadron production at sqrt(s) about 100 TeV, a kinematic region beyond direct accelerator reach, and does so through a small set of physically interpretable variables rather than through ad hoc reweighting of shower simulations."],"fun_headline_variants":["Proton-air showers hint at 2.8–4.6x hadronic boost for muon excess","Auger shower depth implies hadronic boost to explain muon puzzle","Shower depth pins down hadronic boost factor for muon puzzle","Muon excess tied to 2.8–4.6x hadronic boost from shower depth","Shower depth demands 2.8–4.6x hadronic boost for muon excess"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the nearly linear relations between predicted shower depth and the mean first-interaction production variables, fitted with only six model versions at a single energy and zenith angle, are valid in Nature, so that the entire observed X_max shift is imputed to the first proton–air interaction rather than to later pion–air interactions, exotic physics, or an alternative primary mass composition.","fun_headline_variants_meta":{"raw":{"variants":["Proton-air showers hint at 2.8–4.6x hadronic boost for muon excess","Auger shower depth implies hadronic boost to explain muon puzzle","Shower depth pins down hadronic boost factor for muon puzzle","Muon excess tied to 2.8–4.6x hadronic boost from shower depth","Shower depth demands 2.8–4.6x hadronic boost for muon excess"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001658,"raw_usage":{"total_tokens":6435,"prompt_tokens":779,"completion_tokens":5656,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":5542}},"tokens_in":523,"tokens_out":5656,"duration_ms":31213,"temperature":1.0,"reasoning_tokens":5542,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:05:58.551023+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsifier would be a hadronic interaction model, or a dedicated accelerator measurement, that reproduces the same mean first-interaction production variables at sqrt(s) about 100 TeV but predicts an <X_max> that deviates from the universal calibration curve by more than the observed residual scatter of about 5 g/cm^2; alternatively, an air-shower measurement in which the zenith-angle dependence of the muon excess and the X_max shift require different amplification factors g_mod would break the proposed link between the two observables.","supporting_citations":[],"review_version":1}