{"id":"fc195243-c927-4694-94f4-72614afd8b13","arxiv_id":"2507.23552","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"FASER's updated forward neutrino measurements are broadly consistent with current hadronic interaction models, with tensions in the 0.3-0.6 TeV range and above 1 TeV.","lead":"FASER's forward neutrino detectors at the LHC measured electron and muon neutrino rates from proton collisions, and compared them with the newest cosmic-ray hadron interaction models. The data mostly agree with the models, but show a small excess of neutrino interactions around 300-600 GeV and fewer high-energy neutrinos than models with charm production predict.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed 300–600 GeV excess and >1 TeV deficit rest on small event counts and are never assigned a statistical significance; without a quantitative test, the model-discrimination implications are not load-bearing.","rationale":"The paper is a self-described preliminary update: 33% and 3% of the available data are analyzed, and the conclusions are worded cautiously ('generally consistent', 'some discrepancies'). The strongest claim, however, goes beyond the data: the abstract and Section 5 connect the observed bin-level differences to constraints on pion, kaon, and charm production and to the muon puzzle. What would have to be true for that claim to hold is that the bin-level discrepancies are larger than expected from statistical and systematic fluctuations. The paper never demonstrates this. The reader's weakest assumption concerned the model dependence of the single scale factor mu in Eq. 2; that is a real but secondary issue for the FASERnu total rates and is explicitly tested with nine flux models. The energy-binned comparison in Fig. 1 (center) comes from the electronic detector, whose flux extraction is not based on the hadronic model shape in the same way; the more pressing problem is that no significance is attached to the 300–600 GeV excess or the >1 TeV deficit. With 36 mu-neutrino interactions total, bin counts are small, and the 'excess in all generator comparisons' does not protect against a common offset in that energy region. I therefore propose a quantitative binned-likelihood test as the settlement. If the discrepancies survive with a global significance above, say, 3sigma, the paper's implications would be strengthened; if not, the paper should restrict its conclusion to reporting the differences without claiming model discrimination. Since the reader already assigned CONDITIONAL and the paper is honest about its preliminary nature, I do not move the verdict.","tokens_in":9476,"tokens_out":6738,"duration_ms":73853,"concrete_test":"Reproduce the central panel of Fig. 1 numerically: recover the five energy-bin counts and their covariance from the FASER publications (Refs. [12,13]) and the generator predictions from the CRMC/chromo outputs used here. For each generator, compute the binned Poisson likelihood ratio for the observed counts, including the published statistical and correlated systematic uncertainties; quote the p-value for the 0.3–0.6 TeV excess and the >1 TeV deficit, with a trial factor for scanning bins and generators. If the global p-value exceeds 0.05, revise the conclusion to state that the discrepancies are not statistically significant.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 states 'In all cases, there is an excess of neutrino events (but not anti-neutrino events) at energies between 300 GeV to 600 GeV' and 'generators including charm hadron production predict more high-energy neutrinos than observed,' and Section 5 concludes that these discrepancies motivate improvements to hadronic models. The load-bearing step is the inference from the binned event counts in Fig. 1 (center) to genuine model discrepancies. No significance is computed anywhere. The samples are small: §3.2 reports 5 nu_e+antinu_e and 19 nu_mu+antinu_mu candidates; §3.3 reports N_int(nu_mu+antinu_mu)=36.0 +16.1/-13.2. The electronic spectrum is divided into five energy bins, so the 0.3–0.6 TeV bin and the >1 TeV bin contain at most a handful of events. With Poisson statistics and large correlated systematic uncertainties (detection efficiency dominates, per §3.3), a few-event excess in one bin is plausibly a fluctuation. The statement that the excess appears 'in all cases' across generators does not increase significance, since the generator predictions in that bin are highly correlated (all are pion/kaon-dominated forward models). Without a p-value or a likelihood-ratio test that includes the published covariance, the data do not rule out consistency of all generators. Thus the 'implications for forward hadron production' are not yet established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the latest FASER neutrino measurements and compares them with hadronic interaction models relevant to forward particle production and the cosmic-ray muon puzzle. Using 9.5 fb^-1 of FASERnu emulsion data, the authors extract electron- and muon-neutrino interaction rates by fitting a single energy-averaged signal strength mu to the observed event counts, with nuisance parameters handled through a Bayesian MCMC likelihood. They combine these rates with previously published FASER electronic-detector measurements of the muon-neutrino energy and pseudorapidity spectra, and compare all results with EPOS-LHC(r), SIBYLL 2.3d/e, QGSJET 2.04/3, and the FASER baseline model. The paper finds that predictions are generally consistent with the measurements, but reports an apparent excess of neutrino (not antineutrino) events around 300-600 GeV and a tendency for charm-including generators to overpredict neutrinos above 1 TeV, interpreting these as motivations for improved forward hadron production models.","tokens_in":9730,"tokens_out":5712,"duration_ms":65252,"significance":"The paper is a useful conference-style summary that brings together the latest FASER neutrino interaction rates and differential spectra and places them side by side with the most recent hadronic interaction models. Its strengths are the clear description of the likelihood and MCMC procedure, the use of official FASER results, and the explicit bias study with nine alternative flux models. If the reported discrepancies were quantitatively established, the comparison would provide valuable validation of pion, kaon, and charm production in forward kinematics and could inform the cosmic-ray muon puzzle. However, the paper's new contribution is largely qualitative: the discrepancies are identified visually in figures and stated in words, but no statistical significance, likelihood-ratio test, or covariance-based comparison is provided. The paper itself acknowledges the limited analyzed dataset and the preliminary character of the results, which tempers the conclusions, but the title's promise of 'implications' needs quantitative support.","major_comments":[{"comment":"The central claims that 'in all cases, there is an excess of neutrino events at energies between 300 GeV to 600 GeV' and that 'generators including charm hadron production predict more high-energy neutrinos than observed' are presented without any quantitative statistical test. No p-value, confidence level, or likelihood-ratio statistic is given for any generator or energy bin. Because the number of events in individual bins is limited and the systematic uncertainties are substantial, a visual excess in one or two bins can easily arise from a fluctuation. The fact that the excess appears for several generators does not add independent statistical power, since the generator predictions in those bins are highly correlated. A quantitative comparison, including the published covariance of the data points and the bin-to-bin correlations, is required to support the claimed model discrimination. This is load-bearing because Section 5 bases its motivation on these discrepancies.","section":"Section 4"},{"comment":"The FASERnu interaction rate N_int is computed by fitting a single scale factor mu to the total observed event count and then multiplying the baseline-model flux by this factor. Consequently, N_int is a normalization measurement whose energy and rapidity shapes are inherited from the baseline model. The paper's bias check with nine flux models is useful, but those models are all built from the same family of hadronic generators and do not cover alternative spectral shapes such as a different pion-to-kaon ratio. The comparison of the left panel of Fig. 1 to generators should therefore be described as a test of the overall normalization only, not as a spectral constraint. This does not invalidate the differential energy and pseudorapidity spectra from the electronic detector, which are separate measurements, but the manuscript should be explicit about the limited interpretation of N_int.","section":"Section 3.3, Eq. (2)"},{"comment":"The caption notes that the data points in the center and right panels are correlated because they use the same dataset, yet no covariance information is used when comparing the measurements with generator predictions. The statement that a discrepancy appears 'in all cases' across generators is a visual observation, not a statistical one; it does not account for the strong correlations among the model predictions. To make the claimed excess and deficit operational, the authors should provide a quantitative goodness-of-fit measure, such as a chi-square or profile-likelihood ratio, for each generator using the full covariance matrix, along with the resulting p-values. Without this, the implications for forward hadron production are not yet established.","section":"Fig. 1 and Section 4"}],"minor_comments":[{"comment":"The abstract states that 'the latest measurements of electron and muon neutrino fluxes are presented,' but the paper mostly reinterprets previously reported FASERnu event rates and combines them with previously published FASER electronic-detector results. The wording should distinguish new results from a new interpretation of existing data.","section":"Abstract and Section 1"},{"comment":"The likelihood in Eq. (3) includes Poisson priors P_k for background Monte Carlo fluctuations, but the text does not describe how these are implemented (e.g., as scaling factors with Poisson constraints). A short explanation would improve reproducibility.","section":"Eq. (3) and Section 3.3"},{"comment":"The 'other syst' uncertainty is the dominant source for muon neutrinos in Table 1, but Section 3.3 later states that the neutrino detection efficiency uncertainty dominates. The relation between 'other syst' and the detection-efficiency uncertainty should be clarified.","section":"Table 1"},{"comment":"The y-axis labels in the three panels of Fig. 1 are inconsistent: the left panel uses 'Number of Interactions' while the center and right panels use 'Number of Interactions per cm^2.' The units and target volumes should be stated consistently in the caption and axes.","section":"Fig. 1"},{"comment":"The text says that EPOS-LHCr underestimates events at 300-600 GeV where 'pi+ and kaons contribute.' This should be expanded to 'charged pions and charged/neutral kaons' to match the production-mode decomposition in Fig. 2.","section":"Section 4 and Fig. 2"},{"comment":"References [8] and [27] cite 'in this proceedings' without article numbers. For an ICRC proceedings paper, the PoS article numbers (e.g., PoS(ICRC2025)358 and PoS(ICRC2025)1182) should be included in the reference list.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is essentially a conference proceedings contribution that combines already-published FASER results with generator predictions. The main new element is a qualitative comparison plot, and the paper's title and conclusions go somewhat beyond what the current statistical analysis supports. The missing significance calculation is a well-defined addition that would substantially strengthen the paper, so I recommend major revision rather than rejection. The editor may also wish to consider whether the 'implications' framing is appropriate for the present dataset or whether a more descriptive title would be more suitable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a short ICRC conference proceedings, not a full paper. What is actually new: the FASERν emulsion result now uses 2.45 times more target mass than the first result, and this is the first public comparison of FASER neutrino data with the newly released EPOS-LHCr, SIBYLL 2.3e, and QGSJET 3. The experimental method is established and clearly described—standard likelihood/MCMC with nuisance parameters, and the analysis is honest about backgrounds and small data fractions.\n\nThe paper does several things well. The overall message is appropriately cautious: predictions are generally consistent with the measured fluxes, and the discrepancies are called \"motivating\" rather than definitive. The single-figure comparison of many generators is useful, and the authors explicitly flag which generators lack charm production. They also check the shape-dependence of their scale-factor extraction against nine flux models and include that as a systematic. That is a good-faith effort.\n\nThe soft spots are as follows. The claimed 300–600 GeV excess and >1 TeV deficit in the muon neutrino spectrum rest on a handful of events, and no significance is computed anywhere. With Poisson statistics and large correlated systematics, those bins could easily fluctuate. The statement that the excess appears \"in all cases\" across generators does not add significance because the generator predictions in that bin are highly correlated. Also, the N_int numbers from FASERν are obtained by fitting a single scale factor to the total event count and then scaling the baseline model; the differential spectrum in the center panel does come from the electronic detector, so it is not affected by that particular circularity, but the left-panel totals inherit the baseline model's flavor and energy mix. The nine-model cross-check is good but stays within the same family of generators.\n\nNone of this is fatal. The paper's main claim—that current generators are broadly consistent with FASER data—holds up. The discrepancies are suggestive and appropriately deferred to future data. What is missing is a quantitative statement of how surprising the discrepancies are; that should have been included even in a proceedings.\n\nWho is this for? Anyone working on forward hadron production, cosmic-ray muon puzzles, or generator tuning. It is a useful status update, not a landmark result. I would bring it to a reading group and would cite it as a reference for the new model comparisons. If this were submitted to a journal, I would send it to referees, but I would ask them to either add a significance estimate or explicitly soften the implication claims.","headline":"A solid, honest FASER status update with useful new generator comparisons, but the energy-bin discrepancies are qualitative and lack significance.","tokens_in":10914,"tokens_out":2694,"would_cite":true,"duration_ms":31058,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FASER's forward neutrino measurements at the LHC find a 300-600 GeV excess of neutrino events over all hadronic generators, and a charm-driven deficit above 1 TeV, giving the first direct LHC constraints on the forward pion, kaon, and…","keywords":["forward neutrino flux","FASER","hadronic interaction models","muon puzzle","cosmic-ray air showers","charm hadron production","pion and kaon production","LHC forward physics"],"falsifier":"Re-analyze the full Run-3 emulsion and electronic-detector data (about 150 inverse femtobarns) with neutrinos and antineutrinos binned separately: if the 300-600 GeV excess persists and the above-1 TeV deficit narrows as the sample grows, the paper's interpretation is supported; if the excess is absorbed by updated cross-section or flux-shape systematics, or becomes consistent once the nine-model shape bias is recalculated with an independent flux model, the claim would be weakened.","tokens_in":9266,"feed_emoji":"⚛️","tokens_out":12091,"duration_ms":116660,"temperature":0.7,"pith_summary":"FASER, a detector 480 m downstream of the LHC collision point, catches neutrinos produced by forward decays of charged pions, kaons, and charm mesons, a region no other LHC experiment sees directly. Using 9.5 inverse femtobarns of emulsion-detector data and 65.6 inverse femtobarns of electronic-detector data, the paper compares measured electron- and muon-neutrino interaction rates with the latest hadronic generators, including EPOS-LHCr, SIBYLL 2.3e, and QGSJET 3. Overall the models agree with the measured fluxes, but in every model the neutrino events in the 300-600 GeV energy window are under-predicted, while antineutrino events are not, and models that include charm production over-predict neutrinos above 1 TeV. These discrepancies matter because forward pion, kaon, and charm yields are the same physics behind the long-standing excess of muons in ultra-high-energy cosmic-ray air showers.","feed_headline":"FASER neutrino excess no hadronic generator predicts","feed_subtitle":"Forward neutrinos at 300-600 GeV outnumber all generators, while charm-driven neutrinos above 1 TeV come up short.","key_machinery":"The measurement uses charged-current neutrino interactions in a 324.1 kg tungsten-emulsion target as the detector, selecting neutral vertices with an electromagnetic shower for electron-neutrino candidates or a penetrating muon track for muon-neutrino candidates. Because only a handful of events is available, the flux is not unfolded directly: a single energy-averaged scale factor $\\mu$ (Eq. 2) multiplies the predicted neutrino and antineutrino fluxes, and its posterior is sampled by a Bayesian Markov-chain Monte Carlo with nuisance parameters for neutral-hadron and neutral-current backgrounds. That scale factor converts the observed event count into an interaction-rate spectrum while preserving the energy and rapidity shape of the baseline Monte Carlo flux, so the energy-binned comparison is only as good as that assumed shape.","core_discovery":"FASER's forward neutrino measurements at the LHC are reaching the precision where they can begin to discriminate among hadronic interaction models. The paper reports $N^{\\rm int}(\\nu_e+\\bar\\nu_e)=12.2^{+8.7}_{-6.4}$ and $N^{\\rm int}(\\nu_\\mu+\\bar\\nu_\\mu)=36.0^{+16.1}_{-13.2}$ measured interactions in the FASER$\\nu$ target, and when these are binned in energy the data sit above every generator prediction in the 300-600 GeV range while falling below the charm-inclusive predictions above 1 TeV. The paper notes that electron-neutrino flux measurements can uniquely constrain kaon and charm hadron production, so the discrepancies are read as the first direct LHC constraints on forward pion, kaon, and charm yields, the same quantities implicated in the cosmic-ray muon puzzle.","pith_inferences":["If the 300-600 GeV excess persists with full statistics, adjusting generators to raise forward pion and kaon yields in that energy window would also change predicted muon rates in air showers, plausibly in the direction needed for the muon puzzle.","The shape-blind scale-factor method is tested against nine flux models from the same generator family; a more severe test would use data-driven pion-to-kaon ratios or a model with a genuinely different pion-to-kaon ratio.","Because the emulsion sample is only 3% of collected data, the statistical reach grows by roughly a factor of 30 soon, so the discrepancy could sharpen, move, or disappear with that sample.","The rapidity spectrum shows no strong tension, suggesting the disagreement is primarily in energy dependence; separate neutrino and antineutrino rapidity spectra would localize whether the excess is a pion versus kaon effect."],"forward_implications":["The overall agreement of EPOS-LHCr, SIBYLL 2.3e, and QGSJET 3 with the measured fluxes means that forward hadron production models are broadly correct in total rate while missing details in specific energy ranges.","The 300-600 GeV excess of neutrino events common to all models indicates under-predicted forward pion and kaon production in that energy window, and it does not come from antineutrinos.","Charm-inclusive generators over-produce neutrinos above 1 TeV, so FASER data prefer models without charm at those energies or require reduced forward charm production.","These comparisons provide a new, model-discriminating input for tuning hadronic interaction models used in cosmic-ray air-shower simulations and for testing the kaon-enhancement scenario of the muon puzzle.","Only 33% of electronic and 3% of emulsion data are used; the stated expectation is that the full Run-3 dataset and later the Forward Physics Facility will refine these comparisons with finer bins and multi-differential analyses."],"supporting_citations":[{"why":"Supplies the 5 electron-neutrino and 19 muon-neutrino candidates and the updated emulsion-detector analysis the fluxes are based on.","marker":"[15]"},{"why":"Defines the standard FASER neutrino flux and cross-section prediction methodology used in Eq. (1).","marker":"[19]"},{"why":"Provides the event selections, detection efficiencies, and neutral-hadron background estimates carried over to this analysis.","marker":"[14]"},{"why":"Provides the electronic-detector measurement of the muon-neutrino flux as a function of energy used in the comparison.","marker":"[12]"},{"why":"Provides the electronic-detector muon-neutrino flux as a function of rapidity shown in the comparison figure.","marker":"[13]"},{"why":"The updated EPOS-LHCr generator that overestimates the above-1 TeV muon-neutrino rate and underestimates the 300-600 GeV rate.","marker":"[7]"},{"why":"The QGSJET-III model without charm production that agrees better with the above-1 TeV data.","marker":"[9]"},{"why":"Articulates the kaon-enhancement scenario for the muon puzzle that the FASER neutrino comparisons can test.","marker":"[11]"}],"fun_headline_variants":["FASER neutrinos show forward excess, high-energy shortfall","FASER neutrino data hint at hadron model gaps","LHC forward neutrinos challenge hadronic models","FASER neutrino fluxes diverge from generator predictions","FASER's neutrino excess hints at pion and kaon yields"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measurement assumes that the energy and rapidity shape of the neutrino flux is the one predicted by the baseline simulation, so a single scale factor can turn the observed event count into an interaction-rate spectrum; if the real pion-to-kaon ratio differs from that simulation, the energy-binned comparisons would be biased even if the total rate came out right.","fun_headline_variants_meta":{"raw":{"variants":["FASER neutrinos show forward excess, high-energy shortfall","FASER neutrino data hint at hadron model gaps","LHC forward neutrinos challenge hadronic models","FASER neutrino fluxes diverge from generator predictions","FASER's neutrino excess hints at pion and kaon yields"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000584,"raw_usage":{"total_tokens":2776,"prompt_tokens":1002,"completion_tokens":1774,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":1692}},"tokens_in":618,"tokens_out":1774,"duration_ms":15506,"temperature":1.0,"reasoning_tokens":1692,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:41:20.044635+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the full Run-3 emulsion and electronic-detector data (about 150 inverse femtobarns) with neutrinos and antineutrinos binned separately: if the 300-600 GeV excess persists and the above-1 TeV deficit narrows as the sample grows, the paper's interpretation is supported; if the excess is absorbed by updated cross-section or flux-shape systematics, or becomes consistent once the nine-model shape bias is recalculated with an independent flux model, the claim would be weakened.","supporting_citations":[{"cited_title":"Pierog and K","cited_arxiv_id":null,"evidence_quote":"The updated EPOS-LHCr generator that overestimates the above-1 TeV muon-neutrino rate and underestimates the 300-600 GeV rate."}],"review_version":1}