{"id":"1fc7a7f3-ea9d-4c86-9ca9-8e3c96a2a935","arxiv_id":"1908.01750","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Average very forward energy in 13 TeV proton-proton collisions rises with central track multiplicity, and every tested event generator overestimates the hadronic energy fraction compared with CMS data.","lead":"The CMS collaboration measured the average energy far forward in proton-proton collisions at 13 TeV, split into electromagnetic and hadronic parts, and compared it with the number of charged particles produced near the collision center. All tested Monte Carlo event generators put too much energy into hadrons, a result relevant for simulations of cosmic-ray air showers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Universal 'all generators overestimate hadron fraction' claim is not shown to survive the quoted intercalibration systematic, which shifts the EM/hadronic ratio down by up to ~20% toward the models.","rationale":"The reader correctly identifies the CASTOR electromagnetic/hadronic intercalibration as the weakest assumption, but the reader's rationale contains a factual error about the direction of that systematic: the paper quotes a maximum 8% decrease of electromagnetic energy and 15% increase of hadronic energy, which moves the measured E_em/E_had ratio down by about 20%, toward the model predictions, not away from them. This matters because the abstract's universal statement ('all generators overestimate the fraction of energy going into hadrons') is a central-value comparison that may not survive the quoted systematic shift. The paper does not provide numerical tables of the ratio with the systematic band, so the robustness of the claim cannot be verified from the text alone. The concrete check described above would settle the issue using the publicly available RIVET/HEPData information. Because the physics interpretation (the muon-deficit argument) depends on this discrepancy being larger than the calibration uncertainty, the paper should be accepted with the abstract and summary qualified unless the check confirms that all generator predictions remain below the shifted systematic band. This is a conditional recommendation: no accusation of error, but the central claim's strength needs explicit verification against the paper's own quoted systematic uncertainty.","tokens_in":30114,"tokens_out":8564,"duration_ms":90059,"concrete_test":"Obtain the data points and full systematic band for Fig. 3 from the paper's HEPData record or the RIVET plugin. In each track-multiplicity bin, form the ratio with the electromagnetic energy scaled by 0.92 and the hadronic energy scaled by 1.15, and recompute the lower edge of the systematic band from the stated intercalibration uncertainty. Then test whether every generator prediction (PYTHIA8 CUETP8M1, PYTHIA8 4C+MBR, EPOS LHC, SIBYLL 2.1, SIBYLL 2.3c, QGSJET-II.04, PYTHIA8 CP5, HERWIG 7.1) remains below that lower edge in all bins. If any model enters the shifted band, the abstract's 'all generators' claim must be qualified to state that the excess is within the intercalibration systematic uncertainty.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract's central claim is that every tested generator overestimates the hadronic energy fraction. In the data this corresponds to a higher electromagnetic-to-hadronic reconstructed-energy ratio, E_em/E_had, than any model prediction (Fig. 3). The load-bearing assumption is the relative calibration of the CASTOR electromagnetic and hadronic sections. Section 3 states that dedicated simulation studies allow a maximum decrease of the electromagnetic energy by 8% and a corresponding increase of the hadronic energy by 15%, included as systematic uncertainties. Because these shifts are anticorrelated, they lower the measured E_em/E_had ratio by up to roughly 20% (0.92/1.15). This shift is in exactly the direction of all the model predictions, which lie below the data. The paper states that 'all model predictions are lower than the data' without explicitly confirming whether this remains true after applying this systematic shift. If any generator falls inside the shifted systematic band, the universal wording of the abstract overstates the result: the correct statement would be that the central values of all generators overestimate the hadron fraction, while some or all are consistent with data within the intercalibration uncertainty. Since the paper uses this discrepancy to argue that the models 'cannot explain the muon deficit' in air showers, the robustness of the claim to this calibration uncertainty is decisive. The reader's rationale incorrectly asserts that the intercalibration systematic shifts the data away from the models; the quoted numbers shift it toward them.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a CMS measurement of the average total, electromagnetic, and hadronic energy reconstructed in the CASTOR calorimeter at pseudorapidities -6.6 < eta < -5.2, as a function of the charged-track multiplicity at |eta| < 2, in proton-proton collisions at sqrt(s) = 13 TeV. The analysis uses 0.22 inverse nanobarns of low-luminosity 2015 data, an unbiased bunch-crossing trigger, and a modified pixel tracking algorithm for the solenoid-off configuration. The data are compared with PYTHIA 8 tunes, EPOS LHC, SIBYLL 2.1/2.3c, QGSJET-II.04, and HERWIG 7.1. For models without full detector simulation, the paper introduces a forward-folding method based on four-dimensional migration matrices and makes the matrices available in a RIVET plugin. The central physics claim is that all tested generators overestimate the fraction of energy going into hadrons, and the paper connects this to the muon deficit in ultra-high-energy air shower simulations.","tokens_in":30296,"tokens_out":4851,"duration_ms":54609,"significance":"If the result holds, it is a valuable and novel measurement: it is the first to correlate very-forward CASTOR energy with central charged multiplicity at 13 TeV, and it provides new constraints on underlying-event modelling, forward particle production, and cosmic-ray interaction models. The analysis has clear strengths: an unbiased trigger, quantified vertex and pileup rejection with residual uncertainties, a detector-level presentation that avoids unfolding, a forward-folding procedure validated against full detector simulation to better than 1%, and a public RIVET plugin with the migration matrices. The main physics conclusion about the hadronic energy fraction, however, depends on the relative electromagnetic/hadronic calibration of CASTOR, and the paper does not currently demonstrate the robustness of the universal 'all generators' statement to that systematic uncertainty.","major_comments":[{"comment":"The universal claim that all generators overestimate the fraction of energy going into hadrons is not shown to survive the quoted intercalibration systematic uncertainty. Section 3 states that dedicated simulation studies allow a maximum decrease of the electromagnetic energy by 8% and a corresponding increase of the hadronic energy by 15%, and that these shifts are anticorrelated. Applied together, they lower the measured E_em/E_had ratio by about 20% (0.92/1.15), i.e., in the direction of all model predictions. The paper states in Section 5 that 'all model predictions are lower than the data' and in Section 6 that the models 'cannot explain the muon deficit', but it does not show whether every generator remains below the data after applying this correlated shift. Please quantify the comparison after applying the intercalibration shifts, or show residuals relative to the full asymmetric systematic band, and qualify the abstract and Section 6 statements if any model becomes consistent with data within this uncertainty.","section":"Section 3, Section 5, Fig. 3, Abstract"},{"comment":"The claim that all generators overestimate the hadronic energy fraction is presented without a quantitative significance or goodness-of-fit measure. Since the abstract makes a universal statement about every tested generator, the paper should report, for each generator and multiplicity bin, the size of the discrepancy relative to the total systematic uncertainty, and state explicitly whether the discrepancy is dominated by the intercalibration uncertainty or by the model predictions themselves. This would make the central claim reproducible and would clarify how much weight the muon-deficit conclusion can bear.","section":"Section 5, Fig. 3"}],"minor_comments":[{"comment":"The phrase 'fraction of energy going into hadrons' refers to the energy reconstructed in the hadronic section of a non-compensating calorimeter, not directly to the particle-level hadronic energy. Because the comparison is forward-folded this is internally consistent, but the wording risks overinterpretation; please state explicitly in the abstract or in Section 5 that this is a detector-level quantity.","section":"Abstract"},{"comment":"The axis label of Fig. 3 should unambiguously define the ratio, preferably as <E_em^reco>/<E_had^reco>, directly on the axis; the current layout is ambiguous and the caption alone does not remove the ambiguity.","section":"Fig. 3"},{"comment":"The ranges of the pileup-rejection uncertainty (1-8% for total and electromagnetic energies, 1-10% for hadronic energy) and the resulting total uncertainties (18-19%, 18-20%, 20-26%) are not fully explained; a sentence on why the hadronic component is more affected would be helpful.","section":"Table 1"},{"comment":"The notation k_ij^lm for the migration matrix is introduced compactly; a brief definition of the four indices in one place would improve readability, since the same indices are used in Eq. (1) and in the bin description.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The measurement itself appears carefully done and the forward-folding tool is a genuine asset. The main issue is that the abstract and summary make a stronger statement than the quoted intercalibration systematic currently supports; this should be fixable with a quantitative check and revised wording, not new data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a genuinely new measurement, the first correlation of very forward energy in CASTOR (-6.6 < eta < -5.2) with central track multiplicity (|eta| < 2) at 13 TeV, and it ships a reusable forward-folding technique plus a RIVET plugin. The analysis is careful: minimal-bias trigger, explicit pileup rejection with quantified residuals, track efficiency corrections, and a forward-folding matrix validated against full detector simulation to better than 1%. Statistical uncertainties are small; the dominant 17% CASTOR energy scale is honestly stated. The paper does what a measurement paper should: it presents the data in a way that future generator tunes can use directly.\n\nThe main physical result is that average total, electromagnetic, and hadronic energy in CASTOR all increase with central multiplicity, and that most generators describe the total energy but predict a smaller electromagnetic-to-hadronic ratio than the data. That discrepancy is interesting and potentially relevant to the air-shower muon puzzle.\n\nThe soft spot is the abstract's universal wording: 'All generators considered overestimate the fraction of energy going into hadrons.' The intercalibration systematic is -8% on electromagnetic energy and +15% on hadronic energy, and these are anticorrelated, lowering the measured EM/hadronic ratio by up to roughly 20%. That shift is in exactly the direction of all the model predictions. The gray band in Fig. 3 presumably includes this uncertainty, but the paper does not explicitly state whether any generator falls inside the band after applying the shift. The reader's summary incorrectly claimed the shift moves the data away from the models; it moves the data toward them. So the honest statement is that the central values of all generators lie below the data, while some models may be consistent within the intercalibration uncertainty. The most recent tunes, PYTHIA 8 CP5 and SIBYLL 2.3c, do disagree significantly, and the paper's conclusions about those tunes hold up. This is a wording/robustness issue rather than a fatal flaw: the measurement itself is not circular, no parameters are fit to data, and the forward-folding model dependence is evaluated with four different migration matrices.\n\nThis paper deserves a serious referee. It is aimed at people working on underlying event modeling, forward physics, and cosmic-ray air shower generators. I would cite it if I worked in those areas. Recommend acceptance after a minor revision that softens the 'all generators overestimate' claim to match the intercalibration uncertainty, and ideally a table showing the ratio with the systematic shift applied.","headline":"A solid first measurement correlating CASTOR very-forward energy with central multiplicity at 13 TeV; the central hadron-fraction claim is real but somewhat overstated relative to the quoted intercalibration systematic.","tokens_in":30889,"tokens_out":2412,"would_cite":true,"duration_ms":26443,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that in 13 TeV proton-proton collisions, the average very forward energy rises with central track multiplicity, and while every tested Monte Carlo generator reproduces the total energy, all of them overestimate the…","keywords":["very forward energy","CASTOR calorimeter","underlying event","track multiplicity","hadronic interaction models","forward folding","cosmic-ray air showers","proton-proton collisions"],"falsifier":"Shifting the CASTOR intercalibration to its quoted extremes—electromagnetic energy down 8% and hadronic up 15%—and re-extracting the hadronic fraction would show whether data still disagree with all generators, and a dedicated test-beam comparison of electron-versus-pion response against GEANT4 would identify which calibration is right.","tokens_in":29844,"feed_emoji":"⚛️","tokens_out":11259,"duration_ms":105498,"temperature":0.7,"pith_summary":"At 13 TeV proton-proton collisions, this paper correlates, for the first time, the energy flowing into the very forward pseudorapidity region $-6.6 < \\eta < -5.2$ with the number of charged tracks in the central region $|\\eta| < 2$, separating the forward energy into electromagnetic and hadronic components. The central result is that the average total forward energy grows with central track multiplicity and is reproduced reasonably well by all tested Monte Carlo generators, but the electromagnetic-to-hadronic split is not: the data carry a larger electromagnetic fraction than any generator predicts. In other words, every generator considered overestimates the fraction of very forward energy going into hadrons, with the largest deviations for the newest tunes. If this holds, it sharpens the known problem that cosmic-ray air-shower simulations underestimate muon production, because the data show that even less energy is available for hadronic processes than the models assume.","feed_headline":"All generators put too much energy into forward hadrons","feed_subtitle":"New CASTOR data show more electromagnetic energy than any generator predicts, a clue to the muon deficit.","key_machinery":"The measurement is carried by two pieces of apparatus-level machinery. The first is the CASTOR calorimeter, a quartz-tungsten sampling calorimeter whose first two channels form a 20-radiation-length electromagnetic section and whose remaining twelve channels form a hadronic section of 10 interaction lengths total; its electron and hadron response, including noncompensation, was measured in test beams and is simulated with GEANT4. The second is a 'forward folding' procedure: four-dimensional migration matrices, constructed from full detector simulations of PYTHIA 8 CUETP8M1, PYTHIA 8 4C+MBR, EPOS LHC, and SIBYLL 2.1, map every generator's true central multiplicity and forward energy onto reconstructed track multiplicity and CASTOR energy, so that any model can be compared to the data at detector level without unfolding. Five matrices are provided, four generator-specific plus one averaged, and the spread among them is assigned as a systematic uncertainty.","core_discovery":"The paper establishes a detector-level measurement of the average total, electromagnetic, and hadronic energy in the CASTOR calorimeter at $-6.6 < \\eta < -5.2$ as a function of the reconstructed charged-track multiplicity at $|\\eta| < 2$, using an unbiased low-luminosity 13 TeV dataset. The total energy rises steeply at low multiplicity and flattens at high multiplicity; all generators reproduce this rise, with SIBYLL 2.1 the closest. The electromagnetic component is generally well described, while most generators overestimate the hadronic component. The ratio of hadronic to electromagnetic energy is lower in data than in all tested models, and this shortfall is largest for PYTHIA 8 CP5 and SIBYLL 2.3c, the most recent tunes. The ratio is approximately independent of track multiplicity, and the paper argues that this disagreement rules out the simplest explanation of the air-shower muon deficit, namely that models simply need more hadronic energy in the forward region.","pith_inferences":["The ratio of hadronic to electromagnetic energy is roughly constant across track multiplicities, so a single energy-independent adjustment to the neutral-to-charged pion ratio in fragmentation might bring models into agreement with data.","Extending this measurement to other collision energies (e.g., 5.02 or 7 TeV) would test whether the hadronic overestimate grows with energy, which would sharpen its relevance for ultra-high-energy cosmic-ray modeling.","The forward-folding matrices could be applied to generators with different parton-shower or color-reconnection schemes to isolate which modeling ingredient controls the hadronic energy fraction.","A particle-level measurement using unfolding, rather than forward folding, would confirm that the effect is not introduced by the detector-response correction itself."],"forward_implications":["The average total very forward energy rises with central track multiplicity, and all tested generators reproduce this rise reasonably well, showing that underlying-event tunes from central rapidity extrapolate to the far forward region for the total energy.","All tested generators overestimate the hadronic fraction of the very forward energy, meaning they produce too many long-lived hadrons relative to neutral pions and photons in this phase space.","The two newest model tunes, SIBYLL 2.3c and PYTHIA 8 CP5, disagree most strongly with the measured electromagnetic-to-hadronic ratio, so retuning them on these data is a direct next step.","Because the data put even more energy into the electromagnetic channel than the models, the cosmic-ray muon deficit cannot be cured by simply shifting more energy into hadronic production in the forward region.","The forward-folding matrices are published in a RIVET plugin, so any future generator or tune can be compared to this dataset without a full detector simulation."],"supporting_citations":[{"why":"Supplies the test-beam measurements of shower containment and noncompensation that set the CASTOR electromagnetic/hadronic response model.","marker":"[7]"},{"why":"Establishes the CASTOR energy calibration and the estimators for particle-level electromagnetic and hadronic energies used here.","marker":"[9]"},{"why":"Provides the connection between the electromagnetic-to-hadronic energy fraction and muon production in air showers, the basis for the physics interpretation.","marker":"[16]"},{"why":"Defines the PYTHIA 8 generator that supplies the central tunes compared to the data.","marker":"[20]"},{"why":"Provides the EPOS LHC predictions, an alternative collective-hadronization model compared to the data.","marker":"[23]"},{"why":"Provides the SIBYLL 2.1 predictions, the cosmic-ray generator found to best describe the total-energy multiplicity dependence.","marker":"[24]"},{"why":"The GEANT4 simulation toolkit used to model CASTOR shower response, on which the forward-folding matrices and intercalibration systematics depend.","marker":"[25]"},{"why":"The RIVET tool through which the forward-folding migration matrices are released for comparison of any model to the data.","marker":"[35]"}],"fun_headline_variants":["CMS forward energy: all models overestimate hadrons","Hadronic to EM ratio in forward region lower than predicted","New CASTOR measurement rules out simple muon fix","Forward hadronic energy lower than generators predict","CMS: hadronic fraction too high in all Monte Carlos"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result depends on the assumption that the CASTOR intercalibration corrections (electromagnetic energy down by up to 8%, hadronic up by up to 15%) correctly convert the measured section energies into the true particle-level electromagnetic and hadronic energies, so if the shower simulation behind those corrections is wrong, the claimed excess of electromagnetic energy in data could change or disappear.","fun_headline_variants_meta":{"raw":{"variants":["CMS forward energy: all models overestimate hadrons","Hadronic to EM ratio in forward region lower than predicted","New CASTOR measurement rules out simple muon fix","Forward hadronic energy lower than generators predict","CMS: hadronic fraction too high in all Monte Carlos"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000177,"raw_usage":{"total_tokens":1269,"prompt_tokens":894,"completion_tokens":375,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":299}},"tokens_in":510,"tokens_out":375,"duration_ms":4483,"temperature":1.0,"reasoning_tokens":299,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:04:11.061793+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Shifting the CASTOR intercalibration to its quoted extremes—electromagnetic energy down 8% and hadronic up 15%—and re-extracting the hadronic fraction would show whether data still disagree with all generators, and a dedicated test-beam comparison of electron-versus-pion response against GEANT4 would identify which calibration is right.","supporting_citations":[{"cited_title":"Performance studies of a full-length prototype for the CASTOR forward calorimeter at the CMS experiment","cited_arxiv_id":null,"evidence_quote":"Supplies the test-beam measurements of shower containment and noncompensation that set the CASTOR electromagnetic/hadronic response model."},{"cited_title":"Measurement of the inclusive energy spectrum in the very forward direction in proton-proton collisions at sqrt(s) = 13 TeV","cited_arxiv_id":"1701.08695","evidence_quote":"Establishes the CASTOR energy calibration and the estimators for particle-level electromagnetic and hadronic energies used here."}],"review_version":1}