{"id":"2accd898-47b2-4815-b4c3-617001661207","arxiv_id":"2607.10637","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"A matched-simulation detector-level pipeline reconstructs 60–85 degree air showers from RF-chain voltage traces with 10.75–11.04% energy scatter and near-zero mean residual, without a fitted global energy scale.","lead":"Radio detectors that record raw voltage traces—not clean electric fields—can still recover cosmic-ray shower direction and energy inside a simulation: mean energy offsets are below one percent with about 11% event-to-event scatter for 60–85 degree showers. The test is explicitly a matched simulation closure, not deployed-array resolution or an independent energy calibration.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The azimuth-collapsed reference-energy template assumption is the load-bearing simplification; the 85–90° breakdown already shows it fails at the edge, and the 60–85° claim inherits the risk.","rationale":"The reader correctly identifies the template assumption as the weakest link. My concern sharpens that: the 85–90° breakdown is strong evidence that the one-dimensional template fails under exactly the conditions present at the upper edge of the primary domain, and the internal zenith trend in Table 6 suggests that the 60–85° aggregate mean masks residual structure. The OOF training on the same simulation framework is a secondary but real circularity risk. This is not an internal contradiction, but it is a correctness risk: the central claim is conditional on an unvalidated azimuth/energy invariance. The paper is honest about boundaries, but the primary domain claim would be materially strengthened or weakened by the proposed test. Hence CONDITIONAL: acceptance is justified only if the template invariance check passes; if not, the central claim needs substantial qualification.","tokens_in":16188,"tokens_out":1613,"duration_ms":19224,"concrete_test":"Split the 693-event primary analysis sample by true azimuth quadrants and by true energy quartiles, and recompute the mean and scatter of δE in each subset for the hard-gated denoised branch. If mean residuals shift by more than ~3 percentage points across azimuth or energy subsets, the template invariance assumption is the cause; if they stay within ~1–2 percentage points, the aggregate closure is more robust. Also run the fit with a template generated at a different reference azimuth (e.g., 0°) and a different reference energy (e.g., 1 EeV) and compare residuals.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—near-zero mean residual and ~11% scatter with no fitted scale over 60–85°—depends critically on the azimuth-collapsed four-arm template (Eq. 3.5), which is generated only at azimuth 45°, reference energy 0.316 EeV, and two primaries, then applied to all events at all energies and azimuths. The paper's own upper-boundary result (+14.2% mean, 19.9% scatter, Table 6) demonstrates that this template model breaks at 85–90°. Within the 60–85° primary domain, Table 6 shows a systematic zenith trend: mean residuals go from -5.1% (60–70°) to +7.1% (80–85°), a 12-percentage-point swing. This is not random noise—it is structured bias that happens to partially cancel in the aggregate. The azimuth-collapsed symmetry assumption (Eq. 3.5) also removes any azimuth dependence of the footprint, even though the geomagnetic angle and charge-excess interference vary with azimuth; the paper does not test whether the closure holds per azimuth. The OOF interpolation weight p_p is trained on primary labels from the same ZHAireS/QGSJET-II-04 framework, so the composition-averaged closure partly reflects the training framework's own templates. The quoted near-zero aggregate mean may therefore be a balance of compensating zenith/azimuth biases rather than evidence that the template family is correct.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a detector-level simulation-closure test for radio air-shower reconstruction. 972 ZHAireS event packages (QGSJET-II-04, proton/iron, 45–89 deg, 0.001–4 EeV) are propagated through an RF chain and trigger; reconstruction uses a voltage-domain ADF plus joint timing–amplitude axis fit, native-grid FFT inversion to the 50–200 MHz electric field, a peak |E_xπ| observable along the shower-plane v×B axis, and fixed symmetric four-arm iron/proton templates at reference energy 0.316 EeV. A nested five-fold out-of-fold logistic-regression weight interpolates between the Fe and p endpoint energies. On a common-quality sample of 725 events (693 with reconstructed zenith 60–85 deg), the three waveform branches (raw-noisy weighted, hard-gated denoised equal-weight, clean equal-weight) achieve mean energy residuals of -0.05%, -0.30%, -0.66% with event-to-event scatter 10.75–11.04%, and a median angular separation of about 0.052 deg. Boundary zenith ranges are reported separately, with the 85–90 deg bin showing a mean residual of +14.2% and scatter 19.9%. The paper explicitly restricts its claims to matched-simulation, cross-fitted closure and does not claim deployed-array calibration or independent energy resolution.","tokens_in":16544,"tokens_out":6671,"duration_ms":73033,"significance":"If the claims hold, this is a useful methodological benchmark: it shows that a trigger-to-energy reconstruction chain can operate without a global calibration constant in a controlled simulation, with reproducible data/scripts, bootstrap intervals, and an honest report of boundary failures. The strengths are the transparent cross-fitting design, the explicit statement of limitations, the separate reporting of edge domains, and the availability of code/data to regenerate figures and tables. However, the central 'near-zero mean residual' result currently rests on an aggregate mean that averages over a strong within-domain zenith trend, and the 'no fitted multiplicative energy scale' wording is contradicted by the per-event amplitude fit in Eq. (3.6). These issues are fixable but require reworking the presentation and the main interpretive claims.","major_comments":[{"comment":"Eq. (3.6) performs a per-event least-squares amplitude fit â_k and sets E_k = 0.316 â_k EeV. The text immediately after says 'No global, primary-dependent, event-dependent or zenith-dependent amplitude scale is fitted,' which is internally inconsistent: â_k is an event-dependent multiplicative amplitude scale. The abstract's claim of doing reconstruction 'without a fitted multiplicative energy scale' is therefore misleading. What the analysis actually avoids is a global calibration constant or a post-fit multiplicative correction. Please reword this central claim precisely.","section":"§3.5, Eq. (3.6); Abstract"},{"comment":"The headline mean residuals of -0.05% to -0.66% are aggregates over the 60–85 deg domain. Table 6 shows a systematic zenith trend within this domain: -5.1% (60–70), -1.4% (70–75), +5.3% (75–80), +7.1% (80–85). The near-zero aggregate mean is thus a cancellation of opposing biases whose relative weights depend on the trigger-selected zenith distribution, not evidence of per-zenith template validity. Please present zenith-differential residuals as the primary closure statistic (or reweight to a stated zenith distribution) and temper the abstract's emphasis on the aggregate mean.","section":"§4.4, Table 6; Abstract"},{"comment":"The template family is azimuth-collapsed: generated only at azimuth 45°, at one reference energy, and averaged symmetrically over four arms. No residual-versus-true-azimuth analysis is provided, although geomagnetic-angle and charge-excess interference vary with azimuth. The upper-boundary failure and the +7.1% mean at 80–85 deg suggest the one-dimensional v×B peak profile already loses accuracy inside the primary domain. Add a diagnostic of residuals versus true azimuth (and ideally true energy) to support the claim that this template family is adequate for the entire 60–85 deg primary domain.","section":"§3.4, Eq. (3.5); §4.4"}],"minor_comments":[{"comment":"'The trigger configuration uses the horizontalychannel' appears to be a typo; it should be 'horizontal y channel'.","section":"§2.2"},{"comment":"The symbol p is used for the integration variable in Eq. (3.4) and also as the proton-endpoint label (p) and the parameter p = n_src sinθ. This triple use is confusing; consider renaming the endpoint label or the parameter.","section":"Eq. (3.4) and Sec. 3.4"},{"comment":"The term 'GP80-like' is not defined or referenced. Provide a citation or a brief description of GP80 so the reader can understand the intended array concept.","section":"§1"},{"comment":"The panel titles report true zenith for display, but the text should state explicitly why reconstructed zenith is not used in the titles, to avoid confusion with the primary-domain definition in Eq. (3.9).","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well-structured and unusually transparent about limitations, and the reproducible-package claim is a strong point. However, the internal contradiction about the fitted amplitude scale and the aggregate-mean cancellation over the zenith trend are substantive; they affect how the headline numbers should be read. I recommend major revision rather than rejection, because the underlying simulation and analysis methodology are sound and the concerns can be addressed by re-presentation and additional diagnostics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, well-scoped simulation closure—probably the first to show that sparse RF-chain-triggered ADC traces can be inverted back to fields and reconstructed end-to-end with ~11% energy scatter and sub-0.1° direction. The authors deserve credit for the unglamorous but necessary work: native-grid inversion, a three-branch comparison, a common quality sample defined before fitting, OOF separation, bootstrap intervals, and separate reporting of boundary zenith bins where the method fails. The paper is also honest about scope: it explicitly says these numbers are matched-simulation closure, not deployed-array resolution or independent calibration.\n\nThe soft spot is the one the stress-test flags, and it is real: the aggregate mean near zero is a balance of structured zenith biases. Their own Table 6 shows mean residuals from -5.1% in the 60-70 bin to +7.1% in the 80-85 bin, a twelve-point swing, and +14.2% at 85-90. So the headline mean of -0.05% depends on the sample's mix of angles and primaries. A reader quoting only the abstract mean would be overreading it. They do report this structure in the text, so it is not hidden—but the abstract could be read as if the estimator is globally unbiased, and it isn't.\n\nThe template side is the bigger underlying assumption: azimuth-collapsed, generated at one azimuth and one reference energy, applied to all events at all azimuths and energies. There is no per-azimuth closure test, and the OOF interpolation weight is trained on primary labels from the same ZHAireS framework that produced the test events. That keeps this a matched-simulation closure rather than an independent validation. These are limitations, not fatal flaws; they are disclosed, and the discussion section says the obvious things about the need for 2D footprints and external validation.\n\nOn balance, I think the paper holds up as a careful detector-level simulation study within its stated domain. It advances the practical methodology for GRAND-like sparse radio reconstruction. The main referee-level push should be on the template's azimuth/energy extrapolation and on presenting the zenith trend as the primary energy result rather than the aggregate mean.","headline":"A careful, honest matched-simulation closure for RF-chain radio reconstruction; read the zenith-resolved tables rather than the abstract mean.","tokens_in":17036,"tokens_out":3285,"would_cite":true,"duration_ms":37608,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Trigger-selected radio voltage traces alone yield air-shower energies with about 11% scatter and near-zero mean bias in simulated events, with no fitted scale.","keywords":["ultra-high-energy cosmic rays","extensive air showers","radio detection","air-shower reconstruction","RF-chain inversion","detector-level simulation closure","energy reconstruction","out-of-fold validation"],"falsifier":"Split the 693-event primary sample by true azimuth: if mean energy residuals vary systematically across azimuth, especially far from the template azimuth, beyond the reported bootstrap intervals, the azimuth-collapsed template assumption fails. A cleaner test would be to run the same pipeline with an independent shower-simulation code and check whether the near-zero mean residual persists or shifts.","tokens_in":16045,"feed_emoji":"📡","tokens_out":6230,"duration_ms":64200,"temperature":0.7,"pith_summary":"The paper attempts to show that an autonomous radio array can reconstruct cosmic-ray air-shower direction and energy starting from the same kind of input a real detector would deliver—trigger-selected ADC voltage traces—rather than from idealized electric-field footprints. In a matched simulation, 693 quality-selected events with reconstructed zenith between 60° and 85° have mean energy residuals between -0.66% and -0.05% and event-to-event scatter between 10.75% and 11.04% across three waveform branches: raw-noisy weighted, hard-gated denoised equal-weight, and clean equal-weight reference. The hard-gated branch has a median angular separation of 0.052°. These numbers are produced by inverting the radio-frequency chain on its native frequency grid and blending iron- and proton-endpoint template fits with a cross-fitted interpolation weight, all without a fitted multiplicative energy scale. The paper is careful that this is cross-fitted closure within one simulation framework, not deployed-array resolution or an independent calibration; the upper-zenith boundary (85–90°) fails with +14.2% mean residual and 19.9% scatter.","feed_headline":"Noisy radio traces alone close shower energies to ~11%","feed_subtitle":"Simulated events in the 60–85° zenith band reconstruct with mean bias under 0.7% and about 11% spread, using no fitted scale.","key_machinery":"The load-bearing mechanism is native RF-chain inversion on each event's own discrete Fourier grid, restricted to a 50–200 MHz passband, which recovers electric-field waveforms from digitized voltages and yields the station observable |Exπ|, the peak amplitude projected along the shower-plane v×B axis. Energy estimation then uses symmetric four-arm lateral-distribution templates built from simulated iron and proton footprints at one reference energy, radio-core recentered by a refractive-shift proxy; a nested five-fold out-of-fold logistic-regression weight interpolates between the two endpoint energies so that no event is predicted by a model trained on it. The four-arm profile is the key si","core_discovery":"The central claim is that the detector-level inverse problem for radio air showers is solvable: from L1-triggered packages of noisy ADC traces, a pipeline of voltage-domain geometry, native-grid RF-chain inversion, 50–200 MHz v×B peak-amplitude extraction, symmetric four-arm iron/proton template fits, and nested five-fold out-of-fold endpoint interpolation reconstructs energy and direction with small mean bias and about 11% scatter across 693 common events in the 60–85° zenith band, without fitting any amplitude or energy scale. The paper also claims that the raw-noisy and hard-gated branches match the clean reference to within a few tenths of a percent in paired mean difference, meaning the","pith_inferences":["Inference: Because the template family is generated at a single azimuth and one reference energy, the closure likely degrades with distance from that azimuth; a true-azimuth split of the 693-event sample would reveal whether charge-excess asymmetry not captured by the azimuth-collapsed profile is hiding in the 11% scatter.","Inference: The same out-of-fold interpolation architecture could be retrained on two-dimensional radio-footprint templates with explicit energy and shower-development dependence; if that eliminates the upper-boundary offset without a fitted scale, the approach would extend to near-horizontal showers and potentially yield composition-sensitive energy estimates.","Inference: Applying the pipeline to an independent shower generator or a different site's radio-frequency chain would test transferability; a shifted mean residual would quantify how much of the closure is specific to the matched simulation framework.","Inference: The near-zero mean residual across all three waveform branches suggests that a fluence-based energy estimator, once its signal window and noise subtraction are calibrated, may achieve comparable or better scatter; comparing peak-amplitude and fluence estimators within this detector-level closure would settle which observable is more noise-tolerant."],"forward_implications":["If the closure is representative, autonomous arrays can reconstruct shower energy and direction from raw triggered traces without requiring ideal electric-field footprints, a fitted energy scale, or per-event composition labels.","The small paired branch differences (0.24–0.60 percentage points relative to clean) imply that denoising and noise-weighting do not materially change the energy result; the roughly 11% scatter is dominated by shower-to-shower and template-model effects, not by electronic noise.","Fixed iron-only or proton-only templates carry opposite mean biases (+8.5% and -8.3% in the denoised branch); out-of-fold interpolation reduces the composition-averaged offset to near zero, so the interpolation mechanism is what makes the single estimator usable over the full angular range.","The sharp upper-boundary failure (85–90°: +14.2% mean, 19.9% scatter) sets an explicit validity limit: the same pipeline and templates should not be applied beyond 85° without a two-dimensional or physically corrected footprint model.","The median angular separation of about 0.05° is a matched-simulation closure value; adding deployed-array timing, positioning, and calibration systematics will degrade it, so it should not be read as field angular resolution."],"fun_headline_variants":["Radio showers: 11% energy spread from raw voltage traces","No fitted scale: air-shower energies closed to 11%","Shower energy from noise: 11% scatter, sub-0.7% bias","Detector-level radio reconstruction hits 11% energy resolution","Radio air showers: unbiased energies, 11% scatter, no scale fit"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The energy closure rests on the assumption that one azimuth-collapsed, four-arm v×B peak-amplitude template family—generated at a single reference azimuth and energy, with a refractive radio-core shift proxy and no charge-excess correction—adequately describes every event in the 60–85° band, and the paper's own +14.2% upper-boundary residual marks where that assumption begins to break.","fun_headline_variants_meta":{"raw":{"variants":["Radio showers: 11% energy spread from raw voltage traces","No fitted scale: air-shower energies closed to 11%","Shower energy from noise: 11% scatter, sub-0.7% bias","Detector-level radio reconstruction hits 11% energy resolution","Radio air showers: unbiased energies, 11% scatter, no scale fit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000215,"raw_usage":{"total_tokens":1309,"prompt_tokens":834,"completion_tokens":475,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":390}},"tokens_in":578,"tokens_out":475,"duration_ms":4734,"temperature":1.0,"reasoning_tokens":390,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T07:10:35.761656+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Split the 693-event primary sample by true azimuth: if mean energy residuals vary systematically across azimuth, especially far from the template azimuth, beyond the reported bootstrap intervals, the azimuth-collapsed template assumption fails. A cleaner test would be to run the same pipeline with an independent shower-simulation code and check whether the near-zero mean residual persists or shifts.","supporting_citations":[],"review_version":2}