{"id":"334c11b8-81f6-40e0-83c3-46b6370ca66b","arxiv_id":"2507.06695","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"GRANDProto300's filtering pipeline selects 26 solid cosmic-ray candidate events from 533,466 radio triggers recorded with 46 antennas between December 2024 and March 2025.","lead":"GRANDProto300, a prototype radio array in the Gobi desert, reports 26 solid cosmic-ray candidate events from its first stable run using an automated signal-filtering pipeline. The result is a key step toward validating the radio technique for the future GRAND experiment and studying where cosmic rays come from.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manual visual cuts with an admitted northern-azimuth bias are the load-bearing step: without a blinded re-review or objective rescoring, the 26-candidate count is not robust.","rationale":"The paper's central claim is that the pipeline yielded 26 solid cosmic-ray candidates from GP300 data. The most fragile link in the chain of cuts is the final manual selection (Section 2.5), which is explicitly biased toward the northern azimuths (Section 3). If the visual reviewers were more permissive for northern events, the candidate list and count would not be reproducible. The reader's weakest assumption identifies the same issue, so I agree. A blinded re-review is the natural test. The reader's CONDITIONAL verdict is appropriate; my concern does not require changing it. A secondary presentation issue is that the 'Full pipeline' efficiency in Fig. 5 (right) is listed as '7×10^-5%' while the caption says these are excluded fractions; this value is inconsistent with the product of the individual cut efficiencies and with the 41-candidate count, suggesting a units or labeling error, but it does not affect the validity of the candidates themselves.","tokens_in":9785,"tokens_out":11830,"duration_ms":135162,"concrete_test":"Perform a blinded re-evaluation: have at least two independent analysts apply the Section 2.5 visual cuts to the full set of events that pass the systematic cuts, with event azimuths masked. Compare the resulting candidate lists and counts. If the analysts' accepted lists disagree by more than a few events, or if unmasking reveals that the final 26 are concentrated in φ=0–45° and 315–360° while the pre-cut events are not, the manual-cut bias is confirmed and the 26-candidate claim is analyst-dependent. As an objective alternative, encode the visual criteria as automated cuts (e.g., pulse width <100 ns, coherent timing residuals <0.5°, footprint convex hull area, spectral index) and check whether a comparable candidate set survives without human input.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The final count of 26 solid candidates depends critically on the manual visual cuts described in Section 2.5 (trace shape, footprint compactness, timing consistency). These criteria are not reduced to quantitative, reproducible rules, and Section 3 explicitly admits that 'particular attention was given to the Northern part of our dataset when performing manual cuts (from azimuths φ=0–45° and 315–360°).' If the visual reviewers' judgment was influenced by this stated azimuthal preference, the candidate set could contain northern events that would have been rejected had they arrived from other directions, and the number 26 would not be a stable output of the pipeline. The paper provides no inter-reviewer consistency check, no blinded re-review, and does not show the candidate list or an arrival-direction map for the 26 events, so the reported count cannot be independently audited. This is load-bearing because the central claim is precisely that these 26 events are solid cosmic-ray candidates; if the manual selection is biased, the claim reduces to an unquantified human judgment.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents the first cosmic-ray search pipeline for GRANDProto300, applied to 533,466 coincident events recorded between December 2024 and March 2025 with 46 antennas. The pipeline combines coincidence, clustering, polarization, and quality cuts (antenna number, zenith, PWF error, SNR, RMS) followed by manual visual cuts and a reconstruction-quality selection. The authors report 41 candidate events after the systematic pipeline and 26 \"solid\" candidates after manual and reconstruction-quality cuts, with energy reconstructions consistent with the expected range.","tokens_in":9988,"tokens_out":7326,"duration_ms":69885,"significance":"If the 26 candidates are genuine cosmic rays, the work demonstrates the autonomous radio-detection technique at the prototype scale and supports the GRAND design. The paper is transparent about its preliminary nature and provides efficiency estimates for each cut based on ZHAireS simulations with data-based noise. However, the load-bearing manual visual cuts are subjective and the admitted northern-azimuth bias in Section 3 undermines the robustness of the reported candidate count without additional checks.","major_comments":[{"comment":"The manual visual cuts (trace shape, footprint compactness, timing consistency) are not defined by quantitative thresholds, and Section 3 states that \"particular attention was given to the Northern part of our dataset when performing manual cuts.\" Because the final count of 26 candidates is obtained after these manual cuts, the central claim is not reproducible and may be biased; the paper provides no inter-reviewer consistency check, blinded re-review, or test of the azimuthal sensitivity of the selection. This needs to be addressed, e.g., by specifying measurable criteria (pulse width, spectral slope, footprint ellipticity, timing-residual limits), performing a blinded re-review, and comparing the candidate azimuthal distribution to the detector exposure.","section":"§2.5, §3"},{"comment":"The polarization cut is calibrated using a noise sample consisting only of events from the plane and transformer directions, but it is then applied to the full dataset; the reported 38% noise rejection may not hold for the broader noise population. Moreover, the noise distribution used to choose (eV·eB)_cut is drawn from the same dataset that is later filtered, introducing a partial circularity that should be mitigated with a disjoint calibration sample or a hold-out period.","section":"§2.3"},{"comment":"The cut-efficiency table reports percentages without uncertainties, and the final line \"Full pipeline 7×10^{-5}%\" is ambiguous: 41/533,466 ≈ 7.7×10^{-5} is the retained fraction, not the excluded fraction, and it is inconsistent with the column header. Uncertainties on the retained fraction, ideally including the manual-cut step, should be provided to support the claimed purity of the sample.","section":"Fig. 5 (right)"},{"comment":"The SNR cut (SNR_Y ≥ 5) is described as \"arbitrary\" and the PWF-error cut as \"more artificial.\" The paper does not demonstrate that the final candidate list is stable under reasonable variations of these thresholds; a one-at-a-time sensitivity scan showing the number of candidates as a function of each threshold would substantiate the claim that the 26 candidates are not an artifact of the specific cut choices.","section":"§2.4"}],"minor_comments":[{"comment":"The abstract contains a typo: \"Between November 2024 u to May 2025\" should read \"November 2024 and May 2025.\"","section":"Abstract"},{"comment":"The phrase \"transient noise induces events\" would be clearer as \"transient noise produces events.\"","section":"§2.1"},{"comment":"The notation for the b-ratio (e𝑉· e𝐵) has inconsistent spacing between the vector dot; please use consistent math formatting throughout.","section":"§2.3"},{"comment":"The sentence \"26 solid cosmic-ray candidates, displayed in [1], Figure 5\" appears to refer to a companion paper's figure, but this manuscript contains its own Figure 5; please clarify the intended cross-reference.","section":"§3"},{"comment":"The final row \"7×10^{-5}%\" should be rewritten as 7.7×10^{-5} (the retained fraction) or 7.7×10^{-3}% to avoid confusion with the column header indicating the excluded proportion.","section":"Fig. 5 (right)"},{"comment":"The definition of angular distance Δ𝜗 = min(Δ𝜃, Δ𝜙) is confusing because Δ𝜃 and Δ𝜙 have different units; either define the minimum of the absolute zenith difference and the wrapped azimuth difference, or use an angular separation on the sky.","section":"§2.2"},{"comment":"The statement that cosmic-ray traces are characterized by \"low-frequency signals\" should be quantified (e.g., peak frequency below 100 MHz) to make the visual cut reproducible.","section":"§2.5"}],"recommendation":"major_revision","confidential_remarks":"This is an ICRC proceedings contribution that fits the journal's scope and is honest about its preliminary status. The principal concern is the reliance on subjective manual cuts without auditing; I believe this is fixable with additional analyses. I did not find evidence of citation manipulation; the reference to \"[1], Figure 5\" is likely a typo."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First: the paper reports the first batch of cosmic-ray candidate events from GRANDProto300, 26 in number, found with a pipeline that is mostly a recombination of previously published techniques (clustering, polarization, footprint, PWF). That is a legitimate detector milestone, and the authors deserve credit for stating their biases openly, including the manual visual cuts and the northern-azimuth preference in Section 3. The polarization cut is grounded in ZHAireS simulations with data-based noise, which is a reasonable way to set this kind of threshold at prototype level. The cut-efficiency table (Fig. 5) is useful, and the pipeline description is sufficiently concrete to be replicated in spirit.\n\nNow the soft spots, in proportion. The main one is that the 26 candidate count is not robust against the manual visual cuts. Section 2.5 lists three criteria — trace shape, footprint compactness, timing consistency — but none is reduced to a quantitative rule, and Section 3 states that 'particular attention' was given to northern azimuths during those cuts. Without an inter-reviewer consistency check or a blinded re-review, there is no way to know whether the number 26 is a stable pipeline output or a reflection of the reviewers' prior. The candidate list and arrival-direction map are not shown, so the claim cannot be independently audited from the text. The stress-test note is right: this is load-bearing, not a side concern.\n\nSeveral smaller issues: cut efficiencies are quoted without uncertainties; the SNR threshold of ≥5 on the Y channel is called 'arbitrary' by the authors themselves; the zenith cut to 60–88° means the sample is restricted, which is fine but should be stated when interpreting the arrival directions. The polarization cut uses noise events from the same dataset that is later filtered; this has some circular flavor, but the signal model is independent, so I consider it a minor concern. The paper is a proceedings, so the absence of a full candidate table is expected, but it does limit the strength of the claim.\n\nVerdict: conditional, as the reader says. This is a promising prototype result, not a physics discovery. The honest writing and transparent pipeline make it worth a serious referee — if submitted to a journal, I would send it out, with the requirement that the manual cuts be replaced or augmented by a blinded or quantitative scoring, and that the 26 candidates be listed with their directions and basic signal parameters. Without that, the central number 26 is a claim about human judgment, not about the detector.","headline":"A transparent prototype result: first GP300 CR candidates, but the count of 26 rests on manual visual cuts with an admitted northern bias, so treat the number as provisional.","tokens_in":10560,"tokens_out":2867,"would_cite":false,"duration_ms":30134,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Six filtering stages applied to 533,466 coincidence-triggered events yield 41 cosmic-ray candidates, of which 26 survive reconstruction-quality checks and are presented as solid detections from the 46-antenna GRANDProto300 prototype.","keywords":["cosmic rays","radio detection","GRANDProto300","air showers","pipeline","polarization cut","b-ratio","prototype array"],"falsifier":"Re-run the pipeline on time-scrambled or phase-randomized versions of the coincident-event data: if a comparable number of candidates survive the visual cuts, the selection is not isolating physical air showers; alternatively, repeating the manual cuts with azimuth hidden would reveal whether the 26 candidates disappear or concentrate toward the transformer and aircraft directions.","tokens_in":9572,"feed_emoji":"📡","tokens_out":9470,"duration_ms":98128,"temperature":0.7,"pith_summary":"The paper reports the first cosmic-ray search in data from GRANDProto300, a 300-antenna radio array prototype under deployment in the Gobi Desert. With only 46 antennas collecting stably between November 2024 and May 2025, the authors build and run a six-step filtering pipeline—coincidence triggering, planar-wavefront direction reconstruction, clustering removal of noise bursts, a polarization cut, footprint-size selection, and systematic plus visual quality cuts—over 533,466 coincident events. The pipeline returns 41 candidate events, and a reconstruction-quality stage reduces these to 26 solid cosmic-ray candidates. The aim is to show that autonomous radio detection can identify cosmic rays in the $10^{17}$–$10^{18.5}$ eV range, the energy window spanning the transition from Galactic to extragalactic sources, while the full array is still being built.","feed_headline":"26 cosmic-ray candidates found in first GRANDProto300 data","feed_subtitle":"Six pipeline stages cut 533,466 noisy triggers to 26 solid events, validating autonomous radio detection.","key_machinery":"The load-bearing estimator is the b-ratio, $e_V\\cdot e_B$: the cosine between the voltage vector measured by an antenna and the local geomagnetic field direction. Because geomagnetic emission from an air shower is polarized along $\\mathbf{k}\\times\\mathbf{B}$, a genuine cosmic-ray signal should have $e_V\\cdot e_B\\approx 0$, whereas randomly polarized noise is spread over all values; the cut keeps events with median $e_V\\cdot e_B < 0.25$. Around this estimator, the clustering cut uses Planar Wave Front (PWF) arrival-direction reconstruction to reject bursts that repeat within $\\Delta\\theta = 5^\\circ$ and $\\Delta t = 5$ s, and the footprint and timing cuts rely on the expectation that cosmic-ray events trigger a compact cluster of 5–10 antennas over a 1–10 km$^2$ footprint.","core_discovery":"The paper's central claim is that the GRANDProto300 pipeline, applied to the first stable dataset, isolates cosmic-ray events from a radio background dominated by transient noise. Starting from 533,466 coincidence-triggered events, a clustering cut removes 79% of events as bursts from an electric transformer and hovering aircraft; a polarization cut based on the b-ratio $e_V\\cdot e_B$ (the component of the measured electric field along the geomagnetic field) removes another 22%; further cuts on footprint size ($\\ge 5$ antennas), zenith angle ($60^\\circ$–$88^\\circ$), planar-wavefront reconstruction error, signal-to-noise ratio, and RMS noise remove most of the rest. The surviving events are inspected visually for short low-frequency pulses, a compact footprint of 5–10 antennas, and timing consistent with a plane wave front, yielding 41 candidates. Three independent reconstruction methods—lateral distribution function, angular distribution function, and graph neural networks—then confirm 26 of these as solid cosmic-ray candidates, with reconstructed energies consistent with the array's exposure calculations.","pith_inferences":["A testable extension the paper does not pursue: re-running the manual cuts with azimuth information blinded would measure how much of the documented northern-sky preference biases the surviving candidate list and could convert the 26 candidates into an unbiased sample.","Because the paper notes that the real background b-ratio distribution is non-uniform, a future pipeline could model the actual noise polarization density instead of assuming uniform noise; that should improve background rejection beyond the reported level and recover candidates currently lost.","The visual cuts—pulse widths of 50–100 ns, compact 5–10 antenna footprint, PWF-timing consistency—could be codified as quantitative features and automated, which is the natural bridge to the machine-learning cuts the authors list as a next step.","If the 26 candidates are genuine, their rate with 46 antennas over the studied months can be combined with the published exposure to produce a first flux estimate in the $10^{17}$–$10^{18.5}$ eV band, a step the paper defers until the dataset is larger."],"forward_implications":["The 26 solid candidates demonstrate that the 46-antenna prototype can already pick out cosmic-ray-like events from a noisy radio environment, validating the autonomous-detection concept before the array is complete.","The cut-efficiency table provides a quantitative noise budget: clustering alone removes 79% of coincident events, so most of GP300's current trigger rate is transient local interference rather than astrophysical signal.","The polarization cut keeps more than 99% of simulated cosmic-ray signals while rejecting roughly a fifth of the data, showing that the b-ratio is a workable positive identifier even before antenna polarity calibration is fully settled.","The reconstruction-quality stage is deliberately conservative, keeping only 26 of 41 candidates, and this purity-over-efficiency choice sets the template for the next stage of the analysis.","Once the full 300-antenna configuration is deployed in 2026, the authors expect around 130 cosmic-ray events per day, which would make the Galactic-to-extragalactic transition region accessible with the same pipeline."],"supporting_citations":[{"why":"Describes the GP300 array layout, deployment status, and antenna characteristics that define the dataset.","marker":"[2]"},{"why":"Provides the exposure calculations and expected event rates used to cross-check candidate energies and to forecast full-array sensitivity.","marker":"[4]"},{"why":"The GRANDlib software framework used for offline processing of data and for simulating the radio-frequency chain for cosmic-ray signals.","marker":"[5]"},{"why":"The Planar Wave Front arrival-direction reconstruction method that underlies the clustering cut and the timing visual cut.","marker":"[8]"},{"why":"Introduces the b-ratio ($e_V\\cdot e_B$) polarization estimator and its efficiency benchmark, on which the polarization cut is based.","marker":"[9]"},{"why":"Predicts the cosmic-ray footprint size of 1–10 km$^2$ (5–10 triggered antennas) used in the footprint-size and visual footprint cuts.","marker":"[11]"},{"why":"Lateral distribution function reconstruction from deconvolved electric fields, one of the three methods in the reconstruction-quality cut.","marker":"[12]"},{"why":"Angular distribution function reconstruction from voltage traces, including the chi-squared criterion used in the final quality cut.","marker":"[13]"},{"why":"Graph neural network reconstruction directly from voltage traces, the third method used to confirm candidates.","marker":"[15]"}],"fun_headline_variants":["26 cosmic-ray candidates from 533k noise triggers","Radio array distills 26 cosmic rays from 533k noise events","First GRANDProto300 data: 26 cosmic-ray candidates","From 533k radio noise events to 26 cosmic-ray candidates","26 cosmic rays emerge from 533k noise triggers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The manual visual cuts—trace shape, footprint compactness, and timing consistency—are assumed to separate real cosmic-ray signals from noise, and the reviewers' stated preference for northern azimuths is assumed not to bias which events survive.","fun_headline_variants_meta":{"raw":{"variants":["26 cosmic-ray candidates from 533k noise triggers","Radio array distills 26 cosmic rays from 533k noise events","First GRANDProto300 data: 26 cosmic-ray candidates","From 533k radio noise events to 26 cosmic-ray candidates","26 cosmic rays emerge from 533k noise triggers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000979,"raw_usage":{"total_tokens":4161,"prompt_tokens":951,"completion_tokens":3210,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":3125}},"tokens_in":567,"tokens_out":3210,"duration_ms":22598,"temperature":1.0,"reasoning_tokens":3125,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:56:36.100131+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the pipeline on time-scrambled or phase-randomized versions of the coincident-event data: if a comparable number of candidates survive the visual cuts, the selection is not isolating physical air showers; alternatively, repeating the manual cuts with azimuth hidden would reveal whether the 26 candidates disappear or concentrate toward the transformer and aircraft directions.","supporting_citations":[{"cited_title":"Grand prototypes: Grandproto300 and grand@auger,","cited_arxiv_id":null,"evidence_quote":"Describes the GP300 array layout, deployment status, and antenna characteristics that define the dataset."},{"cited_title":"KatoPoS ICRC2025(2025) 298","cited_arxiv_id":null,"evidence_quote":"Provides the exposure calculations and expected event rates used to cross-check candidate energies and to forecast full-array sensitivity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The GRANDlib software framework used for offline processing of data and for simulating the radio-frequency chain for cosmic-ray signals."},{"cited_title":"Ferrière, S","cited_arxiv_id":null,"evidence_quote":"The Planar Wave Front arrival-direction reconstruction method that underlies the clustering cut and the timing visual cut."},{"cited_title":"Chiche, K","cited_arxiv_id":null,"evidence_quote":"Introduces the b-ratio ($e_V\\cdot e_B$) polarization estimator and its efficiency benchmark, on which the polarization cut is based."},{"cited_title":"Benoit-Lévy, K","cited_arxiv_id":null,"evidence_quote":"Predicts the cosmic-ray footprint size of 1–10 km$^2$ (5–10 triggered antennas) used in the footprint-size and visual footprint cuts."},{"cited_title":"GülzowPoS ICRC2025(2025) 283","cited_arxiv_id":null,"evidence_quote":"Lateral distribution function reconstruction from deconvolved electric fields, one of the three methods in the reconstruction-quality cut."},{"cited_title":"GuelfandPoS ICRC2025(2025) 278","cited_arxiv_id":null,"evidence_quote":"Angular distribution function reconstruction from voltage traces, including the chi-squared criterion used in the final quality cut."},{"cited_title":"Emergences","cited_arxiv_id":null,"evidence_quote":"Graph neural network reconstruction directly from voltage traces, the third method used to confirm candidates."}],"review_version":1}