{"id":"552d5eba-ae2e-4c37-87cc-e596d5a1635f","arxiv_id":"2508.17461","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Layered water Cherenkov detectors can be calibrated on atmospheric muons, and a triangular GCOS array needs about 15,000 tanks at roughly 2.2 km spacing for full efficiency above 10 EeV.","lead":"Two layered water Cherenkov tank prototypes at the Pierre Auger site have run continuously for over a decade, and their muon-based calibration matches simulations. The paper also estimates that a future GCOS array covering 60,000 square kilometers would need more than 15,000 tanks spaced about 2.2 km apart to trigger on cosmic rays above 10 EeV.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GCOS spacing and detector count rest on a single-station footprint proxy rather than a full-array trigger simulation; this is the main load-bearing gap.","rationale":"The reader's weakest assumption is the same one I identify: the leap from a single-station 90% trigger footprint to full-array efficiency and maximum spacing. This is genuinely load-bearing because the paper's main quantitative deliverable, the 2.2 km spacing and the resulting 15,000-tank requirement, is derived from it. I agree that the concern is addressable rather than fatal: the prototype has more than a decade of operational data, the calibration procedure for the bottom layer is supported by the muon-peak feature, and the simulation reproduces the measured calibration histograms at least qualitatively. Those parts of the paper give real independent support. The missing piece is an array-level trigger simulation or an equivalent coverage calculation that includes multiplicity, noise, and dead time. Since the current verdict is CONDITIONAL and that condition is exactly this check, my read does not change the verdict. No code or data are provided, so the simulation cannot be inspected independently, but this is a normal state for a design study at ICRC and does not by itself warrant rejection.","tokens_in":6384,"tokens_out":4335,"duration_ms":48309,"concrete_test":"Run a full-array Monte Carlo for a triangular array with spacings d = 2.0, 2.2, 2.5, and 3.0 km, using the same CORSIKA-based station trigger probability maps for proton and iron showers at 10 EeV and zenith angles from 0 to 63 degrees. Draw shower cores uniformly over a unit cell, apply the proposed time-over-threshold trigger with realistic noise and dead time, and compute the fraction of showers that satisfy the array-level trigger condition, first requiring at least 3 stations and then at least 5 stations. If the efficiency at d = 2.2 km is at or above the GCOS requirement (essentially 100% at 10 EeV) and falls below it at d = 3 km, the spacing claim survives; otherwise the maximum spacing must be reduced and the detector count revised upward.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The GCOS array-efficiency claim in Sections 4 and 5 rests on a single-station proxy. The authors compute the lateral trigger probability of one tank as a function of distance from the shower core, define an ellipse at 90% single-station trigger efficiency, and take its minor axis as the maximum spacing on a triangular grid. This step assumes that if every shower core lies within the 90% footprint of at least one station, the array reaches essentially 100% trigger efficiency at 10 EeV. That does not follow: array triggering requires a multiplicity condition (the GCOS strawman demands multiplicity greater than 5 at 30 EeV), and the per-shower efficiency is not simply the maximum single-station probability. A shower whose nearest station has 90% trigger probability would fail about 10% of the time for a single-station trigger, and the failure rate can be larger when several stations must fire in coincidence. Station noise, dead time, and the time-over-threshold trigger logic are not folded into the footprint calculation. No full-array Monte Carlo or coverage integral over shower-core positions in the triangular unit cell is presented, and Figure 4 contains no error bars or systematic checks. The 15,000-tank number is therefore not established by the simulation shown; it is a plausible estimate whose error could be sizable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on two layered water Cherenkov detector (LCD) prototypes operated at the Pierre Auger Observatory site since 2014, focusing on their calibration and on simulations of their response. The authors show that the bottom-layer charge spectrum has a well-defined atmospheric muon peak, which provides a clean calibration reference in vertical equivalent muon units, and they compare simulated charge histograms with the prototype data. They also scan top-layer heights and tank diameters relevant to the GCOS and PEPS designs, and they use simulated shower footprints to estimate the detector spacing and total number of tanks needed for a 60,000 km^2 GCOS surface array, concluding that a 2.2 km spacing and more than 15,000 tanks are required for 100% trigger efficiency above 10 EeV.","tokens_in":6626,"tokens_out":2328,"duration_ms":26454,"significance":"The 10-year prototype dataset and the proposed bottom-layer muon-peak calibration are valuable and credible contributions to the design studies for GCOS, PEPS, and SWGO. The fast, configurable simulation is a practical tool, and the qualitative agreement with the measured charge histograms supports the calibration concept. The paper is also explicit that the top-layer coincidence calibration is not yet implemented in the field, which is an honest statement of scope. The main significance, however, depends on the GCOS array-efficiency estimate: as presented, that estimate rests on a single-station footprint proxy rather than a full-array trigger calculation, so the detector-count claim is not yet established at the level the paper suggests.","major_comments":[{"comment":"The central claim that a 2.2 km spacing gives essentially full trigger efficiency above 10 EeV and that more than 15,000 tanks are needed is derived from the minor axis of an ellipse defined by 90% single-station trigger probability. This is a proxy, not a demonstration of array-level efficiency: the GCOS strawman explicitly requires a multiplicity of trigger detectors larger than 5 at 30 EeV, and a shower whose nearest station triggers with 90% probability will often fail the array-level condition because multiple near-threshold stations must fire in coincidence. Station noise, dead time, and the time-over-threshold logic are not incorporated into the footprint calculation, and no full-array Monte Carlo or analytic coverage integral over shower-core positions in the triangular unit cell is presented. The 15,000-tank number is therefore a plausible estimate whose error could be sizable; the authors should either provide an array-trigger simulation or explicitly restate the conclusion as an upper-limit-style estimate with quantified uncertainty.","section":"Section 4, Fig. 4"},{"comment":"The statement that the simulation 'reproduces the measurements' is supported only by visual comparison. Figure 2 shows simulated histograms but no corresponding data histogram overlaid, and Figure 3 has no error bars, so the dependence of the muon-peak photoelectron yield on top-layer height and diameter is presented without statistical or systematic uncertainties. This is load-bearing because the paper uses this simulation to recommend design choices for PEPS and GCOS, and to justify the calibration method. The authors should add quantitative agreement metrics (e.g., fitted peak positions and widths with uncertainties) and at least statistical error bars to Figure 3.","section":"Section 3, Figs. 2 and 3"},{"comment":"The top-layer calibration via a top-bottom coincidence is validated only in simulation; the authors state in Section 2.1 that the implementation 'is a subject of further studies' and in the conclusion that it 'will be implemented and tested in the near future in the field.' The paper should make clear in the abstract and introduction that the field-validated calibration result concerns the bottom layer only, and that the top-layer coincidence method is a simulation-based proposal, not yet a demonstrated field procedure. This distinction is important for readers planning detector designs based on the presented calibration performance.","section":"Section 2.1 and Section 3"}],"minor_comments":[{"comment":"The coefficients a=0.6 and b=0.4 are taken from Ref. [6] and are not fitted to the present prototype data. This is reasonable, but the paper should state explicitly that these values are assumed from the earlier design study, since the calibration and signal-separation claims in Section 2 rely on them.","section":"Section 2, Eq. (1)"},{"comment":"The transition from the 90% footprint ellipse to the maximum spacing is stated in one sentence without a derivation. A short formula or a diagram showing the triangular-lattice coverage condition would make the argument reproducible and would clarify the role of the minor axis.","section":"Section 4, Fig. 4"},{"comment":"The text alternates between 'detectors' and 'tanks' when quoting the 15,000 number; using one term consistently would avoid ambiguity for the broad ICRC readership.","section":"Abstract and Section 5"},{"comment":"The sentence 'These represent the first studies of the LCD in the context of GCOS and PEPS and are the basis for further performance studies involving full air-shower simulations' is a useful scope statement but sits in the results section; moving it to the introduction or conclusion would improve the narrative flow.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conference proceedings contribution and the prototype-calibration part is solid. My main concern is that the GCOS detector-count claim is presented in the abstract and conclusion as a quantitative result, while the underlying method is a single-station proxy without an array-level trigger simulation. This is fixable within the manuscript's scope by adding an explicit array-coverage calculation and softening the claim, so I recommend major revision rather than rejection. The self-citation to the authors' prior design study [6] for the coefficients a and b is appropriate and not problematic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real meat here is the 10-year operational record of the two layered WCD prototypes and the bottom-layer muon-peak calibration. That part is solid: the seasonal modulation of the calibration constants makes sense, the simulation reproduces the charge histograms qualitatively, and the coincidence method for the top layer is a sensible fix for an otherwise messy peak. The geometry scan for PEPS is a nice practical addition. The paper is a legitimate extension of the earlier proposal and data paper, not a rehash.\n\nThe soft spot is exactly where the reader put it: the GCOS spacing and detector count are built on a single-station footprint proxy. Taking the minor axis of the 90% trigger-efficiency ellipse as the maximum triangular-grid spacing is a covering approximation, and it hasn't been checked against a full-array trigger simulation or even a simple coverage integral over the unit cell. A triangular lattice with spacing d has covering radius d/√3, so if the 90% contour has half-width R, you need d ≤ √3·R, not d = 2R. The paper's 2.2 km/15,000-tank numbers may be roughly right, but they're presented as if they come from a calculation they don't actually show. Station noise, dead time, and the multiplicity requirement are all left out. That's the main load-bearing gap, and it should be flagged clearly.\n\nThe missing error bars in Fig. 3 and the absence of quantitative agreement metrics are minor in a proceedings context, but they'd need to be addressed in a full paper. The coincidence calibration is still only simulated, which the authors admit.\n\nOverall, this deserves a serious referee as a design study / status report for the instrumentation community. The GCOS claim should be softened or backed by a quick array-level Monte Carlo; the calibration material can stand mostly as is. I'd accept it for proceedings with requests for revision, and if it were a journal submission I'd send it back for that array coverage check.","headline":"Useful prototype-calibration status report; the GCOS spacing/count estimate is a heuristic that needs an array-level check.","tokens_in":7143,"tokens_out":2969,"would_cite":false,"duration_ms":31882,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Layered water-Cherenkov detectors calibrate cleanly and need 2.2 km spacing for full 10 EeV efficiency.","keywords":["layered water Cherenkov detector","cosmic-ray air showers","muon content","VEM calibration","GCOS","trigger efficiency","detector spacing"],"falsifier":"Simulate, or build a small test cell of, a triangular array at 2.2 km spacing with full station electronics, dead time, and noise, and count 10 EeV proton and iron showers that trigger with high multiplicity; if the efficiency at 10 EeV is measurably below 100%, the paper's spacing and detector count are over-optimistic.","tokens_in":6206,"feed_emoji":"🔭","tokens_out":7994,"duration_ms":72747,"temperature":0.7,"pith_summary":"This paper makes the case that a layered water-Cherenkov tank, a single detector with an optically separated top and bottom water volume, can serve as the workhorse unit of next-generation cosmic-ray arrays. Using data from two prototype tanks that have run for over a decade, it shows that the bottom layer's charge spectrum has a clean atmospheric-muon peak that provides a stable vertical-equivalent-muon calibration, and that a purpose-built simulation reproduces the measured histograms. The same simulation is then used to size a next-generation observatory with a 60,000 km2 target area: a triangular grid with roughly 2.2 km spacing gives near-100% trigger efficiency for showers above 10 EeV, which means more than 15,000 tanks. The paper does not claim to have built such an array; its contribution is the calibration method and the detector-spacing estimate that follows from it.","feed_headline":"2.2 km spacing gives cosmic-ray array full 10 EeV coverage","feed_subtitle":"Muon-peak calibration is clean, and a 60,000 km2 array would need more than 15,000 detectors.","key_machinery":"The central object is the layered water-Cherenkov tank: a cylindrical water volume split by a reflective barrier into a thin top layer, which absorbs the electromagnetic component (attenuation length about 40 cm), and a thicker bottom layer, which records light from through-going muons. The two measured signals are combined through the linear system $S_{\\rm top}=a\\,S_{\\gamma,e^\\pm}+b\\,S_\\mu$ and $S_{\\rm bottom}=(1-a)\\,S_{\\gamma,e^\\pm}+(1-b)\\,S_\\mu$, with $a\\approx 0.6$ and $b\\approx 0.4$, so the electromagnetic and muon contributions can be recovered station by station. The array-size argument is carried by the single-station trigger probability footprint: the minor axis of the 90%-trigger-efficiency ellipse is taken as the maximum spacing that can be afforded between detectors.","core_discovery":"The central claim is that the layered water-Cherenkov detector separates the electromagnetic and muonic parts of an air shower in a single tank, and that this separation makes the array easy to calibrate and cheap enough to deploy at the scale required by next-generation observatories. In the bottom layer the muon peak is so well defined that the calibration constant can be read directly from a one-minute charge histogram, with only seasonal and settling drift visible over ten years. The top layer's muon peak is hidden by electromagnetic background, but requiring a coincidence with the bottom PMT restores a usable peak. The detector-spacing result follows from simulated single-station trigger footprints: taking the minor axis of the 90%-efficiency ellipse as the maximum allowed spacing gives about 2.2 km, which in a triangular lattice means more than 15,000 detectors for 60,000 km2, and moves full efficiency up to about 30 EeV if the spacing is relaxed to 3 km.","pith_inferences":["The spacing estimate is derived from a single-station footprint proxy rather than from a full array-level trigger simulation; including station dead time, noise, and electronics effects in an array simulation could reduce the achievable spacing and push the detector count above 15,000.","The same footprint argument could be used to design a lower-cost sparse subarray that accepts reduced efficiency at 10 EeV while still meeting science goals at higher energies.","The muon-peak calibration method should transfer to any water-Cherenkov tank whose top layer is thick enough to absorb the electromagnetic cascade; for very thin top layers the bottom-layer muon peak may be contaminated.","The reported photoelectron increase with smaller tank diameter suggests that geometry optimization is not complete, and a full scan over layer heights, diameters, and PMT positions could find a cheaper configuration than the baseline assumed here."],"forward_implications":["A 60,000 km2 observatory built from layered water-Cherenkov tanks requires roughly 15,000 detectors at 2.2 km spacing to be fully efficient above 10 EeV.","Bottom-layer calibration from the atmospheric-muon peak is stable and self-contained, so a large array can be calibrated uniformly without per-station beam calibrations.","Top-layer calibration becomes practical with a top-bottom PMT coincidence trigger, which suppresses the electromagnetic background and makes the muon peak usable.","If spacing is relaxed to 3 km, the array becomes fully efficient only near 30 EeV, so the spacing choice directly sets the energy threshold of the observatory.","The fast, geometry-configurable simulation can scan tank dimensions and PMT layouts for other proposed detector designs before full detector construction."],"supporting_citations":[{"why":"Introduces the layered water-Cherenkov detector concept, the two-layer geometry, and the coefficients $a\\approx0.6$ and $b\\approx0.4$ used to separate electromagnetic and muon signals.","marker":"[6]"},{"why":"Defines the standard VEM calibration procedure for water-Cherenkov detectors that the bottom-layer muon-peak calibration adapts and validates.","marker":"[12]"},{"why":"Supplies the CORSIKA air-shower simulations whose ground-particle output feeds the tank-response simulation used for the calibration histograms and trigger footprints.","marker":"[13]"},{"why":"Sets the target area of 60,000 km2 and the requirement of 100% trigger efficiency above 10 EeV that the spacing calculation is designed to meet.","marker":"[2]"},{"why":"Provides the re-analysis of ten years of prototype data, including daily event rates and calibration constants, that underlies the experimental comparison.","marker":"[11]"},{"why":"Motivates the need for muonic-content measurement by showing hadronic-model systematic uncertainties in depth-of-maximum and ground-signal predictions.","marker":"[5]"}],"fun_headline_variants":["Bottom-layer muon peak enables one-minute calibration","2.2 km spacing: full 10 EeV coverage for gamma and cosmic rays","Layered Cherenkov tanks separate muons, simplify calibration","15,000+ tanks for 60,000 km²: spacing is key","Muon peak in bottom layer: direct calibration constant"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The detector count rests on the assumption that the minor axis of the 90% single-station trigger-efficiency footprint directly sets the maximum spacing on a triangular grid, with no full-array trigger simulation, station dead time, or noise and electronics effects; if the true array-level efficiency at 10 EeV is lower than this footprint proxy, the required spacing would shrink and the number of tanks would rise above 15,000.","fun_headline_variants_meta":{"raw":{"variants":["Bottom-layer muon peak enables one-minute calibration","2.2 km spacing: full 10 EeV coverage for gamma and cosmic rays","Layered Cherenkov tanks separate muons, simplify calibration","15,000+ tanks for 60,000 km²: spacing is key","Muon peak in bottom layer: direct calibration constant"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000291,"raw_usage":{"total_tokens":1724,"prompt_tokens":990,"completion_tokens":734,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":644}},"tokens_in":606,"tokens_out":734,"duration_ms":7818,"temperature":1.0,"reasoning_tokens":644,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:04:02.341804+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate, or build a small test cell of, a triangular array at 2.2 km spacing with full station electronics, dead time, and noise, and count 10 EeV proton and iron showers that trigger with high multiplicity; if the efficiency at 10 EeV is measurably below 100%, the paper's spacing and detector count are over-optimistic.","supporting_citations":[{"cited_title":"Layered water Cherenkov detector for the study of ultra high energy cosmic rays","cited_arxiv_id":"1405.5699","evidence_quote":"Introduces the layered water-Cherenkov detector concept, the two-layer geometry, and the coefficients $a\\approx0.6$ and $b\\approx0.4$ used to separate electromagnetic and muon signals."},{"cited_title":"Heck et al.,CORSIKA: A Monte Carlo code to simulate extensive air showers, Forschungszentrum Karlsruhe Report FZKA-6019(1998)","cited_arxiv_id":null,"evidence_quote":"Supplies the CORSIKA air-shower simulations whose ground-particle output feeds the tank-response simulation used for the calibration histograms and trigger footprints."},{"cited_title":"Flaggs and I.C","cited_arxiv_id":null,"evidence_quote":"Provides the re-analysis of ten years of prototype data, including daily event rates and calibration constants, that underlies the experimental comparison."}],"review_version":2}