{"id":"2a2c498b-9431-4b7d-ad85-f95c47d9fece","arxiv_id":"2505.16440","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"An automated pipeline corrects long-term efficiency drift across the 16 GRAPES-3 muon telescope modules using Bayesian block segmentation, correlation-based reference selection, and Savitzky-Golay smoothing.","lead":"The GRAPES-3 cosmic ray observatory replaces operator-chosen reference detectors with an automated pipeline that uses Bayesian blocks, inter-module correlation, and Savitzky-Golay smoothing to remove long-term efficiency drift from 16 muon telescope modules. The result is a 22-year continuous muon rate stream for solar and space weather studies that correlates somewhat better with independent neutron monitor data than the legacy treatment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'stable reference' premise is untested and the simulation is circular: a common drift in the two selected modules would be absorbed as cosmic-ray signal.","rationale":"The paper is a coherent engineering write-up of a plausible pipeline, and I credit the authors for specifying the rejection, pressure-correction, Bayesian-block, correlation, and SG-filter steps in sufficient detail to be reproduced, and for attempting both a simulation and an external neutron-monitor comparison. The most dangerous point is the identification of the reference rate in Sec. 4.5: the algorithm selects the two modules with the highest mutual correlation after iteratively dropping the least-correlated module. The implicit assumption is that correlation with sibling modules is a proxy for efficiency stability. That is reasonable for a single module with an idiosyncratic drift, but it fails if the two surviving modules share a common instrumental drift, because that drift enters the reference rate and is indistinguishable from genuine solar modulation. The paper offers no calibration pulser, no known-rate source, and no other absolute reference to rule this out.\n\nThe simulation in Sec. 5.1 does not close this gap. The mock data are generated from real M15 data with 'detector effects' defined by r_SG,i, the same Savitzky-Golay-filtered ratios the method estimates. The simulation therefore tests internal consistency, not the accuracy of the efficiency model. It cannot reveal a common drift in the reference modules because such a drift would be part of the injected R(t) and would be treated as physics. The only non-circular evidence is the PSNM comparison. Table 1 shows the new method is better in 6 of 10 years, but 3 years go backwards (2008, 2009, 2013), one ties, and no uncertainties are given; the headline improvement rests on 2016 and 2017. This is suggestive but not enough to support the abstract's claim of a stream 'free of all known instrumental effects over calendar years.' I also note the 70% cut in Sec. 4.3 is never defined, but this affects short-term rejection rather than the central long-term claim.\n\nIf the concern lands, the corrected 22-year omni rate in Fig. 20 could contain a residual long-term instrumental component. A concrete test is to simulate with independently injected efficiency curves, including a common drift in the selected reference pair, and see whether the method recovers the known physical signal. If it does, the concern is retired; if not, the paper needs a reduced claim and an absolute calibration. I therefore keep the reader's CONDITIONAL verdict.","tokens_in":23655,"tokens_out":6684,"duration_ms":59310,"concrete_test":"Re-run the Sec. 5.1 simulation with injected module efficiencies drawn from independent random walks or calibration data (not from SG-filtered ratios), and include a slow common drift in the two modules that the algorithm would select as the reference pair. Run the full Sec. 4 pipeline and compare the final corrected omni rate with the known injected physical signal. If the common drift appears in the corrected rate or biases the recovered signal, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that the two modules with the highest mutual correlation in each Bayesian block are efficiency-stable (Sec. 4.5). This is never validated against an external standard. The Sec. 5.1 simulation cannot validate it: the mock data are generated as R_i(t) = R(t)/r_SG,i, where r_SG,i is the Savitzky-Golay-filtered ratio of M15 to module i, i.e., the same quantity the method estimates. The injected 'true' efficiency variations are therefore defined by the method's own smoothing, so the simulation only shows that the pipeline reproduces its own model. A common drift in the two selected reference modules, whether from shared gas supply, common electronics aging, or a common maintenance pattern, would be absorbed into the reference and appear as cosmic-ray modulation. The neutron monitor comparison is the only independent check, but Table 1 shows improvements in only 2 of 10 years with three years going slightly backwards, and no uncertainties are provided for any correlation coefficient. The abstract's claim of a stream 'free of all known instrumental effects' is stronger than the evidence supports.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an automated algorithmic pipeline for correcting long-term efficiency variations in the 16 independent modules of the GRAPES-3 muon telescope. The method uses Bayesian blocks to segment each module's time series, iteratively selects the two mutually most correlated modules per block as an efficiency-stable reference, stitches block baselines, and applies a Savitzky-Golay filter to each module's ratio to the reference rate to estimate and remove slow efficiency drifts. The result is a 22-year (2001-2022) combined all-direction muon rate. The authors validate the method by comparing it to the legacy operator-selected reference approach on simulated data and by correlating the corrected muon rate with the Princess Sirindhorn Neutron Monitor (PSNM) count rate for 2008-2017.","tokens_in":23850,"tokens_out":2223,"duration_ms":19748,"significance":"If the central claims hold, the paper provides a genuinely useful contribution: a reproducible, operator-independent correction procedure for a long-running muon telescope, with potential applicability to other multi-detector cosmic-ray instruments. The pipeline is described in sufficient detail to be reimplemented, and the final corrected data product (Fig. 20) is of direct value for solar and heliospheric physics studies. The use of an independent neutron monitor for at least partial validation is a real strength, as is the explicit aim of removing human subjectivity from a historically judgment-based calibration step. However, the significance is contingent on demonstrating that the automatically chosen 'stable reference' is not merely self-consistent but actually tracks the true cosmic-ray signal, and on quantifying whether the improvements over the legacy method are statistically meaningful.","major_comments":[{"comment":"The mock-data validation is circular. The simulated rates for the other 15 modules are constructed as Ri(t) = R(t)/rSG,i, where rSG,i is the Savitzky-Golay-filtered ratio of M15 to module i computed from the real data, i.e., exactly the quantity that the new method estimates. The simulation therefore injects the method's own smoothing output as the 'true' efficiency, and the agreement in Fig. 23 only shows that the pipeline can reproduce its own smoothing model. It does not validate the stable-pair selection in Sec. 4.5 or demonstrate that the corrected rate is more accurate than the legacy result, despite the claim in Sec. 5.2 of 'more accurate results'.","section":"Sec. 5.1, Eq. (3)"},{"comment":"The load-bearing premise that the two modules with the highest mutual correlation in each Bayesian block are the most efficiency-stable is never tested against an external standard. If the two selected modules share a common drift—from a common gas-supply pressure change, common electronics aging, or synchronized maintenance—that drift is absorbed into the reference rate and misattributed to cosmic-ray modulation. The neutron-monitor comparison (Sec. 5.3) is the only external check, but it covers only 10 years, shows clear improvement in only two years, and is not used to validate the reference-module selection directly.","section":"Sec. 4.5 and Sec. 4.7"},{"comment":"The quantitative evidence for superiority over the legacy method is weak. In Table 1, the new method improves the correlation coefficient in only 2 of 10 years (2016 and 2017) by a large amount, with modest changes in other years and slight degradations in 2008, 2009, and 2013 (0.62→0.59, 0.15→0.12, 0.56→0.54). No uncertainty or significance level is provided for any of the correlation coefficients, so it is unclear whether the 2016 and 2017 improvements are statistically robust or whether the year-to-year differences are within sampling noise. The abstract and Sec. 6 draw conclusions stronger than what Table 1 supports.","section":"Table 1, Sec. 5.3"},{"comment":"The Savitzky-Golay parameters (window size 10 days, polynomial order 2) are selected by maximizing the post-correction correlation of each module with the reference rate, which is itself constructed from the same correlation-based stable-module selection. This introduces a degree of circularity in the tuning: the metric used to choose parameters is the same internal self-consistency metric used to define the reference. The resulting 'convergence' in Fig. 17 therefore partly measures how well the filter matches the reference construction, not how well either tracks the true cosmic-ray flux. An external metric (e.g., neutron-monitor correlation as a function of these parameters) would strengthen the choice.","section":"Sec. 4.8, Fig. 17"}],"minor_comments":[{"comment":"In the sentence describing the EAS mechanism, 'mechanishm' is a typo for 'mechanism'; please correct.","section":"Sec. 2"},{"comment":"The matrix in Fig. 10 is visually difficult to parse; the long repeated rows of module numbers are not clearly labeled as a function of iteration. A compact tabular or annotated version would improve readability.","section":"Sec. 4.5 (Fig. 10)"},{"comment":"The correlation coefficients in Table 1 are quoted without uncertainties or a statement of the effective number of independent points, despite the strong autocorrelation in daily rates. Please provide bootstrap or effective-degrees-of-freedom uncertainties, particularly for the 2016-2017 comparisons.","section":"Sec. 5.3"},{"comment":"The phrase 'free of all known instrumental effects' (also in the abstract) is too strong given the caveats in Sec. 4.5 and the mixed neutron-monitor results. Rephrase to 'corrected for the known instrumental effects addressed in this work'.","section":"Sec. 4.9 and Sec. 6"},{"comment":"The choice of ±3σ in the bad-packet rejection and ±4σ in the short-unstable-period rejection appears heuristic; please state whether these thresholds are optimized or are chosen from prior experience, and whether the final results are sensitive to them.","section":"Sec. 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a technical methods contribution rather than a new physics result, which is appropriate for astro-ph.IM. The core algorithmic idea is sensible and the long-term corrected data product is potentially valuable. The main risk is overclaiming: the simulation is circular, and the neutron-monitor evidence is mixed. I would encourage the editor to request a revised version that (a) re-runs the simulation with efficiency curves injected independently of the SG-filter-derived ratios, (b) quantifies the significance of the Table 1 correlations, and (c) tones down the 'free of all known instrumental effects' claim. If the authors can address these, the paper could be acceptable. I did not find evidence of problematic citation practices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent data-hygiene methods paper from the GRAPES-3 collaboration. The genuinely new bit is the automatic, per-segment selection of the two most stable reference modules via iterative correlation analysis, replacing the operator-chosen fixed reference used in their legacy correction [11]. That is a real improvement, and the pipeline is described well enough to reimplement. The 22-year corrected muon rate in Fig. 20 will be useful for long-term modulation studies.\n\nWhere it is weaker: the validation does not do what the prose claims. The mock-data test in Sec. 5.1 generates the other 15 modules' rates as R(t)/r_SG,i, where r_SG,i is the Savitzky-Golay-smoothed ratio of M15 to module i from real data. So the injected 'true' efficiency variations are exactly what the method's own smoother produces. The simulation can only show the pipeline is self-consistent, not that it is more accurate than the legacy method. The paper should compare both methods' outputs to the known input R(t) directly.\n\nThe neutron-monitor comparison is the only independent check, and it is not as strong as the abstract implies. Across 2008-2017, the new method improves correlation with PSNM in six years, is essentially unchanged in one, and gets slightly worse in three (so the stress-test line about 'only 2 of 10 years' overstates the case). The large gains are concentrated in 2016 and 2017 (0.44->0.59 and 0.29->0.48). That is suggestive, not conclusive, especially since no uncertainties are given for any correlation coefficient. The abstract's 'free of all known instrumental effects' is stronger than what is demonstrated; Sec. 6 itself says 'expected to be free'.\n\nThe load-bearing premise—that the two highest-correlated modules are the most efficiency-stable—is never directly tested. If those two modules share a drift (common gas supply, electronics aging, maintenance pattern), that drift is absorbed into the reference and misread as cosmic-ray modulation. The neutron monitor gives some guard against this, but only over a decade and without error bars. An injection test that adds a common slow drift to two modules would address this directly; the current simulation cannot.\n\nMinor points: the '70% cut' in Sec. 4.3 is never defined, and the SG parameters are chosen by maximizing post-correction correlation with the reference, which has a mild circularity. Public release of the code would also help.\n\nOverall: the method is a plausible, well-specified improvement over the legacy approach, and the paper deserves a serious referee. It is not ready in its current form—the validation claims need to be scaled back or strengthened, and the stable-reference assumption needs a dedicated test. I would send it to review with a request for major revision, and I would bring it to the reading group to discuss the selection-by-correlation premise.","headline":"A genuinely useful automation of GRAPES-3's efficiency correction with a per-segment reference-selection step, but the validation is weaker than the prose and the stable-reference premise needs a real test.","tokens_in":24546,"tokens_out":3302,"would_cite":true,"duration_ms":27031,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An automated pipeline using Bayesian blocks and inter-module correlations replaces operator-chosen reference detectors, producing a 22-year efficiency-corrected muon rate that agrees better with neutron monitor data than the legacy method.","keywords":["GRAPES-3","muon telescope","efficiency correction","Bayesian blocks","Savitzky-Golay filter","cosmic ray modulation","neutron monitor correlation"],"falsifier":"Extend the neutron-monitor comparison to every overlapping year with quoted uncertainties: if the new method's correlation advantage over the legacy method disappears or reverses when errors are included, the claimed improvement is not established. A sharper check is to inject a common-mode drift into the mock data across all sixteen modules; if the auto-selected reference pair follows the injected drift, the method will mislabel it as physical modulation, directly testing the central assumption.","tokens_in":23440,"feed_emoji":"🔭","tokens_out":4599,"duration_ms":34119,"temperature":0.7,"pith_summary":"The paper claims that long-term efficiency drift in the sixteen independent modules of the GRAPES-3 muon telescope can be corrected without relying on an operator's choice of a reference detector. The pipeline splits each module's pressure-corrected rate into Bayesian blocks, then within each block iteratively discards the least-correlated modules until the two most mutually correlated modules remain, using their average as a reference rate. That reference is stitched across blocks and used with a Savitzky-Golay filter to model and remove each module's slow efficiency decline. On mock data the new method agrees with the legacy approach at correlation 0.995, and against neutron-monitor data it improves the yearly correlation in 2016 from 0.44 to 0.59 and in 2017 from 0.29 to 0.48. The result is a 22-year all-direction muon rate that the authors present as corrected for all known instrumental effects and ready for physics analysis.","feed_headline":"Automated pipeline cleans 22 years of muon-telescope drift","feed_subtitle":"Bayesian blocks plus cross-module correlations replace human-chosen reference detectors and improve neutron-monitor agreement.","key_machinery":"The machinery is a multi-stage pipeline. Bayesian blocks discretize each module's pressure-corrected rate into periods separated by change points, with a p-value threshold of 0.01 and merging of change points closer than 30 days. Within each block, an iterative correlation analysis removes the module with the lowest average correlation with the others until two modules remain; those two are taken as the efficiency-stable reference pair. A rate stability map records each module's correlation with that reference for every block, a stitching step uses the most stable module common to consecutive blocks to align baselines, and a Savitzky-Golay filter with a 10-day window and polynomial order 2 is applied to the ratio of the reference rate to each module to model and correct slow efficiency decline.","core_discovery":"The central claim is that a fully automated, algorithmic correction can remove all known instrumental effects from the G3MT muon data over calendar years, and that the resulting all-direction muon rate is at least as clean as, and in recent years cleaner than, the legacy operator-selected-reference product. The evidence is the corrected 2001-2022 rate, the mock-data agreement with the legacy method, and the neutron-monitor comparison in which the new method raises the yearly correlation substantially in 2016 and 2017 while matching the legacy method in the other years tested.","pith_inferences":["The paper's validation does not test the case where the two 'stable' modules share a common instrumental drift, for example from common gas supply or electronics aging; in that scenario the method would attribute the shared drift to cosmic-ray physics, so a simulation with injected common-mode drift would be a natural next test.","The neutron-monitor comparison is limited to 2008-2017 and is reported without uncertainties; extending it through 2022 and adding error bars would make the claimed improvement in 2016-2017 statistically quantifiable.","Because the reference is built purely from internal correlations, any instrumental effect that is common to all sixteen modules is invisible to the method; the phrase 'free of all known instrumental effects' should be read as 'free of effects visible in the inter-module correlation structure.'","The same pipeline could be applied to other ground-based muon or neutron telescopes with redundant modules, and to retrospective reanalysis of older data that was previously corrected with operator-chosen references."],"forward_implications":["The 22-year all-direction muon rate is corrected for all known instrumental effects and can be used directly for studies of long-term cosmic-ray modulation, spanning two solar cycles and a solar magnetic cycle.","Applying the method separately to each of the 225 direction bins preserves angular information while removing efficiency drift, enabling directional studies without operator-chosen references.","The method transfers to any experiment that monitors long-term natural particle flux with redundant detectors, replacing subjective operator judgment with a reproducible algorithm.","Legacy calendar-year boundaries no longer introduce artificial discontinuities, because Bayesian blocks define the correction periods from the data itself."],"supporting_citations":[{"why":"Defines the legacy method of operator-selected reference-module efficiency correction that this paper replaces and compares against.","marker":"[11]"},{"why":"Supplies the Savitzky-Golay filtering technique used to model and correct each module's slow efficiency variation.","marker":"[22]"},{"why":"Provides the pressure coefficient used to correct muon rates for atmospheric pressure before efficiency correction.","marker":"[19]"},{"why":"Provides the atmospheric temperature coefficient applied to the muon data in the neutron-monitor validation.","marker":"[1]"},{"why":"Characterizes the high-rigidity neutron monitor used as the independent validation dataset for comparing legacy and new corrected muon rates.","marker":"[25]"},{"why":"Describes the NM-64 neutron monitor design underlying the reference neutron data used in the validation.","marker":"[27]"},{"why":"Supplies the atmospheric temperature reanalysis data used for the temperature correction in the validation comparison.","marker":"[28]"},{"why":"Documents the GRAPES-3 muon telescope detector design, establishing the sixteen-module geometry and independent-module assumption the method relies on.","marker":"[13]"}],"fun_headline_variants":["Bayesian blocks clean 22 years of muon data","Algorithm removes detector drift in muon telescope","Automated method scrubs muon telescope drift","New pipeline corrects muon data without human picks","Muon telescope data drift fixed by automated algorithm"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the two modules with the highest mutual correlation in each Bayesian block are the most efficiency-stable, so their average rate is a valid reference for correcting all other modules; if those two drift together from a shared cause, their common drift is absorbed into the reference and mistaken for cosmic-ray modulation.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian blocks clean 22 years of muon data","Algorithm removes detector drift in muon telescope","Automated method scrubs muon telescope drift","New pipeline corrects muon data without human picks","Muon telescope data drift fixed by automated algorithm"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000147,"raw_usage":{"total_tokens":1131,"prompt_tokens":836,"completion_tokens":295,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":452,"completion_tokens_details":{"reasoning_tokens":222}},"tokens_in":452,"tokens_out":295,"duration_ms":2805,"temperature":1.0,"reasoning_tokens":222,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:01:27.896357+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Extend the neutron-monitor comparison to every overlapping year with quoted uncertainties: if the new method's correlation advantage over the legacy method disappears or reverses when errors are included, the claimed improvement is not established. A sharper check is to inject a common-mode drift into the mock data across all sixteen modules; if the auto-selected reference pair follows the injected drift, the method will mislabel it as physical modulation, directly testing the central assumption.","supporting_citations":[{"cited_title":"PoS ICRC2017, 357 (2018) https://doi.org/10.22323/1.301.0357","cited_arxiv_id":null,"evidence_quote":"Defines the legacy method of operator-selected reference-module efficiency correction that this paper replaces and compares against."},{"cited_title":"Astroparticle Physics94, 22–28 (2017) https://doi.org/10.1016/j.astropartphys.2017.07.002","cited_arxiv_id":null,"evidence_quote":"Provides the atmospheric temperature coefficient applied to the muon data in the neutron-monitor validation."},{"cited_title":"Astrophys","cited_arxiv_id":null,"evidence_quote":"Characterizes the high-rigidity neutron monitor used as the independent validation dataset for comparing legacy and new corrected muon rates."},{"cited_title":"12.4/summary","cited_arxiv_id":null,"evidence_quote":"Supplies the atmospheric temperature reanalysis data used for the temperature correction in the validation comparison."},{"cited_title":"Nuclear Instruments and Methods in Physics Research A 545, 643–657 (2005) https://doi.org/10.1016/j.nima.2005.02","cited_arxiv_id":null,"evidence_quote":"Documents the GRAPES-3 muon telescope detector design, establishing the sixteen-module geometry and independent-module assumption the method relies on."}],"review_version":1}