{"id":"93ddc975-9c7a-45f5-9bec-7c2f703acd97","arxiv_id":"2411.14852","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":12,"one_line_summary":"A DBSCAN time-based pixel selection retains more low-charge shower pixels than Tailcuts at matched data-reduction factors in CTAO simulations, with mixed noise-rejection performance.","lead":"This paper tests two ways to shrink the data recorded by gamma-ray telescopes and finds that grouping pixels by when their light arrives keeps more faint shower pixels than the standard charge-only method. The result matters for the CTAO observatory, which must cut its data volume by 10 to 50 times without losing scientific information.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The conclusion's 'lower noise detection' claim is contradicted by Fig. 7, and the signal-efficiency comparison is not protected against selection bias, so the headline superiority is not yet supported.","rationale":"The reader's CONDITIONAL verdict is appropriate. My primary concern is more direct than the reader's weakest assumption: the paper's own Fig. 7 contradicts the 'lower noise detection' clause of the headline conclusion at the parameter sets actually used. The selection-bias issue identified by the reader is also real and compounds the problem, because the signal-efficiency advantage may be inflated by optimizing on the evaluation data. Neither concern by itself proves the method is inferior; it means the strongest claims are not currently supported by the evidence presented. The proposed held-out, joint-objective test would settle whether the advantage survives. Since the reader already conditioned acceptance on fixing the comparison protocol and overstatement, my assessment does not move the verdict.","tokens_in":14714,"tokens_out":4675,"duration_ms":53847,"concrete_test":"Split the simulated events into disjoint tuning and evaluation sets. On the tuning set, optimize each method's parameters using a pre-registered joint objective that rewards signal efficiency and penalizes the NSB-only noise-cluster fraction (for example, maximize efficiency subject to a false-positive fraction below a fixed threshold). Then evaluate both methods on the held-out set. If Clustering no longer beats Tailcuts on both metrics, the conclusion must be weakened. As a minimal internal check, recompute the DVR=40 points in Fig. 7 from the exact parameter sets in Tables 1 and 2 to confirm the noise-cluster fractions quoted in the text.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion states that Clustering achieves 'superior signal efficiency and lower noise detection' relative to Tailcuts for all studied cameras. Section 4.1, Fig. 7, using the parameter sets selected for signal efficiency (Tables 1 and 2), shows the opposite on the noise axis: at DVR=40, Clustering produces false clusters in roughly 40% of NSB-only events versus about 1% for Tailcuts. The authors acknowledge this can be mitigated by re-optimizing parameters for noise, but then the methods are no longer compared under a common optimization objective. Separately, Section 3.1 selects the best point per method on the same simulated data used for evaluation, and Clustering has more free parameters (6 versus 4), so the reported signal-efficiency advantage may partly reflect tuning freedom rather than a true algorithmic gain. No error bars or counts of tested configurations are provided. Together these issues undermine the strongest claim as stated, although the core idea may still be viable after re-analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes and evaluates data volume reduction (DVR) algorithms for CTAO imaging atmospheric Cherenkov telescopes, comparing a Tailcuts baseline with a DBSCAN-based time-space clustering method. Using CORSIKA/sim_telarray simulations of MST-FlashCam and, more briefly, of the other CTAO cameras, the authors measure signal efficiency, noise false-positive rate, robustness to NSB/broken pixels/calibration uncertainty, and angular/energy resolution after cleaning. The central claim is that Clustering achieves superior signal efficiency and lower noise detection than Tailcuts across all studied cameras, with a factor 10–50 reduction in stored data volume. The manuscript also discusses waveform cropping, computational time, and several clustering variants in an appendix.","tokens_in":14932,"tokens_out":3018,"duration_ms":31058,"significance":"If the central comparison is reliable, the result is practically important: CTAO faces a reduction from hundreds of PB/year to a few PB/year, and a robust pixel-selection method that preserves low-charge shower pixels would directly affect the observatory's data pipeline. The paper benefits from detailed Monte Carlo simulations, a concrete algorithmic proposal that is applicable beyond CTAO, and a sensible set of robustness checks. However, the evaluation protocol contains a load-bearing selection-bias problem, and the 'lower noise detection' conclusion is contradicted by the paper's own Fig. 7 under the stated parameter choice. The core idea remains plausible, but the headline superiority claim is not yet established as presented.","major_comments":[{"comment":"The signal-efficiency comparison is vulnerable to selection bias: the text states that for each DVR factor 'only the point with higher signal efficiency is selected and plotted for each method,' meaning that the same simulated data are used both to choose the free parameters and to evaluate the comparison. Clustering has six free parameters (spatial scale, time scale, minPts, noise cut, charge threshold, dilated rows) while Tailcuts has four, and no validation split, no number of tested configurations, and no error bars are reported. The apparent advantage of Clustering may therefore reflect extra tuning freedom rather than an algorithmic property. I request a separate validation sample, a count of the parameter sets tested per DVR factor, and uncertainty estimates on the efficiency curves.","section":"§3.1, Fig. 4"},{"comment":"The Conclusion's statement that Clustering achieves 'lower noise detection compared to Tailcuts' is directly contradicted by Fig. 7: at DVR=40, using the parameters selected for signal efficiency in Tables 1 and 2, Clustering misidentifies clusters in roughly 40% of NSB-only events while Tailcuts does so in about 1%. The paper notes that parameters could be re-optimized for noise rejection, but then the two methods are not compared under a common optimization objective. The claim as written is unsupported; either report false-positive rates at matched signal efficiency, or restrict the conclusion to signal efficiency and reconstruction resolution, with a clearly stated caveat about the noise trade-off.","section":"§4.1, Fig. 7 and §6"},{"comment":"The angular and energy resolution comparison is not controlled in a way that supports the conclusion. The two methods are compared at DVR factors of 120 and 140 rather than at exactly matched reduction values, and the parameters were selected to maximize signal efficiency rather than reconstruction performance, as the authors themselves acknowledge in the text. Because parameter choice strongly affects resolution, the apparent improvement from Clustering could be an artifact of the tuning criterion. I ask for matched DVR factors (e.g., by selecting parameters for each method at the same achieved reduction), parameter selection on validation data, and statistical uncertainties on the resolution curves.","section":"§3.2, Fig. 6"},{"comment":"The blanket claim that Clustering has 'lower noise detection' or 'always detects more signal and less noise' for all cameras is not supported by the presented figures. Fig. 12 shows that for the SST camera, Clustering is slightly worse than Tailcuts at high charge thresholds, and Fig. 11 contains no noise axis at all. Please add the noise false-positive rate for each camera or qualify the claim to the specific cameras and charge ranges shown.","section":"§5, Fig. 12"}],"minor_comments":[{"comment":"There are numerous typographical errors, e.g., 'The CTAO data rates needs' in the abstract and 'di fferent', 'T ailcuts', and 'su ffi ciently' in the body. A careful proofread is needed.","section":"Abstract and throughout"},{"comment":"In Table 2, the boundary threshold for DVR=30 (5.0 PE) is higher than the picture threshold (4.0 PE), which is unusual for Tailcuts; please explain or correct this entry.","section":"Tables 1 and 2"},{"comment":"The y-axis label says 'Relative number of events with at least one cluster' but the caption clarifies normalization to the nominal NSB rate; please make the normalization explicit in the axis label or caption.","section":"Fig. 7"},{"comment":"The text '10 8' and other inline math are not typeset correctly; also, the resolution plots would benefit from error bars or at least a statement of the statistical uncertainty from the finite Monte Carlo sample.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a timely and practical problem for CTAO, and the Monte Carlo study is substantial. However, the evaluation protocol—especially the selection of the best parameter point on the same data used for scoring—needs to be reworked before the central claim can be accepted. The noise contradiction between the conclusion and Fig. 7 should also be resolved explicitly. Given the scope of the necessary re-analysis, I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a practical, well-scoped engineering study for CTAO data volume reduction, and the time-based DBSCAN idea is worth taking seriously. But the conclusion overreaches: the paper claims Clustering achieves 'lower noise detection' relative to Tailcuts, while Fig. 7 shows the opposite at DVR=40 (about 40% of NSB-only events get noise clusters with Clustering versus 1% with Tailcuts). That is a direct contradiction, not a subtle point.\n\nWhat is genuinely new: applying DBSCAN on pixel position plus arrival time for IACT image selection. The building blocks are standard, but the combination and the systematic comparison across the four CTAO cameras (FlashCam, NectarCAM, LST, SST) is the most complete I know of. The Monte Carlo setups are detailed, and the robustness checks under NSB, broken pixels, and calibration uncertainty are sensible. The signal-efficiency advantage at matched DVR factors looks real: Clustering retains more low-charge pixels (above 3 PE) across the full DVR range for both proton and gamma showers, and the angular resolution gain at DVR 120/140 is a useful result.\n\nThe soft spots are concentrated in the evaluation protocol. Section 3.1 says they plot, for each method and DVR factor, only the parameter set with the highest signal efficiency, measured on the same simulated data used for scoring. Clustering has six free parameters versus four for Tailcuts, so part of the reported gain could be tuning freedom. There are no error bars or counts of tested configurations, so we cannot judge variance or overfitting. The noise axis is worse: the conclusion states 'lower noise detection' but Fig. 7, using the signal-efficiency-tuned parameters, shows Clustering producing far more false clusters at DVR=40. The authors acknowledge you could re-optimize for noise, but that changes the objective and the comparison is no longer apples-to-apples.\n\nThe paper also ships no code or data, and the authors admit real-data validation is still needed. Those are acknowledged limitations, not fatal flaws for an engineering paper.\n\nOverall: the central idea is viable, and the study is a useful reference for anyone planning DVR for CTAO or similar arrays. But the strongest claims as written are not yet supported. A serious referee should ask for a validation split (or at least error bars) in the parameter scan, and a rewrite of the conclusion to separate signal retention from noise rejection.\n\nI would send it to peer review, not desk reject. The IACT community needs this kind of systematic comparison, and the issues are fixable.","headline":"Useful DVR study with a genuine new idea, but the noise-superiority claim is contradicted by the paper's own Fig. 7 and the tuning procedure invites selection bias.","tokens_in":15504,"tokens_out":3544,"would_cite":true,"duration_ms":33219,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Time-based clustering can reduce IACT data storage by 10 to 50 times while keeping more signal pixels than Tailcuts.","keywords":["data compression","imaging Cherenkov telescopes","image processing","gamma-ray astronomy","DBSCAN","time-based clustering","data volume reduction","CTAO"],"falsifier":"A test that tunes Tailcuts and clustering to the same false-positive rate on independent data, or measures their signal efficiency with a fixed parameter budget, would settle the claim: if Tailcuts matches clustering's signal efficiency once noise rates are equalised, the central advantage disappears.","tokens_in":14465,"feed_emoji":"🔭","tokens_out":4655,"duration_ms":46008,"temperature":0.7,"pith_summary":"This paper argues that the Cherenkov Telescope Array Observatory (CTAO) can meet its required 10-to-50-fold data volume reduction using pixel selection that exploits the arrival time of Cherenkov light, not just pixel charge. The authors compare a DBSCAN-based time clustering algorithm with the standard Tailcuts cleaning on simulated air showers for four IACT cameras. They find that clustering keeps more true signal pixels, especially low-charge pixels in shower tails, at equal data volume reduction, and leads to better angular and energy resolution for gamma-ray reconstruction. If correct, the result gives CTAO and other IACT arrays a storage strategy that reduces hundreds of petabytes per year to a few petabytes while retaining more scientifically useful information.","feed_headline":"DBSCAN clustering preserves more pixels than Tailcuts at 10-50x data cuts","feed_subtitle":"Time-based clustering keeps more shower pixels and better angular resolution across all CTAO cameras.","key_machinery":"The central object is Time-based Clustering, a DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm applied to the three-dimensional space of pixel $x$/$y$ positions and reconstructed photon arrival time. DBSCAN groups pixels that are close in space and time into clusters and labels isolated points as noise, with additional noise-cut, charge-threshold, and dilation steps. The time dimension is the key mechanism: signal from a shower arrives nearly simultaneously in neighbouring pixels, while noise triggers are random in time, so density in space-time separates signal from noise more effectively than charge thresholds alone.","core_discovery":"The central claim is that a DBSCAN-based data reduction algorithm, operating on pixel positions and reconstructed arrival times, can provide very robust and effective data volume reduction for CTAO and IACT arrays in general. Time-based clustering achieves superior signal efficiency and lower noise detection than Tailcuts for all studied Cherenkov cameras, retaining more low-level signal pixels that carry shower-tail information. In simulations, clustering identifies more than 85% of signal pixels above 3 photo-electrons for proton showers, while Tailcuts drops to 75%; for gamma showers clustering reaches around 95%. At equal reduction factors it improves angular resolution relative to Tailcuts, gives a modest energy-resolution improvement, and remains robust to night-sky background, broken pixels, and calibration uncertainty in simulations.","pith_inferences":["The comparison may overstate clustering's advantage because each method's parameters were chosen to maximize signal efficiency on the same simulation used for evaluation, and the methods have different numbers of free parameters; a fairer test would fix the noise rate or tune on independent data.","Real observational data are required to confirm robustness, since simulations may not capture all detector artefacts that affect time correlations.","The space-time clustering idea could transfer to other pixelated Cherenkov or imaging detectors where signal arrives coherently in time, even outside gamma-ray astronomy.","If adopted for data volume reduction, the same algorithm can serve as an offline image-cleaning step, unifying online reduction and offline analysis."],"forward_implications":["CTAO can reduce raw data from hundreds of petabytes per year to a few petabytes per year while meeting its 10-to-50-fold data reduction target.","Low-charge pixels in shower tails are retained, which matters for template-based reconstruction methods like ImPACT that use tail information.","Angular resolution of gamma-ray showers improves with clustering compared to Tailcuts at equal or higher data reduction, with a modest energy-resolution gain.","The method works across all four CTAO camera designs tested (LST, SST, NectarCAM, FlashCam), not just FlashCam.","Both methods tolerate increasing night-sky background, up to 20% broken pixels, and up to 50% calibration uncertainty with negligible efficiency loss."],"supporting_citations":[{"why":"Supplies the DBSCAN density-based clustering algorithm that is the core of the proposed method.","marker":"[13]"},{"why":"Defines the Tailcuts/Supercuts image-cleaning approach used as the baseline comparison.","marker":"[10]"},{"why":"Provides the ctapipe-based Hillas parameterisation and Random Forest reconstruction used to evaluate array-level performance.","marker":"[11]"},{"why":"Describes ImPACT, the template-based reconstruction whose sensitivity to shower tails motivates retaining low-charge pixels.","marker":"[12]"},{"why":"Supplies the CORSIKA air-shower simulations used to generate the test data.","marker":"[17]"},{"why":"Supplies the sim_telarray telescope and camera simulation used to model the IACT arrays.","marker":"[18]"}],"fun_headline_variants":["DBSCAN clustering beats Tailcuts for IACT data cuts","Time-based clustering preserves more shower pixels","Clustering improves angular resolution at 10x data cuts","DBSCAN retains 95% signal pixels in gamma showers","Robust clustering reduces CTAO data volume efficiently"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that choosing each method's parameter set to maximize signal efficiency on the same simulated data is a fair way to measure algorithmic performance, so part of the reported advantage could come from differences in how many free parameters each method is allowed to tune.","fun_headline_variants_meta":{"raw":{"variants":["DBSCAN clustering beats Tailcuts for IACT data cuts","Time-based clustering preserves more shower pixels","Clustering improves angular resolution at 10x data cuts","DBSCAN retains 95% signal pixels in gamma showers","Robust clustering reduces CTAO data volume efficiently"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00024,"raw_usage":{"total_tokens":1534,"prompt_tokens":976,"completion_tokens":558,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":480}},"tokens_in":592,"tokens_out":558,"duration_ms":6267,"temperature":1.0,"reasoning_tokens":480,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:47:26.765174+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A test that tunes Tailcuts and clustering to the same false-positive rate on independent data, or measures their signal efficiency with a fixed parameter budget, would settle the claim: if Tailcuts matches clustering's signal efficiency once noise rates are equalised, the central advantage disappears.","supporting_citations":[{"cited_title":"Ester, H.-P","cited_arxiv_id":null,"evidence_quote":"Supplies the DBSCAN density-based clustering algorithm that is the core of the proposed method."},{"cited_title":"Punch, C","cited_arxiv_id":null,"evidence_quote":"Defines the Tailcuts/Supercuts image-cleaning approach used as the baseline comparison."},{"cited_title":"Linho ff, S","cited_arxiv_id":null,"evidence_quote":"Provides the ctapipe-based Hillas parameterisation and Random Forest reconstruction used to evaluate array-level performance."},{"cited_title":"Parsons, J","cited_arxiv_id":null,"evidence_quote":"Describes ImPACT, the template-based reconstruction whose sensitivity to shower tails motivates retaining low-charge pixels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CORSIKA air-shower simulations used to generate the test data."},{"cited_title":"Bernl ¨ohr, Simulation of imaging atmospheric Cherenkov telescopes with CORSIKA and sim telarray, Astroparticle Physics 30 (3) (2008) 149–158","cited_arxiv_id":null,"evidence_quote":"Supplies the sim_telarray telescope and camera simulation used to model the IACT arrays."}],"review_version":1}