REVIEW 4 major objections 4 minor 26 references
A Systematic Assessment of Data Volume Reduction for IACTs
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Time-based clustering can reduce IACT data storage by 10 to 50 times while keeping more signal pixels than Tailcuts.
desk verdict Useful DVR study with a genuine new idea, but the noise-superiority claim is contradicted by the paper's own Fig. 7 and the tuning procedure invites selection bias. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is Time-based Clustering, a DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm applied to the three-dimensional space of pixel $x$/$y$ positions and reconstructed photon arrival time. DBSCAN groups pixels that are close in space and time into clusters and labels isolated points as noise, with additional noise-cut, charge-threshold, and dilation steps. The time dimension is the key mechanism: signal from a shower arrives nearly simultaneously in neighbouring pixels, while noise triggers are random in time, so density in space-time separates signal from noise more effectively than charge thresholds alone.
What would settle it
A test that tunes Tailcuts and clustering to the same false-positive rate on independent data, or measures their signal efficiency with a fixed parameter budget, would settle the claim: if Tailcuts matches clustering's signal efficiency once noise rates are equalised, the central advantage disappears.
Extended reading notes
Core claim
The central claim is that a DBSCAN-based data reduction algorithm, operating on pixel positions and reconstructed arrival times, can provide very robust and effective data volume reduction for CTAO and IACT arrays in general. Time-based clustering achieves superior signal efficiency and lower noise detection than Tailcuts for all studied Cherenkov cameras, retaining more low-level signal pixels that carry shower-tail information. In simulations, clustering identifies more than 85% of signal pixels above 3 photo-electrons for proton showers, while Tailcuts drops to 75%; for gamma showers clustering reaches around 95%. At equal reduction factors it improves angular resolution relative to Tailcuts, gives a modest energy-resolution improvement, and remains robust to night-sky background, broken pixels, and calibration uncertainty in simulations.
Load-bearing premise
The comparison assumes that choosing each method's parameter set to maximize signal efficiency on the same simulated data is a fair way to measure algorithmic performance, so part of the reported advantage could come from differences in how many free parameters each method is allowed to tune.
Editorial extensions
If this is right
- CTAO can reduce raw data from hundreds of petabytes per year to a few petabytes per year while meeting its 10-to-50-fold data reduction target.
- Low-charge pixels in shower tails are retained, which matters for template-based reconstruction methods like ImPACT that use tail information.
- Angular resolution of gamma-ray showers improves with clustering compared to Tailcuts at equal or higher data reduction, with a modest energy-resolution gain.
- The method works across all four CTAO camera designs tested (LST, SST, NectarCAM, FlashCam), not just FlashCam.
- Both methods tolerate increasing night-sky background, up to 20% broken pixels, and up to 50% calibration uncertainty with negligible efficiency loss.
Reading between the lines
- The comparison may overstate clustering's advantage because each method's parameters were chosen to maximize signal efficiency on the same simulation used for evaluation, and the methods have different numbers of free parameters; a fairer test would fix the noise rate or tune on independent data.
- Real observational data are required to confirm robustness, since simulations may not capture all detector artefacts that affect time correlations.
- The space-time clustering idea could transfer to other pixelated Cherenkov or imaging detectors where signal arrives coherently in time, even outside gamma-ray astronomy.
- If adopted for data volume reduction, the same algorithm can serve as an offline image-cleaning step, unifying online reduction and offline analysis.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes and evaluates data volume reduction (DVR) algorithms for CTAO imaging atmospheric Cherenkov telescopes, comparing a Tailcuts baseline with a DBSCAN-based time-space clustering method. Using CORSIKA/sim_telarray simulations of MST-FlashCam and, more briefly, of the other CTAO cameras, the authors measure signal efficiency, noise false-positive rate, robustness to NSB/broken pixels/calibration uncertainty, and angular/energy resolution after cleaning. The central claim is that Clustering achieves superior signal efficiency and lower noise detection than Tailcuts across all studied cameras, with a factor 10–50 reduction in stored data volume. The manuscript also discusses waveform cropping, computational time, and several clustering variants in an appendix.
Significance. If the central comparison is reliable, the result is practically important: CTAO faces a reduction from hundreds of PB/year to a few PB/year, and a robust pixel-selection method that preserves low-charge shower pixels would directly affect the observatory's data pipeline. The paper benefits from detailed Monte Carlo simulations, a concrete algorithmic proposal that is applicable beyond CTAO, and a sensible set of robustness checks. However, the evaluation protocol contains a load-bearing selection-bias problem, and the 'lower noise detection' conclusion is contradicted by the paper's own Fig. 7 under the stated parameter choice. The core idea remains plausible, but the headline superiority claim is not yet established as presented.
major comments (4)
- [§3.1, Fig. 4] The signal-efficiency comparison is vulnerable to selection bias: the text states that for each DVR factor 'only the point with higher signal efficiency is selected and plotted for each method,' meaning that the same simulated data are used both to choose the free parameters and to evaluate the comparison. Clustering has six free parameters (spatial scale, time scale, minPts, noise cut, charge threshold, dilated rows) while Tailcuts has four, and no validation split, no number of tested configurations, and no error bars are reported. The apparent advantage of Clustering may therefore reflect extra tuning freedom rather than an algorithmic property. I request a separate validation sample, a count of the parameter sets tested per DVR factor, and uncertainty estimates on the efficiency curves.
- [§4.1, Fig. 7 and §6] The Conclusion's statement that Clustering achieves 'lower noise detection compared to Tailcuts' is directly contradicted by Fig. 7: at DVR=40, using the parameters selected for signal efficiency in Tables 1 and 2, Clustering misidentifies clusters in roughly 40% of NSB-only events while Tailcuts does so in about 1%. The paper notes that parameters could be re-optimized for noise rejection, but then the two methods are not compared under a common optimization objective. The claim as written is unsupported; either report false-positive rates at matched signal efficiency, or restrict the conclusion to signal efficiency and reconstruction resolution, with a clearly stated caveat about the noise trade-off.
- [§3.2, Fig. 6] The angular and energy resolution comparison is not controlled in a way that supports the conclusion. The two methods are compared at DVR factors of 120 and 140 rather than at exactly matched reduction values, and the parameters were selected to maximize signal efficiency rather than reconstruction performance, as the authors themselves acknowledge in the text. Because parameter choice strongly affects resolution, the apparent improvement from Clustering could be an artifact of the tuning criterion. I ask for matched DVR factors (e.g., by selecting parameters for each method at the same achieved reduction), parameter selection on validation data, and statistical uncertainties on the resolution curves.
- [§5, Fig. 12] The blanket claim that Clustering has 'lower noise detection' or 'always detects more signal and less noise' for all cameras is not supported by the presented figures. Fig. 12 shows that for the SST camera, Clustering is slightly worse than Tailcuts at high charge thresholds, and Fig. 11 contains no noise axis at all. Please add the noise false-positive rate for each camera or qualify the claim to the specific cameras and charge ranges shown.
minor comments (4)
- [Abstract and throughout] There are numerous typographical errors, e.g., 'The CTAO data rates needs' in the abstract and 'di fferent', 'T ailcuts', and 'su ffi ciently' in the body. A careful proofread is needed.
- [Tables 1 and 2] In Table 2, the boundary threshold for DVR=30 (5.0 PE) is higher than the picture threshold (4.0 PE), which is unusual for Tailcuts; please explain or correct this entry.
- [Fig. 7] The y-axis label says 'Relative number of events with at least one cluster' but the caption clarifies normalization to the nominal NSB rate; please make the normalization explicit in the axis label or caption.
- [§3.2] The text '10 8' and other inline math are not typeset correctly; also, the resolution plots would benefit from error bars or at least a statement of the statistical uncertainty from the finite Monte Carlo sample.
Circularity Check
No significant circularity: the performance comparison is measured against independent simulations with Tailcuts as an external baseline; parameter selection on the same data is an in-sample tuning concern, not a definitional or self-citational circularity.
full rationale
The paper's central claims are empirical comparisons of two cleaning/reduction algorithms on simulated CTAO data. Signal efficiency is measured as the fraction of true signal pixels above 3 PE that survive cleaning, and the DVR factor is computed directly as the ratio of pixels before and after reduction; neither quantity is defined in terms of the other method's output or in terms of the conclusion being drawn. Tailcuts is an external, independently specified baseline algorithm, and the Time-based Clustering method is an application of the standard DBSCAN algorithm; their relative performance is not forced by construction. The main methodological weakness is that Section 3.1 states 'only the point with higher signal efficiency is selected and plotted for each method' and Tables 1 and 2 give parameters 'selected based on Fig. 4' to maximize signal efficiency on the same simulated data used for evaluation. This is an in-sample selection or overfitting concern, not circular reasoning: the algorithms are not defined in terms of the measured efficiencies, and no equation reduces the comparison to its own inputs. The self-citations present, e.g., to Schwefer et al. for reconstruction methods and to Steinmaßl for the noise-cut concept, are ordinary methodological references and are not load-bearing justifications of the paper's central performance claim. The internal tension between the conclusion's 'lower noise detection' statement and Fig. 7's noise-cluster results is an inconsistency in the interpretation of the measurements, not a circularity. Therefore the derivation chain is self-contained with respect to the circularity failure modes considered here.
Assumptions & free parameters
free parameters (12)
- Spatial scale for DBSCAN (m) =
0.175 (DVR 30/40/120), 0.15 (DVR 140)
- Time scale for DBSCAN (ns) =
5.0, 4.0, 5.0, 5.0
- minPts (DBSCAN minimum points) =
6, 4, 5, 7
- Noise cut (PE) =
2.3, 2.5, 2.3, 2.3
- Charge threshold (PE) =
8.0, 6.0, 10.0, 8.0
- Number of dilated rows =
2, 1, 0, 0
- Tailcuts picture threshold (PE) =
4.0, 6.0, 5.0, 12.0
- Tailcuts boundary threshold (PE) =
5.0, 4.0, 3.0, 2.0
- Tailcuts picture neighbours =
2, 2, 1, 0
- Tailcuts number of rows =
2, 2, 0, 0
- Weighted clustering parameters A and B =
A=2.0, B=4.0
- Waveform integration window =
7 ns (3 ns before peak, 4 ns after)
assumptions (4)
- domain assumption Simulated events from CORSIKA and sim_telarray with the preliminary CTAO Prod-6 configuration faithfully represent the response of real CTAO cameras and the night sky background.
- ad hoc to paper A true signal pixel is adequately defined as containing more than 3 photo-electrons.
- domain assumption The neighbour-based waveform extractor with interpolation and pole-zero deconvolution recovers charge and arrival time accurately enough for time-based clustering.
- domain assumption The DVR factor computed from proton simulations is a fair normalization for comparing gamma-ray signal efficiency.
Cite this review
Pith. "Pith review of A Systematic Assessment of Data Volume Reduction for IACTs." pith.science (2026). https://pith.science/paper/RZKQL3D4
@misc{pith2026241114852,
author = {Pith},
title = {Pith review of: A Systematic Assessment of Data Volume Reduction for IACTs},
year = {2026},
howpublished = {\url{https://pith.science/paper/RZKQL3D4}},
note = {Machine review of arXiv:2411.14852}
}
read the original abstract
High energy cosmic-rays generate air showers when they enter Earth's atmosphere. Ground-based gamma-ray astronomy is possible using either direct detection of shower particles at mountain altitudes, or with arrays of imaging air-Cherenkov telescopes (IACTs). Advances in the technique and larger collection areas have increased the rate at which air-shower events can be captured, and the amount of data produced by modern high-time-resolution Cherenkov cameras. Therefore, Data Volume Reduction (DVR) has become critical for such telescope arrays, ensuring that only useful information is stored long-term. Given the vast amount of raw data, owing to the highest resolution and sensitivity, the upcoming Cherenkov Telescope Array Observatory (CTAO) will need robust data reduction strategies to ensure efficient handling and analysis. The CTAO data rates needs be reduced from hundreds of Petabytes (PB) per year to a few PB/year. This paper presents algorithms tailored for CTAO but also applicable for other arrays, focusing on selecting pixels likely to contain shower light. It describes and evaluates multiple algorithms based on their signal efficiency, noise rejection, and shower reconstruction. With a focus on a time-based clustering algorithm which demonstrates a notable enhancement in the retention of low-level signal pixels. Moreover, the robustness is assessed under different observing conditions, including detector defects. Through testing and analysis, it is shown that these algorithms offer promising solutions for efficient volume reduction in CTAO, addressing the challenges posed by the array's very large data volume and ensuring reliable data storage amidst varying observational conditions and hardware issues.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
T. C. Weekes, The atmospheric Cherenkov technique in very high energy gamma-ray astronomy, Space Science Reviews 75 (1) (1996) 1–15
work page 1996
-
[2]
S. Vercellone, C. Consortium, et al., The next generation Cherenkov Telescope Array observatory: CTA, Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment 766 (2014) 73–77
work page 2014
-
[3]
CTAO Website, layouts for alpha con- figuration, https://www.ctao.org/news/ ctao-releases-layouts-for-alpha-configuration/ , (Ac- cessed 08 November 2024)
work page 2024
-
[4]
CTAO Website, data and computing, https://www.ctao.org/ emission-to-discovery/data-and-computing/ , (Accessed 08 November 2024)
work page 2024
-
[5]
F. Werner, C. Bauer, S. Bernhard, M. Capasso, S. Diebold, F. Eisenkolb, S. Eschbach, D. Florin, C. F ¨ohr, S. Funk, et al., Performance verification of the FlashCam prototype camera for the Cherenkov Telescope Array, Nuclear Instruments and Methods in Physics Research Section A: Accel- erators, Spectrometers, Detectors and Associated Equipment 876 (2017) ...
work page 2017
-
[6]
J. L. Contreras, K. Satalecka, K. Bernl ¨ohr, C. Boisson, J. Bregeon, A. Bul- garelli, G. de Cesare, R. Reyes, V . Fioretti, K. Kosack, et al., Data model issues in the Cherenkov Telescope Array project, arXiv preprint arXiv:1508.07584 (2015)
work page Pith review arXiv 2015
-
[7]
A. M. Hillas, Cerenkov light images of EAS produced by primary gamma, in: 19th Intern. Cosmic Ray Conf-V ol. 3, no. OG-9.5-3, 1985
work page 1985
-
[8]
G. Schwefer, R. Parsons, J. Hinton, A hybrid approach to event re- construction for atmospheric Cherenkov Telescopes combining machine learning and likelihood fitting, Astroparticle Physics 163 (2024) 103008
work page 2024
Show all 26 references
-
[9]
Nowlin, J
C. Nowlin, J. Blankenship, Elimination of undesirable undershoot in the operation and testing of nuclear pulse amplifiers, Review of Scientific Instruments 36 (12) (1965) 1830–1839
1965
-
[10]
Punch, C
M. Punch, C. Akerlof, M. Cawley, D. Fegan, R. Lamb, M. Lawrence, M. Lang, D. Lewis, D. Meyer, K. O’Flaherty, et al., Supercuts: an im- proved method of selecting gamma-rays, in: Proceedings of the 22nd In- ternational Cosmic Ray Conference. 11-23 August, 1991. Dublin, Ire- lan...
1991
-
[11]
Linho ff, S
M. Linho ff, S. Bhattacharyya, J. P ´erez Romero, S. Stani ˇc, V . V odeb, S. V orobiov, D. Zavrtanik, M. Zavrtanik, M. ˇZivec, ctapipe-prototype open event reconstruction pipeline for the Cherenkov Telescope Array, in: 38th International Cosmic Ray Conference [also] ICRC2023,...
2023
-
[12]
Parsons, J
R. Parsons, J. Hinton, A Monte Carlo template based analysis for air- Cherenkov arrays, Astroparticle Physics 56 (2014) 26–34
2014
-
[13]
Ester, H.-P
M. Ester, H.-P. Kriegel, J. Sander, X. Xu, et al., A density-based algorithm for discovering clusters in large spatial databases with noise, in: kdd, V ol. 96, 1996, pp. 226–231
1996
-
[14]
R. J. Campello, D. Moulavi, J. Sander, Density-based clustering based on hierarchical density estimates, in: Pacific-Asia conference on knowledge discovery and data mining, Springer, 2013, pp. 160–172
2013
-
[15]
Ankerst, M
M. Ankerst, M. M. Breunig, H.-P. Kriegel, J. Sander, OPTICS: Ordering points to identify the clustering structure, ACM Sigmod record 28 (2) (1999) 49–60
1999
-
[16]
Steinmaßl, Probing particle acceleration in stellar binary systems using gamma-ray observations, Ph.D
S. Steinmaßl, Probing particle acceleration in stellar binary systems using gamma-ray observations, Ph.D. thesis, Ruprecht-Karls-Universit ¨at Hei- delberg (2023)
2023
-
[17]
D. Heck, J. Knapp, J. Capdevielle, G. Schatz, T. Thouw, et al., CORSIKA: A Monte Carlo code to simulate extensive air showers (1998)
1998
-
[18]
Bernl ¨ohr, Simulation of imaging atmospheric Cherenkov telescopes with CORSIKA and sim telarray, Astroparticle Physics 30 (3) (2008) 149–158
K. Bernl ¨ohr, Simulation of imaging atmospheric Cherenkov telescopes with CORSIKA and sim telarray, Astroparticle Physics 30 (3) (2008) 149–158
2008
-
[19]
Zanin, CTA–the World’s largest ground-based gamma-ray observa- tory, in: Proceedings of 37th International Cosmic Ray Conference (ICRC2021), V ol
R. Zanin, CTA–the World’s largest ground-based gamma-ray observa- tory, in: Proceedings of 37th International Cosmic Ray Conference (ICRC2021), V ol. 395, SISSA, 2022, p. 005
2022
-
[21]
F. A. Aharonian, W. Hofmann, A. Konopelko, H. V ¨olk, The potential of ground based arrays of imaging atmospheric Cherenkov telescopes. I. Determination of shower parameters, Astroparticle Physics 6 (3-4) (1997) 343–368
1997
-
[22]
Donath, C
A. Donath, C. Deil, M. P. Arribas, J. King, E. Owen, R. Terrier, I. Re- ichardt, J. Harris, R. B¨uhler, S. Klepser, Gammapy-A Python package for gamma-ray astronomy, arXiv preprint arXiv:1509.07408 (2015)
2015 arXiv
-
[23]
Van Rossum, F
G. Van Rossum, F. L. Drake, et al., Python reference manual, V ol. 111, Centrum voor Wiskunde en Informatica Amsterdam, 1995
1995
-
[24]
Cornils, S
R. Cornils, S. Gillessen, I. Jung, W. Hofmann, M. Beilicke, K. Bernl ¨ohr, O. Carrol, S. Elfahem, G. Heinzelmann, G. Hermann, et al., The optical system of the HESS imaging atmospheric Cherenkov telescopes. Part II: mirror alignment and point spread function, Astroparticle Phy...
2003
-
[25]
Depaoli, Status of the SST camera for the Cherenkov Telescope Array, arXiv preprint arXiv:2310.01183 (2023)
D. Depaoli, Status of the SST camera for the Cherenkov Telescope Array, arXiv preprint arXiv:2310.01183 (2023)
2023 arXiv
-
[26]
Tavernier, J.-F
T. Tavernier, J.-F. Glicenstein, F. Brun, Status and performance results from NectarCAM–a camera for CTA medium sized telescopes, arXiv preprint arXiv:1909.01969 (2019)
2019 arXiv
-
[27]
Kobayashi, A
Y . Kobayashi, A. Okumura, F. Cassol, H. Katagiri, J. Sitarek, P. Gliwny, S. Nozaki, Y . Nogami, Camera Calibration of the CTA-LST prototype, arXiv preprint arXiv:2108.05035 (2021). 12
2021 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.