REVIEW 3 major objections 4 minor 13 references
To clean or not to clean? Influence of pixel removal on event reconstruction using deep learning in CTAO
T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Removing 88 percent of the pixels through the production-like data-volume-reduction mask leaves the deep-learning reconstruction of gamma-ray events essentially unchanged, while two standard cleaning masks degrade it, especially below 100…
desk verdict Useful MC-based comparison showing DVR cleaning preserves γ-PhysNet performance, but the production-pipeline claim needs more evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is γ-PhysNet, a convolutional neural network with an attention mechanism that performs multi-task event reconstruction directly from DL1 images. The comparison is organized by three pixel-selection masks: tailcut cleaning, a two-threshold procedure with thresholds {8, 4, 2}; lstchain cleaning, the default LST-1 mask consisting of the same tailcut followed by time-consistency filtering on the time map; and data-volume reduction, a tailcut with parameters {8, 4, 0} followed by three morphological dilations. The yardstick is a set of instrument response functions computed on energy bins: energy resolution, energy bias, angular resolution, and the area under the classification curve.
What would settle it
Run the identical γ-PhysNet training and evaluation on real LST-1 Crab Nebula data comparing DVR-masked and unmasked images, and compare energy resolution and angular resolution per energy bin; a difference larger than the statistical uncertainties, especially below 100 GeV, would refute the paper's conclusion.
Extended reading notes
Core claim
On the paper's own terms, the finding is that the data-volume-reduction mask used in production for LST-1 removes 88% of pixels yet leaves γ-PhysNet's energy resolution, energy bias, angular resolution, and gamma-versus-proton classification essentially unchanged across all energy bins, whereas tailcut and lstchain cleaning remove 98% of pixels and produce a significant drop in all these metrics, concentrated below 100 GeV. The same qualitative behavior appears when the simulations are run with a lower night-sky background, so the result is not an artifact of the high-noise setting.
Load-bearing premise
The conclusion rests on Monte-Carlo simulated images standing in for real LST-1 data, and on the study's DVR variant being equivalent to the production pre-processing; if either assumption fails, the no-impact result may not transfer to real observations.
Editorial extensions
If this is right
- The production DVR pipeline for LST-1 can be kept in place without expecting the γ-PhysNet reconstruction to lose performance.
- Switching to lstchain or tailcut cleaning before feeding images to the network would degrade energy and angular resolution and classification, with the largest losses below 100 GeV.
- Aggressive cleaning also discards about 20% of events entirely, so the performance drop is compounded by a loss of usable data.
- The roughly 88% pixel reduction achieved by DVR shows that substantial data compression can coexist with deep-learning event reconstruction.
Reading between the lines
- If a DVR mask can remove 88% of pixels with no measured impact, the information-bearing part of the image may be much smaller than the raw image; a saliency or ablation study on γ-PhysNet could map the minimum pixel set needed for reconstruction.
- Because the three masks all share a tailcut core, a systematic scan of thresholds and dilation counts could locate the boundary where reconstruction performance starts to degrade, potentially enabling even more aggressive compression.
- The study uses one source model (Crab-like), one zenith angle, one telescope, and one network architecture, so extending to other spectra, zenith angles, and telescope types is needed before the no-impact conclusion is generalized to the full CTAO array.
- If real-data Crab observations confirm the Monte-Carlo result, the pipeline can separate concerns: DVR for storage and transfer, with aggressive astrophysical cleaning applied only outside the deep-learning branch.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies how three pixel-removal cleaning procedures (tailcut, lstchain, and a data-volume-reduction (DVR) mask) affect the performance of the γ-PhysNet deep-learning event reconstruction for the LST-1 telescope. Using Monte Carlo simulations tuned for the Crab Nebula with high night-sky background at 20° zenith, the authors compare energy resolution, energy bias, angular resolution, and gamma/proton classification AUC against a no-mask baseline. They report that DVR yields nearly identical performance to no mask, while tailcut and lstchain cleaning cause significant degradation, particularly below 100 GeV. The paper concludes that the DVR preprocessing used in production for LST-1 should not impact the analysis pipeline and notes that real-data validation is left to future work.
Significance. If the main result holds, it is practically important for CTAO: it suggests that the mandatory data-volume-reduction step can be applied to DL1 images without degrading the deep-learning reconstruction, whereas aggressive traditional cleaning masks (tailcut/lstchain) should be avoided when using γ-PhysNet. The study is useful because it provides an apples-to-apples comparison on standard IRF metrics with a no-mask reference, and the authors use public tools (gammalearn, glearn_irfs) that make the pipeline reproducible. The main caveats are that the quantitative claims are not accompanied by statistical uncertainties, the analysis does not account for the 20% of events fully removed by tailcut/lstchain cleaning, and the extrapolation to the production DVR rests on an unverified 'similar to' statement rather than an identical implementation.
major comments (3)
- [Section 3, Figure 3] The IRF plots and the qualitative statements 'very similar response' and 'significant drop' are shown without any statistical uncertainties or confidence intervals. Since the comparisons are made from Monte Carlo samples with finite statistics, especially in the low-energy bins, the reader cannot determine whether the observed differences are statistically meaningful. Please add bootstrap errors, Poisson uncertainties, or another quantitative measure of the spread, and use them when making claims about the magnitude of the impact.
- [Section 2, Dataset; Section 3, Results] The paper states that tailcut and lstchain cleaning remove all pixels in 20% of the images and that these events are discarded from the dataset, but the subsequent IRFs are computed only on the surviving events. In a real analysis pipeline, discarding 20% of gamma-like events changes the effective collection area and the event selection, which is itself part of the 'impact on the analysis pipeline' claimed in the conclusion. The current comparison therefore does not capture the full effect of the cleaning methods. Please either include an efficiency or effective-area metric, or explicitly limit the conclusion to reconstruction performance on events that survive the cleaning.
- [Section 2, DVR paragraph; Section 4, Conclusion] The conclusion that 'the pre-processing used in production for LST-1 should have no impact on the analysis pipeline' relies on the assertion that the DVR method used here is 'similar to the pre-processing used in production.' The paper provides no evidence of functional equivalence: there is no quantitative mask-overlap comparison, no test of the production algorithm on the same events, and no real-data validation (which is deferred to future work). If the production DVR differs in threshold values, number of dilations, ordering of operations, or includes additional time/dead-pixel filtering, then the extrapolation from the tested DVR to the production pipeline is not justified. Please either verify equivalence or soften the conclusion to apply only to the specific DVR implementation tested.
minor comments (4)
- [Section 2, first paragraph] There is a typo: 'distringuish' should be 'distinguish.'
- [Section 2, paragraph after cleaning definitions] The text says 'see Fig. 1' when referring to the effect of masks on images; this should be Fig. 2, since Figure 1 shows the network architecture and Figure 2 shows the mask applications.
- [Section 3, last paragraph] The lower-noise MC experiment is only described with the sentence 'We observe the same behavior.' This is a qualitative claim with no figure or quantitative support; please provide the corresponding plot or at least a summary of the metric differences.
- [Section 4, Conclusion] The study covers only one source model (Crab-like with high NSB), one zenith angle, and one network architecture, yet the conclusion is phrased generally. Adding a caveat that the result has been demonstrated for this configuration would improve the accuracy of the statement.
Circularity Check
No circularity found: the study is an empirical comparison of fixed cleaning masks against a no-mask baseline, and the production-pipeline extrapolation rests on an unverified similarity assumption rather than on a definitional reduction.
full rationale
The paper performs a measurement: it applies three fixed pixel-removal methods (tailcut, lstchain cleaning, and DVR) to Monte Carlo images and compares the resulting γ-PhysNet reconstruction metrics against a 'no mask' baseline. There is no fitted parameter that is later renamed as a prediction: the mask parameters ({8,4,2}, {8,4,2} with time filtering, {8,4,0} with three dilations) are stated a priori from standard or production practice, and the baseline is explicitly external ('We compare the three masking methods with the classical approach, named no mask'). The central result, that DVR and no mask have 'a very similar response for every metric', is an empirical outcome, not a construction. The self-citations to γ-PhysNet and to Vuillaume et al. (2021) are references to the network architecture and to prior sensitivity studies; they do not supply the result that DVR preserves performance. The concluding generalization that 'the pre-processing used in production for LST-1 should have no impact on the analysis pipeline' does depend on the unverified assertion that the DVR method is 'similar to the pre-processing used in production', but this is a limitation of external validity or an unproven equivalence, not a circular derivation: the paper never defines DVR in terms of the conclusion, nor does it fit anything to the IRFs it later reports. The lack of statistical uncertainties on Figure 3 is a reporting weakness, not a circularity. Accordingly, no step satisfies the requirement of exhibiting a specific reduction of the claimed result to its own inputs.
Assumptions & free parameters
free parameters (3)
- tailcut cleaning thresholds =
picture=8, boundary=4, min_neighbors=2
- lstchain cleaning thresholds =
tailcut {8,4,2} plus time consistency filter
- DVR cleaning thresholds and dilations =
tailcut {8,4,0} plus 3 dilations
assumptions (2)
- domain assumption Monte Carlo simulations accurately reproduce LST-1 detector response and night-sky background noise.
- domain assumption The gamma-PhysNet architecture and its training procedure from prior work are appropriate for this comparison.
Cite this review
Pith. "Pith review of To clean or not to clean? Influence of pixel removal on event reconstruction using deep learning in CTAO." pith.science (2026). https://pith.science/paper/JWMKXDIQ
@misc{pith2026250207643,
author = {Pith},
title = {Pith review of: To clean or not to clean? Influence of pixel removal on event reconstruction using deep learning in CTAO},
year = {2026},
howpublished = {\url{https://pith.science/paper/JWMKXDIQ}},
note = {Machine review of arXiv:2502.07643}
}
read the original abstract
The Cherenkov Telescope Array Observatory (CTAO) is the next generation of ground-based observatories employing the imaging air Cherenkov technique for the study of very high energy gamma rays. The software Gammalearn proposes to apply Deep Learning as a part of the CTAO data analysis to reconstruct event parameters directly from images captured by the telescopes with minimal pre-processing to maximize the information conserved. In CTAO, the data analysis will include a data volume reduction that will definitely remove pixels. This step is necessary for data transfer and storage but could also involve information loss that could be used by sensitive algorithms such as neural networks (NN). In this work, we evaluate the performance of the gamma-PhysNet when applying different cleaning masks on images from Monte-Carlo simulations from the first Large-Sized Telescope. This study is critical to assess the impact of pixel removal in the data processing, mainly motivated by data compression.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archive author booktitle chapter collaboration edition editor eid eprint howpublished institution journal key month note number numpages organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.con...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in " " * FUNCTION format....
-
[3]
A., Antonelli, L., Aramo, C., Arbet-Engels, A., Arcaro, C., et al
Abe, H., Abe, K., Abe, S., Aguasca-Cabot, A., Agudo, I., Crespo, N. A., Antonelli, L., Aramo, C., Arbet-Engels, A., Arcaro, C., et al. 2023 a , The Astrophysical Journal, 956, 80
work page 2023
-
[4]
A., Antonelli, L., Aramo, C., Arbet-Engels, A., Artero, M., Asano, K., Aubert, P., et al
Abe, S., Aguasca-Cabot, A., Agudo, I., Crespo, N. A., Antonelli, L., Aramo, C., Arbet-Engels, A., Artero, M., Asano, K., Aubert, P., et al. 2023 b , Astronomy & astrophysics, 673, A75
work page 2023
-
[5]
De, S., Maitra, W., Rentala, V., & Thalapillil, A. M. 2023, Physical Review D, 107, 083026
work page 2023
-
[6]
2010, Astroparticle Physics, 34, 25
Fiasson, A., Dubois, F., Lamanna, G., Masbou, J., & Rosier-Lees, S. 2010, Astroparticle Physics, 34, 25
work page 2010
-
[7]
2021, in International Conference on Pattern Recognition (Springer), 174
Jacquemont, M., Vuillaume, T., Benoit, A., Maurin, G., & Lambert, P. 2021, in International Conference on Pattern Recognition (Springer), 174
work page 2021
-
[8]
2023, in Proceedings, 38th International Cosmic Ray Conference, vol
Linhoff, M., Beiske, L., Biederbeck, N., Fröse, S., Kosack, K., & Nickel, L. 2023, in Proceedings, 38th International Cosmic Ray Conference, vol. 444-703
work page 2023
Show all 13 references
-
[9]
2021, arXiv preprint arXiv:2109.03515
Lopez-Coto, R., Moralejo, A., Artero, M., Baquero, A., Bernardos, M., Contreras, J., Di Pierro, F., Garc \' a, E., Kerszberg, D., L \'o pez-Moya, M., et al. 2021, arXiv preprint arXiv:2109.03515
2021 arXiv
-
[10]
(CTA, LST Project) 2022, in ASP Conf
L\'opez-Coto, R., et al. (CTA, LST Project) 2022, in ASP Conf. Ser., vol. 532, 357
2022
-
[11]
2021, arXiv preprint arXiv:2112.01828
Miener, T., L \'o pez-Coto, R., Contreras, J., Green, J., Green, D., Mariotti, E., Nieto, D., Romanato, L., & Yadav, S. 2021, arXiv preprint arXiv:2112.01828
2021 arXiv
-
[12]
D., & Ohm, S
Parsons, R. D., & Ohm, S. 2020, The European Physical Journal C, 80, 1
2020
-
[13]
A., Poireau, V., Maurin, G., Benoit, A., Lambert, P., Lamanna, G., & Project, C.-L
Vuillaume, T., Jacquemont, M., de Bony de Lavergne, M., Sanchez, D. A., Poireau, V., Maurin, G., Benoit, A., Lambert, P., Lamanna, G., & Project, C.-L. 2021, Analysis of the cherenkov telescope array first large-sized telescope real data using convolutional neural networks. 21...
2021 arXiv
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.