Pith. sign in

REVIEW 3 major objections 4 minor 13 references

To clean or not to clean? Influence of pixel removal on event reconstruction using deep learning in CTAO

T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Removing 88 percent of the pixels through the production-like data-volume-reduction mask leaves the deep-learning reconstruction of gamma-ray events essentially unchanged, while two standard cleaning masks degrade it, especially below 100…

desk verdict Useful MC-based comparison showing DVR cleaning preserves γ-PhysNet performance, but the production-pipeline claim needs more evidence. read the letter →

arxiv 2502.07643 v1 pith:JWMKXDIQ submitted 2025-02-11 astro-ph.IM

classification astro-ph.IM
keywords imagingatmosphericCherenkovtelescopesCTAOLST-1gamma-rayastronomydeeplearningeventreconstructionpixelcleaningdatavolumereduction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether the pixel-removal step built into CTAO data handling, which is needed for storage and transfer, silently throws away information that a deep-learning event reconstructor relies on. It compares three cleaning masks against no mask using Monte-Carlo simulations of the first Large-Sized Telescope and the γ-PhysNet network, which reconstructs energy, direction, and particle type from telescope images. The answer is asymmetric: the production-like data-volume-reduction mask, which drops 88% of pixels through a tailcut plus three dilations, leaves every performance metric essentially unchanged, while tailcut and lstchain cleanings degrade performance sharply, most visibly below 100 GeV. The authors conclude that the DVR pre-processing used in production for LST-1 should have no impact on the analysis pipeline.

What carries the argument

The central object is γ-PhysNet, a convolutional neural network with an attention mechanism that performs multi-task event reconstruction directly from DL1 images. The comparison is organized by three pixel-selection masks: tailcut cleaning, a two-threshold procedure with thresholds {8, 4, 2}; lstchain cleaning, the default LST-1 mask consisting of the same tailcut followed by time-consistency filtering on the time map; and data-volume reduction, a tailcut with parameters {8, 4, 0} followed by three morphological dilations. The yardstick is a set of instrument response functions computed on energy bins: energy resolution, energy bias, angular resolution, and the area under the classification curve.

What would settle it

Run the identical γ-PhysNet training and evaluation on real LST-1 Crab Nebula data comparing DVR-masked and unmasked images, and compare energy resolution and angular resolution per energy bin; a difference larger than the statistical uncertainties, especially below 100 GeV, would refute the paper's conclusion.

Watch

Extended reading notes

Core claim

On the paper's own terms, the finding is that the data-volume-reduction mask used in production for LST-1 removes 88% of pixels yet leaves γ-PhysNet's energy resolution, energy bias, angular resolution, and gamma-versus-proton classification essentially unchanged across all energy bins, whereas tailcut and lstchain cleaning remove 98% of pixels and produce a significant drop in all these metrics, concentrated below 100 GeV. The same qualitative behavior appears when the simulations are run with a lower night-sky background, so the result is not an artifact of the high-noise setting.

Load-bearing premise

The conclusion rests on Monte-Carlo simulated images standing in for real LST-1 data, and on the study's DVR variant being equivalent to the production pre-processing; if either assumption fails, the no-impact result may not transfer to real observations.

Editorial extensions

If this is right

  • The production DVR pipeline for LST-1 can be kept in place without expecting the γ-PhysNet reconstruction to lose performance.
  • Switching to lstchain or tailcut cleaning before feeding images to the network would degrade energy and angular resolution and classification, with the largest losses below 100 GeV.
  • Aggressive cleaning also discards about 20% of events entirely, so the performance drop is compounded by a loss of usable data.
  • The roughly 88% pixel reduction achieved by DVR shows that substantial data compression can coexist with deep-learning event reconstruction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If a DVR mask can remove 88% of pixels with no measured impact, the information-bearing part of the image may be much smaller than the raw image; a saliency or ablation study on γ-PhysNet could map the minimum pixel set needed for reconstruction.
  • Because the three masks all share a tailcut core, a systematic scan of thresholds and dilation counts could locate the boundary where reconstruction performance starts to degrade, potentially enabling even more aggressive compression.
  • The study uses one source model (Crab-like), one zenith angle, one telescope, and one network architecture, so extending to other spectra, zenith angles, and telescope types is needed before the no-impact conclusion is generalized to the full CTAO array.
  • If real-data Crab observations confirm the Monte-Carlo result, the pipeline can separate concerns: DVR for storage and transfer, with aggressive astrophysical cleaning applied only outside the deep-learning branch.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies how three pixel-removal cleaning procedures (tailcut, lstchain, and a data-volume-reduction (DVR) mask) affect the performance of the γ-PhysNet deep-learning event reconstruction for the LST-1 telescope. Using Monte Carlo simulations tuned for the Crab Nebula with high night-sky background at 20° zenith, the authors compare energy resolution, energy bias, angular resolution, and gamma/proton classification AUC against a no-mask baseline. They report that DVR yields nearly identical performance to no mask, while tailcut and lstchain cleaning cause significant degradation, particularly below 100 GeV. The paper concludes that the DVR preprocessing used in production for LST-1 should not impact the analysis pipeline and notes that real-data validation is left to future work.

Significance. If the main result holds, it is practically important for CTAO: it suggests that the mandatory data-volume-reduction step can be applied to DL1 images without degrading the deep-learning reconstruction, whereas aggressive traditional cleaning masks (tailcut/lstchain) should be avoided when using γ-PhysNet. The study is useful because it provides an apples-to-apples comparison on standard IRF metrics with a no-mask reference, and the authors use public tools (gammalearn, glearn_irfs) that make the pipeline reproducible. The main caveats are that the quantitative claims are not accompanied by statistical uncertainties, the analysis does not account for the 20% of events fully removed by tailcut/lstchain cleaning, and the extrapolation to the production DVR rests on an unverified 'similar to' statement rather than an identical implementation.

major comments (3)
  1. [Section 3, Figure 3] The IRF plots and the qualitative statements 'very similar response' and 'significant drop' are shown without any statistical uncertainties or confidence intervals. Since the comparisons are made from Monte Carlo samples with finite statistics, especially in the low-energy bins, the reader cannot determine whether the observed differences are statistically meaningful. Please add bootstrap errors, Poisson uncertainties, or another quantitative measure of the spread, and use them when making claims about the magnitude of the impact.
  2. [Section 2, Dataset; Section 3, Results] The paper states that tailcut and lstchain cleaning remove all pixels in 20% of the images and that these events are discarded from the dataset, but the subsequent IRFs are computed only on the surviving events. In a real analysis pipeline, discarding 20% of gamma-like events changes the effective collection area and the event selection, which is itself part of the 'impact on the analysis pipeline' claimed in the conclusion. The current comparison therefore does not capture the full effect of the cleaning methods. Please either include an efficiency or effective-area metric, or explicitly limit the conclusion to reconstruction performance on events that survive the cleaning.
  3. [Section 2, DVR paragraph; Section 4, Conclusion] The conclusion that 'the pre-processing used in production for LST-1 should have no impact on the analysis pipeline' relies on the assertion that the DVR method used here is 'similar to the pre-processing used in production.' The paper provides no evidence of functional equivalence: there is no quantitative mask-overlap comparison, no test of the production algorithm on the same events, and no real-data validation (which is deferred to future work). If the production DVR differs in threshold values, number of dilations, ordering of operations, or includes additional time/dead-pixel filtering, then the extrapolation from the tested DVR to the production pipeline is not justified. Please either verify equivalence or soften the conclusion to apply only to the specific DVR implementation tested.
minor comments (4)
  1. [Section 2, first paragraph] There is a typo: 'distringuish' should be 'distinguish.'
  2. [Section 2, paragraph after cleaning definitions] The text says 'see Fig. 1' when referring to the effect of masks on images; this should be Fig. 2, since Figure 1 shows the network architecture and Figure 2 shows the mask applications.
  3. [Section 3, last paragraph] The lower-noise MC experiment is only described with the sentence 'We observe the same behavior.' This is a qualitative claim with no figure or quantitative support; please provide the corresponding plot or at least a summary of the metric differences.
  4. [Section 4, Conclusion] The study covers only one source model (Crab-like with high NSB), one zenith angle, and one network architecture, yet the conclusion is phrased generally. Adding a caveat that the result has been demonstrated for this configuration would improve the accuracy of the statement.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the study is an empirical comparison of fixed cleaning masks against a no-mask baseline, and the production-pipeline extrapolation rests on an unverified similarity assumption rather than on a definitional reduction.

full rationale

The paper performs a measurement: it applies three fixed pixel-removal methods (tailcut, lstchain cleaning, and DVR) to Monte Carlo images and compares the resulting γ-PhysNet reconstruction metrics against a 'no mask' baseline. There is no fitted parameter that is later renamed as a prediction: the mask parameters ({8,4,2}, {8,4,2} with time filtering, {8,4,0} with three dilations) are stated a priori from standard or production practice, and the baseline is explicitly external ('We compare the three masking methods with the classical approach, named no mask'). The central result, that DVR and no mask have 'a very similar response for every metric', is an empirical outcome, not a construction. The self-citations to γ-PhysNet and to Vuillaume et al. (2021) are references to the network architecture and to prior sensitivity studies; they do not supply the result that DVR preserves performance. The concluding generalization that 'the pre-processing used in production for LST-1 should have no impact on the analysis pipeline' does depend on the unverified assertion that the DVR method is 'similar to the pre-processing used in production', but this is a limitation of external validity or an unproven equivalence, not a circular derivation: the paper never defines DVR in terms of the conclusion, nor does it fit anything to the IRFs it later reports. The lack of statistical uncertainties on Figure 3 is a reporting weakness, not a circularity. Accordingly, no step satisfies the requirement of exhibiting a specific reduction of the claimed result to its own inputs.

Assumptions & free parameters 3 free parameters · 2 assumptions · 0 invented entities

No new physical entities are introduced. The free parameters are all cleaning thresholds chosen by hand, not fitted to data. The axiomatic assumptions are reasonable for an MC-based instrument response study, but they limit the generality of the pipeline-level conclusion.

free parameters (3)
  • tailcut cleaning thresholds = picture=8, boundary=4, min_neighbors=2
    Hand-chosen values from the ctapipe implementation; the comparison result may depend on these specific thresholds.
  • lstchain cleaning thresholds = tailcut {8,4,2} plus time consistency filter
    Default LST-1 analysis mask; hand-chosen as in prior work.
  • DVR cleaning thresholds and dilations = tailcut {8,4,0} plus 3 dilations
    Production-like data volume reduction; hand-chosen, and the result is specific to this configuration.
assumptions (2)
  • domain assumption Monte Carlo simulations accurately reproduce LST-1 detector response and night-sky background noise.
    The entire evaluation rests on simulated data; real data verification is deferred to future work, as stated in the Conclusion.
  • domain assumption The gamma-PhysNet architecture and its training procedure from prior work are appropriate for this comparison.
    The paper reuses the network without re-tuning; the interaction between the network and mask could differ for other architectures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of To clean or not to clean? Influence of pixel removal on event reconstruction using deep learning in CTAO." pith.science (2026). https://pith.science/paper/JWMKXDIQ

@misc{pith2026250207643,
  author       = {Pith},
  title        = {Pith review of: To clean or not to clean? Influence of pixel removal on event reconstruction using deep learning in CTAO},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JWMKXDIQ}},
  note         = {Machine review of arXiv:2502.07643}
}
read the original abstract

The Cherenkov Telescope Array Observatory (CTAO) is the next generation of ground-based observatories employing the imaging air Cherenkov technique for the study of very high energy gamma rays. The software Gammalearn proposes to apply Deep Learning as a part of the CTAO data analysis to reconstruct event parameters directly from images captured by the telescopes with minimal pre-processing to maximize the information conserved. In CTAO, the data analysis will include a data volume reduction that will definitely remove pixels. This step is necessary for data transfer and storage but could also involve information loss that could be used by sensitive algorithms such as neural networks (NN). In this work, we evaluate the performance of the gamma-PhysNet when applying different cleaning masks on images from Monte-Carlo simulations from the first Large-Sized Telescope. This study is critical to assess the impact of pixel removal in the data processing, mainly motivated by data compression.

Figures

Figures reproduced from arXiv: 2502.07643 by the authors.

Figure 1
Figure 1. γ-PhysNet (Jacquemont et al. 2021) architecture: multi-task event re￾construction 2. Methods We distringuish 3 pixel selection methods: tailcut cleaning: It consists in a two-threshold tail-cuts procedure as implemented in the ctapipe method (Linhoff et al. 2023). The method has 3 parameters : picture threshold, which is the threshold above which all pixels are retained, boundary thresh￾old, which is the threshold a… view at source ↗
Figure 2
Figure 2. Visualisations of the binary masks application us [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. IRFs for the cleaning methods. From left to right: E [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 11 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archive author booktitle chapter collaboration edition editor eid eprint howpublished institution journal key month note number numpages organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.con...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in " " * FUNCTION format....

  3. [3]

    A., Antonelli, L., Aramo, C., Arbet-Engels, A., Arcaro, C., et al

    Abe, H., Abe, K., Abe, S., Aguasca-Cabot, A., Agudo, I., Crespo, N. A., Antonelli, L., Aramo, C., Arbet-Engels, A., Arcaro, C., et al. 2023 a , The Astrophysical Journal, 956, 80

  4. [4]

    A., Antonelli, L., Aramo, C., Arbet-Engels, A., Artero, M., Asano, K., Aubert, P., et al

    Abe, S., Aguasca-Cabot, A., Agudo, I., Crespo, N. A., Antonelli, L., Aramo, C., Arbet-Engels, A., Artero, M., Asano, K., Aubert, P., et al. 2023 b , Astronomy & astrophysics, 673, A75

  5. [5]

    De, S., Maitra, W., Rentala, V., & Thalapillil, A. M. 2023, Physical Review D, 107, 083026

  6. [6]

    2010, Astroparticle Physics, 34, 25

    Fiasson, A., Dubois, F., Lamanna, G., Masbou, J., & Rosier-Lees, S. 2010, Astroparticle Physics, 34, 25

  7. [7]

    2021, in International Conference on Pattern Recognition (Springer), 174

    Jacquemont, M., Vuillaume, T., Benoit, A., Maurin, G., & Lambert, P. 2021, in International Conference on Pattern Recognition (Springer), 174

  8. [8]

    2023, in Proceedings, 38th International Cosmic Ray Conference, vol

    Linhoff, M., Beiske, L., Biederbeck, N., Fröse, S., Kosack, K., & Nickel, L. 2023, in Proceedings, 38th International Cosmic Ray Conference, vol. 444-703

Show all 13 references
  1. [9]

    2021, arXiv preprint arXiv:2109.03515

    Lopez-Coto, R., Moralejo, A., Artero, M., Baquero, A., Bernardos, M., Contreras, J., Di Pierro, F., Garc \' a, E., Kerszberg, D., L \'o pez-Moya, M., et al. 2021, arXiv preprint arXiv:2109.03515

  2. [10]

    (CTA, LST Project) 2022, in ASP Conf

    L\'opez-Coto, R., et al. (CTA, LST Project) 2022, in ASP Conf. Ser., vol. 532, 357

  3. [11]

    2021, arXiv preprint arXiv:2112.01828

    Miener, T., L \'o pez-Coto, R., Contreras, J., Green, J., Green, D., Mariotti, E., Nieto, D., Romanato, L., & Yadav, S. 2021, arXiv preprint arXiv:2112.01828

  4. [12]

    D., & Ohm, S

    Parsons, R. D., & Ohm, S. 2020, The European Physical Journal C, 80, 1

  5. [13]

    A., Poireau, V., Maurin, G., Benoit, A., Lambert, P., Lamanna, G., & Project, C.-L

    Vuillaume, T., Jacquemont, M., de Bony de Lavergne, M., Sanchez, D. A., Poireau, V., Maurin, G., Benoit, A., Lambert, P., Lamanna, G., & Project, C.-L. 2021, Analysis of the cherenkov telescope array first large-sized telescope real data using convolutional neural networks. 21...

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.