REVIEW 3 major objections 5 minor 21 references
Gamma/hadron separation in the TAIGA experiment with neural network methods
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Convolutional neural networks can separate gamma-ray from hadron-induced air showers in TAIGA imaging Cherenkov telescope data, recovering the Crab Nebula at more than 5.5 sigma in 21 hours of observations.
desk verdict A useful incremental validation of CNN-based gamma/hadron separation at TAIGA, but the headline 6σ Crab signal rests on a train/test split the paper never documents. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the CNN gamma-score classifier: a block of convolutional feature-extraction layers followed by fully connected layers, with ReLU activations, a sigmoid output, and binary cross-entropy training, mapping each preprocessed telescope image to a probability in $[0,1]$ that the primary particle was a gamma. Around it sits a physics-driven pre-selection on the Hillas angle $\alpha$: the angle between the shower image's major axis and the source direction is sharply peaked near zero for gamma events and flat for isotropic hadrons, so requiring $\alpha<20^\circ$ (soft) or $\alpha<6^\circ$ (strict) together with $\mathrm{Size}>120$ photoelectrons already suppresses the hadron background by roughly a factor of 10 before the network runs. The network's role is to refine the surviving sample, and the paper uses the resulting gamma score, thresholded at the value where 50% of validation gammas are lost, as the final classifier.
What would settle it
Inspect the training-data provenance: if any of the 40,000 experimental hadron images were recorded on the same six nights used for the test, retrain the CNN on hadrons from other nights and recompute the LiMa significance; a drop to or below the Hillas value of $5.84\sigma$ would show that the claimed CNN advantage came from seeing the test conditions.
Extended reading notes
Core claim
The central claim is that a CNN, after cleaning, Wobble correction, hexagonal-to-square image transformation, and logarithmic scaling, can be trained to distinguish gamma-ray from hadron-induced air showers with enough purity to extract a TeV source signal from modern IACT data. Trained on 38,400 Monte Carlo gamma images with a $-2.6$ spectral index and 40,000 experimental hadron images (expanded by $60^\circ$ rotations to 470,400 samples), the network assigns each image a gamma score; thresholding at 0.9965 yields 99.93% classification accuracy and hadron suppression $B \approx 3000$ on validation, and combined with a $\mathrm{Size}>120$ phe cut and an angle cut $\alpha<20^\circ$ the effective suppression reaches $B\approx 10000$ with a loss of about half the gamma events. On the 21-hour Crab data, the soft selection gives an excess of 91.2 events and $6.26\sigma$, the strict selection ($\alpha<6^\circ$) gives 68.7 events and $6.48\sigma$, and the standard Hillas cuts give 66 events and $5.84\sigma$. The interpretation the authors defend is that deep learning performs comparably to the classical analysis and deserves continued development, rather than that it fully solves gamma/hadron separation.
Load-bearing premise
The load-bearing premise is that the experimental hadron images used for training were not drawn from the same events or nights as the 21-hour test data; the paper does not state this split, and if they overlap, the reported significance is inflated.
Editorial extensions
If this is right
- A CNN can be deployed alongside the Hillas-parameter method as an independent check on source detections, since the two selections share only a fraction of their accepted events.
- For a fixed observation campaign, strict CNN selection delivers higher statistical significance per hour than the standard cuts when the background rate is the limiting factor, at the cost of retaining fewer gamma events.
- Because the classifier was trained on Monte Carlo gammas with a single power-law index ($-2.6$), applying it to sources with different spectral shapes requires re-evaluating the score threshold.
- The same image features used for classification can support the planned next step of energy-spectrum reconstruction from TAIGA-IACT data.
- The validation numbers ($B\approx 10000$ at 50% gamma loss) set a target for future IACT processing pipelines, but only if the train/test independence is confirmed.
Reading between the lines
- The headline significance is mostly a statement about the pre-selection plus CNN, not about the CNN alone: the $\alpha$ angle already suppresses hadrons by roughly 10, and the CNN's marginal gain over Hillas cuts is about 0.4 to 0.6 sigma.
- If the training hadrons were drawn from the same six nights as the test, the measured $>5.5\sigma$ would likely be optimistic; a clean temporal split is the first thing to check.
- The single anomalous night (29 November) suggests the classifier is sensitive to nightly atmospheric or hardware conditions, so per-night normalization or domain adaptation may improve robustness.
- An interesting extension not explored in the paper: train the CNN without any $\alpha$ cut and compare against Hillas cuts with the same $\mathrm{Size}$ threshold, isolating how much of the separation comes from image morphology rather than from the source-direction prior.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes a convolutional neural network (CNN) approach for gamma/hadron separation in the TAIGA-IACT experiment, trained on Monte Carlo gamma showers and experimental hadron events, and tested on 21 hours of Crab Nebula observations from six nights in November–December 2019. The CNN is compared with the standard Hillas parameter cuts. The authors report LiMa significances of 6.26σ (soft selection) and 6.48σ (strict selection) for the CNN method, versus 5.84σ for the standard Hillas cuts, and conclude that the CNN achieves a performance level comparable to the standard method. The paper also discusses the selection thresholds, the validation procedure, and the distributions of Hillas parameters for the events selected by both methods.
Significance. If the reported result is valid, it provides a useful demonstration that a CNN can serve as an alternative to classical Hillas cuts for IACT data analysis in a high-background regime (gamma-to-hadron ratio around 1:10^4). The paper is valuable for the TAIGA collaboration and for the broader IACT community because it compares the two methods on the same experimental data, uses a standard LiMa significance calculation, and is candid about the limitations, including the fact that the CNN does not outperform the standard method and that fine-tuning is still required. The main potential contribution is the empirical evidence on the achievable background suppression with a CNN on real telescope data. However, the strength of this evidence depends critically on the independence between the experimental hadron events used for training/validation and the test nights, as well as on whether the CNN score threshold was actually applied in the test analysis.
major comments (3)
- [II.1, II.2, and III] The manuscript never states whether the 40,000 training hadrons and 9,600 validation hadrons are disjoint in time or field position from the six Crab nights used as the test set. Since the training and validation hadrons are explicitly chosen to include 'noise effects related with the operation of the telescope,' the reported 6.26σ/6.48σ significances in Table I can be inflated by night-specific leakage if any of those events come from the same nights. Please state the observation nights and ON/OFF regions of the training/validation hadrons, and, if necessary, repeat the analysis with a strictly disjoint temporal split.
- [II.2 and Table I] The validation set gives a class threshold of 0.9965, leaving 3 hadrons out of 9,600 (B≈3000). If this threshold were applied to the test data, the CNN soft row's average OFF count of 161.8 would be reduced to about 0.05; instead Table I reports 161.8, which is exactly the level expected from the Size>120, α<20 pre-cuts alone. This suggests the CNN score threshold was not actually applied in the test, or the table rows report only the pre-CNN selection. Please clarify the exact event-selection chain used for each row and report the post-threshold ON and OFF counts.
- [Abstract, Conclusion, and Table I] The abstract states a signal 'higher than 5.5σ,' the conclusion states 6.5σ, and Table I reports 6.26σ (soft) and 6.48σ (strict). The sentence excluding 29 November reports 6.3σ for Hillas cuts, which is higher than the full-sample 5.84σ in Table I; these discrepancies are not explained. Please reconcile the reported significances and specify which selection and data sample each number refers to.
minor comments (5)
- [Throughout] There are numerous typos and grammatical errors, e.g., 'Hilas' for Hillas, 'handron' for hadron, 'It's demonstrated' instead of 'It is demonstrated,' and 'interdependence of Hills parameters' instead of 'Hillas parameters.' A careful proofreading pass is needed.
- [II.1] The description of the image preprocessing would benefit from a brief explanation of why the logarithmic scaling specifically maps pixel amplitudes to [0,1] and how the axial method handles the hexagonal-to-square transformation, as these details affect reproducibility.
- [III] The sentence 'In the meantime, 46 events were selected simultaneously by the standard method and CNN (soft selection conditions) in ON point' is unclear: does this mean 46 events passed both selections, and what is the total number of selected events in each method for that comparison?
- [III and Figure 5] The caption and text mention 'interdependence of Hillas parameters' but the figure shows distributions and scatter plots; please clarify what is plotted in each panel and how the CNN-selected events are defined if the CNN threshold is applied.
- [IV] The conclusion says '68 gamma events were obtained during this period of time,' but Table I lists an excess of 68.7 for the strict CNN selection; please state whether this is the rounded value and whether the number refers to the excess or to the ON count after background subtraction.
Circularity Check
No significant circularity: the CNN signal significance is a held-out test result, and the classifier's threshold and cuts are set on separate training/validation data, not on the Crab excess.
full rationale
The paper's central claim—Crab signal at 6.26σ/6.48σ with the CNN—is obtained by applying a classifier trained on 38,400 Monte Carlo gamma events and 40,000 experimental hadron events, then validated on 9,500 MC gamma events and 9,600 experimental hadrons, to six test nights (November–December 2019) that are not used for training or validation. The α and Size pre-selections are physics-driven cuts derived from Monte Carlo distributions (retaining 90% of gammas and 20% of hadrons for soft, 50% and <10% for strict), and the 0.9965 class threshold is set on the validation set at 50% gamma loss, not on the Crab ON/OFF excess. The LiMa significances in Table I are measured outcomes on the held-out test data and are compared against an independent Hillas-cut analysis (Eq. 1, cited to prior work [21]); no equation in the paper redefines a fitted parameter as a prediction, and no load-bearing argument reduces to a self-citation. The undocumented temporal/geometric split between the experimental hadron training/validation samples and the test nights is a legitimate statistical-independence risk (potential leakage), but that is a correctness and experimental-reporting concern, not circularity, because the reported significance is not constructed to equal any training target. Consequently, no specific circular step can be exhibited from the paper's own derivation chain.
Assumptions & free parameters
free parameters (6)
- Size threshold =
120 phe
- soft alpha cut =
20 degrees
- strict alpha cut =
6 degrees
- CNN class threshold =
0.9965
- cleaning thresholds =
7 phe low, 14 phe high
- MC spectral index =
-2.6
assumptions (4)
- domain assumption CORSIKA Monte Carlo and the TAIGA detector response simulation accurately model gamma-ray induced air showers and telescope images.
- domain assumption The experimental hadron events used for class 0 training are representative of the hadron background in the test observations and do not include the test signal events.
- domain assumption The Wobble pointing mode and the ten OFF-point averaging correctly estimate the hadron background under the ON source region.
- standard math The LiMa significance formula is applicable to these count statistics.
Cite this review
Pith. "Pith review of Gamma/hadron separation in the TAIGA experiment with neural network methods." pith.science (2026). https://pith.science/paper/OMVNQRIU
@misc{pith2026250201500,
author = {Pith},
title = {Pith review of: Gamma/hadron separation in the TAIGA experiment with neural network methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/OMVNQRIU}},
note = {Machine review of arXiv:2502.01500}
}
read the original abstract
In this work, the ability of rare VHE gamma ray selection with neural network methods is investigated in the case when cosmic radiation flux strongly prevails (ratio up to {10^4} over the gamma radiation flux from a point source). This ratio is valid for the Crab Nebula in the TeV energy range, since the Crab is a well-studied source for calibration and test of various methods and installations in gamma astronomy. The part of TAIGA experiment which includes three Imaging Atmospheric Cherenkov Telescopes observes this gamma-source too. Cherenkov telescopes obtain images of Extensive Air Showers. Hillas parameters can be used to analyse images in standard processing method, or images can be processed with convolutional neural networks. In this work we would like to describe the main steps and results obtained in the gamma/hadron separation task from the Crab Nebula with neural network methods. The results obtained are compared with standard processing method applied in the TAIGA collaboration and using Hillas parameter cuts. It is demonstrated that a signal was received at the level of higher than 5.5{\sigma} in 21 hours of Crab Nebula observations after processing the experimental data with the neural network method.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Sitarek, Galaxies 10, 21 (2022)
J. Sitarek, Galaxies 10, 21 (2022). https://doi.org/10.3390/galaxies10010021
-
[2]
V. S. Murzin, Astrophysics of cosmic rays: Textbook for universities (Logos, Moscow, 2006). (In Russian)
work page 2006
-
[3]
J. Aleksi´ c, S. Ansoldi, L. A. Antonelli, et al., Journal of High Energy Astrophysics 5, 30 (2015). https://doi.org/10.1016/j.jheap.2015.01.002
-
[4]
F. Aharonian, A. G. Akhperjanian, A. R. Bazer-Bachi et al., Astron. Astrophys. 457, 899 (2006). https://doi.org/10.1051/0004-6361:20065351
-
[5]
J. Holder, R. W. Atkins, H. M. Badran, et al. Astropart. Phys. 25, 391 (2006). https://doi.org/10.1016/j.astropartphys.2006.04.002
-
[6]
The CTA Consortium, Science with the Cherenkov Telescope Array (World Scientific, Singapore, 2019). https://doi.org/10.1142/10986
doi:10.1142/10986 2019
-
[7]
N. M. Budnev, L. Kuzmichev, R. Mirzoyan, et al., PoS ICRC2021, 731 (2021). https://doi.org/10.22323/1.395.0731
-
[8]
A. M. Hillas, PoS ICRC1985, 445 (1985)
work page 1985
Show all 21 references
-
[9]
T. C. Weekes, M. F. Cawley, D. J. Fegan, et al. Astrophys. J. 342, 379 (1989). https://doi.org/10.1086/167599
1989 doi
-
[10]
Postnikov, A
E. Postnikov, A. Kryukov, S. Polyakov and D. Zhurov, CEUR-WS Proceedings 2406, 90 (2019). https://ceur-ws.org/Vol-2406/paper11.pdf
2019
-
[11]
Gres and A
E. Gres and A. Kryukov, PoS DLCP2021, 015 (2021). https://doi.org/10.22323/1.410.0015
2021 doi
-
[12]
E. O. Gres, A. P. Kryukov, A. P. Demichev, et al. Moscow Univ. Phys. 78 (Suppl 1), S45–S51 (2023). https://doi.org/10.3103/S002713492307010X
2023 doi
-
[13]
Goodfellow, Y
I. Goodfellow, Y. Bengio, A. Courville, Deep Learning (Adaptive Computation and Machine Learning series) (The MIT Press, USA, 2016)
2016
-
[14]
Hornik, M
K. Hornik, M. Stinchcombe and H. White, Neural Networks 2, 359 (1989). https://doi.org/10.1016/0893-6080(89)90020-8
1989 doi
-
[15]
Grinyuk, E
A. Grinyuk, E. Postnikov and L. Sveshnikova, Phys. Atom. Nuclei 83, 262–267 (2020). https://doi.org/10.1134/S106377882002012X
2020 doi
-
[16]
D. Heck, J. Knapp, J. N. Capdevielle, et al., report FZKA-6019 (1998)
1998
-
[17]
E. B. Postnikov, I. I. Astapov, P. A. Bezyazeekov, et al., Bull. Russ. Acad. Sci.: Phys. 83, 955 (2019). https://doi.org/10.3103/S1062873819080331
2019 doi
-
[18]
V. P. Fomin, A. Stepanian, R. Lamb, et al., Astropart. Phys. 2, 137 (1994)
1994
-
[19]
Hoogeboom, J
E. Hoogeboom, J. W. T. Peters, T. S. Cohen and M. Welling, International Conference on Learning Representations (2018). https://openreview.net/forum?id=r1vuQG-CW
2018
-
[20]
Li and Y.-Q
T.-P. Li and Y.-Q. Ma, Astrophys. J. 272, 317 (1983). https://doi.org/10.1086/161295
1983 doi
-
[21]
L. G. Sveshnikova, P. A. Volchugov, E. B. Postnikov, et al. Bull. Russ. Acad. Sci. Phys. 87, 904–909 (2023). https://doi.org/10.3103/S1062873823702738 7
2023 doi
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.