{"id":"ecfeb94d-bd6c-4db2-bd41-36f859dc23dc","arxiv_id":"2502.01500","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A convolutional neural network with alpha and Size pre-cuts separates gamma-ray events from cosmic-ray background in TAIGA-IACT data at a level comparable to the standard Hillas method, yielding a Crab Nebula signal at about 6 sigma in 21 hours.","lead":"The authors apply a convolutional neural network to gamma/hadron separation in the TAIGA imaging atmospheric Cherenkov telescope data, achieving a LiMa signal significance around 6 sigma for the Crab Nebula in 21 hours. The result is comparable to the standard Hillas-parameter cut method, suggesting CNNs are a viable independent analysis route for this experiment.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 6σ Crab signal is load-bearing on undocumented independence between the experimental hadron training/validation samples and the 21-hour test nights; if training hadrons come from the same nights, the CNN's background suppression is potentially inflated by night-specific leakage.","rationale":"The reader's weakest_assumption—undocumented train/test independence for the experimental hadron sample—is the same concern I identify as most load-bearing. I verified from the text that Section II.1 describes class 0 as 'experimental data (hadrons)' and Section III defines the test period, but no statement separates the training/validation hadrons from the test nights. The paper's own justification for using experimental hadrons, namely capturing telescope noise effects, makes night-specific overfitting a real risk rather than a pedantic worry. If the training hadrons include events or conditions from the same six nights, the CNN could suppress background more effectively on those nights than on truly unseen data, directly inflating the reported 6.26σ/6.48σ significances. This concern is central because the paper's headline claim is exactly that significance level. I did not elevate the LiMa α inconsistency or the 29 November anomaly to primary status: those could stem from a definitional convention or a single-night systematic, whereas the independence assumption, if violated, invalidates the whole comparison. The result may well survive a provenance check, so the appropriate verdict remains CONDITIONAL pending that check; I therefore leave the reader's verdict unchanged.","tokens_in":7348,"tokens_out":8471,"duration_ms":73056,"concrete_test":"Request the run/night provenance of the class-0 training and validation hadrons and check whether any run ID overlaps the six test nights (Nov–Dec 2019). If overlap exists, retrain the CNN from scratch on hadrons drawn exclusively from nights not in the test set (same MC gamma sample, same Size/alpha pre-cuts, same preprocessing) and recompute the LiMa significances in Table I. A drop below ~5σ would demonstrate leakage; a stable ≥5.5σ would clear the concern. Also report whether training hadrons were taken only from OFF regions, so that gamma-ray events in the ON region were never labeled as hadrons.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—Crab Nebula signal at 6.26σ/6.48σ with the CNN (Table I; abstract: 'higher than 5.5σ')—rests on the classifier having been evaluated on data it did not see during training. Section II.1 states that class 0 consists of 'experimental data (hadrons)' and the validation set also uses 9,600 experimental hadrons, but the manuscript never specifies which observation nights or ON/OFF regions those hadrons come from. Section III defines the test set as six nights in November–December 2019. If any of the 40,000 training hadrons (or the validation hadrons used to set the 0.9965 class threshold) were drawn from those same nights, the CNN can exploit night-specific detector conditions or even memorize background events, inflating the ON/OFF excess. This is not a remote possibility: the authors justify using experimental hadrons precisely because they contain 'noise effects related with the operation of the telescope,' so such effects are exactly what the network is trained on. Without a documented temporal/geometric split, the reported 6σ excess cannot be attributed to gamma/hadron separation rather than to training/test leakage. The 'anomalous' night of 29 November and the apparent inconsistency in the LiMa α convention are secondary; the independence assumption is the load-bearing one.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a convolutional neural network (CNN) approach for gamma/hadron separation in the TAIGA-IACT experiment, trained on Monte Carlo gamma showers and experimental hadron events, and tested on 21 hours of Crab Nebula observations from six nights in November–December 2019. The CNN is compared with the standard Hillas parameter cuts. The authors report LiMa significances of 6.26σ (soft selection) and 6.48σ (strict selection) for the CNN method, versus 5.84σ for the standard Hillas cuts, and conclude that the CNN achieves a performance level comparable to the standard method. The paper also discusses the selection thresholds, the validation procedure, and the distributions of Hillas parameters for the events selected by both methods.","tokens_in":7531,"tokens_out":9052,"duration_ms":78399,"significance":"If the reported result is valid, it provides a useful demonstration that a CNN can serve as an alternative to classical Hillas cuts for IACT data analysis in a high-background regime (gamma-to-hadron ratio around 1:10^4). The paper is valuable for the TAIGA collaboration and for the broader IACT community because it compares the two methods on the same experimental data, uses a standard LiMa significance calculation, and is candid about the limitations, including the fact that the CNN does not outperform the standard method and that fine-tuning is still required. The main potential contribution is the empirical evidence on the achievable background suppression with a CNN on real telescope data. However, the strength of this evidence depends critically on the independence between the experimental hadron events used for training/validation and the test nights, as well as on whether the CNN score threshold was actually applied in the test analysis.","major_comments":[{"comment":"The manuscript never states whether the 40,000 training hadrons and 9,600 validation hadrons are disjoint in time or field position from the six Crab nights used as the test set. Since the training and validation hadrons are explicitly chosen to include 'noise effects related with the operation of the telescope,' the reported 6.26σ/6.48σ significances in Table I can be inflated by night-specific leakage if any of those events come from the same nights. Please state the observation nights and ON/OFF regions of the training/validation hadrons, and, if necessary, repeat the analysis with a strictly disjoint temporal split.","section":"II.1, II.2, and III"},{"comment":"The validation set gives a class threshold of 0.9965, leaving 3 hadrons out of 9,600 (B≈3000). If this threshold were applied to the test data, the CNN soft row's average OFF count of 161.8 would be reduced to about 0.05; instead Table I reports 161.8, which is exactly the level expected from the Size>120, α<20 pre-cuts alone. This suggests the CNN score threshold was not actually applied in the test, or the table rows report only the pre-CNN selection. Please clarify the exact event-selection chain used for each row and report the post-threshold ON and OFF counts.","section":"II.2 and Table I"},{"comment":"The abstract states a signal 'higher than 5.5σ,' the conclusion states 6.5σ, and Table I reports 6.26σ (soft) and 6.48σ (strict). The sentence excluding 29 November reports 6.3σ for Hillas cuts, which is higher than the full-sample 5.84σ in Table I; these discrepancies are not explained. Please reconcile the reported significances and specify which selection and data sample each number refers to.","section":"Abstract, Conclusion, and Table I"}],"minor_comments":[{"comment":"There are numerous typos and grammatical errors, e.g., 'Hilas' for Hillas, 'handron' for hadron, 'It's demonstrated' instead of 'It is demonstrated,' and 'interdependence of Hills parameters' instead of 'Hillas parameters.' A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The description of the image preprocessing would benefit from a brief explanation of why the logarithmic scaling specifically maps pixel amplitudes to [0,1] and how the axial method handles the hexagonal-to-square transformation, as these details affect reproducibility.","section":"II.1"},{"comment":"The sentence 'In the meantime, 46 events were selected simultaneously by the standard method and CNN (soft selection conditions) in ON point' is unclear: does this mean 46 events passed both selections, and what is the total number of selected events in each method for that comparison?","section":"III"},{"comment":"The caption and text mention 'interdependence of Hillas parameters' but the figure shows distributions and scatter plots; please clarify what is plotted in each panel and how the CNN-selected events are defined if the CNN threshold is applied.","section":"III and Figure 5"},{"comment":"The conclusion says '68 gamma events were obtained during this period of time,' but Table I lists an excess of 68.7 for the strict CNN selection; please state whether this is the rounded value and whether the number refers to the excess or to the ON count after background subtraction.","section":"IV"}],"recommendation":"major_revision","confidential_remarks":"The most important issue is the lack of a documented split between the training/validation hadron samples and the test nights. If the hadrons used for training come from the same nights, the reported significance could be inflated, and this cannot be assessed from the current text. In addition, the numbers in Table I appear incompatible with the stated validation background suppression, suggesting that the CNN threshold may not have been applied to the test data; this needs to be clarified before the paper can be considered for publication. The topic is within the scope of the journal, and the manuscript is potentially salvageable after these points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a credible incremental step in a program the authors have been running for years. They take the CNN classifier from a simulated 1:1000 imbalance to 1:10000 and, for the first time in this line of work, show it on real Crab data with a LiMa significance compared directly against the standard Hillas cuts on the same nights. That comparison is the paper's real value: same 21 hours, same events, two methods, one table. The CNN at 6.26/6.48σ comes out slightly ahead of Hillas at 5.84σ, and the authors are honest that complete separation is not achieved and the model needs fine-tuning.\n\nWhat is good: the LiMa calculation is standard, the use of experimental hadrons for training is sensible, the preprocessing and Wobble handling are described in enough detail to be reproducible, and the paper acknowledges the anomalous night and the dependence of the conclusions on it. The citation pattern is fine; the self-citations are to their own prior work and to the simulation software, which is appropriate.\n\nSoft spots, in order of seriousness. First, the manuscript never says whether the 40,000 training hadrons and 9,600 validation hadrons are disjoint from the six test nights. If they overlap, the CNN can learn night-specific detector conditions or even memorize background events, and the reported 6σ is inflated. This is not an obscure point: the authors justify experimental hadrons precisely because they contain real telescope noise, so the network is being trained on the thing that could leak. The fix is a sentence specifying the temporal/geometric split; without it, the headline number is not yet trustworthy.\n\nSecond, the paper reports 5.5σ in the abstract, 6.26/6.48σ in Table I, and 6.5σ in the conclusion. These need to be reconciled. Third, the post hoc exclusion of 29 November cuts the CNN significance from 6.26/6.48 to 5.1/5.5 while raising Hillas to 6.3. That tells me the CNN result is fragile on one night, and the paper should lead with the full-data and the excluded-night numbers together rather than burying the exclusion in the text.\n\nThe train/test issue is load-bearing and fixable. I would send this to a serious referee: the experiment is real, the comparison is useful, and the missing documentation is exactly what a referee should force the authors to provide. I would not cite it in my own work right now, but I would put it in front of a reading group focused on ML in astro.","headline":"A useful incremental validation of CNN-based gamma/hadron separation at TAIGA, but the headline 6σ Crab signal rests on a train/test split the paper never documents.","tokens_in":8206,"tokens_out":2918,"would_cite":false,"duration_ms":25556,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["07.05.Mh","29.40.Ka","95.55.Ka","95.85.Pw","98.70.Rz"],"model":"deepseek-v4-flash","headline":"Convolutional neural networks can separate gamma-ray from hadron-induced air showers in TAIGA imaging Cherenkov telescope data, recovering the Crab Nebula at more than 5.5 sigma in 21 hours of observations.","keywords":["gamma-ray astronomy","IACT","image recognition","convolutional neural networks","Crab Nebula","gamma/hadron separation","extensive air showers","LiMa significance"],"falsifier":"Inspect the training-data provenance: if any of the 40,000 experimental hadron images were recorded on the same six nights used for the test, retrain the CNN on hadrons from other nights and recompute the LiMa significance; a drop to or below the Hillas value of $5.84\\sigma$ would show that the claimed CNN advantage came from seeing the test conditions.","tokens_in":7085,"feed_emoji":"🔭","tokens_out":8433,"duration_ms":67502,"temperature":0.7,"pith_summary":"This paper aims to establish that a convolutional neural network (CNN) can serve as an independent gamma/hadron separator for the TAIGA imaging atmospheric Cherenkov telescopes in a regime where cosmic-ray background outnumbers point-source gamma photons by about $10^4$ to 1. The authors train the CNN on simulated gamma showers and experimental hadron images, then test it on 21 hours of Crab Nebula observations. They report LiMa significances of $6.26\\sigma$ with a soft pre-selection and $6.48\\sigma$ with a strict pre-selection, versus $5.84\\sigma$ from the standard Hillas-parameter cuts on the same data. The paper's conclusion is that CNN classification quality is on the same level as the standard method, and that it is a viable independent selection approach, not a replacement.","feed_headline":"CNN recovers Crab Nebula at 6.5 sigma in 21 hours","feed_subtitle":"A convolutional network trained on simulated gammas and real hadrons rivals Hillas cuts on the same TAIGA dataset.","key_machinery":"The carrying object is the CNN gamma-score classifier: a block of convolutional feature-extraction layers followed by fully connected layers, with ReLU activations, a sigmoid output, and binary cross-entropy training, mapping each preprocessed telescope image to a probability in $[0,1]$ that the primary particle was a gamma. Around it sits a physics-driven pre-selection on the Hillas angle $\\alpha$: the angle between the shower image's major axis and the source direction is sharply peaked near zero for gamma events and flat for isotropic hadrons, so requiring $\\alpha<20^\\circ$ (soft) or $\\alpha<6^\\circ$ (strict) together with $\\mathrm{Size}>120$ photoelectrons already suppresses the hadron background by roughly a factor of 10 before the network runs. The network's role is to refine the surviving sample, and the paper uses the resulting gamma score, thresholded at the value where 50% of validation gammas are lost, as the final classifier.","core_discovery":"The central claim is that a CNN, after cleaning, Wobble correction, hexagonal-to-square image transformation, and logarithmic scaling, can be trained to distinguish gamma-ray from hadron-induced air showers with enough purity to extract a TeV source signal from modern IACT data. Trained on 38,400 Monte Carlo gamma images with a $-2.6$ spectral index and 40,000 experimental hadron images (expanded by $60^\\circ$ rotations to 470,400 samples), the network assigns each image a gamma score; thresholding at 0.9965 yields 99.93% classification accuracy and hadron suppression $B \\approx 3000$ on validation, and combined with a $\\mathrm{Size}>120$ phe cut and an angle cut $\\alpha<20^\\circ$ the effective suppression reaches $B\\approx 10000$ with a loss of about half the gamma events. On the 21-hour Crab data, the soft selection gives an excess of 91.2 events and $6.26\\sigma$, the strict selection ($\\alpha<6^\\circ$) gives 68.7 events and $6.48\\sigma$, and the standard Hillas cuts give 66 events and $5.84\\sigma$. The interpretation the authors defend is that deep learning performs comparably to the classical analysis and deserves continued development, rather than that it fully solves gamma/hadron separation.","pith_inferences":["The headline significance is mostly a statement about the pre-selection plus CNN, not about the CNN alone: the $\\alpha$ angle already suppresses hadrons by roughly 10, and the CNN's marginal gain over Hillas cuts is about 0.4 to 0.6 sigma.","If the training hadrons were drawn from the same six nights as the test, the measured $>5.5\\sigma$ would likely be optimistic; a clean temporal split is the first thing to check.","The single anomalous night (29 November) suggests the classifier is sensitive to nightly atmospheric or hardware conditions, so per-night normalization or domain adaptation may improve robustness.","An interesting extension not explored in the paper: train the CNN without any $\\alpha$ cut and compare against Hillas cuts with the same $\\mathrm{Size}$ threshold, isolating how much of the separation comes from image morphology rather than from the source-direction prior."],"forward_implications":["A CNN can be deployed alongside the Hillas-parameter method as an independent check on source detections, since the two selections share only a fraction of their accepted events.","For a fixed observation campaign, strict CNN selection delivers higher statistical significance per hour than the standard cuts when the background rate is the limiting factor, at the cost of retaining fewer gamma events.","Because the classifier was trained on Monte Carlo gammas with a single power-law index ($-2.6$), applying it to sources with different spectral shapes requires re-evaluating the score threshold.","The same image features used for classification can support the planned next step of energy-spectrum reconstruction from TAIGA-IACT data.","The validation numbers ($B\\approx 10000$ at 50% gamma loss) set a target for future IACT processing pipelines, but only if the train/test independence is confirmed."],"supporting_citations":[{"why":"Introduces the Hillas image moments that define the standard analysis parameters and the angle $\\alpha$ pre-selection used here.","marker":"[8]"},{"why":"Early IACT application of Hillas parameters to gamma-ray source detection, establishing the classical analysis method this paper compares against.","marker":"[9]"},{"why":"Previous TAIGA deep-learning studies that reported an achievable hadron suppression limit of about 1000 and flagged the Wobble source-position effect, motivating the pre-selection used here.","marker":"[10–12]"},{"why":"Provides the Monte Carlo gamma-ray simulation set with a $-2.6$ spectral index that forms the gamma class of the training data.","marker":"[15]"},{"why":"CORSIKA software used to simulate the extensive air showers from which gamma images were generated.","marker":"[16]"},{"why":"Simulation of the TAIGA-IACT camera response that turns simulated showers into telescope images for training.","marker":"[17]"},{"why":"Introduces the Wobble pointing mode, which the paper applies in image preprocessing and uses to estimate OFF-source background.","marker":"[18]"},{"why":"Provides the axial hexagonal-to-square transformation needed to feed hexagonal camera images into a CNN.","marker":"[19]"},{"why":"Li and Ma's significance formula, used to compute the signal significances reported in Table I.","marker":"[20]"},{"why":"Source of the standard Hillas-parameter cuts used as the baseline method for comparison.","marker":"[21]"}],"fun_headline_variants":["TAIGA CNN sees Crab at 6.5σ in 21 hours","CNN matches Hillas cuts for Crab in TAIGA","Neural nets rival standard gamma/hadron cuts","Deep learning finds Crab signal in TAIGA data","TAIGA neural net matches Hillas for Crab"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the experimental hadron images used for training were not drawn from the same events or nights as the 21-hour test data; the paper does not state this split, and if they overlap, the reported significance is inflated.","fun_headline_variants_meta":{"raw":{"variants":["TAIGA CNN sees Crab at 6.5σ in 21 hours","CNN matches Hillas cuts for Crab in TAIGA","Neural nets rival standard gamma/hadron cuts","Deep learning finds Crab signal in TAIGA data","TAIGA neural net matches Hillas for Crab"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000735,"raw_usage":{"total_tokens":3336,"prompt_tokens":1048,"completion_tokens":2288,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":2205}},"tokens_in":664,"tokens_out":2288,"duration_ms":16539,"temperature":1.0,"reasoning_tokens":2205,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T15:05:36.044305+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the training-data provenance: if any of the 40,000 experimental hadron images were recorded on the same six nights used for the test, retrain the CNN on hadrons from other nights and recompute the LiMa significance; a drop to or below the Hillas value of $5.84\\sigma$ would show that the claimed CNN advantage came from seeing the test conditions.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Hillas image moments that define the standard analysis parameters and the angle $\\alpha$ pre-selection used here."},{"cited_title":"Grinyuk, E","cited_arxiv_id":null,"evidence_quote":"Provides the Monte Carlo gamma-ray simulation set with a $-2.6$ spectral index that forms the gamma class of the training data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CORSIKA software used to simulate the extensive air showers from which gamma images were generated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Simulation of the TAIGA-IACT camera response that turns simulated showers into telescope images for training."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Wobble pointing mode, which the paper applies in image preprocessing and uses to estimate OFF-source background."},{"cited_title":"Hoogeboom, J","cited_arxiv_id":null,"evidence_quote":"Provides the axial hexagonal-to-square transformation needed to feed hexagonal camera images into a CNN."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the standard Hillas-parameter cuts used as the baseline method for comparison."}],"review_version":1}