{"id":"cc5a2566-6875-4520-83b2-965ae923627e","arxiv_id":"2502.07643","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Data volume reduction with a tailcut and three dilations leaves gamma-PhysNet reconstruction virtually unchanged, while tailcut and lstchain cleanings degrade performance at low energies.","lead":"This study tests whether removing pixels from telescope images before feeding them to a neural network harms gamma-ray event reconstruction. It finds that the data volume reduction mask used for CTAO's LST-1 preserves performance, while harsher cleanings degrade it, especially below 100 GeV.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The production-pipeline claim rests on an unverified 'similar to' DVR implementation; exact mask equivalence and real-data validation are missing.","rationale":"The manuscript is a compact empirical study with a clear protocol: three masks versus no mask under fixed MC conditions. I found the central comparison plausible and the conclusion useful if properly scoped. The single most load-bearing gap is not the MC-to-real extrapolation generally, but the specific assertion that the DVR implementation tested equals the production LST-1 pre-processing. Section 2 describes DVR as 'similar to' production with no quantitative mask-equivalence check, and Section 4 defers real-data validation. Since the title conclusion is about production, this equivalence is the hinge. If the production DVR has extra time filtering, different dilation count, or different thresholds, the reported 'no impact' cannot be transferred. I also note the absence of IRF error bars in Figure 3, which would be needed to claim equivalence rather than mere visual agreement, but this is secondary to the unverified production mapping. These issues are addressable and do not overturn the in-sample findings; thus the reader's CONDITIONAL verdict remains appropriate.","tokens_in":3312,"tokens_out":5430,"duration_ms":49524,"concrete_test":"Use the exact production LST-1 DVR code (or a bit-exact implementation) to preprocess the same MC DL1 events studied here, compare the binary masks per event (e.g., Jaccard index and fraction of discarded events) with the paper's DVR definition in §2, and recompute the four IRFs in Figure 3 on the production-masked images. If the masks differ materially or the recomputed curves separate from the reported DVR curves by more than the MC uncertainty, the conclusion should be restricted to the studied DVR variant, not the production pipeline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing condition for the title conclusion is that the DVR method in §2 is functionally identical to the LST-1 production DVR. The paper only states that DVR is 'similar to the pre-processing used in production' and gives no quantitative mask-overlap check, no comparison against the production algorithm on identical events, and no real-data validation (deferred to future work in §4). The strongest empirical claim — DVR preserves γ-PhysNet performance — is demonstrated only for this one MC configuration (Crab-like high NSB, 20° zenith) and for one network. If the production DVR differs in thresholds, number of dilations, ordering, or includes additional time/dead-pixel filtering, conclusion 'pre-processing used in production should have no impact' does not follow. A secondary aggravating factor is that Figure 3 has no statistical uncertainties, so 'very limited impact' and 'significant drop' are not quantified; however the primary gap is the unverified DVR-production equivalence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how three pixel-removal cleaning procedures (tailcut, lstchain, and a data-volume-reduction (DVR) mask) affect the performance of the γ-PhysNet deep-learning event reconstruction for the LST-1 telescope. Using Monte Carlo simulations tuned for the Crab Nebula with high night-sky background at 20° zenith, the authors compare energy resolution, energy bias, angular resolution, and gamma/proton classification AUC against a no-mask baseline. They report that DVR yields nearly identical performance to no mask, while tailcut and lstchain cleaning cause significant degradation, particularly below 100 GeV. The paper concludes that the DVR preprocessing used in production for LST-1 should not impact the analysis pipeline and notes that real-data validation is left to future work.","tokens_in":3480,"tokens_out":4766,"duration_ms":44481,"significance":"If the main result holds, it is practically important for CTAO: it suggests that the mandatory data-volume-reduction step can be applied to DL1 images without degrading the deep-learning reconstruction, whereas aggressive traditional cleaning masks (tailcut/lstchain) should be avoided when using γ-PhysNet. The study is useful because it provides an apples-to-apples comparison on standard IRF metrics with a no-mask reference, and the authors use public tools (gammalearn, glearn_irfs) that make the pipeline reproducible. The main caveats are that the quantitative claims are not accompanied by statistical uncertainties, the analysis does not account for the 20% of events fully removed by tailcut/lstchain cleaning, and the extrapolation to the production DVR rests on an unverified 'similar to' statement rather than an identical implementation.","major_comments":[{"comment":"The IRF plots and the qualitative statements 'very similar response' and 'significant drop' are shown without any statistical uncertainties or confidence intervals. Since the comparisons are made from Monte Carlo samples with finite statistics, especially in the low-energy bins, the reader cannot determine whether the observed differences are statistically meaningful. Please add bootstrap errors, Poisson uncertainties, or another quantitative measure of the spread, and use them when making claims about the magnitude of the impact.","section":"Section 3, Figure 3"},{"comment":"The paper states that tailcut and lstchain cleaning remove all pixels in 20% of the images and that these events are discarded from the dataset, but the subsequent IRFs are computed only on the surviving events. In a real analysis pipeline, discarding 20% of gamma-like events changes the effective collection area and the event selection, which is itself part of the 'impact on the analysis pipeline' claimed in the conclusion. The current comparison therefore does not capture the full effect of the cleaning methods. Please either include an efficiency or effective-area metric, or explicitly limit the conclusion to reconstruction performance on events that survive the cleaning.","section":"Section 2, Dataset; Section 3, Results"},{"comment":"The conclusion that 'the pre-processing used in production for LST-1 should have no impact on the analysis pipeline' relies on the assertion that the DVR method used here is 'similar to the pre-processing used in production.' The paper provides no evidence of functional equivalence: there is no quantitative mask-overlap comparison, no test of the production algorithm on the same events, and no real-data validation (which is deferred to future work). If the production DVR differs in threshold values, number of dilations, ordering of operations, or includes additional time/dead-pixel filtering, then the extrapolation from the tested DVR to the production pipeline is not justified. Please either verify equivalence or soften the conclusion to apply only to the specific DVR implementation tested.","section":"Section 2, DVR paragraph; Section 4, Conclusion"}],"minor_comments":[{"comment":"There is a typo: 'distringuish' should be 'distinguish.'","section":"Section 2, first paragraph"},{"comment":"The text says 'see Fig. 1' when referring to the effect of masks on images; this should be Fig. 2, since Figure 1 shows the network architecture and Figure 2 shows the mask applications.","section":"Section 2, paragraph after cleaning definitions"},{"comment":"The lower-noise MC experiment is only described with the sentence 'We observe the same behavior.' This is a qualitative claim with no figure or quantitative support; please provide the corresponding plot or at least a summary of the metric differences.","section":"Section 3, last paragraph"},{"comment":"The study covers only one source model (Crab-like with high NSB), one zenith angle, and one network architecture, yet the conclusion is phrased generally. Adding a caveat that the result has been demonstrated for this configuration would improve the accuracy of the statement.","section":"Section 4, Conclusion"}],"recommendation":"major_revision","confidential_remarks":"This is a short proceedings-style paper with a clear empirical design and a potentially useful practical message. The main concerns are quantitative: missing uncertainties, the unaccounted event-loss effect, and the overreach from the tested DVR to the production pipeline. These are fixable within the scope of the manuscript (adding error bars, discussing efficiency, and softening or validating the production claim). I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you're in the IACT-DL space. The paper gives the first direct comparison I know of between three pixel-cleaning masks (DVR, tailcut, lstchain) applied to the γ-PhysNet network on LST-1 MC, and the result is clean: DVR matches no-mask across energy resolution, bias, angular resolution, and AUC, while tailcut and lstchain cost several tens of percent in the low-energy bins. That's a useful, concrete data point for CTAO pipeline decisions, and the authors lay out the masks and thresholds clearly enough to reproduce the setup.\n\nThe main message survives my reading. The IRF comparison is internally consistent, and the relative plots make the effect easy to see. It is also honest that the lower-noise MC experiment is only described qualitatively, and the authors flag real-data validation as future work.\n\nThe soft spots are real but manageable. There are no error bars or statistical uncertainties on the IRFs, so 'significant drop' and 'very limited impact' are not quantified. More importantly, the DVR implementation is described as 'similar to' the production LST-1 preprocessing, but exact equivalence is not verified—no mask-overlap check, no side-by-side on identical events. The conclusion that 'the pre-processing used in production should have no impact' is therefore an extrapolation from a close-but-not-demonstrated-identical surrogate. That doesn't break the paper's core comparison, but it should be softened or backed up. The single-source, single-zenith, single-network setup also limits generalization, though it is an honest first step.\n\nOverall: a solid engineering validation with a clear, modest claim. It deserves full peer review; the fixes are straightforward (add uncertainties, clarify or demonstrate DVR equivalence, temper the production claim). I'd send it to referees.","headline":"Useful MC-based comparison showing DVR cleaning preserves γ-PhysNet performance, but the production-pipeline claim needs more evidence.","tokens_in":4006,"tokens_out":2473,"would_cite":true,"duration_ms":22954,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Removing 88 percent of the pixels through the production-like data-volume-reduction mask leaves the deep-learning reconstruction of gamma-ray events essentially unchanged, while two standard cleaning masks degrade it, especially below 100…","keywords":["imaging atmospheric Cherenkov telescopes","CTAO","LST-1","gamma-ray astronomy","deep learning","event reconstruction","pixel cleaning","data volume reduction"],"falsifier":"Run the identical γ-PhysNet training and evaluation on real LST-1 Crab Nebula data comparing DVR-masked and unmasked images, and compare energy resolution and angular resolution per energy bin; a difference larger than the statistical uncertainties, especially below 100 GeV, would refute the paper's conclusion.","tokens_in":3119,"feed_emoji":"🔭","tokens_out":9121,"duration_ms":78047,"temperature":0.7,"pith_summary":"The paper asks whether the pixel-removal step built into CTAO data handling, which is needed for storage and transfer, silently throws away information that a deep-learning event reconstructor relies on. It compares three cleaning masks against no mask using Monte-Carlo simulations of the first Large-Sized Telescope and the γ-PhysNet network, which reconstructs energy, direction, and particle type from telescope images. The answer is asymmetric: the production-like data-volume-reduction mask, which drops 88% of pixels through a tailcut plus three dilations, leaves every performance metric essentially unchanged, while tailcut and lstchain cleanings degrade performance sharply, most visibly below 100 GeV. The authors conclude that the DVR pre-processing used in production for LST-1 should have no impact on the analysis pipeline.","feed_headline":"Pixel stripping for storage barely hurts gamma-event reconstruction","feed_subtitle":"Conservative DVR drops 88% of pixels with no performance loss; standard cleanings lose low-energy events.","key_machinery":"The central object is γ-PhysNet, a convolutional neural network with an attention mechanism that performs multi-task event reconstruction directly from DL1 images. The comparison is organized by three pixel-selection masks: tailcut cleaning, a two-threshold procedure with thresholds {8, 4, 2}; lstchain cleaning, the default LST-1 mask consisting of the same tailcut followed by time-consistency filtering on the time map; and data-volume reduction, a tailcut with parameters {8, 4, 0} followed by three morphological dilations. The yardstick is a set of instrument response functions computed on energy bins: energy resolution, energy bias, angular resolution, and the area under the classification curve.","core_discovery":"On the paper's own terms, the finding is that the data-volume-reduction mask used in production for LST-1 removes 88% of pixels yet leaves γ-PhysNet's energy resolution, energy bias, angular resolution, and gamma-versus-proton classification essentially unchanged across all energy bins, whereas tailcut and lstchain cleaning remove 98% of pixels and produce a significant drop in all these metrics, concentrated below 100 GeV. The same qualitative behavior appears when the simulations are run with a lower night-sky background, so the result is not an artifact of the high-noise setting.","pith_inferences":["If a DVR mask can remove 88% of pixels with no measured impact, the information-bearing part of the image may be much smaller than the raw image; a saliency or ablation study on γ-PhysNet could map the minimum pixel set needed for reconstruction.","Because the three masks all share a tailcut core, a systematic scan of thresholds and dilation counts could locate the boundary where reconstruction performance starts to degrade, potentially enabling even more aggressive compression.","The study uses one source model (Crab-like), one zenith angle, one telescope, and one network architecture, so extending to other spectra, zenith angles, and telescope types is needed before the no-impact conclusion is generalized to the full CTAO array.","If real-data Crab observations confirm the Monte-Carlo result, the pipeline can separate concerns: DVR for storage and transfer, with aggressive astrophysical cleaning applied only outside the deep-learning branch."],"forward_implications":["The production DVR pipeline for LST-1 can be kept in place without expecting the γ-PhysNet reconstruction to lose performance.","Switching to lstchain or tailcut cleaning before feeding images to the network would degrade energy and angular resolution and classification, with the largest losses below 100 GeV.","Aggressive cleaning also discards about 20% of events entirely, so the performance drop is compounded by a loss of usable data.","The roughly 88% pixel reduction achieved by DVR shows that substantial data compression can coexist with deep-learning event reconstruction."],"supporting_citations":[{"why":"It supplies the γ-PhysNet architecture and the reconstruction method whose performance is being measured.","marker":"Jacquemont et al. 2021"},{"why":"It establishes the method's extra sensitivity to observation conditions, which motivates studying the impact of pixel removal.","marker":"Vuillaume et al. 2021"},{"why":"It supplies the tailcut procedure used as one of the cleaning baselines.","marker":"Linhoff et al. 2023"},{"why":"It defines the lstchain cleaning mask, the default LST-1 analysis mask that includes time-consistency filtering.","marker":"Abe et al. 2023a"},{"why":"It provides the time-map and time-consistency filtering step that distinguishes the lstchain mask from the plain tailcut.","marker":"López-Coto et al. 2022"}],"fun_headline_variants":["88% pixel removal? Gamma AI keeps its cool","Aggressive pixel pruning barely dents gamma reconstruction","Deep learning survives 88% pixel wipe for gamma events","Pixel purge for storage: gamma AI unaffected","Light cleaning fine, heavy cleaning hurts gamma AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on Monte-Carlo simulated images standing in for real LST-1 data, and on the study's DVR variant being equivalent to the production pre-processing; if either assumption fails, the no-impact result may not transfer to real observations.","fun_headline_variants_meta":{"raw":{"variants":["88% pixel removal? Gamma AI keeps its cool","Aggressive pixel pruning barely dents gamma reconstruction","Deep learning survives 88% pixel wipe for gamma events","Pixel purge for storage: gamma AI unaffected","Light cleaning fine, heavy cleaning hurts gamma AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000222,"raw_usage":{"total_tokens":1392,"prompt_tokens":822,"completion_tokens":570,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":438,"completion_tokens_details":{"reasoning_tokens":497}},"tokens_in":438,"tokens_out":570,"duration_ms":5958,"temperature":1.0,"reasoning_tokens":497,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T12:01:34.209692+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical γ-PhysNet training and evaluation on real LST-1 Crab Nebula data comparing DVR-masked and unmasked images, and compare energy resolution and angular resolution per energy bin; a difference larger than the statistical uncertainties, especially below 100 GeV, would refute the paper's conclusion.","supporting_citations":[{"cited_title":"2021, in International Conference on Pattern Recognition (Springer), 174","cited_arxiv_id":null,"evidence_quote":"It supplies the γ-PhysNet architecture and the reconstruction method whose performance is being measured."},{"cited_title":"Analysis of the Cherenkov Telescope Array first Large-Sized Telescope real data using convolutional neural networks","cited_arxiv_id":"2108.04130","evidence_quote":"It establishes the method's extra sensitivity to observation conditions, which motivates studying the impact of pixel removal."},{"cited_title":"2023, in Proceedings, 38th International Cosmic Ray Conference, vol","cited_arxiv_id":null,"evidence_quote":"It supplies the tailcut procedure used as one of the cleaning baselines."},{"cited_title":"(CTA, LST Project) 2022, in ASP Conf","cited_arxiv_id":null,"evidence_quote":"It provides the time-map and time-consistency filtering step that distinguishes the lstchain mask from the plain tailcut."}],"review_version":1}