{"id":"697dab87-afbf-41e7-b0d7-7451c0dee4be","arxiv_id":"2412.11208","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"DeepSets and graph attention networks improve simulated pion energy resolution and particle identification in a digital hadronic calorimeter, with little loss when pads are quadrupled in size.","lead":"The paper tests neural networks that treat particle showers as 3D point clouds to reconstruct hadron energy and type in a digital calorimeter. It reports better energy resolution than earlier algorithms and suggests larger readout pads may suffice, which could cut detector costs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed improvement over baseline algorithms in Fig. 2 is confounded by the unstated shower-start preselection in the first 10 layers, so the central result is not yet established.","rationale":"The reader's weakest assumption identified exactly this combined confound: external baselines from different detector technologies plus the uncontrolled pre-selection. My stress-test narrows the load-bearing part to the pre-selection because it is internal to the paper's own methodology and can be directly corrected by the authors. Selecting only showers that start in the first 10 layers is a physically motivated cut to reduce leakage, but it changes the population being measured; comparing that subset with published baseline resolutions that include all events conflates algorithm performance with event selection. The paper is clearly written and the ML engineering is plausible, but the central quantitative claim depends on a comparison that is not apples-to-apples. A controlled re-evaluation would settle the question; until then the conditional verdict is appropriate. I do not see a reason to change the reader's verdict: the methodological issue is addressable, so conditional (not reject) is right, and the paper's internal logic otherwise hangs together. My agreement rating is 'agree' because the reader's weakest assumption already names the same concern; the only difference is emphasis, with the pre-selection being the more decisive and testable element.","tokens_in":4355,"tokens_out":3657,"duration_ms":34206,"concrete_test":"Recompute Figure 2 under controlled conditions: take the same simulated RPWELL DHCAL events and run (a) the DeepSets model, (b) the hit-count/hit-density algorithm from Ref. [3], and (c) the CALICE [4] approach if available, on both the full sample and the shower-start-in-first-10-layers preselected sample, using identical energy bins, incident angle (0 deg), particle type (charged pions), and the same resolution metric. If the DeepSets advantage over Ref. [3] disappears or shrinks materially when both are evaluated on the full unselected sample (or when both use the same preselection), the claimed improvement over baselines is an artifact of the event selection and the central conclusion fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline result—DeepSets pion energy resolution improving on the RPWELL [3] and CALICE RPC [4] baselines—rests on comparing a simulated, pre-selected sample against external baselines whose event selection is not matched. In the Methods section the authors state: 'To minimize the longitudinal leakage of the shower, the events are pre-selected for the analysis with the identified shower start in the first 10 layers of the calorimeter.' This cut removes late-starting and high-leakage events, which are exactly the events that most degrade energy resolution. The magenta curve in Fig. 2 is therefore the resolution of an early-shower subset, while the black and green curves from Refs. [3,4] were not produced with this same cut (and in the case of [4] are for a different RPC-based DHCAL technology). Without knowing how much of the gap is due to the selection, the central claim that GNN reconstruction outperforms traditional algorithms is not supported. The correct test is to apply the same event selection to both the proposed model and the baseline algorithms, or to evaluate both on the full unselected sample, and to state the resolution definition and statistical uncertainties explicitly.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a simulation study of point-cloud deep learning for particle shower reconstruction in a proposed RPWELL-based Digital Hadronic Calorimeter (DHCAL). The authors simulate hadronic showers (pions, kaons, protons, neutrons) with GEANT4, pre-select events with shower start in the first 10 layers, and train DeepSets and Graph Attention Transformer (GAT) models for energy regression and particle identification. They report energy resolution curves versus energy, angle, and pad size, and claim that DeepSets outperforms traditional algorithms from Refs. [3] and [4], that GAT performs similarly at higher computational cost, and that increasing the pad size from 1x1 cm^2 to 2x2 cm^2 does not significantly degrade energy resolution.","tokens_in":4541,"tokens_out":6405,"duration_ms":58949,"significance":"If the central claims are supported, this work would be a useful step toward applying GNN-based algorithms to digital calorimeters, with potential implications for detector design and cost reduction. The paper has clear strengths: it uses a relatively large simulated sample, describes two neural architectures, and studies several dependencies (particle type, angle, pad size). However, the headline improvement over published baselines rests on an uncontrolled comparison, and the quantitative claims are not accompanied by uncertainties. The result is therefore interesting but not yet established.","major_comments":[{"comment":"The headline comparison in Fig. 2 is not controlled. The DeepSets curves are evaluated on events pre-selected to have the shower start in the first 10 layers, as stated in the Methods: “To minimize the longitudinal leakage of the shower, the events are pre-selected for the analysis with the identified shower start in the first 10 layers of the calorimeter.” The baseline curves taken from Refs. [3] and [4] are for different detector technologies (RPWELL prototype versus RPC-based CALICE DHCAL), and no statement is made about whether the same shower-start selection, energy-resolution definition, or event selection was applied to them. Because an early-shower-start cut preferentially removes late-developing and high-leakage events, it can improve the measured resolution independently of the reconstruction method. The claim that the magenta curve in Fig. 2 “outperforms traditional algorithms” is therefore not established. The authors should apply the baseline algorithms to the same pre-selected simulated sample, or evaluate their method on the full unselected sample, and should state the resolution definition (e.g., sigma/mean from a Gaussian fit versus RMS/mean) used for all curves. The shower-start identification algorithm should also be described explicitly.","section":"Methods and Results, Fig. 2"},{"comment":"No statistical or systematic uncertainties are reported for the energy-resolution curves or for the confusion-matrix entries in Table 1. Energy resolution is estimated from a finite simulated sample, and the differences discussed in the text, such as the claim in Fig. 4 that 2x2 cm pads “does not degrade the performance significantly,” are small and could be within sample fluctuations. Without error bars, confidence intervals, or at least the number of events per energy bin, these comparisons are not quantitatively meaningful.","section":"Results, Figs. 3-4 and Table 1"},{"comment":"The simulated detector response is not validated against the test-beam results in Ref. [3], although the paper states that “the expected performance of the DHCAL was evaluated based on past measurements with smaller RPWELL prototypes [3].” No comparison of simulated hit multiplicity, MIP detection efficiency, shower profiles, or other observables to the measurements is shown. Without such validation, the absolute resolution values and their comparison to the external baselines in Fig. 2 rest on an unverified simulation model. A validation plot or a quantitative statement of the model agreement is needed.","section":"Methods and Results, Fig. 2"},{"comment":"The claimed improvement over baselines is demonstrated only for the DeepSets architecture: Fig. 2 is explicitly for DeepSets, and no GAT energy-resolution curve is provided. The abstract states that the combination of GATs and DeepSets “results in an improvement over existing baseline techniques,” but the energy-resolution evidence in the paper is limited to DeepSets. The GAT results are presented only in terms of computational cost and particle identification; its energy-resolution performance is not quantified. The authors should either provide the GAT energy-resolution curve or limit the claim accordingly.","section":"Results and Abstract"}],"minor_comments":[{"comment":"The event counts after pre-selection are ambiguous: the text says 1.2M showers were simulated and then “for each particle type, the data set contains 600k events,” which could mean 600k per particle type or 600k in total. The number of events per energy bin should also be stated.","section":"Methods"},{"comment":"Hyper-parameters such as learning rate and batch size are given, but the number of layers, hidden dimensions, activation functions, and the procedure for choosing the GAT attention radius are not reported, which is insufficient for reproducibility.","section":"Methods"},{"comment":"Ref. [7] is listed as “Attention Is All You Need” with arXiv:1706.03762 (2023), but the graph attention architecture used in the paper should be cited to Veličković et al., “Graph Attention Networks” (2018). The current citation appears to be to the Transformer paper.","section":"References"},{"comment":"The sentence “The probabilities are obtained by applying the Softmax function 4.1 [10]” appears to refer to a nonexistent equation number; either label the equation or remove the number.","section":"Methods"},{"comment":"There are several typos and grammatical issues, including “GEANT41” instead of GEANT4, “the production of of neutral pions,” and “the shape and a shower development vary.” The paper would benefit from a careful proofreading pass.","section":"General"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Maryna and team have written a clear, short study applying DeepSets and GAT point-cloud methods to a simulated RPWELL DHCAL. The useful piece is the granularity study: they report that 2x2 cm pads give roughly the same pion energy resolution as 1x1 cm pads, which, if it holds under a matched analysis, is a practical result for reducing channel count. The angular-dependence plot is also a nice sanity check. The paper reads honestly and does not overclaim in the text; the ML setup (masked attention, batch size, MSE loss) is described well enough to reproduce.\n\nThe soft spot is exactly the one the stress-test note flags. The authors pre-select events with shower start in the first ten layers 'to minimize longitudinal leakage,' then compare the resulting DeepSets resolution to published baselines from Refs. [3] and [4]. Those baselines were not produced with that cut, and [4] is a different RPC technology. So the magenta vs black vs green gap in Fig. 2 is not a controlled comparison. The preselection removes the late-starting, high-leakage events that dominate the poor-resolution tail, so part of the apparent gain could simply be the selection. The correct fix is to run the same baseline algorithms on the same simulated sample, or at least to characterize how much resolution changes with the cut. I also note there are no statistical uncertainties on the resolution curves, and no mention of prior ML-for-calorimetry work, which makes the literature picture incomplete.\n\nNone of this is fatal to the internal logic. The DeepSets and GAT implementations are coherent, the PID confusion matrix is plausible, and the claim about pad size is independent of the baseline comparison. But the central headline claim—that GNN reconstruction outperforms traditional algorithms—is not yet supported. The paper is a legitimate proof-of-concept for a serious referee, not a definitive result. I would send it to review with the expectation of major revision: matched baselines, error bars, and a decision about what to claim. I would not cite it in its current form.","headline":"Sensible ML application and a useful granularity result, but the headline improvement over baselines is not established because the event preselection is uncontrolled and the comparators come from different detectors.","tokens_in":5122,"tokens_out":2157,"would_cite":false,"duration_ms":21178,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Point-cloud neural networks improve hadron energy resolution in a digital calorimeter.","keywords":["digital hadronic calorimeter","point cloud deep learning","graph neural networks","DeepSets","graph attention transformers","energy resolution","particle identification","calorimeter simulation"],"falsifier":"Measure the energy resolution of pions of known momenta in a test-beam run of an actual RPWELL-based DHCAL module with both 1x1 and 2x2 cm2 pads, and apply the same DeepSets and GAT models trained on that detector's data. If the measured pion resolution is no better than the traditional hit-counting algorithm, or if the 2x2 cm2 pad size degrades resolution by more than the statistical uncertainty, the paper's central claim is refuted.","tokens_in":4138,"feed_emoji":"⚛️","tokens_out":7454,"duration_ms":65211,"temperature":0.7,"pith_summary":"This paper argues that treating hadronic showers in a digital hadronic calorimeter as point clouds—unordered sets of fired pad positions—and processing them with deep set and graph-network architectures gives better energy resolution and particle identification than the traditional hit-counting baselines. The authors train DeepSets and Graph Attention Transformer networks on hundreds of thousands of simulated pion, kaon, proton, and neutron showers between 1 and 60 GeV, then compare the predicted pion energy resolution with published RPWELL- and RPC-based DHCAL baselines. They report that DeepSets outperforms both baselines, and that enlarging the readout pads from 1x1 to 2x2 cm2—four times fewer channels—does not significantly degrade the resolution. If this holds, ML-based reconstruction could improve hadronic energy measurements at future colliders and allow cheaper, coarser calorimeter readout.","feed_headline":"Point-cloud networks beat hit counting for hadron energy","feed_subtitle":"A GNN shower reconstruction improves pion energy resolution and tolerates four times larger readout pads.","key_machinery":"The point-cloud representation carries the argument: each fired calorimeter pad is a three-dimensional point, so a shower is an unordered set of points, and the network must be invariant to how the set is ordered. DeepSets passes each point through a multilayer perceptron and combines the resulting feature vectors with average pooling, capturing set-level information without explicit spatial relations. The Graph Attention Transformer adds edges between pads within a cone of radius 0.1 and uses masked attention to share information only among geometrically close neighbours, so the model sees local shower structure while retaining global context. This combination—permutation-invariant set processing plus locality-aware graph attention—is what lets the networks learn shower shape rather than counting hits.","core_discovery":"The central claim is that a neural network which learns from the spatial pattern of fired pads, rather than just their number, can reconstruct the energy of hadronic showers more accurately than the algorithms used for existing digital hadronic calorimeters. For charged pions, the DeepSets energy-resolution curve lies below both the RPWELL-DHCAL and the RPC-based DHCAL baseline curves. The same architecture holds its resolution when the pad size is increased by a factor of four, suggesting that the limiting information is not transverse granularity alone. A Graph Attention Transformer reaches similar energy performance at three times the computational cost and ten times the memory, but it is the variant that separates particle types, with the most reliable identification for protons and kaons, which the authors trace to baryon- and strangeness-number conservation suppressing event-to-event fluctuations in the electromagnetic fraction of the shower.","pith_inferences":["Beyond the paper: the same point-cloud pipeline could be applied to full jets and to two-shower separation, where the graph structure may help disentangle overlapping showers in ways that global pooling cannot.","Beyond the paper: the insensitivity to pad size suggests that transverse granularity is not the dominant resolution term, so the argument may transfer to other gaseous digital calorimeter technologies, not only RPWELL, as long as the longitudinal sampling is preserved.","Beyond the paper: because the baselines come from different detectors with different selections, an unambiguous gain estimate requires a same-detector comparison; this is a testable follow-up rather than a flaw in the simulation study itself.","Beyond the paper: the masked-attention radius of 0.1 may encode a physical scale; testing whether the optimal radius tracks the calorimeter's transverse shower width would give a transfer rule for applying graph attention to other absorbers."],"forward_implications":["If correct, the 2x2 cm2 pad result means a DHCAL with four times fewer readout channels can match 1x1 cm2 performance for single-hadron energy measurement.","If correct, point-cloud networks become a credible replacement for hit-counting energy reconstruction, improving hadronic and possibly jet energy resolution in particle-flow experiments.","If correct, GAT-based particle identification from calorimeter patterns alone could support background rejection and event classification without relying on a separate PID system.","The observed degradation for incidence angles beyond about 20 degrees implies practical deployments need either angular-specific training or larger, more diverse training samples."],"supporting_citations":[{"why":"Supplies the RPWELL-based DHCAL detector model, MIP efficiency and multiplicity assumptions, and the baseline resolution that the DeepSets result is claimed to outperform.","marker":"[3]"},{"why":"Supplies the RPC-based DHCAL test-beam resolution curve used as the second baseline.","marker":"[4]"},{"why":"Supplies the Monte Carlo simulation toolkit used to generate the 1.2M hadronic showers the networks are trained and tested on.","marker":"[5]"},{"why":"Supplies the DeepSets architecture, the permutation-invariant set-processing method used for energy regression.","marker":"[6]"},{"why":"Supplies the attention mechanism that the GAT model uses with masked local neighbourhoods.","marker":"[7]"}],"fun_headline_variants":["GNN point clouds sharpen hadron energy in DHCAL","DeepSets cuts error, GAT IDs particles in digital calorimeter","Point-cloud nets tolerate 4x larger pads for hadron energy","Shower point clouds beat hit counting for pion energy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the new networks outperform the traditional algorithms assumes that the published RPWELL-DHCAL and RPC-based DHCAL baseline curves are directly comparable to the simulated RPWELL-DHCAL resolution, even though the detectors, geometries, and event selections differ, including a preselection that requires the shower to begin within the first ten layers.","fun_headline_variants_meta":{"raw":{"variants":["GNN point clouds sharpen hadron energy in DHCAL","DeepSets cuts error, GAT IDs particles in digital calorimeter","Point-cloud nets tolerate 4x larger pads for hadron energy","Shower point clouds beat hit counting for pion energy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00016,"raw_usage":{"total_tokens":1195,"prompt_tokens":869,"completion_tokens":326,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":255}},"tokens_in":485,"tokens_out":326,"duration_ms":3628,"temperature":1.0,"reasoning_tokens":255,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:10:09.002663+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the energy resolution of pions of known momenta in a test-beam run of an actual RPWELL-based DHCAL module with both 1x1 and 2x2 cm2 pads, and apply the same DeepSets and GAT models trained on that detector's data. If the measured pion resolution is no better than the traditional hit-counting algorithm, or if the 2x2 cm2 pad size degrades resolution by more than the statistical uncertainty, the paper's central claim is refuted.","supporting_citations":[{"cited_title":"Shaked-Renous et al., Test-beam and simulation studies towards RPWELL-based DHCAL","cited_arxiv_id":null,"evidence_quote":"Supplies the RPWELL-based DHCAL detector model, MIP efficiency and multiplicity assumptions, and the baseline resolution that the DeepSets result is claimed to outperform."},{"cited_title":"NIM A 939, 89–105(2019)","cited_arxiv_id":null,"evidence_quote":"Supplies the RPC-based DHCAL test-beam resolution curve used as the second baseline."},{"cited_title":"Agostinelli et al., GEANT4: A Simulation toolkit","cited_arxiv_id":null,"evidence_quote":"Supplies the Monte Carlo simulation toolkit used to generate the 1.2M hadronic showers the networks are trained and tested on."}],"review_version":1}