{"id":"771ae2c1-d100-423e-b8d3-8954fd3a96c7","arxiv_id":"2601.05289","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A vision-transformer flow-matching model generates calorimeter showers across regular and irregular detector geometries at millisecond speeds, and pretraining plus fine-tuning cuts training cost by about half.","lead":"This paper trains a vision-transformer network to imitate how particle-detector calorimeters record electron, photon, and pion showers, testing it on six detector geometries with generation times of tens of milliseconds on a GPU, and shows that pretraining on one detector set and fine-tuning on another roughly halves training cost. If the claims hold, it could make high-energy-physics simulation dramatically cheaper and easier to adapt to future detectors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fine-tuning benefits are only shown for near-identical detectors; the claimed latent-space proximity is untested, so the universal transfer-learning savings may not generalize.","rationale":"The central claim has two components: fidelity of generated showers and transfer-learning savings. The fidelity overstatement is real (ResNet AUCs of 0.68–0.80 contradict 'indistinguishable'), but it is an overclaim in wording and can be corrected without changing the method's substance. The transfer-learning claim is more load-bearing because the 'universal' and cost-saving aspects of the paper depend on it. The paper's own experiments only probe near-identical detectors (same geometry or same detector family), and the one distinct geometry (CaloHadronic) is tested only at full dataset size, where the gain is small and no data-efficiency evidence is offered. The proposed test (data-efficiency scan on a truly different geometry) directly targets whether the claimed generalizable savings exist. The reader's conditional verdict already accounts for this uncertainty; our concern does not shift the verdict but sharpens the condition under which the paper would be accepted.","tokens_in":27834,"tokens_out":9244,"duration_ms":93407,"concrete_test":"Repeat the data-efficiency scan of Fig. 9 (left) for the CaloHadronic dataset: fine-tune the LEMURS-pretrained backbone on training subsets of 10k/25k/50k/80k showers and compare the ResNet/high-level AUC against from-scratch training on the identical subsets. If fine-tuned AUC is not significantly better at small sample sizes, the data-efficiency claim does not generalize across detector types. Independently, measure the Wasserstein-2 distance between the distributions of the pretrained network's conditional embedding vector on source and target datasets; if it is large relative to within-dataset variance, the 'small shift' assumption is falsified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 states 'We posit that general features, e.g the sparsity, of calorimeter showers are encoded in the large ViT backbone during pretraining' and attributes fine-tuning gains to 'a smaller shift in how the data are mapped to the Gaussian latent space.' This premise is not tested. The transfer-learning evidence (Section 6) is confined to near-identical detectors: LEMURS→ds2 changes particle type and conditions in the same Par04 geometry; ds2→ds3 is a superresolution of the same detector. The only truly different target, CaloHadronic, is fine-tuned on the full dataset with a small AUC improvement (0.614→0.600, Fig. 9) and no data-efficiency scan is provided. If a target detector's preprocessed showers are not nearby in latent space, the claimed universal training-cost reductions and data-efficiency gains do not follow; the headline 'universal vision transformer' is then unsupported beyond the specific geometries tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a vision-transformer (ViT) based fast calorimeter simulation architecture built upon CaloDREAM, extended to irregular geometries via flexible patching and a 3D positional embedding. The authors benchmark this 'CaloDREAM++' model on CaloChallenge datasets ds1/ds2/ds3, the multi-detector LEMURS dataset, and a down-sampled CaloHadronic dataset, reporting generation times of O(10–100) ms per shower on GPU. They further propose a transfer-learning strategy: pretraining on LEMURS or ds2 and fine-tuning on target detectors, claiming reduced training iterations, improved data efficiency, and in some cases better fidelity. The paper includes code and data releases.","tokens_in":28202,"tokens_out":6713,"duration_ms":62910,"significance":"If the claims are correct, the paper is a strong step toward a single ViT backbone usable for multiple detector geometries and particle types, with practical training-cost savings. The public code/data release, the use of external Geant4 references, and the broad evaluation across regular, irregular, and full-detector geometries are notable strengths. The transfer-learning results, particularly the ds2 data-efficiency scan, are valuable. However, the abstract and conclusions overstate the indistinguishability of generated showers, and the FPD numbers in Table 3 are inconsistent with the metric definition and Table 1, undermining some quantitative claims. The universality of the transfer-learning benefit is not yet demonstrated for genuinely different geometries.","major_comments":[{"comment":"The abstract and conclusions state that generated showers are 'statistically indistinguishable from Geant4 in multiple evaluation metrics.' This is contradicted by the paper's own ResNet classifier AUCs: Table 1 shows ds2 AUC=0.683(9) and ds3=0.799(9); Table 3 shows FCCeeALLEGRO AUC=0.688(16); Fig. 5 shows ds1-pi low-level AUC=0.633(3). These are far from the 0.5 expected for indistinguishable samples. The text in Section 4 even says 'only the more advanced ResNet extracts mismodeled generated features.' The claim should be qualified to high-level features or explicitly acknowledge low-level discrepancies.","section":"Abstract / Section 7"},{"comment":"The FPD column in Table 3 reports values near 1.0033 for both Geant4 and ViT-CFM across all detectors. Table 1 reports FPD multiplied by 10^3, with Geant4 ds2 at 10.7(8), i.e. FPD ~0.0107. A distance metric between a sample and itself should be near zero, not ~1.003. These numbers are inconsistent and make the statement that 'the FPD metric scores all networks as Geant4-like' uninterpretable. Please correct the definition/units or explain the discrepancy before the FPD results can be used as evidence.","section":"Table 3"},{"comment":"The transfer-learning evidence is not sufficient for the claimed universality. Fine-tuning is demonstrated for LEMURS→ds2, where the target is the same Par04 geometry (the one-hot encoding is set to Par04SiW), and for ds2→ds3, a superresolution of the same geometry. The only geometrically different target, CaloHadronic, shows a small high-level AUC improvement (0.614(3)→0.600(4)) and no data-efficiency scan is provided. Section 2.2 posits that general features are encoded in the backbone and that fine-tuning needs only 'a smaller shift' in latent space, but this premise is not tested. The abstract's 'higher data efficiency' claim is therefore only demonstrated for near-identical detectors. Please either provide a data-efficiency scan on a genuinely different geometry or temper the universality claim.","section":"Section 6 / Figs. 7-9"},{"comment":"For ds3, the fine-tuned ResNet AUC (0.789(9)) is not significantly different from the from-scratch value (0.799(9)) — the difference is within uncertainties. The paper acknowledges 'The high-granular ds3 shows a smaller gain,' but the abstract's phrase 'or altogether improves the fidelity of generated showers' is not supported by this result. The only significant fidelity improvement appears in ds2, so the wording should be softened accordingly.","section":"Figure 9 (right)"}],"minor_comments":[{"comment":"The claim that 'a third detector with the Par04 geometry can be encoded as a weighted combination of the two materials already seen during training' is speculative and not tested. Please mark it as a hypothesis or remove it unless an experiment supports it.","section":"Section 6"},{"comment":"The sentence 'The CaloDREAM++ samples for ds2 are statistically indistinguishable from the Geant4 reference' is directly contradicted by the ResNet AUC 0.683(9) in Table 1. Rephrase to refer to high-level features or to the selected metrics.","section":"Section 4"},{"comment":"The caption says 'from a total of 200k showers.' Clarify whether this is per detector or across all detectors, and whether the FPD is computed on the same 200k sample as the classifiers.","section":"Table 3 caption"},{"comment":"The target velocity for the linear trajectory x(t)=(1-t)ϵ + t x0 is x0−ϵ; the notation (x−ϵ) in Eq. (5) is clear only after the expectation over x∼pdata is understood. Consider writing (x0−ϵ) explicitly for readability.","section":"Eq. (5)"}],"recommendation":"major_revision","confidential_remarks":"The FPD inconsistency in Table 3 is serious enough that it may indicate a systematic issue in the evaluation pipeline; please ask the authors to clarify whether the reported values are an unusual definition or a typo. The abstract overclaim should be corrected regardless, as the ResNet AUCs themselves contradict 'statistically indistinguishable.' The transfer-learning section would also benefit from a more carefully qualified scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First things first: this is a real, reproducible contribution. The patching scheme for irregular geometries plus the 3D positional embedding lets a ViT handle CaloHadronic's EM+HCAL layout, and they ship code and data. The baseline tables on ds1/ds2/ds3 and LEMURS are carefully done, and the paper is unusually honest in places—it admits the FCCeeALLEGRO ResNet AUC is 0.688 and the ds1-pi low-level AUC is 0.633. That is exactly the right way to report.\n\nThe problem is the abstract. 'Statistically indistinguishable from Geant4' is not supported by their own numbers: several ResNet AUCs are clearly above 0.5 (ds2 0.683, ds3 0.799, FCCeeALLEGRO 0.688, ds1-pi 0.633). The paper even acknowledges the ResNet finds differences. So the abstract overclaims, and the conclusion repeats it.\n\nThere are two other soft spots. First, the FPD values in Table 3 are around 1.003 for both Geant4 and the model; that is not a distance metric as defined in the text. Either the definition has changed or the table is mislabeled, and a referee should ask for the formula. Second, the transfer-learning evidence is narrower than the 'universal vision transformer' framing suggests. The convincing gains are LEMURS->ds2 and ds2->ds3; both are the same Par04 geometry family. The only genuinely different detector, CaloHadronic, shows a small high-level AUC improvement (0.614 to 0.600) with no data-efficiency scan. So the claim that pretraining 'leads to reduced training costs and higher data efficiency' across arbitrary geometries is not yet demonstrated. The stress-test note got this right.\n\nThe comparison in Section 5 says the ds1 results 'improve over all the other submissions to the CaloChallenge,' but I could not find a table with those other submissions. Either add it or drop the sentence.\n\nGood things: the evaluation is external (Geant4), the metric discussion about ResNet vs FPD is useful, and the release of code and data is real evidence. Self-citations to CaloDREAM are legitimate starting points, not proof.\n\nVerdict: worth sending to peer review. The method is credible and the flaws are fixable in revision—mostly by aligning claims with numbers and tightening the transfer-learning conclusions. I would cite it for the irregular-geometry patching, but I would not repeat the 'indistinguishable' phrasing.","headline":"Solid extension of CaloDREAM to irregular geometries with honest evaluations, but the abstract overstates fidelity and the transfer-learning evidence is thinner than the 'universal' claim.","tokens_in":28614,"tokens_out":2228,"would_cite":true,"duration_ms":22167,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Vision transformers can emulate Geant4 for both electromagnetic and hadronic calorimeter showers across arbitrary geometries, in tens of milliseconds on one GPU, and transfer learning cuts training time in half.","keywords":["fast calorimeter simulation","vision transformer","conditional flow matching","generative networks","transfer learning","Geant4 emulation","irregular geometries","detector simulation"],"falsifier":"Fine-tune the LEMURS-pretrained backbone on a detector that has a substantially different material composition and voxel layout (for example, a liquid-argon calorimeter or an ECal with much finer lateral granularity) and compare the ResNet AUC after the same number of iterations against a from-scratch run. If the fine-tuned network does not reach from-scratch performance in approximately half the iterations, the claimed transfer-learning advantage does not hold.","tokens_in":27782,"feed_emoji":"⚛️","tokens_out":4335,"duration_ms":39277,"temperature":0.7,"pith_summary":"Fast calorimeter simulation is a bottleneck at the LHC and future colliders. This paper shows that a vision transformer trained with conditional flow matching can generate both electromagnetic and hadronic showers that are statistically indistinguishable from Geant4 across multiple detectors and geometries, in tens of milliseconds on one GPU. It also shows that pretraining the transformer on a large multi-detector dataset and fine-tuning on a target geometry roughly halves training time and improves data efficiency. The authors present the first patch-based simulation of a full detector with both electromagnetic and hadronic calorimeters.","feed_headline":"One vision transformer emulates Geant4 showers in 100 ms","feed_subtitle":"The same network spans regular and irregular detectors, electrons and hadrons, and learns new geometries with half the training.","key_machinery":"The central object is a patch-based vision transformer (ViT) shape network trained with conditional flow matching: a universal backbone of transformer blocks acts on patch embeddings of the voxelized shower, while detector-specific embedding and head layers are reinitialized or interpolated during fine-tuning. This split—universal backbone plus light detector-specific layers—is what lets one network handle irregular geometries and enables transfer learning, because the backbone is expected to encode general shower features such as sparsity and dynamic range.","core_discovery":"The paper's central claim is that a single vision-transformer architecture, extended to arbitrary voxelized geometries via patching, can emulate the Geant4 detector response across electromagnetic and hadronic showers with statistically negligible deviations. On the LEMURS multi-detector dataset, five detectors with different granularities are covered by one shape network, and the same backbone fine-tunes to CaloChallenge datasets and a high-granular ILC calorimeter. Generation stays at 10–100 ms per shower on a single NVIDIA A100 GPU. Fine-tuning from the LEMURS pretraining reaches the accuracy of a from-scratch network in roughly half the iterations and improves generalization when trainin","pith_inferences":["If the latent-space proximity assumption generalizes, a publicly released pretrained calorimeter backbone would let experiments fine-tune on their own detector geometry with a fraction of the current computational and data budget.","A natural stress test is to fine-tune on a detector with very different materials or segmentation (e.g., liquid argon or a much finer ECal) and measure whether the training-time savings persist; the paper only tested similar cylindrical and cartesian pad geometries.","The metric discrepancy between FPD and ResNet AUC suggests current benchmarks may overstate fidelity; a classifier-based 'distance to Geant4' on the same low-level voxels is a stricter target.","The same universal-backbone recipe could plausibly extend to other sparse, high-dynamic-range detector signals such as tracker hits or Cherenkov images, since the patching only needs a voxel-to-patch map."],"forward_implications":["Geant4-quality shower generation at 10–100 ms per shower on a single GPU, versus hours of CPU simulation per event.","The same trained backbone transfers to new detectors and particle types, halving the iterations needed to match from-scratch quality.","Fine-tuned networks generalize better than from-scratch networks at fixed training-data size, reducing Geant4 data production needs.","Irregular and high-granular geometries, including a full detector with electromagnetic and hadronic sections, are handled without mapping to a regular grid.","Metric choices matter: ResNet AUC and FPD can disagree, so a faithful evaluation should use the most discriminative classifier."],"fun_headline_variants":["One vision transformer covers all calorimeter geometries","Universal ViT cuts training cost for new detectors","ViT matches Geant4 showers on regular and irregular grids","Pretrained vision transformer speeds up shower simulation","Single model emulates Geant4 in 10-100 ms per shower"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The transfer-learning savings rely on the assumption that showers from different detectors and particle types are already close in the Gaussian latent space, so a pretrained backbone only needs to learn a small shift; the paper states this as a posit without an ablation.","fun_headline_variants_meta":{"raw":{"variants":["One vision transformer covers all calorimeter geometries","Universal ViT cuts training cost for new detectors","ViT matches Geant4 showers on regular and irregular grids","Pretrained vision transformer speeds up shower simulation","Single model emulates Geant4 in 10-100 ms per shower"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000251,"raw_usage":{"total_tokens":1348,"prompt_tokens":649,"completion_tokens":699,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":393,"completion_tokens_details":{"reasoning_tokens":621}},"tokens_in":393,"tokens_out":699,"duration_ms":7221,"temperature":1.0,"reasoning_tokens":621,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T12:03:38.674787+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fine-tune the LEMURS-pretrained backbone on a detector that has a substantially different material composition and voxel layout (for example, a liquid-argon calorimeter or an ECal with much finer lateral granularity) and compare the ResNet AUC after the same number of iterations against a from-scratch run. If the fine-tuned network does not reach from-scratch performance in approximately half the iterations, the claimed transfer-learning advantage does not hold.","supporting_citations":[],"review_version":1}