{"id":"28bceaeb-4b48-431a-a1be-4c844cfaf1d2","arxiv_id":"2607.13107","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"DeepCormack integrates a CNN, an MLP, and a UNet into the modified Cormack reconstruction method, improving simulated TPMD reconstruction PSNR by about 8.5 dB and enabling lower-count (faster) measurements.","lead":"A team applies deep learning to reconstruct three-dimensional electron momentum densities from sparse positron-annihilation measurements, improving simulated reconstruction quality by about 8.5 dB PSNR over the standard method and remaining stable at 10x lower photon counts. The gains depend strongly on matching training data to the target material, so the authors recommend generating sample-specific synthetic training data from density-functional theory.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 8.5 dB gain is measured on synthetic in-distribution data; the noise model is hand-calibrated to one ZrZn2 experiment and no real-data validation is presented, so the practical speed-up claim is conditional.","rationale":"I read the paper as an honest, well-structured proposal: the operator formulation of MCM is useful, the synthetic data pipeline is described in detail, and the OOD degradation on ZrZn2 is reported rather than hidden. The strongest claim as stated is about reconstruction quality on test data, and within the synthetic DMD test set it is supported. The weakest point is exactly the distribution-match assumption: both the TPMD diversity and the noise model are derived from a single copper reference and a single hand-calibrated ZrZn2 experiment, so the reported gains have not been shown to transfer to real measurements. The paper's own recommendation—pair DeepCormack with per-sample DFT training—is a sensible mitigation, but it remains untested on experimental data. This does not undermine the synthetic-data claim, but it means the practical benefit (weeks of saved scanning time) is conditional on the distribution-match assumption and on the calibration of the noise simulator. The reader's CONDITIONAL verdict is appropriate; I do not see a reason to move to REJECT or UNVERDICTED, as the method is clearly specified and the in-distribution results are reproducible in principle. The concrete test I propose would settle whether the central practical claim survives on real data.","tokens_in":19279,"tokens_out":9582,"duration_ms":100137,"concrete_test":"Apply the recommended per-sample DFT workflow to a real experimental ACAR dataset with an independently known Fermi surface (e.g., ZrZn2 from Major et al., with dHvA/ARPES cross-checks): compute a DFT TPMD for that exact material, generate training data with the Section 2.3 simulator, train DCU/DCCMU, reconstruct from real projections at 10M and 200M counts (subsampling the real event list if needed), and compare recovered Fermi surface features (extremal areas, topology) against the known values. If the 10M-count learned reconstruction meets or exceeds the 200M-count MCM on those Fermi-surface metrics, the practical claim holds; if not, the improvement is confined to synthetic in-distribution data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the synthetic training distribution being representative of real ACAR measurements. Two linked components are least secure. (1) TPMD geometry: all training volumes derive from a single copper DFT reference via SVD/DMD (Section 2.3), so the DMD test set and the copper test set are in-distribution by construction; the OOD ZrZn2 test is still simulated with the same noise pipeline, and the best combined model DCCMU falls below MCM on it (32.51 vs 33.96 dB, Table 1). (2) Noise model: Section 3.1 calibrates the simulator to one ZrZn2 experiment by hand-adjusting two parameters (+60% convolution widths, 40% effective counts); no independent material or second dataset validates this calibration. Because the headline numbers (e.g., DCCMU 38.22 dB at 10M vs MCM 32.38 dB at 200M, Table 2) come from this same simulator, the model may be learning simulator-specific artifacts rather than generalizable physics. The paper explicitly recommends per-sample DFT training, but that workflow is not tested on experimental data, so the statement 'enabling significantly faster acquisition' is an extrapolation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DeepCormack, a family of learned reconstruction modules (1D CNN, MLP, 2D UNet, optionally FiLM-conditioned) inserted at different stages of the modified Cormack method (MCM) for 2D-ACAR Fermi surface tomography. To obtain training data without months of measurements, the authors generate synthetic 3D TPMD volumes from a single DFT copper reference: SVD sampling of Chebyshev coefficients produces central slices, and DMD/Koopman evolution produces slice-to-slice variation; measurement effects (detector blur, momentum sampling function, Poisson noise) are then simulated. The method is evaluated on synthetic DMD test data, on the copper reference, and on out-of-distribution ZrZn2 data, at count levels from 10M to 200M, using PSNR/SSIM and qualitative Fermi-surface images. The headline result is that the best combined model DCCMU reaches 38.22 dB PSNR at 10M counts on the DMD test set, exceeding MCM at 200M (32.38 dB), a gain of about 8.5 dB at 200M. However, on ZrZn2 the same configuration falls below MCM (32.51–32.82 dB vs. 33.96 dB at 200M), and the paper recommends per-sample DFT training rather than demonstrating transfer to experimental data.","tokens_in":19628,"tokens_out":6283,"duration_ms":66863,"significance":"If the reported gains transferred to real ACAR measurements, DeepCormack could substantially reduce acquisition times (weeks to months) or improve reconstruction quality for Fermi-surface studies. The operator-based reformulation of MCM and the SVD/DMD data-generation pipeline from a single DFT reference are useful and potentially transferable contributions. The paper is also commendably honest: it reports all configurations, including failures on out-of-distribution ZrZn2, and explicitly discusses the distribution-matching limitation. That said, the central practical claim is currently supported only by in-distribution synthetic experiments with a hand-calibrated noise model; the generalization evidence is negative for the flagship configuration on the only truly OOD sample.","major_comments":[{"comment":"The central speed-up claim is contradicted by the paper's own OOD experiment. In Table 4 (ZrZn2), DCCMU at 10M gives 32.32 dB PSNR, below MCM at 200M (33.96 dB), and at 200M DCCMU gives 32.82 dB vs. 33.96 dB for MCM. The abstract's statements that DeepCormack 'remains stable at reduced counts' and 'enables significantly faster acquisition times' hold for the in-distribution DMD set and for Cu, but not for a sample outside the training distribution. Since real ACAR materials are necessarily outside any training set unless a per-sample DFT is computed—a workflow recommended but not tested here—the practical claim needs to be rephrased as conditional and supported by real-data validation.","section":"§3.3.3, Tables 1 and 4"},{"comment":"The noise simulator is calibrated by hand to a single ZrZn2 experiment using +60% convolution widths and 40% effective counts. No second experiment or independent material is used to check whether these parameters generalize; every training sample and every test measurement (DMD, Cu, ZrZn2) is passed through this same simulator. The learned models may therefore be fitting simulator-specific artifacts (e.g., the particular MSF, blur, and effective count rescaling) rather than the physics of ACAR. A real-data reconstruction, or at minimum a held-out experimental noise calibration, is needed before the reported dB gains can be translated into expected experimental performance.","section":"§3.1, noise simulator calibration"},{"comment":"The training/test protocol makes the headline result in-distribution by construction. Section 2.3 generates all synthetic volumes by SVD/DMD from one DFT copper reference, and the DMD test set is drawn from the same generative pipeline. The copper test is the same reference density used to construct the SVD/DMD models. Hence the +8.5 dB PSNR gain over MCM at 200M (Table 2) demonstrates that the networks reproduce the generative model's manifold, not that they generalize to unseen physics. The ZrZn2 experiment is the only OOD test, and there the flagship DCCMU falls below MCM; this should be stated prominently in the abstract and conclusion.","section":"§2.3 and §3.2"},{"comment":"The quantitative evaluation is in p-space TPMD using PSNR/SSIM, not the downstream Fermi surface in k-space, as acknowledged in §4. Because the LCW step and gradient extraction are nonlinear, a 8.5 dB PSNR gain does not guarantee a proportionally improved Fermi surface. Qualitative Fermi-surface figures are shown, but no quantitative downstream metric is provided. Adding a metric defined on the recovered Fermi surface (e.g., error in high-symmetry plane contours or extracted Fermi-surface features) would materially strengthen the claim that DeepCormack improves Fermiology, not only image quality.","section":"§4, evaluation metrics"}],"minor_comments":[{"comment":"Tables 1 and 4 report different PSNR values for the same DCCMU/ZrZn2/200M condition (32.51 dB vs. 32.82 dB). Please clarify the evaluation protocol (e.g., slice range, test realization) or explain the discrepancy.","section":"Tables 1 and 4"},{"comment":"Figure captions 2–4 contain 'Workflow of the The...' typos; the Discussion contains 'wile' (for 'while') and 'desiged' (for 'designed'). A copyedit pass is needed.","section":"Figure captions and Discussion"},{"comment":"The text says 20 ideal projections in [0°,45°] are used to approximate the ground truth copper density, while the experimental pipeline uses 5 projections. Clarify that the 20 projections are used only for SVD/DMD coefficient extraction and not for the learned reconstruction experiments.","section":"§2.3"},{"comment":"Figure 6 uses 165M counts for the experimental comparison, while Tables 1–4 use 200M. State whether 165M is the experimental total and how the 40% effective-count factor is applied to the simulated data at each count level.","section":"§3.1 and Figure 6"},{"comment":"The captions for Tables 3 and 4 both refer to 'Figure 9' for the PSNR/SSIM plots; the ZrZn2 results appear to be shown in Figure 11. Check all cross-references.","section":"Tables 3 and 4 captions"},{"comment":"The configuration names such as 'CNNPT + MLPPT → UNet' are difficult to parse. Define the arrow notation and the acronyms (e.g., 'PT' = pre-trained) once at first use, and consider a small table mapping configuration names to pipeline stages.","section":"Notation, §3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its limitations, which is a genuine strength, but the abstract and conclusion overstate the practical impact relative to the evidence. I would be willing to accept after a revision that either adds a real-data validation or carefully conditions all practical claims to in-distribution/sample-specific training. The discrepancy between Tables 1 and 4 should also be resolved before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper applies supervised deep learning to the modified Cormack method for reconstructing Fermi surfaces from positron annihilation data. That has not been done before, and the synthetic data pipeline they introduce — SVD and DMD from a single DFT density — is a clever solution to the lack of experimental training data. The experiments are thorough: multiple network placements (projection space, radial density space, image space), count levels down to 10M, and an honest out-of-distribution test on ZrZn2 that shows the best combined model falling below the classical baseline while the standalone UNet still holds up. They also disclose the hand-calibration of the noise simulator to one experiment. That is good scientific practice.\n\nThe soft spot is the one the authors themselves name: the headline 8.5 dB gain is measured on synthetic data drawn from the same distribution used for training. The copper 'real' test uses the same DFT reference that generated the training data, so the only genuine OOD evaluation is ZrZn2, and there the gains are smaller and model-dependent. The noise model is calibrated to a single ZrZn2 measurement, so the claim that measurement time can drop from months to weeks rests on the assumption that this simulator captures real experimental degradation. That assumption is reasonable but unvalidated on a second material. No code or data are released, which makes independent verification harder.\n\nNone of this undermines the core contribution: the method works better than the classical baseline across the board, and the framework is flexible and reproducible in principle. The paper would benefit from a real-data test on at least one additional material, and from releasing the synthetic data generator. This is exactly the kind of work that should go through peer review and get the community to attempt replication.","headline":"A well-executed learned reconstruction for a niche but important materials characterization problem; the headline speed-up is real on synthetic data but remains conditional until tested on real ACAR experiments.","tokens_in":20082,"tokens_out":1979,"would_cite":false,"duration_ms":20417,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["71.18.+y","78.70.Bj"],"model":"deepseek-v4-flash","headline":"Augmenting the modified Cormack method with neural networks at three stages recovers Fermi-surface momentum densities from twenty times fewer annihilation counts, with an SVD/DMD generator supplying training data from one DFT reference.","keywords":["Fermi surface","2D-ACAR","two-photon momentum density","modified Cormack method","deep learning reconstruction","dynamic mode decomposition","synthetic data generation","positron annihilation"],"falsifier":"Run DeepCormack, trained on a copper-derived DMD dataset, on real 10M-count 2D-ACAR projections of a material whose Fermi surface is known from other methods and whose DFT reference was not used in training; if the reconstruction does not match or beat MCM on 200M-count data at the level claimed here, the method's practical value fails. A cheaper falsification: expand the paper's own out-of-distribution test to several materials with multiple Fermi-surface sheets and check whether the degradation seen for ZrZn2 is systematic.","tokens_in":19184,"feed_emoji":"🔬","tokens_out":10184,"duration_ms":92766,"temperature":0.7,"pith_summary":"This paper claims that a learned, model-based upgrade to the standard modified Cormack reconstruction can make 2D-ACAR Fermi-surface measurements far faster without sacrificing reconstruction quality. In the synthetic and simulated experiments shown, the full DeepCormack stack at 10M counts per projection reconstructs momentum densities better than plain MCM at 200M counts, implying scan times could drop from months to weeks. Because real experimental training sets are impractical, the paper also introduces a data generator: it decomposes a single DFT copper momentum density into Chebyshev coefficients, samples new central slices with truncated SVD, and evolves them into full 3D volumes using dynamic mode decomposition. The key caveat stated by the authors is distribution dependence—copper-trained models degrade sharply on out-of-distribution ZrZn2 data—so the recommended use is sample-specific training from a DFT calculation of the target material.","feed_headline":"Rebuild Fermi surfaces from 20x fewer scan counts","feed_subtitle":"Adding neural networks to the modified Cormack method recovers sharper Fermi surfaces and cuts months-long scans to weeks.","key_machinery":"The engine is the modified Cormack method (MCM) itself, written as an operator chain O Z W S^-1 P^-1 C: C interpolates projections onto polar coordinates, P is a symmetry-truncated Fourier transform keeping only harmonics multiple of the crystal's rotational order (4 for FCC), S expands into Chebyshev sine series, W weights the coefficients, Z builds Zernike polynomials, and O recombines radial density functions into the two-photon momentum density. DeepCormack inserts learned components around this operator—a 1D CNN before the MCM, an MLP after the radial-density step, and a UNet on the final Euclidean image, optionally conditioned on the log count level to signal noise. The other load-bear","core_discovery":"On its own terms, the paper establishes that the modified Cormack method can be wrapped in a supervised learning stack—a 1D CNN that denoises the measured projections, an MLP that refines the radial density functions, and a UNet that corrects the reconstructed TPMD image—and that the combined model remains accurate at count levels an order of magnitude below standard practice. The discovery that makes this trainable is the SVD/DMD synthetic-data generator: it turns one DFT-derived reference volume into arbitrarily many paired ground-truth / degraded-measurement training examples. The authors report 40.69 dB PSNR for the best configuration on synthetic test data at 200M counts versus 32.38 dB","pith_inferences":["If the speed-up transfers to real samples, 2D-ACAR could shift from a months-long, one-material measurement to a screening tool for a series of alloys or dopings: the per-sample overhead is reduced to one DFT reference and a shorter scan.","The SVD/DMD generator is a general recipe: any tomography problem with a known reference volume and a smooth slice-to-slice evolution could use the same two-step sampling-plus-evolution scheme to generate training data, not just Fermi surface reconstruction.","A directly testable extension is to train on multiple DFT references simultaneously and measure out-of-distribution robustness; the paper's own OOD results imply generalization should improve with diversity, but they do not test it.","The choice of PSNR/SSIM as evaluation metrics may understate or misstate the real goal (Fermi surface geometry); retraining or at least evaluation with a Fermi-surface-aware metric such as gradient of the LCW occupation could change the relative ranking of configurations."],"forward_implications":["A 10M-count per projection ACAR scan is sufficient for DeepCormack to reconstruct a momentum density comparable to or better than MCM at 200M counts, so acquisition can be cut by roughly an order of magnitude while keeping or improving quality.","Even at standard 200M counts, DeepCormack improves PSNR by about 8.5 dB on in-distribution synthetic data, so Fermi-surface features—especially at high momentum away from the slice centre—should be recovered more faithfully.","Because the synthetic training data comes from a DFT reference, the practical workflow becomes: run a DFT calculation of the target material (about a day), generate training data, then collect a shorter experimental measurement; the paper recommends this sample-specific pairing rather than a generic pretrained model.","With shorter scans, collecting more than 5 projections at proportionally lower counts becomes viable; the paper argues this richer angular sampling could further improve reconstruction quality, benefiting configurations that exploit multiple input channels.","The operator formulation of MCM links ACAR tomography to the standard inverse-problems literature, so improvements from that community (e.g., learned regularisation, uncertainty quantification) can be applied directly to Fermi surface reconstruction."],"fun_headline_variants":["Neural nets sharpen Fermi surface maps from sparse scans","Deep learning reconstructs Fermi surfaces with 20x fewer counts","AI tomography: Fermi surface maps in weeks, not months","One DFT volume trains deep model for faster Fermi surface scans","DeepCormack: neural nets improve Fermi surface reconstruction by 8 dB"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline gains rest on the assumption that the synthetic training distribution—built from one copper DFT reference plus a hand-calibrated noise model (+60% convolution width, 40% of true counts matched to one ZrZn2 experiment)—is representative enough of the target material's true momentum density and measurement noise that the learned reconstruction transfers to real data.","fun_headline_variants_meta":{"raw":{"variants":["Neural nets sharpen Fermi surface maps from sparse scans","Deep learning reconstructs Fermi surfaces with 20x fewer counts","AI tomography: Fermi surface maps in weeks, not months","One DFT volume trains deep model for faster Fermi surface scans","DeepCormack: neural nets improve Fermi surface reconstruction by 8 dB"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1379,"prompt_tokens":865,"completion_tokens":514,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":439}},"tokens_in":609,"tokens_out":514,"duration_ms":5877,"temperature":1.0,"reasoning_tokens":439,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:23:33.327031+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run DeepCormack, trained on a copper-derived DMD dataset, on real 10M-count 2D-ACAR projections of a material whose Fermi surface is known from other methods and whose DFT reference was not used in training; if the reconstruction does not match or beat MCM on 200M-count data at the level claimed here, the method's practical value fails. A cheaper falsification: expand the paper's own out-of-distribution test to several materials with multiple Fermi-surface sheets and check whether the degradation seen for ZrZn2 is systematic.","supporting_citations":[],"review_version":1}