{"id":"68ad778d-e511-444d-b501-6a2459439dab","arxiv_id":"2608.01601","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A sine-activated PINN with elastic-equilibrium and compatibility priors reconstructs 4D-STEM strain maps from 10% of probe positions on one experimental dataset (R2(εxx) = 0.80), with explicit caveats about single-mask evaluation and a mild information leak.","lead":"A physics-informed neural network was used to reconstruct 2D strain maps of a thermoelectric material from as few as 10% of the probe positions a dense 4D-STEM scan would use, recovering the chevron strain-band pattern with R2 = 0.80. The same framework attaches per-pixel uncertainty estimates, making it a candidate protocol for dose-reduced strain mapping of beam-sensitive specimens.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single random mask per sampling fraction leaves all headline metrics without error bars; repeated-mask runs are required before the quantitative claims can be trusted.","rationale":"The reader's weakest assumption—that one random mask per fraction cannot support the precise quantitative claims—is the most load-bearing concern in the paper. All headline numbers derive from a single realization, and the paper itself explicitly acknowledges that repeated masks are needed for error bars. This concern is concrete, addressable, and does not require reinterpreting the physics or the architecture; it is a matter of experimental protocol and statistical reporting. The information leak from full-map standardization and the capped GP baseline are real but secondary: the standardization leak is acknowledged and a fix is proposed, and the GP cap only affects one of the two baseline comparisons. The single-mask issue affects every quantitative claim, including the absolute reconstruction quality, the baseline margins, and the ablation percentages. Thus the paper should remain CONDITIONAL pending repeated-mask experiments, exactly as the reader concluded. My independent read does not change that verdict.","tokens_in":18504,"tokens_out":4919,"duration_ms":53915,"concrete_test":"For each sampling fraction, generate 10 independent random masks with different seeds, retrain the PINN (and at 10% sampling also CS and GP) with identical hyperparameters, and report mean ± one standard deviation for R2, MAE, and RMSE on εxx, plus the pairwise differences PINN−CS and PINN−GP. If the standard deviation of R2 at 10% exceeds ~0.02, or the 95% confidence interval for the MAE reduction vs GP includes zero, the headline quantitative claims are not robust to mask choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claims—R2=0.80 at 10% sampling, MAE reductions of ~26% vs CS and ~22% vs GP, ablation percentages, and the saturation plateau—are all point estimates from exactly one random sampling mask per fraction. Section 4.4 concedes this: 'The reported study uses one dataset and one random mask per sampling fraction; repeated mask realizations would attach error bars to Table 1.' This is a missing support rather than a hidden flaw, but it is load-bearing because the strain map contains extended chevron bands; different masks sample these structured features differently, especially at 1–10% sampling where a band may or may not be intersected. If mask-to-mask variance is comparable to the reported effect sizes (e.g., ±0.05 in R2 or ±10% in MAE reduction), the headline numbers become unstable and the comparative claims (including the 26%/22% margins) could reverse. The reader's weakest assumption correctly identifies this as the primary issue; the standardization information leak and GP data cap are secondary, but they compound the uncertainty in the same quantitative comparisons.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a physics-informed neural network (PINN) architecture for reconstructing 2D strain fields from sparsely sampled 4D-STEM data. The architecture embeds elastic equilibrium and Saint-Venant compatibility as soft constraints in the training loss, using a coordinate-based SIREN backbone with sine activations, frozen residual-scale normalization, an exponential physics-weight ramp, adaptive collocation, and two Bayesian variants (Monte Carlo dropout and mean-field variational inference) for per-pixel uncertainty. The method is demonstrated on a single experimental 180×400-pixel strain map of PbGeSnSe1.5Te1.5. Across sampling fractions of 1–75% (720–54,000 probe positions), the paper reports R²(ε_xx) rising from 0.44 at 1% to 0.80 at 10% and saturating near 0.86, MAE reductions of ~26% versus compressed sensing and ~22% versus Gaussian-process regression at 10% sampling, and an ablation showing that the physics prior improves accuracy only at extreme sparsity while always improving the imposed residual-based self-consistency metrics. The central claims are explicitly conditional on one dataset and one random mask per sampling fraction, and the paper discloses several secondary caveats.","tokens_in":18728,"tokens_out":6900,"duration_ms":76614,"significance":"If the quantitative claims are robust, the paper offers a practical dose-reduction protocol for 4D-STEM strain mapping, a topic of genuine interest for beam-sensitive materials. The architectural choices are well motivated by the structure of the second-order PDE constraints, and the paper is unusually honest: it reports the regime where the physics prior hurts accuracy, provides a reproducible notebook and scripts, and compares against external ground truth (the dense experimental map) rather than only fitting artifacts. The main weakness is that the headline numbers—R²=0.80 at 10%, the 26%/22% comparative margins, and the ablation percentages—are point estimates from one random mask per sampling fraction, with no error bars. The GP baseline is also capped at 1,500 training points, which disadvantages it in the very comparison used for the headline margin. These issues are fixable within the manuscript's scope, but they are load-bearing for the central quantitative claims.","major_comments":[{"comment":"All headline quantitative claims (R²=0.80 at 10%, MAE reductions of ~26% vs CS and ~22% vs GP, ablation percentages, saturation plateau) are point estimates from a single random sampling mask per fraction. The paper itself states in §4.4 that 'repeated mask realizations would attach error bars to Table 1.' Because the target field contains extended chevron strain bands, mask-to-mask variation at 1–10% sampling may be comparable to the reported effect sizes, and the comparative and ablation claims are not yet supported. Please run multiple mask realizations per sampling fraction and report means with confidence intervals for Table 1 and Tables S2–S4.","section":"§3.3, Table 1, §4.4"},{"comment":"The claim that the physics prior 'always improves physical self-consistency' is an evaluation of the training objective itself. The PINN explicitly minimizes Requil and Rcompat through the composite loss, while the data-only SIREN is never trained to minimize those residuals. Lower residuals for the PINN are therefore expected by construction and do not constitute independent evidence of physical consistency. Please either reframe these metrics as training-objective diagnostics or evaluate physical consistency with quantities not directly optimized, e.g., displacement compatibility obtained by integrating the reconstructed strain, or equilibrium residuals evaluated in physical units without per-component standardization.","section":"§3.6, Eq. (11), Table S4"},{"comment":"The GP baseline is capped at 1,500 training points, while the 10% sampling condition supplies 7,200 retained probes. Thus the headline 22% MAE reduction relative to GP compares against a GP that saw at most 1,500 of the available points. The text limits the caveat to 'GP comparisons at ≥25% sampling understate its performance,' but the same cap binds at 10%. Please acknowledge this for the 10% comparison, or use a scalable GP approximation that uses the full retained set, or match the number of training points across methods.","section":"§4.4, Table S2"},{"comment":"Per-component standardization statistics are estimated from the full dense reference map and used to normalize both data and physics residuals. This is a mild information leak relative to a strictly prospective sparse acquisition, as the paper concedes. Because the reported R², MAE, and ablation percentages are obtained under this leak, they do not directly establish the prospective protocol. Please quantify the sensitivity by recomputing standardization statistics from retained measurements only (or from training splits) and report the change in the headline metrics.","section":"§3.2–3.3, §4.4"}],"minor_comments":[{"comment":"The abstract states R² 'saturates near 0.86 by 25%,' but Table 1 reports R²(ε_xx)=0.84 at 25%, 0.85 at 50%, and 0.86 at 75%. Please adjust the wording to agree with the table.","section":"Abstract, §3.5, Table 1"},{"comment":"The text says 'Fig. 3(b) shows loss trajectories and data/physics breakdown at 50% sampling,' but the decomposition is in panel (c). Please correct the cross-reference.","section":"§3.3, Figure 3"},{"comment":"The early-stopping validation split is not described. Please specify how validation points are chosen (e.g., held-out fraction of the sampled positions) and whether the validation set is excluded from the data term.","section":"§2.5"},{"comment":"The uncertainty-error correlations (ρ≈0.32 for MFVI, ρ≈0.27 for MCD) are point estimates from a single mask. Given the single-mask design, report intervals or at least note the single-run nature in the caption.","section":"§3.7"},{"comment":"The hyperparameters λ_data, α_eq, α_co, λ_max_phys, τ, ω0, and the SIREN width/depth are given in prose. A consolidated hyperparameter table would improve reproducibility.","section":"§2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually transparent about its limitations, and the missing error bars and GP cap are exactly the load-bearing issues that require revision. The single worked example limits generality, but the reproducible code means the required repeated-mask and standardization-sensitivity experiments are straightforward to run within the scope of a revision. The editor may also wish to consider whether the methodological novelty relative to standard PINN practice is sufficient for the journal's readership; the five design requirements are well-motivated but individually known in the PINN literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuinely useful methods paper, and one of the more honest PINN-for-STEM evaluations I've seen. It ships code, data, and a notebook that regenerates every table, which is more than most papers in this area do. The central claim—that a SIREN-based PINN recovers the chevron strain morphology from 10% of probe positions—is supported by the reported numbers and by the supplementary tables. The ablation is also done honestly: the data-only SIREN beats the PINN once sampling fraction is above ~5%, and the authors say so plainly rather than spinning it. That alone earns credit.\n\nWhat's actually new is the design analysis connecting the two PDE priors to five concrete architectural requirements, especially the frozen residual normalization. The empirical finding that the physics prior helps accuracy only at extreme sparsity, while always improving physical self-consistency, is the kind of result that should shape future work. The uncertainty maps are a nice practical addition, though not deeply novel.\n\nThe soft spots are real but specific. First: every headline metric—R2=0.80 at 10%, the 26%/22% MAE reductions, the ablation percentages—comes from exactly one random mask per sampling fraction. The paper itself concedes this in Section 4.4. With extended chevron bands in the map, mask-to-mask variance could be material; if a band is or isn't intersected at 1% sampling, the numbers could shift. Repeated masks with error bars are needed before the quantitative claims can be taken at face value. Second: the standardization statistics for the physics residuals are estimated from the dense reference map, an acknowledged information leak relative to a prospective sparse acquisition. That's a fair caveat but it means the reported absolute errors are optimistic. Third: the GP baseline is capped at 1,500 training points, so its comparison at 25% and above understates GP; the comparison at 10% is the only one that really counts. None of these are fatal, but all three apply to the same set of headline numbers, so the uncertainty compounds.\n\nThe math and data handling look sound to me. The physics residuals are imposed, not derived from the data, so the circularity burden is low. The citation pattern is appropriate; the paper builds on standard PINN and STEM literature and credits it.\n\nFor whom is this paper? Anyone working on sparse reconstruction in electron microscopy, or on PINNs for inverse problems with second-order priors. It deserves a serious referee. I'd send it out, but I'd tell the referee to focus on the missing error bars and the information leak, and I'd expect a revision that either supplies repeated-mask experiments with a prospective standardization scheme or softens the quantitative claims accordingly.","headline":"Honest, reproducible methods paper with a real empirical finding, but every headline number rests on one random mask per sampling fraction, so the quantitative claims need error bars before they can be trusted.","tokens_in":19260,"tokens_out":1250,"would_cite":true,"duration_ms":17052,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a physics-informed neural network can reconstruct a continuous 2D strain tensor from 10% of the probe positions of a 4D-STEM scan, recovering the large-scale strain morphology while cutting nominal electron dose about","keywords":["4D-STEM","strain mapping","physics-informed neural networks","sparse reconstruction","automatic differentiation","Saint-Venant compatibility","Bayesian uncertainty","dose reduction"],"falsifier":"Repeat the 10% sampling benchmark with, say, 20 independently drawn random masks on the same strain map and compute the spread of R², MAE, and the differences between the physics-informed network, compressed sensing, and Gaussian-process regression. If the spread overlaps the claimed 26% and 22% margins, the comparative claims are not settled. A second check: re-estimate the per-channel standardization statistics from the retained training pixels only, retrain, and see whether R² at 10% drops materially below 0.80; a large drop would show the reported edge partly reflects leakage from the dens","tokens_in":18366,"feed_emoji":"🔬","tokens_out":7039,"duration_ms":72568,"temperature":0.7,"pith_summary":"The paper claims that a physics-informed neural network can reconstruct a full 2D strain tensor from 10% of the probe positions of a conventional 4D-STEM scan, cutting the nominal electron dose about tenfold while reproducing the large-scale strain-band morphology. On an experimental 72,000-pixel strain map of a domain-structured thermoelectric material, the method reaches R² of 0.80 on the principal strain component at 10% sampling and saturates near 0.86 by 25%. At 10% sampling it reduces mean absolute error by roughly 26% over compressed sensing and 22% over Gaussian-process regression on the same sampling masks. The key mechanism is embedding two continuum-mechanics constraints, elastic equilibrium and Saint-Venant compatibility, into the network's training loss through automatic differentiation. The paper also shows that the physics prior improves accuracy only in the extreme-sparsity regime, always improves physical self-consistency, and that Bayesian variants give per-pixel uncertainty maps correlated with reconstruction error.","feed_headline":"Physics-aware network rebuilds strain maps from a tenth of the probes","feed_subtitle":"Sparse 4D-STEM reconstruction hits R² 0.80 on the main strain component at 10% sampling, with per-pixel uncertainty maps.","key_machinery":"The load-bearing element is the composite physics-plus-data loss of Eq. (10). Elastic equilibrium residuals (from ∇·σ = 0 under isotropic plane stress) and the 2D Saint-Venant residual ∂²y εxx + ∂²x εyy − 2∂x∂y εxy = 0 are evaluated by second-order automatic differentiation of the network at fresh collocation points. The named architecture is SIREN—a sine-activated implicit neural representation whose non-vanishing second derivatives make it compatible with these second-order priors. The remaining components—frozen exponential-moving-average normalization (freezing after a 200-step warmup prevents the physics term from driving the field to the trivial uniform solution), an exponential physic","core_discovery":"Central claim: a physics-informed neural network reconstructs a continuous strain tensor from sparse 4D-STEM measurements, recovering chevron strain bands from 10% of probe positions with R² ≈ 0.80 on εxx at that fraction. The network is an implicit coordinate field (SIREN) outputting (εxx, εyy, εxy, θ); the loss is a masked data term plus residuals of elastic equilibrium and Saint-Venant compatibility, evaluated by automatic differentiation. Five architectural choices—implicit representation, sine activations, frozen residual normalization, an exponential physics-weight ramp, and residual-based adaptive collocation—are argued necessary for the priors to train. The paper also claims the phys","pith_inferences":["If the single-mask caveat is resolved favourably, the natural next test is a fully prospective workflow in which per-channel standardization statistics are estimated only from the retained measurements; the paper concedes the current retrospective benchmark leaks information from the dense reference map.","The finding that the homogeneous equilibrium prior becomes a bias when data are abundant suggests that an eigenstrain-augmented or per-pixel learnable constitutive closure could make the physics term beneficial across the full sampling range rather than only at extreme sparsity.","The same second-order PDE-embedding recipe may extend to other 4D-STEM modalities, such as differential phase contrast or ptychography, by replacing the elasticity residuals with the appropriate forward-model residuals.","A practical falsifiable roadmap: apply the trained protocol to a fresh beam-sensitive specimen and compare its sparse reconstructions against a full-dose map of the same region, checking whether the 10% uncertainty maps identify where a targeted second pass yields the largest error reduction."],"forward_implications":["Experimental 4D-STEM strain maps could be acquired at roughly tenfold lower nominal electron dose by retaining 10% of probe positions, widening the range of beam-sensitive specimens accessible to quantitative strain mapping.","With a constitutive model matched to the specimen, the framework is claimed to transfer to semiconductors, oxides, 2D materials, and battery and thermoelectric phases.","Physics-consistency residuals are improved at every sampling fraction, so downstream derivative quantities such as rotation and shear gradients are more reliable than those from data-only or classical sparse reconstructions.","The Bayesian uncertainty maps provide a target for adaptive second-pass acquisition, concentrating new probes on the regions where the first pass had highest epistemic uncertainty.","The ablation supports a practical protocol: enable physics in proportion to sparsity, and treat a growing gap between physics-informed and data-only accuracy as evidence of unmodeled physics, such as domain eigenstrain."],"supporting_citations":[{"why":"Defines 4D-STEM and frames the dense-versus-sparse acquisition problem that motivates the reconstruction task.","marker":"[Ophus, 2019]"},{"why":"Supplies the strain-measurement method from scanning nanobeam diffraction at subpicometer precision, grounding the data channels.","marker":"[Padgett et al., 2020]"},{"why":"The py4DSTEM pipeline referenced as the standard way the strain maps are produced and validated.","marker":"[Savitzky et al., 2021]"},{"why":"Provides the experimental PbGeSnSe1.5Te1.5 strain map used as the worked example and the domain-eigenstrain interpretation for the ablation result.","marker":"[Liu et al., 2024]"},{"why":"Origin of the physics-informed neural network loss formulation that the paper adapts to elasticity.","marker":"[Raissi et al., 2019]"},{"why":"SIREN sine-activated networks, the backbone whose non-vanishing second derivatives satisfy the network requirement for the second-order priors.","marker":"[Sitzmann et al., 2020]"},{"why":"Source of the residual-based adaptive refinement used to concentrate collocation points on high-residual regions.","marker":"[McClenny and Braga-Neto, 2023]"},{"why":"Analysis of PINN gradient pathologies supporting the exponential physics-weight ramp and the frozen residual normalization.","marker":"[Wang et al., 2022]"},{"why":"Monte Carlo dropout as the Bayesian approximation used for one of the per-pixel uncertainty maps.","marker":"[Gal and Ghahramani, 2016]"},{"why":"Mean-field variational inference used for the second per-pixel uncertainty map.","marker":"[Blundell et al., 2015]"}],"fun_headline_variants":["PINN maps strain from just 10% of 4D-STEM probes","Physics-informed network cuts 4D-STEM sampling to 10%","Strain maps from sparse 4D-STEM via physics-informed network","Neural net with physics priors rebuilds strain at 10% probes","Sparse 4D-STEM strain reconstruction with PINN hits R²=0.80"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The headline numbers rest on the assumption that a single random sampling mask per sampling fraction is representative; the paper itself notes that repeated mask draws would be needed to attach error bars to its Table 1.","fun_headline_variants_meta":{"raw":{"variants":["PINN maps strain from just 10% of 4D-STEM probes","Physics-informed network cuts 4D-STEM sampling to 10%","Strain maps from sparse 4D-STEM via physics-informed network","Neural net with physics priors rebuilds strain at 10% probes","Sparse 4D-STEM strain reconstruction with PINN hits R²=0.80"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1333,"prompt_tokens":855,"completion_tokens":478,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":374}},"tokens_in":599,"tokens_out":478,"duration_ms":5662,"temperature":1.0,"reasoning_tokens":374,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T00:17:46.965500+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the 10% sampling benchmark with, say, 20 independently drawn random masks on the same strain map and compute the spread of R², MAE, and the differences between the physics-informed network, compressed sensing, and Gaussian-process regression. If the spread overlaps the claimed 26% and 22% margins, the comparative claims are not settled. A second check: re-estimate the per-channel standardization statistics from the retained training pixels only, retrain, and see whether R² at 10% drops materially below 0.80; a large drop would show the reported edge partly reflects leakage from the dens","supporting_citations":[{"cited_title":", title =","cited_arxiv_id":null,"evidence_quote":"Defines 4D-STEM and frames the dense-versus-sparse acquisition problem that motivates the reconstruction task."},{"cited_title":"and Holtz, M","cited_arxiv_id":null,"evidence_quote":"Supplies the strain-measurement method from scanning nanobeam diffraction at subpicometer precision, grounding the data channels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The py4DSTEM pipeline referenced as the standard way the strain maps are produced and validated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the residual-based adaptive refinement used to concentrate collocation points on high-residual regions."},{"cited_title":"and Yu, X","cited_arxiv_id":null,"evidence_quote":"Analysis of PINN gradient pathologies supporting the exponential physics-weight ramp and the frozen residual normalization."}],"review_version":1}