{"id":"2ca38da7-fc26-42fc-8309-b43a83068b02","arxiv_id":"1908.00659","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Unregularized linear filters fit to 4D-STEM diffraction data can reproduce arbitrary real-space patterns, while regularized, validated filters yield reconstructions comparable to HAADF imaging, though selection of the regularization strength is not formalized.","lead":"A new study shows that 4D scanning transmission electron microscopy data can be filtered to produce almost any desired image, including artificial or overly sharp structures, if the filter is fit carelessly. The authors argue that regularized regression with validation recovers trustworthy structure images comparable to standard HAADF imaging, but the supporting evidence is visual and lacks rigorous model selection.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The overfitting warning is solid, but the positive claim of HAADF-comparable robust filters is not quantitatively validated: λ is chosen by eye, no held-out metric is reported, and the training target is derived from the same HAADF image used as benchmark.","rationale":"The reader's weakest assumption (linear model, HAADF-derived targets) is related but not identical. The linearity concern is real for thick or dynamically scattering samples, but the paper's cautionary point does not depend on linearity being true: in the p >> n regime, a linear map can overfit any target even when the true scattering is nonlinear. The more decisive gap is that the paper never quantifies the validation it advertises. A λ path shown as images, a hand-picked r, and a visual line profile do not constitute 'careful model selection and validation'. This matters because the abstract and discussion explicitly promise robust filters comparable to HAADF; without a quantitative, pre-registered metric and an independent ground truth, that promise is not testable from the paper as written. However, this is a validation gap rather than a demonstrated error: the overfitting warning is well supported and the regularized reconstructions do appear plausible by eye. The reader's CONDITIONAL verdict already reflects this, so no verdict change is needed.","tokens_in":8047,"tokens_out":5530,"duration_ms":59396,"concrete_test":"Perform nested cross-validation on the existing 4D-STEM data: fix r=5e-5, split the 4-fold-dopant probe positions into three folds, fit Eq. (3) on two folds for a grid of λ values, and choose λ by maximizing a pre-specified quantitative metric (e.g., Pearson correlation or normalized MSE) between the reconstruction and the simultaneously acquired HAADF image on the held-out fold. Then evaluate the selected filter on the independent 3-fold-dopant dataset, reporting the same metric with bootstrap confidence intervals and comparing against a trivial periodic baseline. If the selected λ does not significantly outperform the baseline, the 'comparable to HAADF' claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is twofold: (1) high-dimensional CBED data allow filters that reproduce almost arbitrary training patterns, and (2) regularization plus validation can nevertheless yield robust filters comparable to HAADF imaging. Part (1) is convincingly demonstrated: with p=65,536 CBED pixels and n=1,088 training samples (Fig. 1b area), ordinary least squares can interpolate arbitrary targets, and Figs. 2, 8, and 9 show training/test inconsistencies. Part (2) is not established. Model selection is invoked but never actually performed: Eq. (3) has hyperparameters (λ, r); r is fixed at 5e-5, and λ is scanned along a decreasing path with the 'best' image chosen by visual inspection (Figs. 4-6). No quantitative metric (e.g., MSE, correlation, Fourier ring correlation) is computed on held-out data, no model selection criterion selects λ, and no error bars are given. The line scan in Fig. 6b is normalized and qualitative. Moreover, the training target (Fig. 1b) is constructed from atom sites estimated from the same HAADF image later used as the benchmark, so the comparison is partly circular: training encodes HAADF-derived lattice information and then recovers HAADF-like contrast. Thus the load-bearing assumption underlying the positive claim—that careful, quantitative model selection and validation recover structure robustly—is unsupported as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the risk of overfitting in 4D-STEM structure retrieval by treating real-space image formation as a linear regression on CBED intensities. It demonstrates that with n = 1,088 training samples and p = 65,536 CBED pixels, filters can be constructed via least squares to reproduce almost arbitrary training targets while failing on held-out data. The authors then apply elastic net regularization (Eq. 3) and claim that, with careful model selection and validation, robust filters can be obtained that yield images comparable to HAADF imaging. They also examine segmented detectors and the use of only the bright-field disk. The paper's main contribution is the cautionary demonstration of the flexibility and overfitting risk in high-dimensional diffraction data.","tokens_in":8340,"tokens_out":3662,"duration_ms":35099,"significance":"The negative result is convincing and practically important: the demonstration in Fig. 2 that arbitrary patterns (high-resolution lattices, random atom sites, rings, and the word 'caveat') can be fitted in-sample while failing in held-out regions is a clear warning to the 4D-STEM community. The numerical setup (n = 1,088, p = 65,536) makes the point starkly, and the comparison between pixelated and segmented detectors in Fig. 7 is interesting. However, the positive claim that regularization plus validation yields robust, HAADF-comparable filters is not quantitatively established. If the negative result is the main message, the paper is a useful cautionary note; if the positive claim is meant to be load-bearing, it needs substantially stronger validation.","major_comments":[{"comment":"The paper assumes a fixed linear relationship y = Xw between vectorized CBED intensities and the real-space image, with no discussion of its validity under the experimental conditions. Dynamical diffraction, detector nonlinearity, sample drift, and the fact that the scattering depends on probe position may all violate this assumption. Since the positive claim of robust filter recovery depends on this linear model being at least approximately correct, the assumption needs justification or an explicit statement of its limitations.","section":"Formalism, Eqs. (1)-(2)"},{"comment":"The paper claims to perform 'statistical model selection', but no model selection criterion is actually applied. The mixing parameter r is fixed at 5e-5 without sensitivity analysis, and λ is chosen by visual inspection of the reconstructions in Figs. 4-6 rather than by a quantitative criterion on validation data. The abstract's statement that 'careful choice of model selection and validation' yields robust filters is therefore not demonstrated by the presented analysis.","section":"Demonstrative Results, Eq. (3)"},{"comment":"The training target image (Fig. 1b) is constructed from atom site positions estimated from the HAADF image in Fig. 1a, and the validation benchmark is the same HAADF image (or a HAADF image acquired under the same conditions). The comparison is therefore partly circular: the filter is trained on HAADF-derived lattice information and then evaluated for similarity to HAADF. An independent benchmark, such as known atomic coordinates from a simulation or a separately determined structure, is needed to support the claim of structure retrieval.","section":"Demonstrative Results, Fig. 1b and Figs. 4-6"},{"comment":"The claimed 'comparable contrast to HAADF' is based on visual inspection and normalized line scans without quantitative similarity metrics or error bars. No mean-squared error, correlation coefficient, or Fourier ring correlation is reported on held-out data. Quantitative metrics are needed to support the claim that the reconstructed images are comparable to conventional HAADF imaging.","section":"Demonstrative Results, Fig. 6b"}],"minor_comments":[{"comment":"The abstract contains grammatical issues, such as 'we demonstrate that, it is possible', and the affiliation includes a typo ('Laborotary').","section":"Abstract"},{"comment":"Equation (1) is referenced but missing from the text; the blank after 'the following linear relationship:' should contain the explicit form of the linear model.","section":"Formalism"},{"comment":"The phrase 'ordinary least square estimates' should be clarified, because with n < p the OLS solution is non-unique; the authors presumably mean the minimum-norm least-squares solution.","section":"Formalism"},{"comment":"The statement that varying r 'does not alter conclusions presented here' is not supported by any shown results; a small sensitivity analysis would make this claim credible.","section":"Demonstrative Results"},{"comment":"The reference to Ref. [33] for 'unsupervised learning and exploratory data analysis' should be checked for accuracy, since the cited paper is a manifold-learning study of 4D-STEM and the connection could be stated more precisely.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The overfitting warning is solid and could justify a publication focused on the cautionary message. The positive claim of HAADF-comparable robust filters is currently under-supported and would require quantitative model selection, error bars, and an independent benchmark to be credible. I would encourage the editor to consider whether the negative result alone is sufficient for the journal's scope, or to require the authors to either substantially strengthen the positive claim or reframe the paper as a cautionary methodological note."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Recommend reading this one for the overfitting demonstration. With p=65,536 CBED pixels and n=1,088 training samples, OLS fits arbitrary training targets—including the word 'caveat'—and then fails on held-out data. That is a concrete, useful caution for the 4D-STEM community, and it's the paper's real contribution.\n\nThe elastic net follow-up is less solid. The authors scan λ, pick the images that look best by eye, and report no quantitative metric on held-out data. No MSE, no correlation, no FRC. The training target is built from atom sites estimated from the same HAADF image that later serves as the benchmark, so the comparison is partly circular. The linear model y = Xw is assumed without justification, though for a paper whose message is 'cross-validate before you trust a filter' that assumption is not the main issue. No code or data released either.\n\nTo the paper's credit, it is honest about its failures: outside-BF data doesn't validate, random atom positions don't validate, and the periodicity observation is a nice lead. The central warning stands: unvalidated linear filters can manufacture arbitrary structure. That is important for the growing 4D-STEM field.\n\nThe reader's conditional verdict and the stress-test note both land. The positive claim of HAADF-comparable robust filters is plausible but not established. The paper should be revised with quantitative validation on held-out data and a clearer separation between the cautionary result and the exploratory model-selection results.\n\nMy take: send it to peer review. The overfitting warning deserves wide circulation, and referees can force the authors to either quantify the validation or soften the positive claim. I'd set a high bar on the revision.","headline":"Convincing overfitting warning for 4D-STEM, but the positive regularization claim is under-validated; worth refereeing with a demand for quantitative metrics.","tokens_in":8882,"tokens_out":2092,"would_cite":true,"duration_ms":20908,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J07","62J05","62-07"],"pacs":["07.78.+s","68.37.Ma"],"model":"deepseek-v4-flash","headline":"Unchecked 4D-STEM inversion can fabricate almost any desired structure image; regularization and cross-validation recover HAADF-comparable images.","keywords":["4D-STEM","scanning transmission electron microscopy","convergent beam electron diffraction","elastic net","regularization","overfitting","cross-validation","structure retrieval"],"falsifier":"Run the same elastic-net training-and-validation protocol on simulated 4D-STEM data for an amorphous or non-periodic sample with a known ground-truth structure, generated by multislice or dynamical diffraction simulation; if validated filters still fail to recover the known atomic positions while fitting the training region, then the claim that careful model selection suffices would be disproved. Equivalently, on the experimental instrument, acquire a 4D dataset from a specimen with an asymmetric dopant configuration, train on one half and predict the other; if a non-periodic ground-truth target never validates at any regularization strength, the paper's periodicity conjecture becomes the limiting boundary of the method.","tokens_in":7850,"feed_emoji":"⚛️","tokens_out":7020,"duration_ms":67686,"temperature":0.7,"pith_summary":"The paper demonstrates a serious pitfall in 4D-STEM structure retrieval: because each convergent-beam electron diffraction pattern is hugely redundant, a simple least-squares filter can match almost any user-specified real-space image—high resolution, random spots, words, rings—while failing completely on data not used in training. The authors frame this as an ill-posed inverse problem and show that elastic-net regularization with a sparsity penalty, combined with validation on held-out probe positions, filters out the spurious solutions and produces atomic-scale images with contrast comparable to high-angle annular dark-field (HAADF) imaging. They also show that information outside the bright-field disk alone cannot pass validation on these datasets, and that segmented detectors with more than about 32 segments match pixelated detectors. The takeaway is that apparent resolution or structural features in reconstructed 4D-STEM images can be artifacts unless model selection and validation are used.","feed_headline":"Unvalidated 4D-STEM filters can fake any atomic-scale image","feed_subtitle":"Regularization and held-out validation turn diffraction patterns into trustworthy structure images.","key_machinery":"The central object is the linear filter model $\\boldsymbol{y} = X \\boldsymbol{w}$, where each row of $X$ is the vectorized CBED pattern at one probe position and $\\boldsymbol{y}$ is the desired real-space image intensity. Because the variable dimensionality $p = 256 \\times 256 = 65536$ vastly exceeds the number of training probe positions, the least-squares solution is ill-posed. The paper's tool is elastic-net regularized regression, $$\\hat{\\boldsymbol{w}} = \\arg\\min_{\\boldsymbol{w}} \\frac{1}{2n}\\|\\boldsymbol{y} - X\\boldsymbol{w}\\|^2 + \\$\\lambda$\\left(r\\|\\boldsymbol{w}\\|_1 + \\frac{1-r}{2}\\|\\boldsymbol{w}\\|$_2^{2}$\\right),$$ solved by cyclical coordinate descent with soft-thresholding, with the regularization strength $\\lambda$ controlling filter sparsity and hence the balance between resolution and generalization. The training/validation protocol—training on one half of the scan, predicting the other half, and transferring to a second specimen—is what exposes the overfitting.","core_discovery":"This paper establishes that direct inversion of the linear mapping $\\boldsymbol{y} = X \\boldsymbol{w}$ from vectorized convergent-beam electron diffraction (CBED) patterns to real-space image intensities is not merely noisy but systematically underdetermined: combinations of detector pixels exist that reproduce a prescribed training image with near-perfect fidelity while producing meaningless, non-transferable predictions elsewhere. The authors show this by training filters on artificial targets—excessively high resolution, random atom sites, ring patterns, and the word 'caveat'—and demonstrating that all fit the training half of the field of view yet fail on the other half. With elastic-net regularization over a path of sparsity strengths, they obtain filters that trade training fidelity for generalizable structure: at suitable values of the regularization parameter, reconstructed graphene dumbbells and dopant contrast match simultaneously acquired HAADF images, and the filters transfer to a second dataset with a different silicon dopant. They further find that bright-field disk information is sufficient, that outside-disk information fails validation, and that segmented detectors with 32 or more segments behave like full pixelated detectors.","pith_inferences":["The same overfitting hazard likely applies to black-box machine-learning models that map diffraction patterns to images; the paper's demonstration that a linear filter can reproduce the word 'caveat' is a minimal proof that any unregularized learned mapping can fit arbitrary targets.","For amorphous or aperiodic samples, the paper's periodicity constraint would likely fail; a testable extension is to check whether symmetry-aware regularization, such as group sparsity on lattice sites, restores transferability.","One could operationalize the paper's validation idea as a quantitative 'generalization index'—the ratio of test-area to training-area reconstruction error along the $\\lambda$ path—and use it to select regularization automatically.","Because a pixelated detector with a linear filter is mathematically a continuously segmented detector, the elastic-net filter landscape can be used to propose optimal discrete detector geometries for new STEM imaging modes."],"forward_implications":["Without held-out validation, one can deliberately construct filters that make a 4D-STEM dataset display arbitrarily high resolution, random atom sites, rings, or text-like patterns, so reported resolution claims are not self-certifying.","Filter estimation on half a scan and prediction on the other half is a workable model-selection protocol; a filter that passes it reconstructs graphene structure comparably to HAADF imaging.","For the graphene datasets studied, only the bright-field disk information survives validation; the region outside the disk, though rich in raw intensity, does not yield generalizable structure.","Segmented detectors with about 32 segments or more reconstruct essentially as well as a full pixelated detector, connecting the statistical filter view to detector-design practice.","Periodic training targets (complement of the atom image and individual sublattices) generalize under regularization, suggesting that crystal symmetry is a usable inductive bias for virtual imaging."],"supporting_citations":[{"why":"Establishes the pixelated-detector 4D-STEM modality whose datasets this paper inverts.","marker":"[1]"},{"why":"Supplies the early detector-geometry formulation that the linear filter model generalizes.","marker":"[4]"},{"why":"Together with [4], supplies the differential phase contrast lineage that motivates finding optimal detector weights.","marker":"[5]"},{"why":"Provides the optimized segmented-detector analysis that the paper's segmented-detector results are checked against.","marker":"[16]"},{"why":"Defines the elastic-net penalty that forms the regularized objective in equation (3).","marker":"[25]"},{"why":"Defines the Lasso extreme of the penalty used for sparse filter estimation.","marker":"[26]"},{"why":"Gives the pathwise coordinate-descent strategy for solving the regularization path over $\\lambda$.","marker":"[27]"},{"why":"Supplies the concrete cyclical coordinate-descent algorithm with soft-thresholding used to compute the filters.","marker":"[28]"}],"fun_headline_variants":["4D-STEM filters without checks can fabricate any image","Statistical traps in 4D-STEM: fake structures lurk","Regularization is essential for reliable 4D-STEM images","Unvalidated 4D-STEM filters produce phony atomic maps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that a single fixed linear map from diffraction intensities to real-space image holds across all probe positions and across specimens, although real electron scattering is nonlinear and the training targets are themselves derived from the HAADF image used for validation.","fun_headline_variants_meta":{"raw":{"variants":["4D-STEM filters without checks can fabricate any image","Statistical traps in 4D-STEM: fake structures lurk","Regularization is essential for reliable 4D-STEM images","Unvalidated 4D-STEM filters produce phony atomic maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000523,"raw_usage":{"total_tokens":2500,"prompt_tokens":886,"completion_tokens":1614,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":1540}},"tokens_in":502,"tokens_out":1614,"duration_ms":11621,"temperature":1.0,"reasoning_tokens":1540,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:39:50.910120+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same elastic-net training-and-validation protocol on simulated 4D-STEM data for an amorphous or non-periodic sample with a known ground-truth structure, generated by multislice or dynamical diffraction simulation; if validated filters still fail to recover the known atomic positions while fitting the training region, then the claim that careful model selection suffices would be disproved. Equivalently, on the experimental instrument, acquire a 4D dataset from a specimen with an asymmetric dopant configuration, train on one half and predict the other; if a non-periodic ground-truth target never validates at any regularization strength, the paper's periodicity conjecture becomes the limiting boundary of the method.","supporting_citations":[{"cited_title":"Efficient phase contrast imaging in STEM using a pixelated detector. Part 1: Experimental demonstration at atomic resolution,","cited_arxiv_id":null,"evidence_quote":"Establishes the pixelated-detector 4D-STEM modality whose datasets this paper inverts."},{"cited_title":"PHASE CONTRAST IN SCANNING TRANSMISSION ELECTRON MICROSCOPY.,","cited_arxiv_id":null,"evidence_quote":"Supplies the early detector-geometry formulation that the linear filter model generalizes."},{"cited_title":"DIFFERENTIAL PHASE CONTRAST IN A STEM.,","cited_arxiv_id":null,"evidence_quote":"Together with [4], supplies the differential phase contrast lineage that motivates finding optimal detector weights."},{"cited_title":"Efficient phase contrast imaging in STEM using a pixelated detector. Part II: Optimisation of imaging conditions,","cited_arxiv_id":null,"evidence_quote":"Provides the optimized segmented-detector analysis that the paper's segmented-detector results are checked against."},{"cited_title":"Regularization and variable selection via the elastic net,","cited_arxiv_id":null,"evidence_quote":"Defines the elastic-net penalty that forms the regularized objective in equation (3)."},{"cited_title":"Regression shrinkage and selection via the lasso,","cited_arxiv_id":null,"evidence_quote":"Defines the Lasso extreme of the penalty used for sparse filter estimation."},{"cited_title":"Pathwise coordinate optimization,","cited_arxiv_id":null,"evidence_quote":"Gives the pathwise coordinate-descent strategy for solving the regularization path over $\\lambda$."},{"cited_title":"Regularization paths for generalized linear models via coordinate descent,","cited_arxiv_id":null,"evidence_quote":"Supplies the concrete cyclical coordinate-descent algorithm with soft-thresholding used to compute the filters."}],"review_version":1}