{"id":"a97de3a9-1da5-44e3-bb9c-907679d83b1e","arxiv_id":"2505.12650","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"AutoMat reconstructs crystal structures from simulated STEM images by orchestrating denoising, template retrieval, symmetry-constrained atom fitting, and machine-learned-potential relaxation, and it outperforms existing baselines on the new STEM2Mat benchmark.","lead":"AutoMat is an AI agent that converts electron microscope images of atomic crystals into computer-ready structures and predicted energies by chaining denoising, template lookup, atom fitting, and energy relaxation tools. The paper also presents STEM2Mat-Bench, a simulated benchmark of 450 crystal samples, where AutoMat reports far lower errors than vision-language models and a specialized atom-detection tool.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Template gallery likely contains the 450 test structures, so the reported success may reflect retrieval-plus-refinement rather than single-image reconstruction.","rationale":"The reader's weakest assumption is correct and is the most load-bearing point: the paper's own module description and dataset construction make exact-template leakage likely, and none of the reported numbers control for it. This is not a dispute with the general approach—retrieval from a known database plus local refinement is a legitimate system—but it changes what the benchmark demonstrates. The reported two-order-of-magnitude improvement over AtomAI and factor-of-eight improvement over VLMs on energy MAE are interpretable as reconstruction gains only if the retrieval step lacks privileged access to ground-truth structures. A leak-free rerun, or even a precise statement of gallery construction, would settle this. Secondary issues that should be fixed in revision but are not the main attack: the tier counts in §3.3 sum to 570, not 450 (35 + 456 + 79), and the abstract's '<350 meV/atom' differs from the table's 321.57 and the text's 332 ± 12 meV/atom; these inconsistencies reduce confidence in the reported numbers. The paper has real strengths: it releases code and data, provides a concrete benchmark, and includes a failure analysis; a clean ablation of the LLM orchestration (e.g., fixed sequential pipeline versus agentic rollback) would further strengthen the causal claim that agentic tool use drives the gains. But the gallery overlap is the threshold issue for the central claim, so the paper should remain conditional pending a leak-free evaluation.","tokens_in":11300,"tokens_out":4657,"duration_ms":46592,"concrete_test":"Inspect the released repository and HuggingFace dataset to determine whether the Image Template Matching gallery includes the 450 test CIFs or their simulated projections. Then run the full AutoMat pipeline under three conditions on the same 450 samples: (a) gallery as released; (b) gallery restricted to the 1,693 training/validation structures, explicitly excluding all 450 test structures; (c) gallery augmented with near-duplicate/relaxed variants of the test structures to probe sensitivity. Compare S.S., RMSD, and energy MAE across conditions. If condition (b) materially degrades performance or retrieval success, the headline result reflects retrieval-plus-refinement with leakage, not single-image reconstruction, and the benchmark must be rebuilt or the claim re-scoped.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing vulnerability is the unstated overlap between the Image Template Matching gallery and the test set. The benchmark is built from one 2,143-structure pool (§3.1); the 450 test images are selected from that same pool (§3.3); and the retrieval module 'compares the enhanced image to a pre-stored or externally mounted database of structural templates' (§4.2), with the agent instructed to 'retrieve candidate templates from the structure database' (§4.3). The paper never states that the 450 ground-truth structures are excluded from this gallery. If they are present, then for every test image the exact target CIF (or its simulated projection) is a candidate, and STEM2CIF's role is reduced to selecting the right template and locally refining it. The reported 83.2% S.S. and 0.11 Å RMSD then measure image-retrieval quality plus refinement, not ab initio reconstruction from a single STEM projection. The error analysis in §5.2 compounds this: 39.3% of failures are attributed to template retrieval, confirming the gallery is load-bearing for the headline claim. Without an explicit gallery/test separation, the comparison with VLMs and AtomAI conflates 'retrieval from a known database' with 'reconstruction from the image.' A precise statement of the gallery construction—or a corrective experiment—is therefore necessary before the central claim can be accepted as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AutoMat, an LLM-orchestrated pipeline that converts a single noisy synthetic STEM projection into a simulation-ready CIF file and a formation-energy prediction. The pipeline chains four modules—pattern-adaptive denoising (MOE-DIV AESR), physics-guided image-template retrieval, symmetry-constrained atomic reconstruction (STEM2CIF), and MLIP-based relaxation/property prediction via MatterSim—with rollback-and-retry orchestrated by a DeepSeek-V3 agent. The authors also introduce STEM2Mat-Bench, a synthetic benchmark of 450 (claimed) image–structure–property triples built from 2,143 curated monolayer structures, and report projected lattice RMSD of about 0.11 Å, formation-energy MAE of about 330 meV/atom, and an overall structure success rate of 83.2%, outperforming GPT-4.1mini, Qwen-VL, LLaMA4V, ChemVLM, and AtomAI by an order of magnitude on the reported tables.","tokens_in":11594,"tokens_out":6497,"duration_ms":68823,"significance":"If the central claims hold, AutoMat would be a useful integration of image denoising, template matching, symmetry-constrained reconstruction, and MLIP validation, and STEM2Mat-Bench would provide a reproducible synthetic evaluation suite for a task that currently lacks standardized benchmarks. The paper has concrete strengths: the authors release code and data, the benchmark construction is described in enough detail to be replicated, the metrics are mostly explicit, and the reported order-of-magnitude improvement over general-purpose VLMs is striking. However, the significance is conditional on two load-bearing points: the test structures must be excluded from the template-retrieval gallery, and the contribution of the LLM agent itself must be separated from the contribution of the specialized modules. The current manuscript does not supply either piece of evidence, and the synthetic-only evaluation further limits the strength of the generalization claims.","major_comments":[{"comment":"The manuscript never states that the 450 test ground-truth structures are excluded from the Image Template Matching gallery. Because §3.3 selects the 450 test images from the same 2,143-structure pool used to build the gallery, and §4.2 describes retrieval from 'a pre-stored or externally mounted database of structural templates,' the reported S.S. (83.2%) and RMSD (0.11 Å) could in principle be achieved by retrieving the correct template from the database and locally refining it, rather than by reconstructing the lattice from the image. This is not merely hypothetical: §5.2 attributes 39.3% of failures to template retrieval, confirming that retrieval is load-bearing for the headline result. Please state explicitly whether test structures are excluded from the gallery and, ideally, report a holdout ablation in which the gallery is restricted to the training/validation split.","section":"§3.3, §4.2, §5.2"},{"comment":"The benchmark size and tier arithmetic are internally inconsistent: the text says 450 test samples were retained, but the tier counts in §3.3 (35 + 456 + 79 = 570) sum to 570, contradicting the abstract, introduction, and §3.3's own '450' figure. Moreover, the aggregate energy MAE of 321.57 in Table 1 matches neither the unweighted mean of the tier means (332.43) nor the weighted mean using the stated tier counts (323.49). Please correct the counts, the aggregate, or the definition of 'Avg.'","section":"§3.3, Tables 1–2"},{"comment":"The central claim is that agentic orchestration by an LLM enables the reported performance, but no ablation compares the DeepSeek-V3 controller against a deterministic, fixed-order execution of the same four modules with the same retry logic. Without such an ablation, the reported improvement over the VLMs cannot be attributed to the agent's closed-loop tool use rather than to the specialized vision and physics modules alone. Please add at least one non-agent ablation (e.g., fixed pipeline, no rollback) and report its tier-wise results.","section":"§5.1, Tables 1–2"},{"comment":"The projected lattice RMSD in Eq. (2) compares only the in-plane lattice-vector lengths a and b and ignores the in-plane angle γ. A prediction with a grossly wrong unit-cell angle can therefore receive a small RMSD, which makes the headline 0.11 Å value incomplete as a structural accuracy measure. Please extend Eq. (2) to include the angle (or the full metric tensor), or justify the current definition and report the angle error separately.","section":"§3.4, Eq. (2)"}],"minor_comments":[{"comment":"There are unresolved placeholder references 'Fig. ??' and 'Appendix ?? and ??'; please fill in or remove them before publication.","section":"§3.1, §4.2"},{"comment":"The model name is written inconsistently as 'LLama4V', 'LLaMA4', and 'Llama-4-Maverick'; please use one canonical name throughout.","section":"Tables 1–2, Figure 4"},{"comment":"The discussion states a mean formation-energy MAE of '332±12 meV/atom', while Table 1 reports an average of 321.57 meV/atom; please reconcile the number and define the reported uncertainty.","section":"§5.1"},{"comment":"The definition of Composition Correctness is per-sample binary, but it is reported as a percentage; please state explicitly that the reported values are means over the test set.","section":"§3.4"},{"comment":"All evaluations use abTEM-simulated images rather than experimental STEM micrographs; please add an explicit limitation statement noting that the benchmark currently measures performance on synthetic data only.","section":"§3.2, §6"}],"recommendation":"major_revision","confidential_remarks":"The decisive issue for me is the holdout question. If the authors can document that the 450 test structures are absent from the template-retrieval gallery, the main technical contribution stands, subject to the orchestration ablation and metric fix. If they cannot, the paper should be reframed as retrieval-augmented reconstruction rather than ab initio reconstruction from a single image, and the headline numbers should be recomputed for the setting in which the exact ground-truth structure is not a gallery candidate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper is worth reading but the headline numbers should not be taken at face value. The core idea—an LLM agent that chains denoising, template retrieval, symmetry-constrained reconstruction, and MLIP relaxation to turn a STEM image into a CIF—is a genuine integration, and the STEM2CIF module plus benchmark release are real contributions. They also run sensible baselines.\n\nWhat the paper does well: it is the first end-to-end agentic pipeline I have seen for this task. The authors are honest in §5.2 that 39.3% of failures are template retrieval errors. They release code and data.\n\nThe soft spots are in the evaluation. The biggest one is the likely overlap between the template gallery and the test set. The benchmark is built from one pool of 2,143 structures. The test set of 450 is sampled from that same pool. The retrieval module matches against a \"structure database\" that is never described as excluding the test structures. If the exact ground-truth CIF is in the gallery, then the reported 0.11 Å RMSD and 83.2% S.S. measure retrieval plus local refinement, not single-image reconstruction. The error analysis reinforces this: it treats \"template retrieval\" as a separate failure mode, which only makes sense if retrieval is doing the heavy lifting. The paper must explicitly state the gallery construction and run a held-out experiment.\n\nTwo smaller issues. The tier counts do not add up: Tier 1=35, Tier 2=456, Tier 3=79 sums to 570, not 450. That is either a typo or a deeper mismatch. Also, there is no ablation that isolates the LLM orchestration from a non-agent version of the same modules, so we do not know what the agent actually contributes.\n\nThe evaluation is entirely synthetic. That is acceptable for a first benchmark, but the abstract's claim about \"microscopic characterization\" overstates the evidence. No experimental STEM images are tested.\n\nWho this is for: people building vision-to-structure pipelines in materials informatics will get useful ideas and a reusable benchmark. But the central accuracy claim needs repair. I would send this to peer review—the issue is fixable with a clear gallery/test split statement and a corrected table—but I would not accept the claim as stated now.","headline":"AutoMat is a serious engineering effort with a likely inflated headline: template leakage and inconsistent tier counts undermine the accuracy claims until fixed.","tokens_in":12129,"tokens_out":2073,"would_cite":false,"duration_ms":22483,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AutoMat claims that a text-only language model, armed with denoising, template retrieval, reconstruction, and relaxation tools, can turn a single noisy STEM micrograph into a simulation-ready crystal structure and a formation-energy…","keywords":["STEM image analysis","crystal structure reconstruction","agentic tool use","closed-loop verification","machine-learned interatomic potentials","template retrieval","2D materials","formation energy prediction"],"falsifier":"Run AutoMat's template-matching stage on the STEM2Mat test images after deleting every test structure's ground-truth template from the gallery (or with a gallery of unrelated structures), and check whether the 83.2% success rate and 0.11 Å RMSD collapse toward retrieval-by-chance; also compare against a closed-loop method that receives no template gallery at all.","tokens_in":11136,"feed_emoji":"🔬","tokens_out":6602,"duration_ms":59405,"temperature":0.7,"pith_summary":"This paper claims that a single noisy STEM image of a 2D crystal can be converted automatically into a simulation-ready crystal structure and a formation-energy estimate, without a human annotator. The proposed system, AutoMat, chains four tools—adaptive denoising, physics-guided template retrieval, symmetry-constrained reconstruction into a CIF, and relaxation plus energy prediction with a machine-learned interatomic potential—and lets a text-only language model orchestrate them with rollback-and-retry when a quality check fails. To test this, the authors build STEM2Mat-Bench, 450 simulated image–structure–energy triples drawn from 2,143 curated 2D materials, and report a projected lattice RMSD of about 0.11 Å, a formation-energy MAE near 330 meV/atom, and an 83.2% structure success rate, an order of magnitude better than vision-language model baselines and specialized STEM toolkits. The paper's point is that closed-loop tool use, not a bigger vision model, is what bridges microscopy and atomistic simulation.","feed_headline":"Agent turns noisy STEM images into crystal structures","feed_subtitle":"Tool-using agent reconstructs lattices to 0.11 Å and predicts energies, beating vision-language models by 10x.","key_machinery":"The load-bearing mechanism is the agentic loop around the four tools: a language-model controller makes tool calls, receives structured intermediate results, runs quality checks, and can roll back to a previous stage and retry. The physics-guided template retrieval branch is the crucial state-dependent auxiliary: rather than interpreting the image ab initio, the system matches the denoised image to a library of simulated STEM projections, which supplies a strong prior for lattice type and element assignment; the reconstruction module then refines the candidate under symmetry constraints. The closed-loop verification is what distinguishes this from a one-pass feed-forward reconstruction.","core_discovery":"The central claim is that inference-time hypothesis search with closed-loop verification lets an otherwise text-only LLM outperform vision-language models on a spatially precise inverse problem. AutoMat treats reconstruction as a sequence of tool calls: a pattern-adaptive denoiser enhances the micrograph; an image template matcher proposes candidate structures from a library of simulated projections, filtered by elemental contrast; a reconstruction module detects atomic peaks by clustering, fits the lattice under symmetry constraints, assigns species from the candidate, and writes a CIF; and a machine-learned interatomic potential relaxes the structure and predicts formation energy. The agent monitors intermediate outputs and rolls back to retry failed stages. On STEM2Mat-Bench the system reaches 0.11 ± 0.03 Å in-plane lattice RMSD, 321.6 meV/atom mean formation-energy error, and 83.2% structure success, against a 48.1 meV/atom oracle floor when the true CIF is fed directly to the potential. Error analysis attributes 39.3% of failures to template retrieval and 60.7% to downstream steps such as projection ambiguity or element contrast confusion.","pith_inferences":["The benchmark design leaves the template-gallery overlap question open: the 450 test images are drawn from the same 2,143-structure pool used to build the retrieval library, and the paper never states that test structures are excluded. If they are not, the headline numbers blend retrieval with refinement, and real-world generalization to unseen crystals would likely be lower.","The same closed-loop, rollback-on-failure pattern could transfer to other ill-posed imaging inversions, such as 3D atomic reconstruction from a tilt series, where a controller could compare candidate 3D models against multiple projections rather than committing to one feed-forward prediction.","A direct extension would be to replace or augment template retrieval with a generative structural head that can hypothesize lattices not present in the library; that would turn AutoMat from a lookup-and-refine system into a true ab initio reconstructor and would resolve the main ambiguity in the evaluation."],"forward_implications":["If the reported accuracy holds, laboratories could pipe raw STEM micrographs into an automated system that outputs a CIF ready for simulation, removing a bottleneck in training and validating machine-learned interatomic potentials.","A text-only LLM equipped with domain tools can beat large vision-language models on a visual, spatially precise scientific task, suggesting that tool orchestration matters more than native image reasoning for such problems.","Since 39.3% of failures are attributed to template retrieval, improving retrieval robustness—via uncertainty-aware or multi-candidate matching—would yield the largest immediate gains, while the remaining failures come from projection ambiguity and elemental confusion even with a correct template.","Because the paper attributes most residual energy error to reconstruction rather than to the machine-learned potential, structural fidelity improvements should translate almost directly into better property predictions."],"supporting_citations":[{"why":"Supplies the pretrained machine-learned interatomic potential used for relaxation and formation-energy prediction, the final stage of the pipeline.","marker":"[1]"},{"why":"Provides the pattern-adaptive denoising module (MOE-DIV AESR) that enhances raw STEM images before template matching.","marker":"[12]"},{"why":"Serves as the specialized STEM toolkit baseline that only extracts atomic coordinates, used to show AutoMat's structural reconstruction advantage.","marker":"[28]"},{"why":"One of the structure databases from which the 2,143 curated 2D materials are harvested, forming the source pool for templates and benchmark samples.","marker":"[29]"},{"why":"Another source database for the curated 2D structure pool that underpins the template library and benchmark.","marker":"[30]"},{"why":"Third source of crystal structures used to build the curated pool and template gallery.","marker":"[31]"},{"why":"The simulation engine used to generate synthetic iDPC-STEM images from the structure pool, providing both the template library and the benchmark images.","marker":"[32]"},{"why":"Represents the general-purpose vision-language baseline whose energy and structural errors AutoMat claims to beat by an order of magnitude.","marker":"[26]"}],"fun_headline_variants":["Agentic AI reconstructs crystal lattices from noisy STEM images","Tool-using agent beats vision-language models at 3D structure reconstruction","AutoMat: AI agent with closed-loop verification decodes atomic structures","Agentic search nails crystal structures from single STEM images","Inference-time search turns STEM noise into validated crystal lattices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported accuracy assumes the retrieval library can contain the exact test structure, because the paper does not state that test structures were excluded from the 2,143-template gallery built from the same pool used to create the 450 test images.","fun_headline_variants_meta":{"raw":{"variants":["Agentic AI reconstructs crystal lattices from noisy STEM images","Tool-using agent beats vision-language models at 3D structure reconstruction","AutoMat: AI agent with closed-loop verification decodes atomic structures","Agentic search nails crystal structures from single STEM images","Inference-time search turns STEM noise into validated crystal lattices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1404,"prompt_tokens":980,"completion_tokens":424,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":596,"completion_tokens_details":{"reasoning_tokens":337}},"tokens_in":596,"tokens_out":424,"duration_ms":4591,"temperature":1.0,"reasoning_tokens":337,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:29:18.407585+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run AutoMat's template-matching stage on the STEM2Mat test images after deleting every test structure's ground-truth template from the gallery (or with a gallery of unrelated structures), and check whether the 83.2% success rate and 0.11 Å RMSD collapse toward retrieval-by-chance; also compare against a closed-loop method that receives no template gallery at all.","supporting_citations":[{"cited_title":"Atomai framework for deep learning analysis of image and spectroscopy data in electron and scanning probe microscopy","cited_arxiv_id":null,"evidence_quote":"Serves as the specialized STEM toolkit baseline that only extracts atomic coordinates, used to show AutoMat's structural reconstruction advantage."},{"cited_title":"The computational 2d materials database: high-throughput modeling and discovery of atomically thin crystals","cited_arxiv_id":null,"evidence_quote":"One of the structure databases from which the 2,143 curated 2D materials are harvested, forming the source pool for templates and benchmark samples."},{"cited_title":"abtem: Ab initio transmission electron microscopy image simulation","cited_arxiv_id":null,"evidence_quote":"The simulation engine used to generate synthetic iDPC-STEM images from the structure pool, providing both the template library and the benchmark images."}],"review_version":1}