{"id":"11eeb024-7c16-482a-8f0b-9d2b1f23d949","arxiv_id":"2607.22975","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AEcroscopyWave exposes scanning probe microscopes to agentic AI via MCP tools and supports on-the-fly arbitrary waveform generation; its only case-study evidence is not statistically significant.","lead":"This paper describes AEcroscopyWave, a distributed software platform that lets AI agents control scanning probe microscopes and design custom excitation waveforms in real time. It is a useful step toward autonomous microscopy, though its main demonstration is preliminary and not statistically significant.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central capability claim—AI agents dynamically designing and executing novel waveforms—is not demonstrated: the case study uses human-designed waveforms, and no quantitative LLM-planner success/failure data or baselines are reported.","rationale":"The reader's weakest assumption—that LLM-generated experimental scripts can reliably and safely drive real instruments—is closely related to my concern, but I locate the load-bearing gap specifically in the absence of any demonstration that the LLM planner can generate novel waveforms successfully. The architecture description and the human-designed pulse-train case study do not exercise the AI-generation loop that the central claim depends on. The lack of quantitative success data, failure taxonomy, and baseline comparisons is therefore not a minor omission; it is the difference between a demonstrated capability and a proposed one. I credit the real engineering in the paper: arbitrary waveform upload, distributed agent architecture, MCP integration, and human-in-the-loop approval are meaningful contributions and are not called into question. The non-significant physics results weaken the illustrative value of the case study but do not by themselves invalidate the platform claim. No load-bearing arithmetic error or circular derivation exists. Since the reader's CONDITIONAL verdict already reflects this level of evidence, my stress test does not change the verdict.","tokens_in":9774,"tokens_out":3309,"duration_ms":35742,"concrete_test":"Run a benchmark of N independent natural-language experimental goals (e.g., 50) through the LLM Planner to generate run() scripts targeting the digital twin. Measure: (a) fraction passing digital-twin validation without error; (b) fraction approved by human reviewers; (c) fraction executing successfully on real microscope hardware; (d) fraction producing the intended physical outcome; and (e) the fraction of generated waveforms that are genuinely novel relative to the waveform library vs merely parameter variations. Compare these metrics against AEcroscopy v1 script execution and against manual expert-written scripts for the same goals. Report a failure taxonomy categorized by syntactic error, semantic/logic error, hardware safety violation, and scientific invalidity. If LLM-generated scripts show high success rates and novel waveforms execute without safety incidents, the concern is res","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that AEcroscopyWave lets AI agents dynamically generate and execute novel excitation waveforms rather than tune predefined parameters—requires the LLM Planner, digital-twin validation, and human-approval gate to produce executable and scientifically valid instrument scripts. The paper provides no quantitative evidence for this. The LLM Planner paragraph describes the MCP tool catalog and response-format contract but reports no success rate, no failure taxonomy, and no comparison to AEcroscopy v1 or manual baselines. The Digital Twin Interface subsection explicitly limits validation to 'syntactic verification' ('catches syntax and logic error'), so it does not establish physical safety or experimental validity. The only end-to-end case study uses a pulse-train waveform whose schematic (Fig. 4b) appears human-designed; the paper does not state that the LLM planner generated it, and the headline physical trends are not statistically significant (p=0.117 for switching probability, p=0.078 for area). Thus the central 'creative experimental design by AI' claim is architecturally plausible but empirically undemonstrated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript describes AEcroscopyWave, a distributed software-defined platform for scanning probe microscopy designed to expose instruments to both human users and LLM-based agents via an MCP tool server. It adds programmable arbitrary waveform generation to the prior AEcroscopy framework, with digital-twin syntactic validation and a human-approval gate before instrument execution. The paper presents the architecture, the waveform-generation stack, and a BTO switching case study using a custom pulse-train waveform. The headline claim is that AI agents can dynamically design, validate, and execute novel microscopy experiments; the case study demonstrates flexible waveform execution but not agentic generation, and its statistical trends are not significant.","tokens_in":10033,"tokens_out":4933,"duration_ms":44922,"significance":"If the architecture works as described, AEcroscopyWave is a useful step toward agentic microscopy: it replaces monolithic instrument scripts with REST/MCP-accessible services, separates planning from execution, and introduces a human-approval gate plus digital-twin syntax checking. The platform-level design choices (distributed server-client model, MCP tool catalog, named workflow registry, three-layer waveform module) are clearly described and appear technically plausible. The public user documentation and the integration with multiple commercial AFM/DAQ backends are strengths. However, the paper does not currently provide evidence for its central capability claim: the only end-to-end experiment uses a human-designed waveform, there are no LLM-planner success/failure statistics, and the reported physics trends are below the significance threshold.","major_comments":[{"comment":"The central claim that AI agents can 'dynamically generate, upload, and execute entirely user-defined excitation waveforms' is not demonstrated by the presented experiment. The case study says 'We designed a pulse-train PFM experiment' (authors, not the LLM planner), and no sentence states that the planner generated the waveform. Since the only end-to-end test is human-designed, the capability claim rests on architecture alone. Please either add an end-to-end example where the LLM planner proposes a waveform that passes digital-twin validation and human approval and is executed, or clearly restrict the paper's claim to 'human-designed waveforms executed through a flexible platform, with agentic generation as a planned extension.'","section":"Case Study (Figs. 4–5)"},{"comment":"No quantitative evidence is given for the LLM planner's reliability. The text describes the prompt, tool catalog, and response-format contract, but reports no success rate for producing a valid run() function, no distribution of error types, no correction-loop statistics, and no comparison against AEcroscopy v1 or manual script writing. Such numbers are necessary to support the statement that 'the AI agent can autonomously design both the experimental workflow and the excitation strategies required.' At minimum, report end-to-end planner success rate on a held-out benchmark and a representative failure taxonomy.","section":"LLM Experimental Planner / MCP Tool Server"},{"comment":"The digital-twin validation is explicitly limited to 'syntactic verification' and 'catches syntax and logic error,' as stated in the Figure 3 callout and the corresponding subsection. It therefore does not establish that an LLM-generated script is physically safe or scientifically meaningful. The paper treats the combination of digital-twin checks and human approval as sufficient for safe execution, but provides no evidence about failure modes that pass syntax yet harm the probe, sample, or instrument. Please discuss what the digital twin does and does not validate, and temper the safety claims or add safeguards (e.g., explicit parameter bounds, hardware interlocks) and tests of those safeguards.","section":"Digital Twin Interface"},{"comment":"The main experimental results are reported as p=0.117 for conditioning-fraction effect on switching probability and p=0.078 for area. The authors themselves call these 'marginal rather than statistically significant.' Summary statements such as 'strongest qualitative trend' and 'suggestive evidence' are appropriate, but the abstract's claim that 'benefits ... is demonstrated by test cases' is not supported by this data. The case study can illustrate the flexible waveform and measurement workflow, but it cannot validate the agentic or scientific-discovery claims without either larger n, pre-registered hypotheses, or independent replication.","section":"Case Study statistics"}],"minor_comments":[{"comment":"Typo/word choice: 'stratified into two distinct regimens' should likely read 'regimes' (unless 'regimens' was intended for a medical analogy, which seems unlikely in context).","section":"Abstract"},{"comment":"'switched domain radius area' is awkward; consider 'domain radius' or 'domain area' depending on what is extracted in Fig. 5(b).","section":"Figure 4 caption"},{"comment":"The sentence 'The experiment was performed at two different sites per image, for different images for a total of four distinct points' is ambiguous. Clarify the number of images, sites per image, and total measurement points.","section":"Case Study"},{"comment":"It is not clear whether waveform output from WaveVI (LabVIEW/NI DAQ) is synchronized with AFM scan/pulse triggers in the VERO/Cypher workflows or only used stand-alone. A short timing/synchronization note would help.","section":"Hardware-Software Interface"},{"comment":"The documentation URL is provided, but not a source-code repository or version tag. Since the paper is about a software platform, please include the repository (or state availability conditions) to enable reproducibility.","section":"Data and Code Availability"},{"comment":"Some references are to arXiv preprints and future-dated venues; consider adding the peer-reviewed version when available, and double-check all reference metadata (e.g., year/volume) before production.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The architecture is plausible and clearly presented, but the paper overstates what is demonstrated. The most important issue is the gap between the agentic-AI framing and the human-designed case study with non-significant statistics. This can be addressed in revision either by adding an agentic end-to-end demonstration and planner-reliability metrics, or by repositioning the manuscript as a platform/architecture paper with explicitly scoped claims. If the authors cannot supply the missing evidence, I would lean toward rejecting the strong 'creative experimental design' claims rather than the platform itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nShort version: this is a real piece of engineering that moves SPM automation forward, but the headline capability—AI agents designing and executing novel waveforms—is not actually demonstrated in the paper. The case study uses waveforms the authors designed themselves, and the one end-to-end experiment is honestly labeled not statistically significant. That doesn't sink the platform, but it should be reflected in how the results are framed.\n\nWhat is genuinely new: the move from fixed waveform libraries to dynamic, arbitrary waveform upload is a meaningful step beyond AEcroscopy v1. The architecture is well thought out—distributed server/client with FastAPI, job queue, instrument registry, and a digital twin that does syntactic validation without touching hardware. The MCP tool server is a sensible way to expose microscope control to LLMs, and the three-layer waveform module (primitives, registry, MCP callable) is practical. The paper does a good job describing the design tradeoffs, and the writing is clear. I also appreciate that the non-significant p-values are reported rather than smoothed over.\n\nWhere it's soft: the central agentic claim needs numbers. No success rate for the LLM planner, no failure taxonomy, no comparison against v1 or manual operation, no evaluation of the digital twin's ability to catch actually dangerous scripts. The case study itself is a conventional factorial screen with a human-designed pulse train; it demonstrates that the waveform upload works, not that an AI agent designed anything creative. The 'data and code availability' line points to user documentation only—no source, no data. For a paper that says 'benefits demonstrated by test cases,' that's a gap.\n\nThat said, the physical case study on BTO is interesting in its own right, and the authors are appropriately cautious about the marginal trends. Nothing here is wrong on the architecture side; the shortfall is all in validation evidence.\n\nBottom line: this is a 'Towards' paper that would be better if it were framed as one. It deserves a serious referee—the architecture and engineering are worth careful scrutiny—but I'd want the authors to either provide quantitative LLM planner evaluation or explicitly limit the claims to platform capabilities, not demonstrated agentic creativity. I'd take it to a reading group for the architecture discussion, and I'd probably cite it if I were working on autonomous SPM.","headline":"New SPM automation engineering with a genuinely useful MCP/arbitrary-waveform architecture, but the claim that AI agents design novel waveforms is unsupported by the presented evidence.","tokens_in":10515,"tokens_out":2934,"would_cite":true,"duration_ms":28087,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AEcroscopyWave is a software-defined microscopy platform that lets AI agents design, validate, and execute arbitrary new experiments—moving beyond tuning preset parameters—while keeping a human in the approval loop.","keywords":["automated microscopy","agentic AI","scanning probe microscopy","waveform generation","model context protocol","digital twin validation","autonomous experimentation","software-defined instrumentation"],"falsifier":"Run a defined benchmark where LLM-generated experimental scripts are executed on a real microscope against a known ground-truth sample, measuring success rate, hardware incidents, and deviation from the intended protocol; if success is not clearly better than random or if safety incidents occur despite validation, the central claim is undermined.","tokens_in":9716,"feed_emoji":"🔬","tokens_out":3185,"duration_ms":31397,"temperature":0.7,"pith_summary":"The paper argues that the next step for autonomous microscopy is not smarter optimization within fixed parameter spaces but giving AI agents the creative freedom to invent entirely new excitation waveforms and experimental procedures. To make that possible, the authors built AEcroscopyWave, a distributed software layer that exposes scanning probe microscopes to large language model (LLM) agents through a flat set of callable tools, validates newly generated code against a simulated 'digital twin' microscope before it touches real hardware, and requires human approval before execution. The central demonstration is a pulse-train piezoresponse force microscopy experiment on a ferroelectric film, where the waveform was defined by the platform rather than chosen from a library. If the approach works at scale, it would let AI systems not just tune parameters but propose and test new measurement strategies, making microscopes composable, remotely accessible instruments for human-AI teams.","feed_headline":"AI agents can now invent their own microscope experiments","feed_subtitle":"The platform lets LLMs generate arbitrary excitation waveforms, with digital-twin checks and human sign-off before real instruments.","key_machinery":"The load-bearing mechanism is the combination of a distributed server-client job queue that separates planning from execution, an MCP tool layer that gives LLM agents a flat and discoverable catalog of microscope operations, a three-layer waveform generator (primitives, a named waveform registry, and MCP-callable tools) that allows arbitrary waveforms to be synthesized and uploaded, and a digital-twin interface that syntactically validates custom scripts before a human approves them for the real instrument. The MCP tool layer is what makes the platform 'agent-native': instead of reading code modules, the agent sees a consistent calling convention for every instrument operation.","core_discovery":"The central claim is that a microscopy platform built around programmatic, arbitrary waveform generation and a model context protocol (MCP) tool server turns a scanning probe microscope into a platform where AI agents can propose, validate, and run genuinely new experiments rather than only sampling predefined parameter spaces. The paper demonstrates the architecture with a ferroelectric switching experiment using custom pulse trains, executing a 72-condition factorial screen and reporting trends suggesting intermediate conditioning amplitude may enhance switching, though the results do not reach statistical significance.","pith_inferences":["If LLM planners become reliable enough, the same architecture could extend beyond scanning probe microscopy to synchrotron beamlines, electron microscopes, or any programmable instrument, turning scientific facilities into platforms where AI designs and runs experiments on demand.","The case study's non-significant trends suggest the bottleneck may shift from instrument control to experimental design and sample variability; closed-loop AI may need to plan replicates across sites to extract statistically meaningful conclusions.","The digital twin currently validates syntax and logic, not scientific meaningfulness; a more physics-aware simulator could catch experiment designs that run without error but cannot produce interpretable measurements.","The authors' deliberately modest use of 'digital twin' as simulated control execution implies a gradual path: as twins become more physically faithful, more of the human approval burden could be automated."],"forward_implications":["AI agents can autonomously generate and execute custom excitation waveforms, enabling spectroscopy and switching experiments that would otherwise require specialized human-built setups.","Pre-approved workflows can be registered in a named registry and invoked directly, so routine experiments run safely without generating new code.","The distributed architecture allows microscopes to be operated remotely and coordinated through a central server, with instruments polling for jobs rather than being locally scripted.","Human approval gates combined with digital-twin syntactic checks place a safety boundary between AI-generated code and physical hardware.","Unified interfaces across multiple microscope platforms make heterogeneous instruments accessible and composable through a single agent-facing toolset."],"fun_headline_variants":["AI agents get hands-on with microscopes, design new experiments","Self-driving microscopy: AI agents propose, humans approve","Microscopy platform lets AI agents invent experiments with oversight","Agentic AI takes the wheel in materials characterization"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that an LLM-generated script that passes digital-twin error checks and human review is safe and reliable enough to run on a real microscope, which the paper does not quantitatively test.","fun_headline_variants_meta":{"raw":{"variants":["AI agents get hands-on with microscopes, design new experiments","Self-driving microscopy: AI agents propose, humans approve","Microscopy platform lets AI agents invent experiments with oversight","Agentic AI takes the wheel in materials characterization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001076,"raw_usage":{"total_tokens":4298,"prompt_tokens":660,"completion_tokens":3638,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":404,"completion_tokens_details":{"reasoning_tokens":3573}},"tokens_in":404,"tokens_out":3638,"duration_ms":23833,"temperature":1.0,"reasoning_tokens":3573,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T03:57:57.334506+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a defined benchmark where LLM-generated experimental scripts are executed on a real microscope against a known ground-truth sample, measuring success rate, hardware incidents, and deviation from the intended protocol; if success is not clearly better than random or if safety incidents occur despite validation, the central claim is undermined.","supporting_citations":[],"review_version":1}