{"id":"ab5b9641-ccc3-42b0-b3dc-c81bd4da2fb3","arxiv_id":"2506.08073","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A multi-objective active-learning microscope workflow maps ferroelectric domain-switching Pareto frontiers and predicts switching ease from structural images alone.","lead":"This paper trains a machine-learning system that guides an atomic-force microscope to probe ferroelectric switching at the most informative locations. It maps trade-offs between switching size, symmetry, and local structure in hours instead of months.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central generalization claim rests on in-sample validation: Fig. 5 compares predictions with the same active-learning measurements used to train the model, so reliable prediction at unmeasured locations is not yet demonstrated.","rationale":"The stress-test pass found one genuinely load-bearing gap rather than a fatal flaw. The reader's conditional verdict already hinges on the same issue: the mapping learned from 16x16 patches must transfer to unmeasured locations, and the paper does not test this with held-out data. I agree with that framing. Independent support includes the open-source AESPM interface, a simulation notebook, and the null-control PZTO experiment; these count as real evidence but do not address predictive generalization. The paper's own conclusion admits simpler alternatives were not benchmarked, which is an honest limitation but secondary to the validation gap. A simple held-out or fresh-validation experiment would settle the concern; no correction to theory or code appears necessary. Verdict remains CONDITIONAL; I would not reject the paper.","tokens_in":14215,"tokens_out":4085,"duration_ms":50937,"concrete_test":"Add a validation phase: before the active loop, designate e.g. every 5th patch or a contiguous stripe as held-out, forbid acquisition from these patches, and after training on the remaining measurements predict poled-domain size for the held-out patches and compare against actual post-pulse measurements. Report R^2 and Spearman correlation overall and separately for under-sampled classes (domain interiors, in-plane domains), plus a Fig. 5e-style distance trend computed only from held-out points. If held-out accuracy is near chance or the distance trend differs from the in-sample trend, the claim that sparse measurements can replace dense grids is not supported. A retrospective temporal split (train on early steps, test on later steps) can serve as a weaker first check using existing data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MOBO-DKL can reliably predict complex material responses based solely on structural information from global maps. This is what would justify sparse active learning replacing dense grids. The main quantitative support, Fig. 5e, overlays predicted poled-domain size versus distance to nearest boundary with measured poled-domain sizes; however, those measured values come from the 210 active-learning measurements used to train the model. Predictions for the global map in Figs. 3 and 4 are not compared against withheld ground truth. In addition, qEHVI deliberately oversamples Pareto-relevant boundary and positive-domain locations, so in-sample agreement can be high even if the learned mapping is poor in under-sampled interior and in-plane regions. Because the global map is reacquired every 10 steps, measured patches are partly non-stationary; without a held-out spatial or temporal split, Fig. 5e cannot distinguish genuine transfer from interpolation among training points. This is not an internal inconsistency, but a missing validation step for the paper's central claim. The PZTO null experiment and consistency with known ferroelectric physics provide supportive but not quantitative evidence for out-of-sample predictive accuracy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a multi-objective Bayesian optimization with deep kernel learning (MOBO-DKL) workflow for automated piezoresponse force microscopy (PFM). A convolutional neural network extracts features from 16×16 structural patches of a global PFM map, and separate Gaussian processes map those features to three reward objectives: average piezoresponse, poled-domain size, and poled-domain symmetry. A joint qEHVI acquisition function selects the next measurement location, enabling fully automated active learning. The workflow is demonstrated on two ferroelectric films: PTKTO, where switching behavior depends strongly on local domain structure, and PZTO, a more uniform control sample. The paper reports that the model predicts poled-domain size as a function of distance to the nearest domain boundary, that exploration preferentially targets domain boundaries and positive domains, and that the PZTO null experiment shows random exploration when no structure–property correlation exists. The authors conclude that MOBO-DKL reliably predicts complex material responses from structural maps and efficiently maps Pareto fronts of switching behavior.","tokens_in":14385,"tokens_out":4338,"duration_ms":53719,"significance":"The paper addresses an important practical problem: reducing the number of expensive and destructive PFM measurements needed to map local structure–property relationships in ferroelectrics. The strengths include a fully automated closed-loop experiment that ran 210 steps in about 7 hours, an open-source instrument-control interface (AESPM), a null experiment on uniform PZTO that validates the algorithm's behavior when no correlation exists, and physically plausible mechanistic interpretations (e.g., switching ease ranking boundaries > positive domains > negative domains > in-plane domains). The release of code and data is commendable. However, the central claim of reliable prediction at unmeasured locations rests on the agreement in Fig. 5e, which is an in-sample comparison; the significance of the work would be substantially higher if out-of-sample predictive accuracy were demonstrated.","major_comments":[{"comment":"The validation of predicted versus measured poled-domain size in Fig. 5e uses the same 210 active-learning measurements that were used to train the MOBO-DKL model. This is an in-sample comparison; it cannot establish the central claim that the model reliably predicts responses at unmeasured locations. The authors should provide a held-out spatial or temporal split (e.g., excluding a random subset of the measured locations from training and evaluating predictions on them) and report out-of-sample predictive error. Without this, the agreement in Fig. 5e may reflect interpolation among training points rather than generalization.","section":"II (Figure 5e)"},{"comment":"The claimed efficiency gain of 'approximately 7 hours for 210 total steps' versus 'an estimated three months' for a full-grid measurement is presented without any derivation of the three-month estimate. Since this efficiency claim motivates the entire active-learning approach, the authors should provide a transparent estimate: the number of grid points in a full 256×256 measurement, the per-measurement time (including the pre-pulse and post-pulse scans), and any assumptions about duty cycle. The statement should also acknowledge that a full-grid experiment would be destructive and non-stationary, so the comparison is conceptual rather than a direct time saving.","section":"II (paragraph on experimental time)"},{"comment":"The paper asserts that MOBO-DKL 'efficiently maps the Pareto front' and 'captures the structure–property relationships across multiple reward metrics,' but it provides no quantitative comparison to baselines such as random sampling, single-objective DKL, or a standard Gaussian process with the same CNN features. Without such a comparison, the specific benefit of the multi-objective active-learning strategy over simpler exploration policies is not demonstrated. A simulation using the released data or a post-hoc comparison of the acquired trajectory against random exploration would strengthen this claim.","section":"III (Conclusion)"}],"minor_comments":[{"comment":"In Reward 1, the notation \\(\\exp(-\\vec r/w)\\) is dimensionally inconsistent; it should be \\(\\exp(-|\\vec r|/w)\\), where \\(|\\vec r|\\) is the distance from the pulse center.","section":"II (Reward definitions)"},{"comment":"The description of Reward 3 as 'how far the poled domain has deviated from a perfect cycle' is opposite to the formula \\(\\overline{r_i^0}/\\sigma(r_i^0)\\), which is large for a circularly symmetric domain and small for a distorted one. This should be reworded as a measure of closeness to a perfect circle.","section":"II (Reward definitions)"},{"comment":"The text states 'we release our datasets openly,' but no dataset repository or DOI is provided; only the code notebook and AESPM link are given. Please add a link to the experimental datasets.","section":"III (Data availability)"},{"comment":"The comparison in Fig. 5e would be clearer if the number of measured points, their error bars, and whether any points were excluded in the linecut construction were reported.","section":"Figure 5e"}],"recommendation":"major_revision","confidential_remarks":"The main technical concern is the in-sample validation of the central prediction claim. The paper is otherwise well-structured and provides reproducible code and a valuable null experiment. I recommend major revision rather than rejection because the core workflow is sound and the missing held-out validation can be added with a relatively focused set of experiments or post-hoc analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zoe, here's my read on the MOBO-DKL PFM paper. What's genuinely new: they've wired a multi-objective Bayesian optimization loop (DKL + qEHVI) into a fully automated PFM system and run it on two real materials, PTKTO and PZTO. The PZTO control is good—same algorithm on a uniform sample behaves like random exploration, which tells you the method isn't just clustering structure blindly. The Pareto scatter plots and the predicted distribution maps are also a nice way to visualize the structure–switching relationships, and the physical trends they extract (boundaries switch easiest, positive domains easier than negative, in-plane domains hardest) are consistent with ferroelectric intuition. So as an engineering demonstration, this is solid.\n\nThe soft spot is exactly what the stress-test note says: the quantitative validation in Fig. 5e is in-sample. The measured poled-domain sizes used there are the same 210 active-learning measurements the model was trained on. qEHVI deliberately concentrates queries near boundaries, so interpolation among training points will look great. To support the claim that the model \"can reliably predict complex material responses based solely on structural information,\" you need a held-out spatial split or a temporally separate test set. The paper does not show that.\n\nTwo other things, both acknowledged or minor. First, the authors state in the conclusion that they didn't compare against simpler dimensionality-reduction baselines like PCA + multi-output GP. For a paper whose pitch is \"deep kernel learning\" architecture, that's a real omission, though they're upfront about it. Second, they say \"we release our datasets openly\" but I don't see a link or dataset DOI in the manuscript—just the notebook for a simulated experiment. The efficiency claim (7 hours vs. 3 months) is an estimate; fine as context, but not a measured baseline.\n\nMy overall take: this is a capable group doing careful work, and the paper is honest about several limitations. The missing out-of-sample validation is the one load-bearing gap. It's fixable—retrain on 80% of locations, test on 20%, or run a small withheld grid at the end. Once that's done, the central claim would be properly supported.\n\nI'd send it to peer review. It's a good candidate for a materials-informatics or instrumentation journal, and the referees should push for the held-out test and a baseline. I'd also bring it to our reading group—it's a useful case study in how autonomous-experiment papers should (and shouldn't) validate their predictions.","headline":"A well-executed autonomous PFM demonstration whose headline generalization claim still needs a proper held-out test; worth a serious referee, but the in-sample validation makes me hold off on full endorsement.","tokens_in":686,"tokens_out":1422,"would_cite":true,"duration_ms":31821,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper seeks to establish that a multi-objective Bayesian optimization workflow with deep kernel learning can drive automated piezoresponse force microscopy, selecting measurement sites so that the Pareto front of ferroelectric…","keywords":["multi-objective Bayesian optimization","deep kernel learning","piezoresponse force microscopy","ferroelectric domain switching","active learning","Pareto front","structure-property relationships","automated experimentation"],"falsifier":"Take the set of already-measured locations, train MOBO-DKL on a random 80% of them, and predict the rewards of the remaining 20%; if held-out poled-domain-size predictions do not match measurements to roughly the same accuracy as the linecut agreement shown in the paper, the claim that structure alone determines switching is falsified. A second check would compare measured poled-domain size at the same nominal structure before and after a neighboring pulse, to see whether pulse history changes the outcome beyond the patch content.","tokens_in":13983,"feed_emoji":"⚡","tokens_out":8130,"duration_ms":94068,"temperature":0.7,"pith_summary":"The paper seeks to establish that a multi-objective active-learning workflow, Bayesian optimization with deep kernel learning, can replace dense grid measurements in piezoresponse force microscopy and still recover the structure–property rules behind ferroelectric domain switching. The method crops a global piezoresponse map into small patches, measures three rewards at selected patch centers (local piezoresponse, switched-domain size, and switched-domain symmetry), and trains a convolutional feature extractor together with one Gaussian process per reward so that the next measurement site is chosen to expand the Pareto front. If the central claim is right, experiments that would take months of grid scanning can be compressed into hours of automated measurement, with the added benefit that the measurement is less destructive to the very structure being studied.","feed_headline":"Active learning maps ferroelectric switching in 7 hours, not 3 months","feed_subtitle":"An automated microscope picks the few spots that reveal how domain walls control switching, and predictions match measured poling.","key_machinery":"The load-bearing machinery is the joint CNN-GP deep kernel model, with one convolutional feature extractor and one Gaussian process per reward, trained jointly so that image features align with measured rewards, together with the q-Expected Hypervolume Improvement (qEHVI) acquisition function, which selects the next patch by the expected gain in the volume of the Pareto front. The unit of structure is the 16 by 16 patch cropped from the global piezoresponse map, and rewards are measured at patch centers, so the argument rests on that patch being a sufficient description of the local switching environment. The loop is closed by reacquiring the global map every ten steps to correct for drift and measurement-induced changes.","core_discovery":"The central discovery claimed is that multi-objective deep kernel learning (MOBO-DKL) learns the relationship between local domain-wall configurations and switching behavior well enough to predict, from the before-pulse image alone, how large and how symmetric a poled domain will be after a voltage pulse. In the PTKTO film, the model's predicted poled-domain size as a function of distance to the nearest domain boundary agrees with directly measured values in magnitude and spatial trend, and the predicted maps reproduce the observed progression from boundary shifting to large, symmetric switched domains. The paper also claims a mechanistic ranking of switching ease: domain boundaries switch most easily, followed by positive domains, negative domains, and in-plane domains, consistent with manual segmentation. A control experiment on a nearly uniform PZTO film produces a random-looking exploration trajectory, which the authors interpret as evidence that the algorithm does not invent structure-dependence where none exists.","pith_inferences":["Not claimed in the paper: the same loop should transfer to any setting where a structural image is cheap and the functional measurement is expensive or destructive, since the only requirements are cropped patches and measurable rewards; this could be tested in electron microscopy or molecular discovery.","The patch-based assumption suggests a testable extension: adding pulse history or tip-state features as auxiliary inputs and checking whether held-out prediction improves would reveal whether sequential-poling memory matters.","A natural next step, flagged by the authors, is to replace the scalar symmetry reward with the orientation of the poled domain's principal axis, converting the Pareto map into a directional map that encodes crystallographic anisotropy.","Because the model outputs reward distributions over every patch of the global map, the same workflow could flag rare microstructures whose predicted rewards lie far from the Pareto front, i.e., candidates for unusual switching physics, before any additional measurement is taken."],"forward_implications":["If the central claim holds, a 210-step automated run (about seven hours) can substitute for an estimated three months of full-grid measurements while still mapping the Pareto front of switching behavior.","The trained model's predicted poled-domain size versus distance-to-boundary curves match measured values, so quantitative structure–property trends can be extracted from structural maps alone.","The exploration trajectory and predicted reward distributions sort local structures by switching ease, giving an interpretable mechanistic ranking: domain boundaries, positive domains, negative domains, then in-plane domains.","The averaged-piezoresponse reward behaves as a surrogate for distance to the nearest domain boundary, so an abstract reward can double as a physical descriptor that guides exploration toward informative structures.","On uniformly structured samples, the acquisition becomes effectively random, providing a null-result check that the model exploits structure only when structure actually matters."],"supporting_citations":[{"why":"Establishes the deep kernel learning construction on which the workflow is built: a neural feature extractor feeding a Gaussian process.","marker":"46"},{"why":"Prior automated scanning-probe experiments using single-objective DKL; this paper extends that line to multiple objectives.","marker":"22-24"},{"why":"Supplies the qEHVI acquisition function used to select each next measurement location.","marker":"43"},{"why":"Defines expected hypervolume improvement, the multi-objective selection criterion at the core of the active-learning loop.","marker":"48,49"},{"why":"Provides the closely related automated-AFM study of ferroelectric domain wall pinning that frames the experimental challenge.","marker":"19"},{"why":"Explains the in-plane domain configuration physics used to interpret the different switching of the two stripe domains in PZTO.","marker":"37,38"}],"fun_headline_variants":["AI learns switching rules from PFM images, no manual scans","Multi-objective model predicts poling from one domain snapshot","Automated PFM + deep kernel learning cuts experiments from months to hours","Deep learning maps ferroelectric switching kinetics directly from imaging data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the switching outcome at any location is fully determined by the local pre-pulse structure within a 16 by 16 patch, so rewards measured at patch centers transfer to unmeasured locations; the paper does not test this premise with a held-out split.","fun_headline_variants_meta":{"raw":{"variants":["AI learns switching rules from PFM images, no manual scans","Multi-objective model predicts poling from one domain snapshot","Automated PFM + deep kernel learning cuts experiments from months to hours","Deep learning maps ferroelectric switching kinetics directly from imaging data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1229,"prompt_tokens":915,"completion_tokens":314,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":242}},"tokens_in":531,"tokens_out":314,"duration_ms":5011,"temperature":1.0,"reasoning_tokens":242,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:20:22.188603+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the set of already-measured locations, train MOBO-DKL on a random 80% of them, and predict the rewards of the remaining 20%; if held-out poled-domain-size predictions do not match measurements to roughly the same accuracy as the linecut agreement shown in the paper, the claim that structure alone determines switching is falsified. A second check would compare measured poled-domain size at the same nominal structure before and after a neighboring pulse, to see whether pulse history changes the outcome beyond the patch content.","supporting_citations":[],"review_version":1}