{"id":"86499a1d-084c-45ed-808e-b4f99200797e","arxiv_id":"2608.09104","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":11,"one_line_summary":"A closed-loop multitask Gaussian-process controller on an atomic force microscope selects both measurement location and protocol, transferring information between tapping-mode and DART roughness across an AlScN wafer.","lead":"Scanning probe microscopes normally decide where to measure; this paper also lets them decide which measurement mode to use, guided by a multi-task statistical model. The live system was tested on an AlScN wafer and reconstructed two roughness maps without measuring both modes at the same locations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-task information transfer rests on rho=0.752 estimated from only five paired seeds, with no held-out comparison to single-task GPs; if rho is overestimated, the reported uncertainty reduction does not demonstrate a multitask benefit.","rationale":"The reader's weakest assumption is that the cross-task correlation from five paired seeds is the linchpin. I agree: the only quantitative evidence of information transfer is rho=0.752 and the model-derived uncertainty reduction. Since the paper itself notes the map agreement is partly construction-imposed and task counts are chance-level, the transfer claim is the one thing that must survive. The concrete test I propose directly compares multitask versus independent-task predictions on held-out observations, which would settle whether transfer occurs. If the multitask model does not beat the single-task baseline, the central claim is unsupported and the verdict should move to REJECT; if it does, the conditional is satisfied. The test uses existing data and released code, so it is inexpensive. I do not see an internal inconsistency in the GP formulation, and the engineering demonstration of live closed-loop modality switching is credible. The concern is purely about the strength of the evidence for the cross-task transfer, which is exactly the reader's weakest assumption.","tokens_in":10771,"tokens_out":5317,"duration_ms":56051,"concrete_test":"Using the released multitask-spm code, perform leave-one-out cross-validation on the 19 DART observations: for each held-out point, train the ICM on all 16 tapping observations plus the remaining 18 DART observations, and record the predictive mean and variance at the held-out location. Repeat with a single-task GP trained only on the 18 DART observations. If the multitask model's RMSE and negative log predictive density are not lower than the single-task baseline, the cross-task transfer is not demonstrated. As a sensitivity check, refit the model with each of the five seed pairs removed and record rho; if any removal moves rho below 0.3, the five-seed estimate is not stable enough to support the headline claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Sec. III.E the authors report a learned cross-task correlation rho=B12/(B11 B22)^{1/2}=0.752 and state that the five co-located seed pairs are the primary constraint on rho, with non-overlapping designs recovering it only slowly (Ref. 34). This rho is the mechanism that lets a tapping measurement update the DART posterior and vice versa; the claimed ~80% uncertainty reduction in both landscapes (Sec. III.F) is computed from the fitted model and therefore inherits whatever error is in rho. If the true rho is much smaller than 0.752 -- possible with only five paired measurements, especially given the low measured-versus-predicted scatter r=0.304 that the authors attribute to noise -- then the transfer contribution vanishes and the experiment reduces to two independent GP reconstructions with no multitask advantage. The paper provides no held-out comparison against single-task GPs trained on the same per-task observations, so the central claim that 'the experiment can construct two response landscapes without acquiring both measurement modes at every position' is not yet independently supported. The concern is not that the model is misspecified internally; it is that the single number carrying the transfer claim is identifiable from almost no paired data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces multitask scanning probe microscopy, a closed-loop active-learning workflow in which a multitask Gaussian process with an intrinsic coregionalization model selects both the next measurement location and the next SPM protocol on an operating microscope. The method is demonstrated on a composition-spread AlScN wafer using tapping-mode and DART roughness as two tasks. Five paired seed measurements initialize the task-covariance matrix, after which 25 active-learning scans (17 model-selected, 8 forced-exploration) are executed with one mode per location. The authors report a learned cross-task correlation ρ=0.752, posterior mean maps that co-vary strongly (r=0.984), and an approximately 80% reduction in posterior standard deviation for both landscapes. The paper is transparent about several model-internal aspects of these numbers, noting in the Fig. 3c caption that the map agreement is partly imposed by the shared kernel and should not be read as independent validation.","tokens_in":11120,"tokens_out":3374,"duration_ms":39292,"significance":"If the central claim is supported—that measurements in one SPM mode can transfer information to another mode so that two response landscapes are reconstructed without measuring both modes everywhere—this would be a useful extension of autonomous SPM from spatial sampling to joint location-and-modality allocation, with clear relevance to wafer-scale and combinatorial characterization. The live execution of the loop (35 real scans, automated mode switching, stage motion) and the public availability of the code are concrete strengths. The paper is also unusually candid about the limitations of its own quantitative evidence, explicitly flagging that the posterior scatter is imposed by construction and that the five paired seeds are the primary constraint on ρ. The remaining weakness is that the load-bearing numbers—ρ=0.752, the 80% uncertainty reduction, and the cross-map agreement—are internal to the fitted model and are not checked against any held-out per-task ground truth or against an independent single-task baseline.","major_comments":[{"comment":"The central evidence for information transfer is model-internal. The learned ρ=0.752 is identified primarily from five co-located seed pairs, while the raw measured-versus-predicted cross-task scatter in Fig. 3b has r=0.304, and the posterior-mean scatter r=0.984 in Fig. 3c is, as the authors note, partly imposed by the shared spatial kernel. To support the claim in Sec. III.F that the experiment reconstructs two landscapes without measuring both modes at every position, the paper needs a held-out comparison against a single-task GP (or a diagonal-B ICM) trained on the same per-task observations. Reporting predictive RMSE and log-likelihood on unmeasured locations, or a leave-one-out analysis over the 35 scans, would provide the missing independent check.","section":"Sec. III.E, Fig. 3"},{"comment":"The reported ~80% uncertainty reduction is the decrease in the model's own posterior standard deviation (0.176 nm and 0.235 nm versus prior values of 0.925 nm and 1.098 nm). This quantity depends directly on the fitted task covariance and does not measure actual prediction error. Because an overestimated ρ would produce overconfident transfer, the authors should report actual predictive error at held-out locations—ideally including a small set of paired measurements acquired after the loop—and show the same metric for the independent-task baseline. Without this, the headline uncertainty reduction does not by itself demonstrate a multitask advantage.","section":"Sec. III.F, Fig. 4f"},{"comment":"The uncertainty in the learned cross-task correlation is not quantified. With only five paired seed measurements constraining ρ, the point estimate 0.752 carries large uncertainty. The paper should report a confidence interval or posterior distribution for ρ (e.g., from a bootstrap over the five seed pairs or from the posterior over B), and should include a sensitivity analysis showing how the reconstructed maps and acquisition decisions change when ρ is fixed to smaller values such as 0.3 or 0.5. This is needed to establish that the transfer mechanism is identifiable from the data actually collected.","section":"Sec. III.D and III.E"}],"minor_comments":[{"comment":"The annealing schedule β_eff = β0/(1+0.15 n_step) should specify whether n_step counts only model-driven active steps or also the forced-exploration steps; this affects reproducibility of the acquisition trajectory.","section":"Sec. II.D, Eq. (12)"},{"comment":"The thin-plate-spline height reference from 17 locations is a practical necessity; reporting the typical residual or error of this interpolant against a few validation points would help the reader judge whether height misestimation could bias the roughness measurements.","section":"Sec. III.A"},{"comment":"The roughness descriptor excludes pixels more than five standard deviations from the image mean before computing the standard deviation; for small images or high outlier densities this exclusion changes the estimator. State the image size and the typical number of excluded pixels.","section":"Sec. III.B"},{"comment":"The claim that co-located measurements improve identifiability of the cross-task correlation and that non-overlapping designs recover it only slowly is central to the seed strategy, but Ref. 34 is an arXiv preprint. If the journal requires archival sources or if the claim is not yet peer-reviewed, the relevant analysis should be summarized in the text or an appendix.","section":"Sec. III.E, Ref. 34"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope as a methods demonstration in autonomous scanning probe microscopy. The main concern is not the live loop, which is a real experimental accomplishment, but the quantitative support for the central claim of information transfer. The authors are already aware of the model-internal nature of the evidence, and the requested comparisons against a single-task baseline and a ρ sensitivity analysis are targeted at that gap. I would also check whether Ref. 34, which carries a load-bearing identifiability claim, is published or still a preprint at the time of review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the live loop: a multitask GP actually steering an operating AFM, switching between tapping and DART modes on a wafer-scale grid, with stage motion, mode switching, and scanning all automated. I believe that. The paper is also unusually honest about its own limits—it tells you the posterior-map agreement is partly imposed by construction, that the task counts are chance-level, and that the cross-task correlation rests on five paired seeds. That candor earns credit.\n\nWhat the paper does well: the integration is real, the experimental execution is credible, and the framing of the problem—choosing both where and which protocol—is a fair step beyond prior channel-selection work on pre-acquired data. The ICM machinery is standard, and the authors say so. The writing is clear about what is new (the live closed loop) and what is not (the algorithmic ingredients).\n\nThe soft spot is the one the stress-test note names, and it is load-bearing: the claimed multitask benefit lives in the learned cross-task correlation ρ=0.752, which is primarily pinned down by five paired seed measurements. The measured-versus-predicted cross-task scatter r=0.304 is low enough to make that estimate fragile. The ~80% uncertainty reduction is computed from the fitted joint model, so it inherits whatever error is in ρ. Without a held-out comparison against single-task GPs trained on the same per-task observations, the paper does not actually demonstrate that the transfer helps. A reader can accept that the loop works and still doubt that the multitask part beats a simpler two-independent-GP baseline. The paper does flag the identifiability issue and cites the relevant work (Ref. 34), which makes this a fixable weakness rather than a hidden one.\n\nMinor: the forced-exploration steps are random task assignments, so only 17 of 25 active selections are model-driven; that is fine, but it dilutes the demonstration somewhat. The choice of roughness as the response is a controlled, low-risk demo; the interesting extensions (topography guiding current or spectroscopy) are speculation, not results.\n\nWho this is for: people building autonomous SPM workflows and anyone doing multitask Bayesian optimization on physical instruments. It deserves a serious referee, but the referee should push for a baseline comparison and a sensitivity analysis on the seed positions before publication. If those hold up, this is a useful experimental milestone.","headline":"Genuinely live closed-loop multimodal AFM, but the info-transfer claim leans on five paired seeds and a fitted-model statistic; worth refereeing after a held-out check.","tokens_in":11612,"tokens_out":616,"would_cite":true,"duration_ms":8567,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper demonstrates a live, closed-loop multitask scanning probe microscope in which a multitask Gaussian process with a learned task-covariance matrix decides both the next measurement location and the next imaging protocol, so that a…","keywords":["multitask Gaussian process","intrinsic coregionalization model","autonomous scanning probe microscopy","active learning","measurement modality selection","composition-spread wafer","DART","tapping mode AFM"],"falsifier":"Repeat the experiment with several different sets of paired seed positions (or with 10-20 paired seeds instead of five) and compare the learned $\\rho$ and the final posterior maps against a dense ground-truth map of both modes; if $\\rho$ swings widely across seed sets or the joint model predicts held-out co-located pairs no better than two independent Gaussian processes, the central transfer claim fails.","tokens_in":10558,"feed_emoji":"🔬","tokens_out":6962,"duration_ms":70331,"temperature":0.7,"pith_summary":"Scanning probe microscopy usually maps one property at a time, and multimodal studies either repeat every mode at every grid point or fix the protocol in advance. This paper claims that a microscope can instead treat each measurement mode as a related prediction task and, during the experiment, choose both where and how to measure next. The load-bearing idea is that a multitask Gaussian process can learn the spatial structure of each mode together with a cross-modal covariance, so a measurement in one mode updates the predicted map of the other mode across the entire wafer. The authors demonstrate this on a composition-spread AlScN wafer with tapping-mode and DART roughness as two tasks, using five paired seed measurements to initialize the relationship and 25 subsequent single-mode scans. If the claim holds, wafer-scale multimodal screening can be done with far fewer slow or damaging measurements.","feed_headline":"AFM learns two wafer maps from one scan per site","feed_subtitle":"A multitask model transfers information between imaging protocols, so wafer-wide maps need far fewer slow, damaging measurements.","key_machinery":"The central object is the intrinsic coregionalization model (ICM), a multitask Gaussian-process construction in which the covariance between two location-task pairs factorizes as $K[(x,t),(x',t')] = K_x(x,x')B_{tt'}$, with $B = WW^T + \\mathrm{diag}(\\kappa)$ the task-covariance matrix. $B$ is what transfers information: a measurement of task $t$ at position $x$ updates the posterior for task $t'$ at all positions through the learned cross-task entry $B_{tt'}$. The implementation pairs this with a random-scalarization upper-confidence-bound acquisition function that samples Dirichlet task weights and selects the location-task pair maximizing the weighted UCB score, plus a forced-exploration step every third iteration that picks the farthest unmeasured position with a random task. Paired seed measurements at the start provide the primary constraint on $B$.","core_discovery":"On the paper's own terms, the central discovery is that a multitask Gaussian process with an intrinsic coregionalization model can run live on an operating atomic force microscope and jointly reconstruct two response landscapes from non-coincident observations. The experiment acquired five paired seed measurements at randomly chosen positions, learned a task-covariance matrix whose normalized off-diagonal entry is $\\rho = 0.752$, and then let a random-scalarization upper-confidence-bound acquisition function pick one location--task pair at a time. After 35 total scans (16 tapping, 19 DART), the joint model reduced the grid-averaged posterior standard deviation from roughly 0.925 nm to 0.176 nm for tapping and from 1.098 nm to 0.235 nm for DART, while only one modality was measured at each non-seed location. The paper is careful to note that this uncertainty reduction is a property of the fitted model and that the high posterior correlation between the two maps ($r=0.984$) is partly imposed by the shared spatial kernel.","pith_inferences":["Editorial inference: for modality pairs that are only weakly correlated, five paired seed measurements may be too few to pin down $B$; a practical rule would be to acquire more paired seeds when the expected cross-task correlation is low or the per-task noise is high.","Editorial inference: the same multitask loop could be applied at the level of detector channels or scalar descriptors extracted from spectra (coercive voltage, loop area), turning any multi-observation SPM protocol into an active-learning task space.","Editorial inference: a direct test of the transfer benefit would compare the joint model against two independent GPs on held-out co-located pairs; the paper reports a cross-task scatter of $r=0.304$ against the learned $\\rho=0.752$, so the gap between raw and model-level correlation is where the transfer claim should be validated.","Editorial inference: the cost-aware formulation suggests a natural simulation benchmark: run the closed-loop policy on a pretrained multimodal dataset and compare total information gain per unit 'cost' against fixed-ratio mode allocation."],"forward_implications":["A single measurement at each location can produce posterior maps for two protocols, eliminating the need to acquire both modes at every grid point.","Rapid, weakly perturbative modes can guide the selective deployment of slower contact or spectroscopic measurements, reducing tip wear and sample modification.","Every single-mode measurement improves the predicted landscape of the unmeasured mode, so knowledge accumulates across the whole wafer even when that mode is never run at most positions.","With the cost-aware acquisition function of Eq. 13, the controller can explicitly trade expected information gain against acquisition time, damage, and mode-switching overhead."],"supporting_citations":[{"why":"Introduces the multitask Gaussian process and the intrinsic coregionalization model used as the joint covariance structure.","marker":"[26]"},{"why":"Supplies the broader multitask GP formulation and the linear-model alternatives the paper extends.","marker":"[27]"},{"why":"Documents the identifiability pitfall that motivates the paired-seed initialization; underpins the claim that co-located pairs are needed to learn the cross-task correlation.","marker":"[34]"},{"why":"Provides the ParEGO random scalarization idea on which the acquisition function is built.","marker":"[42]"},{"why":"Supplies the Dirichlet-weighted random scalarization for multi-objective Bayesian optimization used to couple task and spatial selection.","marker":"[43]"},{"why":"Defines the upper-confidence-bound acquisition criterion that scores candidate location-task pairs.","marker":"[44]"},{"why":"Provides the farthest-point maximin criterion used for forced exploration steps.","marker":"[45]"},{"why":"Defines the Dual AC Resonance Tracking (DART) mode used as one of the two measurement tasks.","marker":"[30]"},{"why":"Describes the automated experiment software environment used to close the loop on the microscope.","marker":"[23]"}],"fun_headline_variants":["AI-driven AFM maps two wafer properties with fewer scans","Autonomous AFM picks spots and modes to map wafers faster","Multitask AFM learns cross-modal links to halve required scans","Closed-loop AFM shares data between modes to speed imaging","AFM AI turns one scan into two wafer maps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole information-transfer argument rests on the cross-task correlation being learned reliably from just five paired seed measurements; if those five pairs are unrepresentative or dominated by noise, the transferred predictions and the reported uncertainty reduction would not hold.","fun_headline_variants_meta":{"raw":{"variants":["AI-driven AFM maps two wafer properties with fewer scans","Autonomous AFM picks spots and modes to map wafers faster","Multitask AFM learns cross-modal links to halve required scans","Closed-loop AFM shares data between modes to speed imaging","AFM AI turns one scan into two wafer maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001081,"raw_usage":{"total_tokens":4524,"prompt_tokens":953,"completion_tokens":3571,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":3485}},"tokens_in":569,"tokens_out":3571,"duration_ms":27135,"temperature":1.0,"reasoning_tokens":3485,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:28:17.611918+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the experiment with several different sets of paired seed positions (or with 10-20 paired seeds instead of five) and compare the learned $\\rho$ and the final posterior maps against a dense ground-truth map of both modes; if $\\rho$ swings widely across seed sets or the joint model predicts held-out co-located pairs no better than two independent Gaussian processes, the central transfer claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the multitask Gaussian process and the intrinsic coregionalization model used as the joint covariance structure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the broader multitask GP formulation and the linear-model alternatives the paper extends."},{"cited_title":"Pitfalls and Remedies for Multi-Task Bayesian Optimization","cited_arxiv_id":"2607.09073","evidence_quote":"Documents the identifiability pitfall that motivates the paired-seed initialization; underpins the claim that co-located pairs are needed to learn the cross-task correlation."},{"cited_title":"Knowles, IEEE Transactions on Evolutionary Computation 10 (1), 50-66 (2006)","cited_arxiv_id":null,"evidence_quote":"Provides the ParEGO random scalarization idea on which the acquisition function is built."},{"cited_title":"Paria, K","cited_arxiv_id":null,"evidence_quote":"Supplies the Dirichlet-weighted random scalarization for multi-objective Bayesian optimization used to couple task and spatial selection."},{"cited_title":"Srinivas, A","cited_arxiv_id":null,"evidence_quote":"Defines the upper-confidence-bound acquisition criterion that scores candidate location-task pairs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the farthest-point maximin criterion used for forced exploration steps."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Dual AC Resonance Tracking (DART) mode used as one of the two measurement tasks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the automated experiment software environment used to close the loop on the microscope."}],"review_version":1}