{"id":"c0425a84-7321-46c2-b3a8-db4ac3cf4af5","arxiv_id":"2502.03572","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A multi-branch CNN fed with the trEFM frequency trace and cantilever parameters extracts bi-exponential kinetic parameters (τ1, τ2, A) more accurately and noise-robustly than the prior single-exponential neural network.","lead":"The authors trained a multi-output convolutional neural network to extract two time constants and a mixing coefficient that describe bi-exponential dynamics from time-resolved electrostatic force microscopy signals. The network also takes cantilever properties as inputs and reconstructs experimental data more accurately than the previous single-exponential model, which matters for mapping fast nanoscale electronic and ionic processes in perovskite solar cell materials.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The real-data claim rests on voltage-pulse tests from (likely) a single fine-tuning cantilever and on reconstruction R2 computed with the same forward model, so cross-cantilever generalization and model-transfer accuracy are not yet demonstrated.","rationale":"The reader's weakest assumption was the accuracy of the FFTA forward model. I agree that the forward model is load-bearing, because it generates all simulated training data and is also used to compute reconstruction R2. However, the forward model alone is not the sharpest concern: the voltage-pulse labeled data provide an external check of the forward model for one experimental configuration, and the network is fine-tuned on those labels. The sharper gap is that the labeled experimental validation appears to come from a single cantilever/tip/session, and the paper explicitly avoids multi-cantilever labeled collection because it is too time-consuming. Since the architecture includes k, Q, and omega as inputs specifically to generalize across cantilevers, the central claim of accurate extraction from real trEFM data requires a held-out test on a different cantilever. Without that, the reported parity R2 on voltage-pulse data may reflect fine-tuning to a specific tip state, and the perovskite reconstruction R2 cannot falsify forward-model bias because it uses the same model that produced the training data. This does not overturn the verdict; it reinforces the CONDITIONAL status. The paper has real strengths: the voltage-pulse experiment provides labeled experimental data, the code and data are publicly available, and the SHAP analysis gives interpretability evidence. I therefore recommend keeping the reader's conditional verdict rather than moving to accept or reject. The concrete cross-cantilever test would settle whether the concern lands and would directly inform how the claim should be stated.","tokens_in":16132,"tokens_out":4656,"duration_ms":46822,"concrete_test":"Fine-tune the CNN only on voltage-pulse traces from cantilever A, then evaluate on held-out voltage-pulse traces collected with cantilevers B and C having different k, Q, and omega values (ideally different tips and days). Report parity R2 and MAE for tau1, tau2, A, and <tau> for the cross-cantilever test set. If the <tau> parity R2 drops materially below the reported same-cantilever value of 0.97, the generalization claim fails and the headline should be restricted to per-cantilever calibration. As a complementary check, simulate delta-omega(t) with a deliberately perturbed forward model (e.g., altered Q dependence or added tip-sample capacitance nonlinearity) and test whether the FFTA-trained network's parameter bias is reflected in reconstruction R2; this would show whether high reconstruction R2 actually certifies parameter accuracy when the forward model is misspecified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the CNN accurately extracts tau1, tau2, and A from real trEFM data and reconstructs real signals better than the prior single-exponential network. The evidence for \"real\" accuracy comes from two places: held-out voltage-pulse parity after fine-tuning (Supporting Information Fig. 6) and reconstruction R2 on real perovskite data (0.89 vs. 0.71 in Section 3.5). Neither independently validates the claim for new experimental conditions. Methods 2.4 states that voltage-pulse labeling is time-consuming and that collecting sufficient labeled data on multiple cantilevers would take too long; the 500 labeled traces are used for fine-tuning, but the paper does not report whether the held-out voltage-pulse test set came from a different cantilever, tip, or session than the fine-tuning set. If the test traces share tip state, noise floor, and cantilever transfer function with the fine-tuning traces, the reported parity R2 values partly measure in-distribution memorization rather than generalization to a new cantilever. The perovskite evaluation then uses reconstruction R2 computed by simulating delta-omega(t) with the same FFTA forward model that generated the training data (Methods 2.3 and Section 3.5), so a systematic forward-model bias would inflate the apparent agreement between extracted parameters and real signals. The absence of a cross-cantilever labeled test, and of any independent reconstruction check that does not rely on the training forward model, leaves the \"real experimental data\" portion of the central claim under-supported. This is a concrete gap, not a fatal flaw, because the voltage-pulse experiment is genuine external evidence; however, it is the weakest load-bearing point in the argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a multi-output convolutional neural network (CNN) for extracting bi-exponential surface-potential dynamics (parameters τ1, τ2, and A) from time-resolved electrostatic force microscopy (trEFM) frequency traces. The network is trained on simulated traces generated with the FFTA forward model, fine-tuned on a small set of experimentally labeled voltage-pulse traces, and then evaluated on simulated test data, held-out voltage-pulse data, an artificial image, and real unlabeled perovskite trEFM data. The authors report accurate parameter extraction on the held-out voltage-pulse tests (R2 = 0.74, 0.79, 0.97 for τ1, τ2, A), improved signal reconstruction R2 compared to a previous single-exponential feedforward network (0.97 vs 0.79 on an artificial image; 0.89 vs 0.71 on real perovskite data), and SHAP-based analyses that confirm physically intuitive feature importance. The core claim is that this network enables more accurate and noise-robust extraction of multi-exponential dynamics from trEFM data than existing single-exponential approaches.","tokens_in":16465,"tokens_out":4848,"duration_ms":41739,"significance":"If the claims hold, this work extends quantitative trEFM analysis from single-exponential to bi-exponential dynamics and provides an open-source, reproducible pipeline. The voltage-pulse experiment supplies an external ground truth that partially validates the parameter extraction, which is a genuine strength. However, the real-data reconstruction metric is computed with the same FFTA forward model used to generate the training data, so the improved reconstruction on unlabeled data may reflect improved fitting to the forward model rather than to true experimental physics. Cross-cantilever generalization is also not directly demonstrated. These gaps currently limit the significance of the practical claims, although the methodological contribution and code availability are valuable.","major_comments":[{"comment":"The reconstruction R2 on real perovskite data (0.89) is computed by simulating Δω(t) from the extracted parameters using the FFTA forward model, which is the same model used to generate all simulated training data. This evaluation is self-consistent: the network is trained on FFTA-simulated traces and the 'reconstructed' signal is simulated with the same FFTA model. Consequently, the comparison against the single-exponential network (0.89 vs 0.71) may indicate which set of parameters better reproduces the forward model's output, rather than which more accurately represents the true experimental signal. Please provide an independent validation that does not rely on the training forward model, for example by comparing reconstructed raw cantilever deflection traces against measured ones, or by using labeled voltage-pulse data collected with a different cantilever and tip state.","section":"Section 3.5 and Methods 2.3"},{"comment":"The held-out voltage-pulse test set is obtained by randomly splitting the 500 fine-tuning traces into 350/50/100 for training/validation/testing, but the paper does not report whether the test traces were collected on a different cantilever, tip, or session from the fine-tuning set. If the test traces share the same cantilever transfer function and noise floor as the fine-tuning traces, the reported parity R2 values (0.74, 0.79, 0.97) partly reflect in-distribution memorization rather than generalization to a new cantilever. Please report the cantilever/session identity of the test traces or add a cross-cantilever validation to support the claimed generalizability to new experimental conditions.","section":"Section 3.3 and Supporting Information Fig. 6"},{"comment":"The training and evaluation are confined to τ1 in [1,10] μs, τ2 in [50,500] μs, and A in [0,1]. The paper does not discuss expected behavior for real data with dynamics outside these ranges, and the perovskite data are not reported to lie within the training domain. Because the network regresses parameters within a bounded, normalized output, extrapolation to values outside the training ranges is unlikely to be reliable. Please state this limitation explicitly and, if possible, report the extracted parameter ranges for the perovskite data to confirm they fall within the training distribution.","section":"Section 3.2, Eq. (1), and Supporting Information Fig. 2"}],"minor_comments":[{"comment":"The main text references 'Supporting Information Fig. 9' for SHAP distributions of cantilever parameters, but the supporting information lists Fig. 9 as the noise-evaluation figure and Fig. 10 as the SHAP cantilever-parameter figure; the reference numbers appear to be swapped.","section":"Section 3.4"},{"comment":"The description of the simulation dataset split (7,000/1,500/1,500 out of 10,000 traces) is clear, but it would be helpful to state the analogous split for the 500 voltage-pulse traces in the same paragraph or in Section 2.4 for consistency.","section":"Section 2.3"},{"comment":"The caption gives the FFTA repository as 'https://github.com/rajgiriUW/ffta' while the main text and Data Availability give 'https://github.com/GingerLabUW/FFTA'; please standardize the URL.","section":"Supporting Information Fig. 4 caption"},{"comment":"The caption refers to 'shown in main text and shown in Figure 5' and later to 'Figure 6', but the main text figure is Figure 5; the second reference appears to be a typo.","section":"Supporting Information Fig. 13 caption"},{"comment":"The statement 'we believe this level of accuracy should be sufficient for many if not most imaging applications' is an assertion without a quantitative criterion; consider providing a concrete example of an imaging application and the acceptable error level for that application.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The core methodological contribution (a multi-output CNN for bi-exponential trEFM parameter extraction with open-source code) is sound and likely of interest to the SPM and machine-learning communities. The voltage-pulse parity results provide a meaningful external validation, but the real-data reconstruction claim is undermined by the use of the training forward model as the evaluation metric. The cross-cantilever question is also critical for the generalizability claim. I would like to see the authors address these two points with additional experiments or a clear statement of scope. The manuscript is well written and the supporting information is thorough; the main revisions needed are experimental or analytical additions rather than text-only fixes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper does a real thing. It extends the authors' earlier feedforward network from single-exponential to bi-exponential parameter extraction, using a multi-branch CNN that outputs tau1, tau2, and A and takes cantilever parameters as additional inputs. That is a genuine step beyond ref 18, and the paper ships code and data for both the CNN and the FFTA forward model, which counts as real evidence.\n\nThe strongest evidence is the held-out voltage-pulse parity test: R2 of 0.74, 0.79, and 0.97 for tau1, tau2, and A, with <tau> R2 of 0.97. That is external ground truth, collected with programmed voltage pulses on a conductive substrate, and it supports the extraction claim. The noise-robustness comparison is also quantitative and favors the CNN over the previous network.\n\nThe soft spots are real but mostly addressable. The reconstruction-R2 evaluation on real perovskite data reuses the same FFTA forward model that generated the training data, so high R2 partly reflects self-consistency between the learned inversion and the simulator. The voltage-pulse test set appears to come from the same cantilever and session as the fine-tuning set; nothing demonstrates generalization to a new cantilever, tip state, or noise floor. They also skip a straightforward baseline: standard bi-exponential least-squares fitting on the same voltage-pulse and real data. And the \"physics-informed\" label oversells what is actually parameter concatenation plus SHAP sanity checks.\n\nNone of this breaks the central claim. The voltage-pulse experiment is a genuine external check, the authors are transparent about the cost of collecting more labeled data, and the limitations they state in the text are consistent with what the data actually show. The paper is a methods contribution for a specialized SPM community, not a revolution, but it is honest and useful.\n\nSend it to peer review. A serious referee should ask for a cross-cantilever generalization test and a standard-fit baseline, but the paper is well within the range of a publishable methods paper in an SPM/ML venue. I would not cite it in my own work because the topic is outside my area, but I would bring it to a reading group interested in ML for scanning probe data.","headline":"Genuine step beyond their own single-exponential network, with shipped code and a real voltage-pulse ground truth, but the real-data reconstruction metric is partly self-consistent.","tokens_in":17041,"tokens_out":1668,"would_cite":false,"duration_ms":14633,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multi-output CNN recovers bi-exponential kinetics from time-resolved electrostatic force microscopy data.","keywords":["time-resolved electrostatic force microscopy","convolutional neural network","bi-exponential kinetics","parameter extraction","cantilever physics","signal reconstruction","halide perovskite","physics-informed machine learning"],"falsifier":"Run the trained CNN on trEFM data from a sample whose transient kinetics are independently known to be tri-exponential or stretched-exponential; if the network still reports high reconstruction R2 while its extracted τ1, τ2, and A disagree with the true kinetic moments, then the CNN is fitting the assumed bi-exponential form rather than the actual dynamics.","tokens_in":15898,"feed_emoji":"🔬","tokens_out":7954,"duration_ms":62098,"temperature":0.7,"pith_summary":"Time-resolved electrostatic force microscopy (trEFM) records how a vibrating cantilever's frequency shifts when a laser excites a material, but the cantilever's own physics is tangled with the material's transient kinetics, so recovering the kinetics from the signal is hard. This paper claims that a multi-output convolutional neural network, fed the frequency trace together with the cantilever's spring constant, quality factor, and resonance frequency, can extract the two time constants and the mixing coefficient of a bi-exponential surface-potential response. On simulated, labeled voltage-pulse, and unlabeled perovskite data, the network reconstructs the measured signal more accurately than the earlier single-exponential feedforward network (average R2 of 0.97 versus 0.79 on an artificial image, and 0.89 versus 0.71 on real perovskite data). The practical payoff is that trEFM imaging could characterize multi-exponential ion and carrier dynamics in materials such as halide perovskites without assuming a single time constant.","feed_headline":"Neural net pulls two time constants out of noisy microscopy data","feed_subtitle":"Multi-output CNN beats single-exponential fits, lifting reconstruction R2 from 0.79 to 0.97.","key_machinery":"The central mechanism is the multi-branch convolutional architecture that learns the multi-timescale structure of the trEFM frequency trace: small kernels read the fast initial frequency shift where the kinetic information is concentrated, larger kernels read the slow cantilever-governed relaxation, and a dedicated branch learns the mixing weight $A$. The cantilever parameters $k$, $Q$, and $\\omega$ are concatenated into the dense layers so the network can factor out cantilever physics from the kinetics. A custom loss function weights the squared error of $\\tau_1$ most heavily, reflecting that $\\tau_1$ produces the subtlest signal changes.","core_discovery":"The paper's central claim is that a multi-branched, multi-output CNN trained on simulated trEFM traces and fine-tuned on labeled voltage-pulse experiments can invert the cantilever-obfuscated frequency signal to recover the parameters ($\\tau_1$, $\\tau_2$, $A$) of an underlying bi-exponential surface potential perturbation $y = A e^{-t/\\tau_1} + (1-A) e^{-t/\\tau_2}$. The network takes the normalized $\\Delta\\omega(t)$ trace through three convolutional branches with different kernel sizes to capture fast and slow features, concatenates the latent features with the physical cantilever parameters ($k$, $Q$, $\\omega$), and regresses the three parameters through dense output branches. The authors show that this model separates fast and slow time constants with mean errors of 1.01 $\\mu$s and 35.87 $\\mu$s on fine-tuned experimental data, extracts $A$ with a parity R2 of 0.97, and reconstructs experimental signals with R2 0.89 on unlabeled perovskite images, substantially better than the single-exponential feedforward baseline. They take this as evidence that bi-exponential assumptions and physics-informed inputs make CNN-based parameter extraction practical for trEFM.","pith_inferences":["A natural extension would be to relax the bi-exponential assumption to tri-exponential or stretched-exponential kinetics; the same multi-branch architecture should transfer, but the training data would need to come from a forward model that includes those forms.","Because the network's inputs include cantilever parameters, it may generalize across cantilevers and tip states without retraining, a testable claim the paper does not directly demonstrate.","The reconstruction-R2 metric is computed through the same forward model used to generate training data, so high R2 on experimental data is necessary but not sufficient evidence that the extracted parameters are physically correct; an independent kinetic measurement would close that gap.","A deployment risk is that the network will confidently fit a bi-exponential to data that are not bi-exponential; adding a model-mismatch output or an uncertainty estimate would let users know when the assumed form is inadequate."],"forward_implications":["trEFM data can be analyzed under a bi-exponential model without empirical cantilever calibration, using only the cantilever parameters already measured in the experiment.","The CNN keeps reconstruction R2 above 0.9 at signal-to-noise ratios down to about 7, where the single-exponential network degrades to 0.43-0.8.","The extracted parameter maps on perovskite films reproduce known physics: grain boundaries show slower surface potential equilibration than grain interiors.","SHAP analysis confirms the network relies on the initial frequency shift, not the cantilever relaxation, for the kinetic parameters, and learns that high-Q cantilevers correspond to faster underlying dynamics.","The approach is claimed to be generalizable to other underlying functional forms and other time-resolved scanning probe techniques."],"supporting_citations":[{"why":"The previous single-exponential feedforward neural network; supplies the baseline that the multi-output CNN must beat and the single-exponential limitation being addressed.","marker":"18"},{"why":"Establishes the trEFM technique and the cantilever frequency-shift response that the forward model simulates.","marker":"17"},{"why":"Provides the fast trEFM protocol and demodulation (Hilbert transform) used to obtain the Δω(t) traces.","marker":"31"},{"why":"Documents bi-exponential surface-potential equilibration in halide perovskites and motivates the bi-exponential model.","marker":"4"},{"why":"Supports the role of cantilever Q in the relaxation dynamics, which the network learns via SHAP.","marker":"19"},{"why":"Provides SHAP, the explainability tool used to verify which parts of the signal and which cantilever parameters drive each output.","marker":"47"},{"why":"Supplies the transfer-learning basis for fine-tuning the simulated-data-trained CNN on a small labeled voltage-pulse dataset.","marker":"35"}],"fun_headline_variants":["Multi-output CNN extracts two time constants from trEFM signals","Physics-informed CNN improves bi-exponential fitting in trEFM","Multi-branch CNN beats single-exponential fits for trEFM","AI decodes bi-exponential kinetics from cantilever-obfuscated trEFM"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forward model that converts a bi-exponential surface-potential decay into a simulated cantilever frequency trace must be a faithful description of real trEFM experiments, because that model generates all of the training data and is also reused to score how well the extracted parameters reconstruct the measured signal.","fun_headline_variants_meta":{"raw":{"variants":["Multi-output CNN extracts two time constants from trEFM signals","Physics-informed CNN improves bi-exponential fitting in trEFM","Multi-branch CNN beats single-exponential fits for trEFM","AI decodes bi-exponential kinetics from cantilever-obfuscated trEFM"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00089,"raw_usage":{"total_tokens":3875,"prompt_tokens":1017,"completion_tokens":2858,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":2781}},"tokens_in":633,"tokens_out":2858,"duration_ms":19009,"temperature":1.0,"reasoning_tokens":2781,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T04:28:50.915571+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained CNN on trEFM data from a sample whose transient kinetics are independently known to be tri-exponential or stretched-exponential; if the network still reports high reconstruction R2 while its extracted τ1, τ2, and A disagree with the true kinetic moments, then the CNN is fitting the assumed bi-exponential form rather than the actual dynamics.","supporting_citations":[],"review_version":1}