{"id":"45bd1f83-e05c-449b-8cd8-454aa1580e46","arxiv_id":"2412.02779","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"BayesMulti, a Bayesian-optimized multinomial noise-injection training scheme, aims to make analog DNNs on perovskite memristors robust to non-idealities, but the theoretical bound is not supported as written.","lead":"Bayesian optimization is used to tune perovskite memristor fabrication while a noise-injection training method, BayesMulti, is claimed to make analog neural networks robust to device imperfections. The paper reports large simulation gains and a small 10x10 hardware demo, but the theoretical guarantee and hardware validation are weaker than the claims.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's robustness guarantee is proven only for ternary multiplicative perturbations (δ in {0,0.5,1}), while the paper's own device noise model (Eq. 4) is continuous log-normal, so the central theoretical claim does not cover the actual tested non-idealities.","rationale":"The reader's central concern is correct: the theoretical guarantee is proven for ternary multiplicative perturbations, whereas the paper's own stochastic non-ideality model is continuous log-normal. This is an internal mismatch, not merely a disagreement with existing consensus: the theorem is supposed to justify the robustness observed in experiments whose perturbations are drawn from Eq. 4. The hardware crossbar experiment also uses real conductance drift, which is not constrained to {0,0.5,1}, so it cannot validate the theorem either. I do not find the reader's secondary algebra objection about m=0 and l=Θ-r decisive: since the lower-bound expression increases in both m and l but faster in m, setting m=0 is the correct choice to minimize the worst case, so that step is at least coherent. The decisive problem is the perturbation-set mismatch, which is sufficient on its own to break the central theoretical claim. The empirical comparisons may still be interesting, but without a theorem covering the tested noise model, or a separate derivation bridging log-normal drift to Eq. 2, the paper's headline claim of a theoretically ensured robustness guarantee is not supported. Given that the theorem is load-bearing for the abstract's central assertion and that code and hyperparameters are not provided, retaining the reader's REJECT verdict is appropriate.","tokens_in":29993,"tokens_out":9603,"duration_ms":100547,"concrete_test":"Analytically recompute the DF bound in Supplementary Note 3 for a one-layer 0-1 classifier with πδ induced by δ_i = e^{λ_i}, λ_i ~ N(0, σ²), for a few small Θ. If the discrete ternary formulas (Eqs. 16-19) do not hold for this continuous δ, Theorem 1 is inapplicable to the paper's Eq. 4 noise model. An empirical complement: train a small MLP on MNIST with BayesMulti at reported p1, p2, inject log-normal noise with σ from Eq. 4 at usability = 0.5, and check whether accuracy remains above 0.5 for perturbations that satisfy the theorem's count-based r; failure would confirm that the theoretical guarantee does not govern the experiments.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's advertised guarantee is Theorem 1 (Eq. 2), proved in Supplementary Note 3. The proof's Lemma 1 and the following robustness analysis partition the support of πδ on the discrete set {0,0.5,1}^Θ and express DF in terms of k, l, m, the counts of coordinates where δ_i equals 0, 0.5, or 1. The robustness set B and the bound r are therefore defined for ternary multiplicative disturbances. However, the paper's own model of stochastic memristor non-ideality, Methods Eq. 4 and Supplementary Note 1 Eq. 2, is θ' ← θ e^λ with λ ~ N(0, σ²), a continuous multiplicative perturbation affecting every parameter. A log-normal realization has δ_i outside {0,0.5,1} with probability 1; in that case πδ gives zero mass to essentially all atoms of {0,0.5,1}^Θ, so Eq. 16 does not compute the DF for the tested noise model and Inequality 12 cannot yield Eq. 2. All simulation experiments, the usability metric (Eq. 5), and the hardware demonstration are evaluated under this log-normal drift, not under ternary δ. Thus the central claim that the method 'theoretically ensures' consistent predictions under memristor non-idealities is unsupported for the perturbation class actually tested.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a synergistic methodology for co-optimizing perovskite memristor fabrication and the robustness of analog deep neural networks. Fabrication conditions are selected through Bayesian optimization using a 'usability' metric derived from measured I-V characteristics, and the training method 'BayesMulti' injects multinomial noise whose parameters are tuned via Bayesian optimization. The central theoretical claim is Theorem 1, which asserts that training with the multinomial noise guarantees consistent predictions under multiplicative parameter perturbations within a robustness set B, with a maximum allowable radius r given by Eq. (2). The method is evaluated on image classification, autonomous driving, biological sequence modeling, a large vision-language model, and a 10x10 perovskite memristor crossbar, with reported improvements over empirical risk minimization and an energy-efficiency comparison against a GPU.","tokens_in":30300,"tokens_out":20573,"duration_ms":201124,"significance":"The empirical scope is broad and the hardware demonstration on a real crossbar is a valuable contribution. If the theoretical guarantee were valid for the noise model actually tested, the work would be an important advance in analog co-design. The fabrication Bayesian optimization pipeline and the crossbar validation are strengths, and the paper provides concrete prototypes in Supplementary Note 4. However, the central robustness theorem does not apply to the continuous log-normal noise used in all experiments, and the stated bound appears numerically vacuous for networks of realistic size. The theoretical contribution therefore needs substantial revision before the paper's central claims can be accepted.","major_comments":[{"comment":"The robustness analysis in Supplementary Note 3 (Lemma 1 and the derivation of Eq. (16)) is carried out for perturbations δ whose coordinates take values in {0, 0.5, 1}, with the counts k, l, m fully determining the DF bound. In contrast, the paper's own stochastic non-ideality model (Methods, Eq. (4)) is θ' ← θe^λ with λ ∼ N(0, σ²), a continuous multiplicative perturbation. For a log-normal draw, δ_i ∉ {0, 0.5, 1} for every coordinate with probability one, so the robustness set B (interpreted in the proof as Θ−m−l≤r) is violated for any r<Θ. Consequently, Theorem 1 does not cover the perturbations generated by PerovskiteMemSim and used in all experimental evaluations, leaving the abstract's claim of a theoretical guarantee for memristor non-idealities unsupported for the tested perturbation class.","section":"Supplementary Note 3 / Methods Eq. (4)"},{"comment":"The statement of Theorem 1 is not well-formed: the robustness set is written as B = {δ : δ − 10 + δ − 0.50 − Θ ≤ r}, which is not a meaningful condition; the proof later interprets it as Θ−m−l≤r. In addition, Eq. (2) has the form r ≤ [ln(1.5−fπ0(θ0)) − Θ ln(1−p2)] / [ln p1 − ln(1−p2)]. Since p1 ≤ 1−p2, the denominator is non-positive, and for typical values (e.g., p1=0.3, p2=0.3, Θ=100, fπ0(θ0)=0.8) the right-hand side is negative. As printed, the theorem therefore yields no positive robustness radius for networks of the size used in the experiments, making the quantitative guarantee vacuous.","section":"Theorem 1, Eq. (2)"},{"comment":"The lower bound obtained by setting m=0 and l=Θ−r has the form fπ0(θ0) − 1 + p1^r(1−p2)^{Θ−r}. Because (1−p2)<1, the second term decays exponentially in Θ. For any network with more than a handful of parameters, fπ0(θ0) − 1 + p1^r(1−p2)^{Θ−r} > 0.5 cannot hold even for r=0 unless fπ0(θ0) is extremely close to 1.5, which is impossible since fπ0(θ0) ≤ 1. Thus the proof method, even under its own discrete-perturbation assumptions, cannot certify positive robustness for realistic network sizes; the quantitative claim in Eq. (2) needs to be revisited.","section":"Supplementary Note 3, robustness analysis after Eq. (23)"}],"minor_comments":[{"comment":"The expression for the robustness set B contains obvious typographical errors and should be rewritten in terms of the count variables k, l, m used in the proof.","section":"Theorem 1"},{"comment":"The usability metric depends on the free parameter Require_len, which is set to 35 in the code but not justified in the main text; the authors should state how this threshold is chosen and whether the conclusions are sensitive to it.","section":"Methods Eq. (5) / Supplementary Note 4"},{"comment":"The caption says 'different hardware non-idealities (σ values)' while the horizontal axis is labeled usability; please make the caption consistent with the axis definition.","section":"Figure 6(b) caption"},{"comment":"The prior work in Reference [30] already proposed a Bayes-optimized noise injection approach for analog DNNs; the text should clarify what BayesMulti adds beyond [30] and why the multinomial distribution is essential to the new method.","section":"Introduction and Reference [30]"},{"comment":"The abstract's 'up to 100-fold improvements' should be accompanied by a precise definition of the ratio being reported (e.g., accuracy ratio, error ratio, or energy efficiency), as the metric is not clear without reading the figures.","section":"Abstract and Results"}],"recommendation":"reject","confidential_remarks":"The central theoretical claim is not supportable as written: the theorem's perturbation model is discrete and does not match the continuous log-normal noise used in the evaluations, and the stated bound is numerically vacuous for realistic network sizes. These are load-bearing issues in a manuscript whose abstract and introduction emphasize a theoretical guarantee. The empirical study is broad and the hardware demonstration is useful, but the theoretical contribution would require a major rewrite or a complete change in the evaluation protocol to be salvageable within the paper's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The empirical side is more substantial than the theory. The usability metric and BO-driven fabrication loop are credible: they define a measurable quantity from I-V sweeps, search 8400 configurations in 12 human-in-the-loop iterations, and validate on a real 10x10 perovskite crossbar. That part is useful for the materials community. BayesMulti, however, is an incremental extension of the authors' own Bernoulli noise-injection work, and they never compare against that baseline, which is a real omission.\n\nThe problem is Theorem 1. The proof only handles multiplicative perturbations δ_i in {0, 0.5, 1}, a discrete ternary set. Their own device noise model (Eq. 4) is continuous log-normal, and every experiment uses that model. So the advertised robustness guarantee does not cover the perturbations under test. The theorem statement itself is garbled (the set B is not well-formed), and while the Case 1.2 algebra actually checks out, the mismatch between the proven guarantee and the tested noise is load-bearing—the abstract and conclusion repeatedly say the approach \"theoretically ensures\" robustness.\n\nThe hardware validation is thin: a toy moon dataset on a 10x10 crossbar, one hidden layer. The energy-efficiency number excludes ADC and peripherals, though that is disclosed. The broad empirical sweep across MNIST/CIFAR/KITTI/biology/MiniGPT-4 is nice but relies entirely on PerovskiteMemSim, which is calibrated from the same measured I-V curves. Not circular, but not independent confirmation either. Code is promised but not shipped.\n\nWho gets value? Materials scientists working on perovskite memristors will find the usability metric and BO loop practical. For the ML theory side, the robustness claim needs a complete rework—either prove a bound for log-normal (or bounded continuous) noise, or explicitly restrict the claim to ternary disturbances and justify why that approximates reality.\n\nRecommendation: this deserves a serious referee, not a desk reject. The empirical contribution is real, and the theory is fixable in principle, but the current version overclaims. Send it to review with a request for major revision.","headline":"A solid empirical co-optimization story undermined by a robustness theorem that doesn't match the test-time noise model.","tokens_in":30859,"tokens_out":4293,"would_cite":false,"duration_ms":41449,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that analog neural networks on imperfect perovskite memristors can be made dependable by jointly optimizing the fabrication recipe with Bayesian optimization and training the network with a BO-tuned multinomial noise…","keywords":["perovskite memristors","analog computing","Bayesian optimization","noise injection","deep neural network robustness","memristor crossbar","in-memory computing","device non-idealities"],"falsifier":"Run BayesMulti with a noise model where each weight is multiplied by a continuous random factor whose support lies inside the claimed robustness set, such as log-normal noise with $\\sigma$ matched to device measurements; if any such perturbation pushes the randomized network's score at or below the decision threshold while the perturbation magnitude lies within the radius $r$ from Theorem 1, the claim that prediction outcomes remain consistent is false for the actual noise model. Alternatively, compute the exact worst-case perturbed prediction for a small two-layer network by solving the optimization over the set $B$, and check whether the minimizer really occurs at $m=0$, $l=\\Theta-r$ as the proof assumes.","tokens_in":29820,"feed_emoji":"⚡","tokens_out":10455,"duration_ms":104598,"temperature":0.7,"pith_summary":"Perovskite memristors could make deep learning much more energy-efficient by doing matrix multiplication in analog hardware, but their imperfect, noisy conductance states wreck neural-network accuracy. This paper claims the problem can be solved by optimizing the device and the training algorithm together: a Bayesian-optimization loop first finds fabrication conditions that maximize a measurable 'usability' score, and a second Bayesian loop tunes a multinomial noise-injection scheme, BayesMulti, that makes networks tolerate the residual hardware imperfections. The authors report that the combined approach keeps high accuracy across tasks ranging from MNIST and CIFAR-10 to autonomous-driving point-cloud detection, antibody and glycan classification, and MiniGPT-4, and that a fabricated 10×10 perovskite crossbar loses only about 15% accuracy in hardware inference versus roughly 45% for standard training. They also present a theorem giving a bound on how large a multiplicative weight disturbance can be while the network's prediction stays consistent.","feed_headline":"Tuned noise keeps analog AI accurate on imperfect perovskite memory","feed_subtitle":"Device recipe and training noise are optimized together, cutting hardware accuracy loss from about 45% to 15%.","key_machinery":"The load-bearing objects are the usability score and the BayesMulti noise distribution. Usability is a single scalar that converts a measured I-V curve into a device-quality target: the length of the longest conductance increasing subsequence relative to a required minimum, discounted by $e^{-\\sigma}$ for stochastic cycle-to-cycle variation. BayesMulti is a training-time randomization in which each weight is multiplied by $\\eta$ with density $p_1$ at $\\eta=0$, $p_2$ at $\\eta=0.5$, and $1-p_1-p_2$ at $\\eta=1$; Bayesian optimization with a Gaussian-process surrogate and expected improvement tunes the noise parameters. The argument is carried by Theorem 1, a functional-optimization bound: minimizing the worst-case perturbed prediction over all bounded $[0,1]$-valued functions produces the robustness radius $r$ stated above. The mechanism is that noise injection randomizes over a neighborhood of the parameter point, and the 0.5 multiplier immunizes the network against the same perturbation when it is applied at inference time.","core_discovery":"The central claim is that analog neural networks can be made robust to perovskite memristor non-idealities by jointly optimizing fabrication and training noise. On the device side, the paper defines usability as $\\mathrm{Usability} = (l_{\\mathrm{LCIS}}/\\mathrm{Require\\_len})e^{-\\sigma}$, where $l_{\\mathrm{LCIS}}$ is the length of the longest monotonically increasing segment of the conductance curve, $\\mathrm{Require\\_len}$ is a minimum working range, and $\\sigma$ is the cycle-to-cycle log-normal variability of conductance; Bayesian optimization over perovskite type, nanowire length and diameter, lead electrodeposition time, and Ag thickness raised usability from 0.36 to 0.93 in twelve iterations. On the algorithm side, BayesMulti multiplies each weight by an independent random factor $\\eta$ drawn from $\\{0, 0.5, 1\\}$ with probabilities $p_1$ and $p_2$, with those probabilities chosen by Bayesian optimization. The paper's Theorem 1 states that, for a binary classifier whose network output is $f_{\\pi_0}(\\theta_0) = \\mathbb{E}_{\\eta\\sim\\pi_0}[f(\\theta_0 * \\eta)]$, the maximum allowable multiplicative disturbance $r$ satisfies $r \\le [\\ln(1.5 - f_{\\pi_0}(\\theta_0)) - \\Theta \\ln(1-p_2)]/[\\ln p_1 - \\ln(1-p_2)]$, where $\\Theta$ is the number of parameters; if correct, any perturbation within that set leaves the prediction on the same side of the decision threshold. This is, to the authors' knowledge, the first demonstration of analog-computing inference on a large vision-language model, and the hardware test on a real perovskite crossbar supports the accuracy-stability claim.","pith_inferences":["The theoretical radius $r$ depends on the parameter count $\\Theta$ through the term $-\\Theta\\ln(1-p_2)$, so for very large models the guaranteed perturbation set may shrink; a layer-wise or block-wise noise schedule is a natural extension the paper does not explore.","Because the theorem's proof considers perturbations whose coordinates are exactly 0, 0.5, or 1, while the simulations use continuous log-normal multipliers, the experimental results should be read as evidence for the method's practical robustness rather than as a direct validation of the theorem's bound.","The usability metric, being derived from basic I-V measurements, could serve as a standardized figure of merit for comparing memristive technologies beyond perovskites, provided its correlation with end-task accuracy is validated across independent labs.","The reported 270× energy advantage counts only the crossbar array; including analog-to-digital conversion and peripheral circuits will change the ratio, and the same training recipe could be tested on full-system benchmarks."],"forward_implications":["If Theorem 1 is correct, a BayesMulti-trained binary classifier's prediction cannot flip for any multiplicative disturbance inside the stated set, so analog implementations can be certified against a bounded class of device drift.","BayesMulti-trained analog networks keep usable accuracy at much lower usability values than standard empirical-risk-minimization training, in some KITTI detection settings achieving 10 to 100 times the ERM accuracy.","The BO fabrication loop, starting from one initial configuration, reached a near-optimal perovskite recipe in twelve iterations, including parameter choices outside human expertise such as smaller nanowire diameters.","A real 10×10 perovskite memristor crossbar implementing inference on a moon-shaped dataset loses about 15% accuracy with BayesMulti versus about 45% with ERM, and the array-only energy efficiency is claimed to exceed a Tesla V100 by over 270 times, excluding ADC and peripheral circuitry.","The same noise-injection recipe transfers to deeper and wider networks, including PointPillar, Mason's CNN, SweetNet, and MiniGPT-4, indicating the approach is architecture- and task-agnostic."],"supporting_citations":[{"why":"Supplies the Bayes-optimized noise-injection training strategy that BayesMulti extends from Bernoulli to multinomial noise.","marker":"[30]"},{"why":"Provides the nonideality-aware training baseline for memristive neural networks that BayesMulti is compared against.","marker":"[14]"},{"why":"Establishes the memristor crossbar vector-matrix multiplication model that makes analog DNN inference possible.","marker":"[7]"},{"why":"Demonstrates the three-dimensional perovskite nanowire memristor platform whose fabrication conditions are optimized.","marker":"[46]"},{"why":"Motivates the log-normal model of conductance variation used to simulate stochastic non-ideality in BayesMulti experiments.","marker":"[81]"},{"why":"Documents multi-level RRAM switching whose variability motivates the log-normal drift model.","marker":"[82]"},{"why":"Provides the fully hardware-implemented memristor CNN whose energy-efficiency methodology is used for the 270x comparison.","marker":"[73]"},{"why":"Supplies the PointPillars network architecture used in the autonomous-driving object-detection evaluation.","marker":"[84]"}],"fun_headline_variants":["Perovskite memory gets 100x boost from device-plus-noise tuning","Joint recipe and training fix make imperfect perovskite chips viable","Analog AI thrives on imperfect perovskite with optimized noise","Theoretical shield: Perovskite memristors get robust despite noise","Optimize both hardware and noise to tame perovskite memory"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The theoretical robustness guarantee assumes the only perturbations that matter are coordinate-wise multiplications by exactly 0, 0.5, or 1, while the noisy hardware is modeled with continuous log-normal multiplicative drift, so the guarantee does not literally cover the noise used in the experiments.","fun_headline_variants_meta":{"raw":{"variants":["Perovskite memory gets 100x boost from device-plus-noise tuning","Joint recipe and training fix make imperfect perovskite chips viable","Analog AI thrives on imperfect perovskite with optimized noise","Theoretical shield: Perovskite memristors get robust despite noise","Optimize both hardware and noise to tame perovskite memory"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000629,"raw_usage":{"total_tokens":3042,"prompt_tokens":1212,"completion_tokens":1830,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":828,"completion_tokens_details":{"reasoning_tokens":1747}},"tokens_in":828,"tokens_out":1830,"duration_ms":13054,"temperature":1.0,"reasoning_tokens":1747,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:06:57.075344+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run BayesMulti with a noise model where each weight is multiplied by a continuous random factor whose support lies inside the claimed robustness set, such as log-normal noise with $\\sigma$ matched to device measurements; if any such perturbation pushes the randomized network's score at or below the decision threshold while the perturbation magnitude lies within the radius $r$ from Theorem 1, the claim that prediction outcomes remain consistent is false for the actual noise model. Alternatively, compute the exact worst-case perturbed prediction for a small two-layer network by solving the optimization over the set $B$, and check whether the minimizer really occurs at $m=0$, $l=\\Theta-r$ as the proof assumes.","supporting_citations":[{"cited_title":"Resistive switching mechanism in ZnxCd1−xS nonvolatile memory devices","cited_arxiv_id":null,"evidence_quote":"Supplies the Bayes-optimized noise-injection training strategy that BayesMulti extends from Bernoulli to multinomial noise."},{"cited_title":"The missing memristor found","cited_arxiv_id":null,"evidence_quote":"Provides the nonideality-aware training baseline for memristive neural networks that BayesMulti is compared against."},{"cited_title":"Compliance-free multileveled resistive switching in a transparent 2d perovskite for neuromorphic computing","cited_arxiv_id":null,"evidence_quote":"Demonstrates the three-dimensional perovskite nanowire memristor platform whose fabrication conditions are optimized."},{"cited_title":"Revisiting anodic alumina templates: From fabrication to applications","cited_arxiv_id":null,"evidence_quote":"Motivates the log-normal model of conductance variation used to simulate stochastic non-ideality in BayesMulti experiments."},{"cited_title":"A neuromorphic bionic eye with filter-free color vision using hemispherical perovskite nanowire array retina","cited_arxiv_id":null,"evidence_quote":"Documents multi-level RRAM switching whose variability motivates the log-normal drift model."},{"cited_title":"Optimization of therapeutic antibodies by predicting antigen speci- ficity from antibody sequence via deep learning.Nature Biomedical Engineering, 5(6):600–612, 2021","cited_arxiv_id":null,"evidence_quote":"Provides the fully hardware-implemented memristor CNN whose energy-efficiency methodology is used for the 270x comparison."},{"cited_title":"In-memory learning with analog resistive switching memory: A review and perspective","cited_arxiv_id":null,"evidence_quote":"Supplies the PointPillars network architecture used in the autonomous-driving object-detection evaluation."}],"review_version":1}