{"id":"143737d6-2f03-45cb-b76c-8a1dba1f7079","arxiv_id":"2502.05948","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A measured effective-bit noise model is used to simulate accuracy loss on four workloads; per-module or per-ADC reference tuning helps some tasks, but the drone result contradicts the claimed general effectiveness.","lead":"This paper measures noise in a 40nm RRAM compute-in-memory test chip and builds a simulation that injects those measured errors into neural network workloads. The authors say tuning the ADC reference voltage mitigates the noise, but their own drone-navigation results show the tuning can make performance far worse.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own drone-navigation results contradict the central claim: per-module tuning reduces MSF from 899m to <20m, and per-ADC tuning to 691m, so reference tuning does not 'maintain robustness'.","rationale":"The reader's verdict is REJECT, and I agree that rejection is warranted, but my load-bearing concern differs from the reader's stated weakest assumption. The reader focuses on whether the noise model (measured on 10 modules, 80 ADCs, with Gaussian sampling) is representative of real deployment; that is a valid external-validity concern. However, the most decisive problem is internal: the paper's own Table III shows reference tuning making the drone navigation task dramatically worse, which contradicts the central claim that per-module/per-ADC tuning mitigates noise. This is a logical inconsistency in the argument, not a question of extrapolation. The reader's rationale does mention the drone degradation ('per-module tuning degrades performance drastically'), but this is not identified as the weakest assumption. Thus my agreement is partial. The concrete test of re-running the drone navigation experiment with multiple seeds directly settles whether the contradiction in Table III is robust or an artifact of a single trial; if robust, the central claim is false as stated. The paper's characterization work on RRAM noise and read disturb may still have value, but the stated contribution of demonstrating effective reference tuning is unsupported by the reported data.","tokens_in":8584,"tokens_out":2299,"duration_ms":22163,"concrete_test":"Re-run the drone autonomous navigation evaluation (8-bit quantized C3F2 model, 25-action probabilistic action space) under three conditions: quantized-only, per-module reference tuning, and per-ADC reference tuning, using at least 10 independent random seeds each and reporting mean and standard deviation of MSF. If per-module tuning consistently yields MSF below 20m and per-ADC tuning yields MSF below the quantized-only mean minus one standard deviation, the central claim that reference tuning mitigates CIM noise is falsified for this workload. Even without new runs, recomputing the ratios from Table III (per-module/quantized-only < 0.022, per-ADC/quantized-only = 0.768) shows both tuned variants underperform the quantized-only baseline, directly contradicting the conclusion.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim, stated in the abstract and conclusion, is that per-module/per-ADC reference tuning 'is effective in maintaining system robustness' under CIM noise. Table III directly contradicts this: the drone autonomous navigation task reports baseline MSF 1215.65m, quantize-only MSF 899.38m, per-module tuning MSF <20m, and per-ADC tuning MSF 690.99m. Since MSF is the mean safe flight distance before collision, higher is better; both tuned configurations are worse than the quantize-only baseline, and per-module tuning is catastrophic. The paper itself notes in Section V-A that per-module tuning is inadequate for ResNeXt50-32x4d, so it switches to per-ADC tuning, but for the drone task even per-ADC tuning degrades performance relative to simply quantizing without noise injection. Thus the central claim is internally inconsistent with the presented experimental evidence. No error bars, number of random seeds, or statistical significance are reported for Table III, so the <20m value could be a single unlucky rollout, but the point estimates as written falsify the claim that tuning 'mitigates' noise or 'maintains robustness'. This is not a matter of disagreement with external consensus; it is a direct contradiction within the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a measurement-based noise model for RRAM compute-in-memory (CIM) systems, using a 40nm test chip to extract effective-bit statistics and read-disturb shifts. It then simulates the impact of this noise on four workloads: VGG-8/CIFAR-10, ResNeXt50/ImageNet, a GridWorld MLP, and drone autonomous navigation. The central claim is that per-module or per-ADC reference tuning effectively mitigates CIM noise impacts across time-independent and time-dependent tasks. The paper reports accuracy degradations for supervised tasks and task-success metrics for reinforcement learning tasks, and concludes that per-module/per-ADC tuning maintains system robustness.","tokens_in":8899,"tokens_out":2907,"duration_ms":29429,"significance":"The work provides valuable hardware characterization data from a 40nm RRAM test chip, including per-cell effective-bit distributions and read-disturb trends, and it connects these measurements to application-level simulation. If the noise model were validated and the tuning results were consistently positive, the paper would offer a practical calibration strategy for RRAM CIM accelerators. However, the central claim is directly contradicted by the paper's own drone-navigation results, and the simulation methodology lacks hardware-in-the-loop validation on the evaluated workloads. The paper's contribution is therefore currently more of a measurement methodology and cautionary negative result than a demonstrated mitigation technique.","major_comments":[{"comment":"Table III directly contradicts the paper's conclusion that per-module/per-ADC tuning 'is effective in maintaining system robustness.' For drone autonomous navigation, the quantize-only baseline achieves an MSF of 899.38m, while per-module tuning yields MSF <20m and per-ADC tuning yields 690.99m. Since higher MSF is better, both tuning configurations degrade performance relative to simply quantizing without noise injection, and per-module tuning is catastrophic. The text accompanying Table III states the opposite of the data. This is a load-bearing inconsistency: the central claim of the abstract and conclusion is falsified by the paper's own experimental evidence.","section":"§V-B, Table III"},{"comment":"The noise injection model is built by fitting effective-bit means and standard deviations from 10 test-chip modules (80 ADCs) and then sampling independent Gaussian noise in simulation. The manuscript provides no validation that this fitted noise model reproduces the actual accuracy or task-success degradation for any of the four evaluated workloads. In particular, there is no hardware-in-the-loop measurement on CIFAR-10, ImageNet, GridWorld, or the drone task. Without such validation, the simulated accuracy drops are a propagation of the fitted noise parameters rather than an independent prediction, and the claimed tuning benefits may not transfer to real deployments. The read-disturb analysis is also based on a single module and an accelerated stress condition, leaving the representativeness of that module and condition unaddressed.","section":"§III-C, §IV-B"},{"comment":"The ImageNet results are presented as 'Imagenet-1k / ResNeXT50-32x4d' in the text, but Table I(b) reports a baseline of 83% on a 'reduced-set' and a noise-injected accuracy of 82% on the same reduced set. The abstract and introduction imply full ImageNet-1K evaluation. The manuscript never specifies the size or composition of the reduced subset, nor does it justify that the subset is representative of full ImageNet. This is important because the paper's conclusion about complex supervised workloads rests entirely on this single reduced-set result, and per-module tuning is acknowledged to fail outright for this network, requiring per-ADC tuning instead. The claim that 'per-module/per-ADC tuning is effective' is therefore not supported for this workload without additional details.","section":"§V-A, Table I(b)"},{"comment":"No error bars, confidence intervals, or number of random seeds are reported for any accuracy, win-rate, or MSF measurement. In particular, Table III reports MSF <20m for per-module tuning as a point value without indicating how many drone episodes were evaluated or how much variance exists across episodes. Given the catastrophic drop from 899m to <20m, it is possible that a single unlucky rollout dominates this result. The absence of statistical characterization makes it impossible to assess whether the reported differences, including the claimed benefits of per-ADC tuning on ImageNet and GridWorld, are significant. This is a methodological gap that affects every quantitative conclusion in Section V.","section":"Tables I–III"}],"minor_comments":[{"comment":"The caption describes reinforcement learning as 'unsupervised learning'; reinforcement learning is neither supervised nor unsupervised in the standard taxonomy, and the phrase is misleading.","section":"Fig. 1 caption"},{"comment":"The description of the absolute binning approach and the example in Fig. 3 are difficult to follow; the text should define 'golden value' explicitly and clarify how the dominant-representation rule is applied across the full 16-state ADC output.","section":"§III-A"},{"comment":"The fitting procedure is said to minimize the sum of absolute differences, but the number of fitted parameters per cell and the degrees of freedom of the residual error are not stated. Adding this detail would make the extracted effective-bit statistics more reproducible.","section":"§III-B"},{"comment":"Fig. 7 reports resistance shift under different BL voltages, but the text does not state the number of cells averaged or the measurement conditions (e.g., temperature and cycle count) for each curve. This makes the claim of 'minimal resistance shift' hard to assess.","section":"§IV-A"},{"comment":"Reference [17] is cited for Neurosim, but the modified version with layer scaling and WAGE quantization is not publicly identifiable; the authors should state which version or commit of Neurosim was used and whether the modifications are available.","section":"References"}],"recommendation":"reject","confidential_remarks":"The paper is within the scope of ISCAS, and the hardware characterization data are potentially useful to the CIM community. However, the central claim of noise mitigation by reference tuning is contradicted by the paper's own Table III, and the simulation methodology lacks validation against hardware on the target workloads. Fixing these issues would require either substantially reframing the paper as a negative result for per-module tuning, or adding new hardware-in-the-loop experiments and statistical rigor. Given the page constraints of a conference paper and the load-bearing nature of the discrepancy, I do not see a path to a sound version within the current manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThe useful part of this paper is the characterization: they take a 40nm RRAM CIM macro, measure effective-bit distributions across 80 ADCs, and show that per-module reference tuning is insufficient for some workloads, so they move to per-ADC tuning. That is a real data set, and the read-disturb measurements (HRS shifts, LRS stable) are consistent with prior work. If you work on CIM reliability, the effective-bit extraction and the histogram of correct outputs are worth a look.\n\nThe problem is the central claim. The abstract and conclusion say per-module/per-ADC tuning 'maintains system robustness.' The drone navigation results in Table III say the opposite: quantize-only MSF is 899m, per-module tuning drops below 20m, and per-ADC tuning is 691m. Both are worse than doing nothing to the references. That is not a subtle mismatch; it is a direct contradiction between the paper's stated conclusion and its own experimental table. No error bars or seeds are reported, so a single unlucky rollout could explain some of it, but the point estimates as written falsify the claim for the time-dependent RL case. The text even acknowledges per-module tuning fails for ResNeXt, then switches to per-ADC, but the per-ADC drone number still does not help.\n\nThere are other soft spots. The ImageNet result is on a 'reduced-set' that is never defined. The read-disturb module is chosen because it is 'relatively linear,' which is a selection bias. The noise model is fit to the test chip and then injected into Neurosim, so the simulated accuracy drops are a propagation of the fit rather than an independent prediction; no hardware-in-the-loop validation on any of the four workloads. No code or data for reproduction.\n\nStill, the characterization methodology is not junk. The effective-bit concept with the LAD fit and residual-error handling is a reasonable way to capture ADC + RRAM variance. The per-module vs per-ADC comparison is a useful data point, even if the paper overinterprets it.\n\nAs submitted, I would not accept this. The central claim needs to be revised to match the actual results, or the drone experiment needs to be rerun with proper statistics. If the authors fix that, the characterization part could be a decent workshop paper or a section in a larger journal paper. For now, I would not cite it, and I wouldn't send it to a serious referee without first asking the authors to explain Table III.","headline":"The characterization data is real, but the paper's own drone results contradict its central claim that reference tuning maintains robustness, so the mitigation story does not hold.","tokens_in":9384,"tokens_out":3059,"would_cite":false,"duration_ms":27988,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Per-module and per-ADC reference tuning recovers most accuracy lost to RRAM and ADC noise in compute-in-memory accelerators.","keywords":["compute-in-memory","RRAM","ADC noise","reference tuning","read disturb","effective bits","reinforcement learning","noise injection"],"falsifier":"Run the same four workloads on the actual test chip with the same per-module/per-ADC reference settings and compare end-task accuracy to the simulated numbers, e.g., VGG-8 should land near 87.7% on CIFAR-10; a systematic gap would show the independent-Gaussian effective-bit model is not capturing real noise.","tokens_in":8407,"feed_emoji":"🔧","tokens_out":7625,"duration_ms":64979,"temperature":0.7,"pith_summary":"This paper argues that a large share of the accuracy lost to analog noise in RRAM-based compute-in-memory (CIM) accelerators can be recovered by tuning the reference voltages of the output ADCs, without changing the network or retraining. Using measurements from a 40nm RRAM test chip, the authors build an 'effective bits' model that folds RRAM cell variability and ADC non-ideality into per-cell mean and variance statistics, then inject those statistics as noise into simulations of four workloads: VGG-8 on CIFAR-10, ResNeXt-50 on ImageNet, a GridWorld agent, and a drone-navigation reinforcement-learning agent. They report that per-module reference tuning keeps VGG-8 within about 2.3 points of its clean quantized baseline, while the ImageNet model needs finer per-ADC tuning to stay near its baseline. The paper also characterizes read disturb and shows that a low-voltage read mode limits resistance drift. If the model is representative, the practical consequence is that a simple hardware calibration knob can substitute for more expensive retraining or redundancy in many CIM deployments.","feed_headline":"ADC reference tuning rescues noisy RRAM compute-in-memory","feed_subtitle":"A simple ADC calibration knob keeps neural-network and drone-navigation accuracy close to clean baselines.","key_machinery":"The 'effective bits' model: each RRAM cell's 0/1 contribution is replaced by a normal random variable whose mean and standard deviation are extracted from test-chip measurements, separately for HRS(0) and LRS(1), within 9-cell accumulation groups. A least-absolute-deviation fit over 256 random input vectors assigns an effective bit value and a residual error distribution to every cell; the residual is injected as dynamic noise. ADC reference tuning adjusts offset, step size, and bitline target voltage at global, per-module, or per-ADC granularity, which shifts the effective-bit distributions and therefore the end-task accuracy.","core_discovery":"The paper's claim is that an RRAM CIM macro's end-to-end inference error can be decomposed into a static component (per-cell effective 0/1 bits, measured by fitting 256 random 9-bit vectors per accumulation group with least absolute deviation) and a dynamic residual, and that both components shift with the ADC reference settings. When the ADC references are tuned per module (8 ADCs) or per ADC, the simulated accuracy of supervised classifiers and RL agents recovers most of the gap to their quantization-only baselines; per-ADC tuning is required for the high-precision ImageNet/ResNeXt case. The read-disturb experiments show HRS cells drift toward LRS under stress, which can push CIFAR-10 accuracy from 87% to 66%, and the paper's low-voltage read mode keeps the shifts minimal. In short, the central discovery is that reference tuning is a viable, low-cost mitigation knob for CIM noise.","pith_inferences":["If the effective-bit statistics generalize from this test chip to other RRAM macros, reference tuning could be applied as a per-instance calibration step after fabrication, reducing the need for noise-aware retraining on every chip.","The per-ADC versus per-module distinction suggests a build-time knob: designers can choose ADC reference granularity based on the precision demands of the target workload, with per-ADC only needed for high-precision tasks.","The read-disturb measurements imply a testable design rule: keeping bitline read voltage near 100-300mV should preserve the tuned effective-bit distributions over long inference runs; measuring drift after many read cycles would validate the claimed stability.","The same noise-injection pipeline could be used to rank other mitigation levers, such as write-verify or ECC, against reference tuning on the same four workloads, giving hardware designers a common benchmark."],"forward_implications":["Per-module reference tuning recovers most of the noise-induced accuracy drop on CIFAR-10/VGG-8, with simulated accuracy at 87.7% versus a 90.03% clean baseline.","The ImageNet/ResNeXt-50 workload fails with per-module tuning and requires per-ADC tuning, landing at 82% versus an 83% baseline, so finer-grained tuning is the fallback when module-level settings are insufficient.","In GridWorld, the per-module-tuned noise-injected policy still wins 78.2% of missions compared with a 95.5% win rate for the quantized-only policy, showing time-dependent RL tasks are more sensitive to CIM noise.","Read-disturb stress can drop CIFAR-10 accuracy from 87% to 66%, and operating at lower read voltages keeps resistance shifts small, so low-voltage read is a mitigation for long-horizon deployments."],"supporting_citations":[{"why":"Supplies the 40nm RRAM macro prototype whose modules, ADCs, and reference configurations produced the measured data.","marker":"[20]"},{"why":"Provides the modified simulator in which VGG-8 and ResNeXt inference accuracy under injected noise is evaluated.","marker":"[17]"},{"why":"Documents how read disturb shifts RRAM resistance and degrades inference accuracy, the basis for the read-disturb experiments.","marker":"[28]"},{"why":"Defines the ResNeXt-50 32x4d architecture used for the ImageNet evaluation.","marker":"[29]"},{"why":"Provides the integer-weight/activation quantization scheme used when training the supervised models.","marker":"[30]"}],"fun_headline_variants":["Reference tuning tames ADC noise in RRAM compute-in-memory","Per-ADC reference tuning restores RRAM CIM accuracy","RRAM CIM noise curbed by per-module ADC reference tuning","Low-cost ADC reference tuning quiets RRAM compute-in-memory","ADC reference tuning slashes RRAM CIM inference error"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measured per-cell effective-bit distributions from 10 modules (80 ADCs) on this one 40nm chip are treated as the noise experienced by the simulated networks, with independent Gaussian sampling per bit; if those measurements don't represent real deployment noise, the reported tuning benefits wouldn't transfer.","fun_headline_variants_meta":{"raw":{"variants":["Reference tuning tames ADC noise in RRAM compute-in-memory","Per-ADC reference tuning restores RRAM CIM accuracy","RRAM CIM noise curbed by per-module ADC reference tuning","Low-cost ADC reference tuning quiets RRAM compute-in-memory","ADC reference tuning slashes RRAM CIM inference error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000766,"raw_usage":{"total_tokens":3390,"prompt_tokens":933,"completion_tokens":2457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":2368}},"tokens_in":549,"tokens_out":2457,"duration_ms":17532,"temperature":1.0,"reasoning_tokens":2368,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T17:15:47.919732+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same four workloads on the actual test chip with the same per-module/per-ADC reference settings and compare end-task accuracy to the simulated numbers, e.g., VGG-8 should land near 87.7% on CIFAR-10; a systematic gap would show the independent-Gaussian effective-bit model is not capturing real noise.","supporting_citations":[{"cited_title":"A 40nm 64kb 26.56 tops/w 2.37 mb/mm 2 rram binary/compute-in-memory macro with 4.23x improvement in density and >75% use of sensing dynamic range,","cited_arxiv_id":null,"evidence_quote":"Supplies the 40nm RRAM macro prototype whose modules, ADCs, and reference configurations produced the measured data."},{"cited_title":"Dnn+neurosim v2.0: An end-to-end benchmarking framework for compute-in-memory accelerators for on-chip training,","cited_arxiv_id":null,"evidence_quote":"Provides the modified simulator in which VGG-8 and ResNeXt inference accuracy under injected noise is evaluated."},{"cited_title":"Impact of read disturb on multi- level rram based inference engine: Experiments and model prediction,","cited_arxiv_id":null,"evidence_quote":"Documents how read disturb shifts RRAM resistance and degrades inference accuracy, the basis for the read-disturb experiments."},{"cited_title":"Aggregated resid- ual transformations for deep neural networks,","cited_arxiv_id":null,"evidence_quote":"Defines the ResNeXt-50 32x4d architecture used for the ImageNet evaluation."}],"review_version":1}