{"id":"f6050687-bbc9-47fd-87ba-0323540e0c00","arxiv_id":"2411.18271","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A memristor column programmed with the inverse of a desired nonlinear function lets a ramp ADC output that function directly, eliminating the digital activation processing step in RNN inference.","lead":"This paper shows a way to compute neural network activation functions directly inside an analog memory chip's analog-to-digital converter, using a specially shaped ramp voltage. It could make recurrent neural networks like LSTMs much more energy-efficient on specialized in-memory computing hardware.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 'on-chip' results are software-emulated at the ADC periphery: the integrator and comparator are not fabricated, so the measured transfer functions and 88.5% KWS accuracy do not yet demonstrate the central hardware claim.","rationale":"The reader identified the un-fabricated integrator and comparator as the weakest assumption, and the stress-test concurs with that assessment. The paper's central contribution is the NL-ADC concept, whose mathematical derivation is not circular and whose memristor ramp generation is experimentally demonstrated. However, the most load-bearing part of the claims—that the activation is computed in hardware during digitization and that the system achieves the reported efficiency gains—depends entirely on peripheral analog circuits that are simulated in software. The measured transfer functions and KWS accuracy inherit this dependency. The efficiency advantage is particularly sensitive because the simulated peripheral circuits dominate the macro-level energy budget; even a modest deviation from the assumed comparator and integrator specifications would flip the energy-efficiency comparison against the proposed design. This does not invalidate the concept, but it means the paper should be read as a validated ramp-generation scheme plus a simulated ADC periphery, not as a full hardware demonstration. The reader's CONDITIONAL verdict is appropriate, and no adjustment is needed.","tokens_in":39589,"tokens_out":4831,"duration_ms":48168,"concrete_test":"Fabricate or assemble a full test chip that includes the actual integrator and comparator (in the same 180 nm process as the memristor array, or a well-characterized 16 nm implementation), and measure the end-to-end NL-ADC transfer function and KWS accuracy using the same programmed memristor column. Report the measured area, power, latency, and INL of the integrator and comparator. If the combined integrator+comparator energy per conversion exceeds 1.2X the 357.5 pJ assumed in Tab. S3, or if the measured end-to-end INL exceeds 1 LSB for the 5-bit configuration, then the projected 1.46X energy-efficiency advantage and the claim of on-chip nonlinear activation would not hold as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing claim is that the nonlinear activation is computed during digitization by real analog circuits. The paper explicitly states: 'The fabricated chip does not include the integrators and the comparator which are implemented in software after obtaining the crossbar output.' Therefore, the measured NL-ADC transfer functions (Fig. 3) and the reported on-chip inference accuracy of 88.5% (Fig. 4d) are obtained by applying a software model of the integrator and comparator to measured memristor column currents, not by an end-to-end hardware conversion. The system-level efficiency estimates in Tables S3/S10 assign 324.4 pJ to the integrator and 33.1 pJ to the comparator (out of 557.8 pJ total at macro level), sourced from references [51] and [45] and scaled to 16 nm. The SPICE validation in Supplementary Note S2 only varies integrator DC gain and gain-bandwidth product; it does not model comparator offset, noise, or mismatch, nor the actual layout parasitics of these blocks. Because the integrator plus comparator constitute roughly 64% of the macro energy, a modest underestimate in their power or area would eliminate the claimed 1.46X system-level energy-efficiency advantage over the conventional 5-bit ADC baseline. The mathematical idea of using a memristor-generated inverse-function ramp is sound, and the memristor programming is well characterized, but the central hardware demonstration—nonlinear activation integrated with a ramp ADC at the periphery—remains unbuilt. The abstract and Discussion overstate the demonstration by saying the activation is computed in the ADC and that 'all MAC and nonlinear operations are executed on chip.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes an in-memory nonlinear ADC (NL-ADC) for analog resistive crossbars, in which an extra memristor column is programmed with conductances proportional to the steps of the inverse of a desired activation function. Integration of that column's current produces a nonlinear ramp, and comparison with the MAC result yields a thermometer code that approximates the activation. The authors derive the mapping, experimentally program six activation functions on a memristor array, demonstrate a one-point calibration scheme, show robustness to read-voltage variations, and estimate accuracy and efficiency for keyword spotting (88.5% claimed) and Penn Treebank character prediction through simulations. The fabricated chip includes the memristor array but not the integrator and comparator, which are implemented in software after the crossbar output is obtained.","tokens_in":39876,"tokens_out":4964,"duration_ms":58996,"significance":"If the central hardware claim were fully validated, the NL-ADC would attack a genuine bottleneck in IMC-based RNNs by merging nonlinear activation with digitization and removing the digital activation processor. The mathematical derivation is sound and parameter-free in the sense that the ramp shape is not fitted to data, and the measured conductance-mapped ramps for six activations, the one-point calibration, and the read-voltage-tracking property are valuable experimental results. The public code and the detailed efficiency-estimation methodology are also strengths. However, because the ADC periphery is not fabricated, the paper does not currently demonstrate an integrated NL-ADC; the headline accuracy and efficiency claims are estimates contingent on simulated peripherals.","major_comments":[{"comment":"The paper states, 'The fabricated chip does not include the integrators and the comparator which are implemented in software after obtaining the crossbar output.' This means the transfer functions in Fig. 3 and the KWS accuracy of 88.5% in Fig. 4d combine measured memristor ramp generation with a software model of the analog periphery; they are not end-to-end hardware measurements. The Discussion's claim that the NL-ADC 'removes the need for any digital processor to implement nonlinear activations' is therefore not yet experimentally established. Please either fabricate and characterize the integrator and comparator or clearly re-label all such results as memristor measurements with simulated peripherals and temper the corresponding claims.","section":"Results, 'In-memory Implementation of Nonlinear ADC and Vector Matrix Multiply in a Crossbar Array'"},{"comment":"The macro-level energy budget assigns 324.42 pJ to the integrator and 33.10 pJ to the comparator out of a total of 557.79 pJ (Tab. S3), sourced from references [51] and [45] and scaled to 16 nm; these are not measured values for the present design. The SPICE validation in Supplementary Note S2 varies only integrator DC gain and gain-bandwidth product and does not model comparator offset, noise, mismatch, or layout parasitics. Since the integrator and comparator together dominate the macro energy, the claimed 1.46X system-level energy-efficiency advantage over a conventional 5-bit ADC baseline is not robust unless these peripheral assumptions are bounded by measurement or by a sensitivity analysis.","section":"Supplementary Note S3, Tab. S3"},{"comment":"Table 1 reports 'KWS task on GSCD (Accuracy %): 88.5' without qualification, and the abstract claims 'experimentally demonstrate the implementation of a non-linear activation function integrated with a ramp ADC.' Because the comparator and integrator are software models, this overstates the level of integration actually demonstrated. The comparisons to prior integrated LSTM chips (e.g., 9.9X area efficiency, 4.5X energy efficiency) should be relabeled as estimates or projections rather than measured system performance.","section":"Table 1 and Fig. 4"},{"comment":"The NLP scalability results in Fig. 5 and Tabs. S14–S17 are based on NeuroSim system-level simulations in which the ADC is replaced by the proposed NL-ADC model. These results are therefore simulation-based, not experimentally demonstrated, and the main text should state this explicitly; the current phrasing in the 'Scaling to large RNNs' section, relying on 'experimentally validated nonideality models,' is potentially misleading.","section":"Supplementary Note S4.b and Fig. 5"}],"minor_comments":[{"comment":"The notation 'f () =g−1()' obscures the argument; write f(t) = g^{-1}(t) for clarity.","section":"Eq. (2)"},{"comment":"The phrase 'tasted limited success' in the abstract is awkward; 'had limited success' would be clearer. There is also a typo in the Introduction: 'network ouptut' should be 'network output.'","section":"Abstract and Introduction"},{"comment":"Please state explicitly how the 'Normal ADC' INL results are generated (e.g., a simulated linear ramp with the same Vread sweep), since the comparison is central to the robustness claim.","section":"Fig. 3b"},{"comment":"In the header of Table 1, 'Nature'2332' contains a stray apostrophe, and reference formatting is inconsistent across the paper; please unify the citation style.","section":"Table 1"},{"comment":"The statement that the NL-ADC 'removes the need for any digital processor' should be qualified: non-monotonic activations such as GELU and Swish still require additional logic as described in Supplementary Note S12.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The central issue is the gap between the measured memristor-ramp generation and the simulated ADC periphery. The authors should be required to make this boundary explicit in the title, abstract, and main text, and to provide either a fabricated periphery or a carefully bounded sensitivity analysis of peripheral nonidealities. The technical content is otherwise sound and the paper should not be rejected outright."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The idea here is real: program a memristor column to generate the inverse of the activation function as the ramp in a ramp ADC, so the comparator output gives the activation directly. The math is simple, parameter-free, and correct. Using one memristor per ramp step is a natural fit to analog conductance states, and the one-point calibration is a clean way to correct programming errors. The measured ramp generation is the strongest part: INL around 0.8-1.2 LSB across six activations, and the read-voltage tracking argument is convincing. The KWS demo uses real programmed weights, noise-aware training, and reports 88.5% on 12-class GSCD, which is credible for that configuration.\n\nThe soft spot is exactly what the stress-test flags. The chip does not include the integrators or the comparator; those are emulated in software after reading the crossbar currents. So the \"on-chip nonlinear activation\" claim is only partially demonstrated: the ramp waveform itself is hardware, but the actual analog conversion is simulated. That distinction matters because the system-level energy and area tables assign roughly 64% of macro energy to the integrator plus comparator, with values taken from prior papers and scaled to 16 nm. The SPICE note varies DC gain and gain-bandwidth product but does not cover offset, noise, mismatch, or layout parasitics. A realistic 2x penalty in those blocks would eat most of the claimed 1.46X energy advantage over a conventional ADC baseline. I also think the non-monotonic function extension adds MUX/adder logic that does not clearly appear in the efficiency estimates.\n\nNone of this kills the contribution. The central concept is novel, the memristor programming is well-characterized, and the controlled comparison against a conventional ADC model is the right methodology. The abstract and Discussion simply overstate what was built: \"experimental demonstration\" of a nonlinear ADC should not be used when the comparator and integrator are software. The authors need to separate measured hardware results from simulated/projected system performance and re-derive the efficiency claims with a sensitivity analysis on the unbuilt blocks.\n\nThis paper deserves a serious referee. The idea is worth the community's time, and the measurements are real. With a revision that aligns the claims with the evidence, it would be a solid contribution to analog IMC for RNNs. I would engage with it.","headline":"The nonlinear-ramp ADC concept is sound and the memristor ramp generation is genuinely measured, but the paper oversells the hardware demonstration: the ADC's comparator and integrator are software models, not on-chip circuits.","tokens_in":40515,"tokens_out":2234,"would_cite":true,"duration_ms":25013,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By programming a memristor column with the inverse of a nonlinear activation, a ramp ADC can compute that activation during digitization, eliminating the digital nonlinearity step that bottlenecks recurrent neural network inference…","keywords":["analog in-memory computing","memristor crossbar","nonlinear activation function","ramp ADC","nonlinear ADC","LSTM inference","keyword spotting","hardware-aware training"],"falsifier":"Fabricate a complete macro with real integrators, comparators, and ripple counters; measure the 5-bit sigmoid and tanh transfer functions' INL and the keyword-spotting LSTM accuracy, and compare measured LSTM-layer area and energy against a same-precision conventional ramp-ADC baseline. If the realized INL exceeds roughly 1 LSB, or the measured accuracy falls below the reported 88.5% by more than a few points, or the area/energy ratios do not approach the projected 6.2X and 1.46X, the central hardware claim is refuted.","tokens_in":39374,"feed_emoji":"🧠","tokens_out":17452,"duration_ms":139061,"temperature":0.7,"pith_summary":"Recurrent networks such as LSTM apply sigmoid and tanh after every matrix-vector product, and in analog in-memory computing those nonlinearities are a latency and energy bottleneck. The paper's claim is that these activations can be computed during analog-to-digital conversion—inside a nonlinear ADC (NL-ADC)—rather than in a separate digital processor. The NL-ADC programs an extra memristor column with the inverse of the desired activation, so the ramp voltage is predistorted and the comparator crossing time directly encodes the activated value. On-chip experiments on a 12-class keyword-spotting LSTM reach 88.5% measured accuracy at 5-bit, and system-level estimates show the LSTM layer at roughly 6.2X area efficiency and 1.46X energy efficiency versus a same-precision conventional ramp-ADC baseline. A scaled simulation on a 6.1-million-weight character-prediction LSTM stays within about 0.015 bits per character of software accuracy, indicating the approach carries over to larger recurrent models.","feed_headline":"Memristor ramp ADC computes nonlinear activations during conversion","feed_subtitle":"Folding sigmoid and tanh into digitization removes the digital bottleneck for in-memory LSTM inference.","key_machinery":"The central object is the inverse-function ramp identity $t_{\\mathrm{in}} = g^{-1}(V_{\\mathrm{in}})$: when the ADC ramp $f(t)$ is chosen so that $f = g^{-1}$, the comparator crossing time becomes the activation value $g(V_{\\mathrm{in}})$, so the activation is computed during digitization. The hardware that realizes this is a dedicated memristor column in the same crossbar: each of the $2^b$ ramp steps is a programmable conductance $G_{\\mathrm{adc},k} \\propto \\Delta V_k$, integrated on a feedback capacitor to build the predistorted ramp, with the comparator being the existing sense amplifier. A few bias memristors set the ramp's start voltage and perform a one-point calibration that corrects programming errors and stuck devices, and using the same memristor circuits for MAC and reference makes read-voltage noise common-mode.","core_discovery":"The central claim is that the nonlinear activation of a recurrent network can be made physically part of the analog-to-digital conversion step in a memristive crossbar. The method replaces the linear ramp of a conventional ramp ADC with a predistorted ramp whose shape is the inverse of the desired activation $g^{-1}$; the comparator then trips at a time proportional to $g(V_{\\mathrm{in}})$ instead of $V_{\\mathrm{in}}$. In the crossbar, the ramp is generated by one extra memristor column whose conductances $G_{\\mathrm{adc},k}$ encode the successive step sizes $\\Delta V_k = g^{-1}(t_k) - g^{-1}(t_{k-1})$, with a small set of bias memristors providing a one-point calibration. Because the MAC result and the ADC reference come from the same memristor circuits, read-voltage fluctuations cancel, giving maximum integral nonlinearity (INL) between 0.02 and 0.44 LSB when read voltage varies from 0.15 V to 0.25 V. The paper supports the claim with on-chip inference of a 12-class keyword-spotting LSTM (32 hidden neurons, 9216 memristors) reaching 88.5% accuracy at 5-bit, and with system-level estimates showing roughly 6.2X area efficiency and 1.46X energy efficiency over a same-precision conventional ramp-ADC baseline, plus a scaled 6.1-million-weight character-prediction LSTM that stays within 0.015 bits per character of software.","pith_inferences":["The inverse-ramp construction is not limited to activations; any bijective function of the MAC output, such as a quantizer, a normalizer, or a log-softmax term, could in principle be fused into the same conversion by reprogramming the ramp column.","The decisive next experiment is a full macro that includes the real integrators, comparators, and ripple counters the current chip omits; measuring the LSTM-layer energy, area, and accuracy directly would verify the projected gains rather than relying on estimates.","Because the supplement shows non-monotonic GELU and Swish can be approximated by splitting the inverse into monotonic pieces, the same mechanism is a plausible route to transformer-style workloads, where those activations are common.","Passing the pulse-width-modulated output directly to the next layer, which the paper mentions but does not quantify, would compound the savings in deep RNNs by also removing input DAC or PWM-generation circuits, so the efficiency advantage should grow with depth."],"forward_implications":["LSTM inference can be executed with the nonlinear activations produced directly by the ADC, so the digital processor that normally computes sigmoid and tanh is removed from the critical path.","On the fabricated 9216-memristor crossbar, the 5-bit NL-ADC achieved 88.5% measured accuracy on a 12-class keyword-spotting task; the 4-bit and 3-bit versions achieved 86.6% and 85.2%.","At the system level the LSTM layer is projected to be about 6.2X more area-efficient and 1.46X more energy-efficient than a same-precision conventional ramp-ADC baseline, and roughly 9.9X and 4.5X better than earlier LSTM circuits.","For a 6.1-million-weight character-prediction LSTM, simulated 5-bit NL-ADC reaches 1.349 bits per character under measured write and read noise, close to the 1.334 software baseline.","Because the ramp and the MAC result use the same memristor circuits, read-voltage variation cancels: maximum INL stays between 0.02 and 0.44 LSB across a 0.15-0.25 V read sweep, versus 4.12-5.5 LSB for a conventional ADC."],"supporting_citations":[{"why":"Supplies the crossbar, driver, and sample-and-hold area and energy data used in the system efficiency estimates.","marker":"15"},{"why":"A 64-core IMC chip with shared digital nonlinearity units; it is the main baseline for the character-prediction BPC and efficiency comparisons.","marker":"23"},{"why":"An analog-AI chip for speech recognition whose off-chip nonlinearity processing is the bottleneck this work removes; it is the main accuracy and efficiency comparison point for KWS.","marker":"32"},{"why":"Provides the CORDIC-based digital processor used to estimate the area, energy, and latency of the conventional nonlinearity path.","marker":"36"},{"why":"Provides the defect-aware training method used to make the LSTM models tolerant to memristor programming noise.","marker":"41"},{"why":"Defines the speech-commands keyword-spotting benchmark and its 12-class split used in the on-chip KWS experiments.","marker":"42"},{"why":"Defines the character-prediction corpus used for the large-network scaling study.","marker":"43"},{"why":"A prior in-memory ramp ADC built with SRAM, used as the overhead comparison showing why memristor-programmed steps are more compact.","marker":"45"},{"why":"A fully integrated ReRAM compute-in-memory chip with a conventional 5-bit ramp ADC; it supplies the baseline ADC area, energy, and latency for the efficiency comparisons.","marker":"52"},{"why":"Another recent memristor chip with a conventional ADC, used as a comparison point in the ADC performance table.","marker":"53"}],"fun_headline_variants":["Nonlinear activations baked into memristor ADC conversion","Fold sigmoid and tanh into the ramp ADC itself","One extra column turns ramp ADC into an activation function"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the integrator and comparator, which are not fabricated on the chip and are instead emulated in software after reading the crossbar output, can be physically built with the area, speed, noise, and power assumed in the circuit simulations and system-level estimates; if those real circuits are larger, slower, noisier, or more power-hungry than modeled, the projected efficiency gains and the on-chip inference accuracy would not carry over to a complete system.","fun_headline_variants_meta":{"raw":{"variants":["Nonlinear activations baked into memristor ADC conversion","Fold sigmoid and tanh into the ramp ADC itself","One extra column turns ramp ADC into an activation function"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000555,"raw_usage":{"total_tokens":2710,"prompt_tokens":1080,"completion_tokens":1630,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":696,"completion_tokens_details":{"reasoning_tokens":1578}},"tokens_in":696,"tokens_out":1630,"duration_ms":10629,"temperature":1.0,"reasoning_tokens":1578,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:22:28.081968+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fabricate a complete macro with real integrators, comparators, and ripple counters; measure the 5-bit sigmoid and tanh transfer functions' INL and the keyword-spotting LSTM accuracy, and compare measured LSTM-layer area and energy against a same-precision conventional ramp-ADC baseline. If the realized INL exceeds roughly 1 LSB, or the measured accuracy falls below the reported 88.5% by more than a few points, or the area/energy ratios do not approach the projected 6.2X and 1.46X, the central hardware claim is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the crossbar, driver, and sample-and-hold area and energy data used in the system efficiency estimates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A 64-core IMC chip with shared digital nonlinearity units; it is the main baseline for the character-prediction BPC and efficiency comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CORDIC-based digital processor used to estimate the area, energy, and latency of the conventional nonlinearity path."},{"cited_title":"& Marcinkiewicz, M","cited_arxiv_id":null,"evidence_quote":"Defines the character-prediction corpus used for the large-network scaling study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A prior in-memory ramp ADC built with SRAM, used as the overhead comparison showing why memristor-programmed steps are more compact."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A fully integrated ReRAM compute-in-memory chip with a conventional 5-bit ramp ADC; it supplies the baseline ADC area, energy, and latency for the efficiency comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Another recent memristor chip with a conventional ADC, used as a comparison point in the ADC performance table."}],"review_version":1}