{"id":"db067f4b-0579-4947-abdf-a6dcc499f243","arxiv_id":"2501.19347","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"A 65 nm all-digital Tsetlin machine accelerator chip classifies MNIST images at 8.6 nJ per frame and 97.42% accuracy, the lowest fully digital EPC reported for this benchmark.","lead":"This paper presents a manufactured 65 nm chip that runs Tsetlin machine image classification in purely digital logic, using 8.6 nJ per image on MNIST with 97.42% accuracy. It matters because it gives the Tsetlin machine community its first silicon proof point and shows a logic-based alternative approaching analog in-memory accelerators on energy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 8.6-nJ/frame headline is a core-only measurement; adding the measured 0.76 mW 3.3 V I/O power raises system energy to ~21 nJ/frame, which is no longer below the 12.92 nJ mixed-signal SNN in [21] or near the 3.32 nJ CNN in [20].","rationale":"The paper is a genuine hardware contribution: a fabricated 65 nm CoTM ASIC with convolutional inference, measured classification rates and accuracies across three datasets, public VHDL, and a clearly described architecture. The internal arithmetic is consistent: 0.52 mW at 60.3k images/s gives 8.6 nJ, 1.15 mW gives 19.1 nJ, and the I/O power of 0.76 mW adds about 12.6 nJ per frame. The reader's weakest-assumption analysis correctly identifies the core-only accounting boundary as the most load-bearing issue, because the headline novelty claim is explicitly comparative ('lowest fully digital', 'second lowest overall'). My attack does not question the measured core power, the accuracy numbers, or the architecture; it questions whether the metric chosen for the headline is the right one for the comparative claim being made. Since the authors disclose the I/O power and explain their SoC-integration rationale, the fix is a clear re-scoping of the claim plus a total-power measurement, not a rejection of the engineering result. The conditional verdict already given remains appropriate: the paper should be accepted with the requirement that the EPC claim be stated with its boundary condition and, ideally, with a total chip-power number for the present test chip.","tokens_in":20855,"tokens_out":3102,"duration_ms":34332,"concrete_test":"Measure total power at all supply pins (0.82 V core and 3.3 V I/O) with the Joulescope during continuous classification of the full 10k-image test set at 27.8 MHz, with clock gating and CSRF enabled, and recompute EPC as total power divided by 60.3k images/s. Then re-run the Table IV comparison using this total-power EPC. If the resulting EPC exceeds 12.92 nJ, the 'second lowest overall' claim should be revised to a core-only ranking or accompanied by a clearly labeled SoC-level estimate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the Introduction and repeated in Table IV and Section VII, is that this ASIC achieves the lowest reported energy per classification among fully digital MNIST solutions and the second lowest overall. That claim rests on reporting only accelerator-core power: 0.52 mW at 0.82 V and 27.8 MHz, giving 8.6 nJ at 60.3k images/s. Section V separately reports 0.76 mW consumed by the 3.3 V digital I/O pads in inference mode. Using the same throughput, total measured chip power is 1.28 mW, or 21.2 nJ per classification. That is 1.6x the 12.92 nJ mixed-signal SNN in [21] and far above the 3.32 nJ analog CNN in [20]. The authors justify excluding I/O power by saying a production SoC would optimize the digital interface, but this is an extrapolation, not a measurement on the presented chip. The headline EPC and the comparative ranking are therefore conditional on an accounting boundary that excludes a measured, non-negligible component of the test chip's power. This is the load-bearing assumption in the paper's main contribution: without it, the 'second lowest overall' claim is not supported by the measured data, and even the 'lowest fully digital' claim would need to be re-scoped to 'lowest core-only fully digital' until a comparable total-power number is reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the first fabricated ASIC based on the coalesced Tsetlin machine (CoTM) with convolution, implemented in 65 nm low-leakage CMOS. The accelerator stores all clause weights and Tsetlin-automaton action signals in registers, evaluates 128 clauses in parallel, and performs inference on booleanized 28x28 images with 10 classes. At 27.8 MHz and 0.82 V core supply, the authors measure 60.3k classifications/s, 0.52 mW core power, and 8.6 nJ per classification; test accuracies are 97.42% on MNIST, 84.54% on Fashion-MNIST, and 82.55% on Kuzushiji-MNIST, matching software models. The paper claims this is the lowest reported energy per classification among fully digital MNIST solutions and the second lowest overall, and it includes measured results, a comparison table, and estimates for scaled-up 28 nm and CIFAR-10 variants.","tokens_in":21215,"tokens_out":4411,"duration_ms":41646,"significance":"If the headline energy figure holds under a fair comparison, this is a useful silicon demonstration: it is the first manufactured CoTM ASIC, it is fully digital and tool-flow-compatible, the VHDL is publicly available, and the accuracy numbers are measured on independent test sets rather than simulated. The internal consistency of the measurements (0.52 mW / 60.3 kHz = 8.6 nJ at 0.82 V; 471 cycles at 27.8 MHz matching the reported throughput) is a strength. The significance is however conditional on the energy-accounting boundary used in the comparative claims, because the paper itself reports a separate 0.76 mW I/O power draw that is excluded from the headline figure.","major_comments":[{"comment":"The 8.6 nJ headline is a core-only measurement: Section V reports 0.52 mW accelerator-core power at 0.82 V and separately reports 0.76 mW consumed by the 3.3 V digital I/O pads in inference mode. Adding these gives 1.28 mW total chip power, or about 21.2 nJ per classification at 60.3k images/s, which is above the 12.92 nJ of [21] and far above the 3.32 nJ of [20]. The Introduction, Section VII, and Table IV use the 8.6 nJ value for the 'lowest fully digital' and 'second lowest overall' claims, so the comparison is load-bearing. The paper should report total measured chip EPC alongside the core-only value, or demonstrate that every comparator in Table IV is also core-only and excludes its I/O or interface power; the SoC-integration argument in Section V is a reasonable projection but is not a measurement on this chip.","section":"Section V, Table II, Table IV"},{"comment":"The record claim rests on a single average measurement from one chip at one operating point (0.82 V, 27.8 MHz), with no error bars, repeated-measurement spread, or chip-to-chip variation reported. Because the paper claims the lowest reported EPC for a fully digital MNIST solution, the sensitivity of the EPC to supply voltage, clock frequency, and measurement repeatability should be quantified, and Table IV should state which numbers are single-chip measurements and which are multi-sample or simulated values.","section":"Section V, Table II"},{"comment":"The comparison in Table IV mixes measured and estimated data without a clear accounting boundary for energy. In particular, the 'This work scaled to 28 nm' column reports 4.3 nJ as an estimate based on Dennard scaling and a literal-count assumption from Section VI-A, but it is presented in the same table as measured EPC values from other chips. This makes the 'second lowest overall' claim hard to reproduce. The table should visually separate measured results from projections, or the projections should be moved to the discussion.","section":"Section VII, Table IV"}],"minor_comments":[{"comment":"The abstract states that the chip 'consumes 8.6 nJ per classification' without the qualifier 'accelerator core only'; since Section V explicitly distinguishes core power from I/O power, the abstract and Introduction should carry the same qualification.","section":"Abstract and Section V"},{"comment":"For the 1 MHz, 0.82 V row, the reported 21 uW core power and 2.27k images/s imply 9.25 nJ per classification, but the table lists 9.6 nJ; please reconcile or explain the discrepancy.","section":"Table II"},{"comment":"The text says 'close to what is achieved for the SNN accelerator in [20]', but Table IV identifies [20] as a CNN accelerator, not an SNN; please correct this reference error.","section":"Section VI-A"},{"comment":"The sentence following Eq. (5) writes the position-bit term as '(Y - WY) + (X - WY)', but the equation and the intended symmetry require '(Y - WY) + (X - WX)'; please correct the typo.","section":"Section III-C"},{"comment":"The CSRF feedback is listed as a contribution and is claimed to reduce switching activity, but Section V reports that CSRF alone provides less than 1% measured power reduction; the paper should state explicitly that the benefit was verified mainly in toggling-rate simulation, not in end-to-end energy.","section":"Section IV-D and Section V"}],"recommendation":"major_revision","confidential_remarks":"This is a well-executed chip demonstration with internally consistent measurements and open-source VHDL, which I regard as genuine strengths. The main risk is fairness of the energy comparison: the core-only accounting boundary flips the ranking once the measured I/O power is included. I would ask the authors to report total chip EPC and to make the accounting boundary of every comparator explicit. This is fixable within the scope of the manuscript, so I do not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. This is the first manufactured ASIC for the coalesced Tsetlin machine with convolution, and the measurements look solid. The chip hits 97.42% on MNIST at 60.3k images/s, and the core power arithmetic checks out: 0.52 mW / 60.3 kHz = 8.6 nJ. Public VHDL, three datasets, internal consistency between latency and throughput, and the test accuracy matches the software models. That's real work, and the TM hardware community now has a proper silicon reference point.\n\nThe main caveat is the energy headline. The 8.6 nJ is accelerator-core only; the measured 3.3 V I/O pads add 0.76 mW, which more than doubles system-level energy to roughly 21 nJ/image. The authors report that I/O power and explain why they think an SoC would optimize it away, so it's not hidden. But the 'lowest fully digital / second lowest overall' claim in the introduction and Table IV is only supported under that boundary, and the comparison table doesn't tell us whether the prior chips include their I/O. The claim should be re-scoped to 'lowest core-only fully digital' and the table should state the assumed accounting boundaries. That is a revision request, not a rejection.\n\nThe other new item, the CSRF feedback circuit, delivers less than 1% measured power reduction, so it's a minor addition. The paper admits this. The CIFAR-10 scaling estimates in Section VI are clearly speculative and fine as discussion.\n\nOverall, this is a solid engineering contribution: well reported, reproducible, and the accuracy/throughput claims hold. The energy-efficiency claim is softer than advertised because of the I/O boundary, but the transparency is good. The paper is for people working on low-power edge ML, Tsetlin machines, or hardware/algorithm co-design; they will get real value from the silicon results. I'd send it to peer review and ask for a re-scoped comparison and a total-power number for the test chip. Not a desk reject.","headline":"First silicon CoTM ASIC with honest measurements, but the 8.6-nJ headline is core-only and the energy comparison needs re-scoping.","tokens_in":21788,"tokens_out":3670,"would_cite":true,"duration_ms":37000,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fully digital 65 nm Tsetlin-machine accelerator classifies 28×28 images at 60.3k per second, using 8.6 nJ per classification and matching software accuracy on three datasets.","keywords":["Tsetlin machine","ASIC accelerator","image classification","low-power inference","coalesced Tsetlin machine","convolution","energy per classification","65 nm CMOS"],"falsifier":"Measure total chip power (core plus 3.3 V I/O) during continuous classification at 27.8 MHz and 0.82 V, and divide by the 60,300 images/s rate; the paper's own numbers imply roughly 21 nJ per image on that boundary. If a system-level energy accounting is required for the comparative claims, this measurement refutes them against several mixed-signal comparators, and a second fully digital design reporting under 8.6 nJ on the same core-only metric would refute the 'lowest fully digital' claim directly.","tokens_in":20669,"feed_emoji":"⚡","tokens_out":9696,"duration_ms":82921,"temperature":0.7,"pith_summary":"The paper aims to show that the Tsetlin machine, a learning algorithm based on Boolean logic rather than neural-network arithmetic, can be built as an all-digital accelerator chip that matches the energy efficiency of analog in-memory-computing designs. The manufactured 65 nm test chip classifies 28×28 greyscale images into ten classes at 60,300 images per second, using only 8.6 nJ per classification at a 0.82 V core supply, while reproducing software accuracy on three standard benchmarks. If correct, this puts logic-based machine learning in the same energy class as analog neural accelerators, with the practical advantage that a fully digital design ports to standard digital workflows and any CMOS process. The claim is bounded: the 8.6 nJ figure counts only core power, excluding a 3.3 V I/O block drawing 0.76 mW.","feed_headline":"Tsetlin-machine chip: 8.6 nJ per image classification","feed_subtitle":"The logic-based learner matches analog accelerators' energy efficiency with a fully digital 65-nm ASIC.","key_machinery":"The mechanism that carries the design is a clause pool with sequential-OR convolution: 128 clauses, each a conjunction of literals selected by Tsetlin automaton actions, are evaluated over 361 overlapping 10×10 patches, and a clause counts as fired for an image if it fired in any patch. Because every TA action bit and every signed class weight is held in registers (45,056 bits total), clause evaluation is combinational and the ten class sums are built by multiplexer-and-adder reduction trees with no multiplications, since clause outputs are 0 or 1. Two power-saving mechanisms are specific to this design: the clause switching reduction feedback (CSRF), which ORs the latched clause output into the clause's own combinational path so that a fired clause stops toggling, and separate clock domains with clock gating that stop the model registers (about 90% of the flip-flops) entirely during inference. CSRF cut the simulated combinational toggling rate by roughly half, though the measured total power reduction was under 1%, because the inference core's clock tree and DFFs dominate.","core_discovery":"The paper reports a manufactured 65 nm CMOS accelerator that implements inference for the coalesced Tsetlin machine (CoTM) with a 10×10 sliding convolution window, applied to 28×28 booleanized images with ten classes. With 128 clauses sharing a single clause pool, all model state—34,816 Tsetlin automaton action bits and 10,240 signed 8-bit clause weights—sits in registers, so each of the 361 patches is evaluated by purely combinational clause logic in one clock cycle and class sums are built by addition-only reduction trees. Measured at 27.8 MHz and 0.82 V core supply, the chip classifies 60,300 images/s with 25.4 µs latency and 8.6 nJ per classification, producing test accuracies of 97.42% on MNIST, 84.54% on Fashion-MNIST, and 82.55% on Kuzushiji-MNIST, identical to the software models. The authors present this as the lowest energy per classification reported for a fully digital MNIST accelerator with comparable accuracy, and the second lowest overall among manufactured solutions, behind an analog time-domain CNN at 3.32 nJ. The design adds two power-saving mechanisms: the clause switching reduction feedback (CSRF), which stops re-evaluating clause logic once a clause has fired, and clock gating with separate clock domains for the static model and the active inference core.","pith_inferences":["The comparative claims are sensitive to the energy-accounting boundary: the paper's own measurements put the 3.3 V I/O block at 0.76 mW, which adds roughly 12.6 nJ per image at 60.3k images/s, so under a full-system boundary the 8.6 nJ figure becomes about 21 nJ and the 'lowest fully digital' claim would not survive against all comparators.","The near-zero measured benefit of the CSRF technique (under 1% power reduction) despite a simulated 50% drop in clause-logic toggling implies that the clause combinational logic is not the dominant energy consumer; future power gains would have to come from the clock tree and register storage.","The same 88% of TA actions that are 'exclude' states suggests a sparse model representation with literal addressing, as sketched for a 28 nm version, could cut roughly 47% of core area—a design change the paper estimates but does not fabricate.","The scaling estimates for 28 nm and for CIFAR-10 rest on Dennard scaling and linear area scaling with model size; a fabricated 28 nm port that deviates far from the predicted 4.3 nJ would point to those assumptions rather than to the architecture itself."],"forward_implications":["A Tsetlin-machine classifier now has a silicon-proven, fully digital implementation whose 8.6 nJ per classification is the lowest reported among fully digital MNIST accelerators with matching accuracy, putting logic-based learners in the same energy class as analog mixed-signal designs.","Because the design uses a standard synchronous digital flow, it can be re-targeted to other process nodes; the paper's estimates for a 28 nm port with literal-limited clauses put the core at 0.27 mm² and about 4.3 nJ per classification, close to the 3.32 nJ of the best analog solution.","For larger images, the architecture scales by composing specialized Tsetlin machines; the paper's TM Composites estimate for CIFAR-10 is 0.9 µJ per classification in 65 nm and 0.45 µJ in 28 nm, at an estimated 79% accuracy.","The measured accuracy exactly matches the software models, indicating that the register-based inference path introduces no accuracy degradation."],"supporting_citations":[{"why":"Defines the Tsetlin machine with clauses and Tsetlin automata that the whole accelerator implements.","marker":"[10]"},{"why":"Introduces the convolutional Tsetlin machine, supplying the sliding-window convolution and booleanization used here.","marker":"[13]"},{"why":"Defines the coalesced Tsetlin machine with a shared clause pool that the accelerator's architecture is based on.","marker":"[19]"},{"why":"The FPGA ConvCoTM accelerator whose inference core and training concepts this ASIC builds on; reports 97.6% MNIST accuracy.","marker":"[12]"},{"why":"The first Tsetlin-machine ASIC; supplies the TA structure and the argmax submodule reused in this design.","marker":"[11]"},{"why":"Charge-domain in-memory-computing ternary CNN at 0.18 µJ EPC on MNIST; a key analog comparator for the energy claims.","marker":"[9]"},{"why":"Time-domain analog CNN with 3.32 nJ EPC on MNIST, the lowest reported figure the paper places its 8.6 nJ behind.","marker":"[20]"},{"why":"Mixed-signal neuromorphic SNN with 12.92 nJ EPC on MNIST; another comparator in the energy comparison.","marker":"[21]"},{"why":"The MNIST dataset used as the primary benchmark for accuracy and energy comparison.","marker":"[14]"}],"fun_headline_variants":["All-digital Tsetlin chip hits 8.6 nJ per image","Tsetlin machine ASIC: 8.6 nJ/frame at 60k fps","Logic-based accelerator: 8.6 nJ per MNIST image","65-nm digital Tsetlin chip: 8.6 nJ per classification","Chip runs Tsetlin machine at 8.6 nJ and 97% MNIST accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline energy figure counts only the accelerator core at a single operating point and excludes the test chip's 3.3 V I/O power, which at 0.76 mW adds roughly 12.6 nJ per image and would more than double the energy per classification under a full-system boundary.","fun_headline_variants_meta":{"raw":{"variants":["All-digital Tsetlin chip hits 8.6 nJ per image","Tsetlin machine ASIC: 8.6 nJ/frame at 60k fps","Logic-based accelerator: 8.6 nJ per MNIST image","65-nm digital Tsetlin chip: 8.6 nJ per classification","Chip runs Tsetlin machine at 8.6 nJ and 97% MNIST accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000903,"raw_usage":{"total_tokens":3963,"prompt_tokens":1102,"completion_tokens":2861,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":718,"completion_tokens_details":{"reasoning_tokens":2747}},"tokens_in":718,"tokens_out":2861,"duration_ms":19567,"temperature":1.0,"reasoning_tokens":2747,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T20:28:16.306705+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure total chip power (core plus 3.3 V I/O) during continuous classification at 27.8 MHz and 0.82 V, and divide by the 60,300 images/s rate; the paper's own numbers imply roughly 21 nJ per image on that boundary. If a system-level energy accounting is required for the comparative claims, this measurement refutes them against several mixed-signal comparators, and a second fully digital design reporting under 8.6 nJ on the same core-only metric would refute the 'lowest fully digital' claim directly.","supporting_citations":[{"cited_title":"Tsetlin machine-based image classification FPGA accelerator with on- device training,","cited_arxiv_id":null,"evidence_quote":"The FPGA ConvCoTM accelerator whose inference core and training concepts this ASIC builds on; reports 97.6% MNIST accuracy."},{"cited_title":"Learning automata based energy-efficient AI hardware design for IoT applications,","cited_arxiv_id":null,"evidence_quote":"The first Tsetlin-machine ASIC; supplies the TA structure and the argmax submodule reused in this design."},{"cited_title":"An in-memory-computing charge-domain ternary CNN classifier,","cited_arxiv_id":null,"evidence_quote":"Charge-domain in-memory-computing ternary CNN at 0.18 µJ EPC on MNIST; a key analog comparator for the energy claims."},{"cited_title":"A 28- nm 3.32-nJ/frame compute-in-memory CNN processor with layer fusion for always-on applications,","cited_arxiv_id":null,"evidence_quote":"Time-domain analog CNN with 3.32 nJ EPC on MNIST, the lowest reported figure the paper places its 8.6 nJ behind."},{"cited_title":"A 65 nm 12.92-nJ/inference mixed-signal neuromorphic processor for image classification,","cited_arxiv_id":null,"evidence_quote":"Mixed-signal neuromorphic SNN with 12.92 nJ EPC on MNIST; another comparator in the energy comparison."}],"review_version":1}