{"id":"cf2ebf27-d2a9-4ada-b3bf-174a6d9fa348","arxiv_id":"2509.05532","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"A physically fabricated superconducting spiking neural network chip with 5,822 Josephson junctions classifies a three-digit MNIST subset at 80.07% accuracy after 7x7 downsampling and quantization.","lead":"This paper reports a framework for designing superconducting spiking neural network chips under real fabrication limits, and describes a fabricated chip that classifies a small subset of MNIST digits. A generalist might read it to see whether superconducting neuromorphic hardware can move from simulations to physical chips.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fabricated-chip accuracy is asserted without measured data: no test protocol, raw outputs, or measured clock/power/energy; reported accuracy is internally inconsistent (80.07% vs 86.2% for the same chip).","rationale":"The stress-test target is the strongest claim: fabricated chip classification accuracy. I read the paper in good faith as a design/framework paper claiming a tapeout. The framework contributions could stand even without chip data, but the abstract and Section 4 explicitly claim 'fabricated SuperSNN chip successfully classified...' So the central claim is empirical. To support it, one would need (a) a description of the cryogenic test setup and measurement protocol, (b) a definition of the test set (which 2/3/4 samples, how many), (c) raw chip outputs, and (d) a rule for mapping output pins to predictions. None appears. The only accuracy tables are software evaluations of trained networks (Table VI) and resource comparisons (Table V). The 3.02 GHz and 6.55 fJ numbers are explicitly from simulation/estimation, not measurement. This is the same load-bearing weakness the Reader identified, so agreement is 'agree.' I also note the internal inconsistency: 80.07% vs 86.2% both attributed to the 2/3/4 chip in different places (abstract/Section 4.1/Table VI vs Table V/Section 4.3). That inconsistency is itself a correctness risk, but the deeper issue is the absence of measured evidence. No ad hominem is intended: the concern is about evidence, not intent. If the authors can provide raw test data, the claim could be verified; absent that, a hardware-demonstration claim cannot be accepted. Therefore I do not adjust the Reader's REJECT verdict, though the rejection is based on unverifiability rather than a demonstrated counterexample.","tokens_in":13193,"tokens_out":5393,"duration_ms":53626,"concrete_test":"Ask the authors to release the measured test protocol and raw chip outputs for the fabricated 2/3/4 chip: for each test image, the three output-pin waveforms/DC levels over the ten-cycle prediction window and the resulting label, plus the exact test-set size. Independently recompute classification accuracy from those traces using the paper's criterion (exactly one correct output neuron spikes, others silent) and compare with 80.07% (and 86.2%). If the raw outputs cannot be supplied, or the recomputed accuracy differs by more than the sampling error of the reported test set, the fabricated-chip accuracy claim is unsupported and should be reclassified as a simulation/design estimate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that a fabricated SuperSNN chip 'successfully classified' digits 2/3/4 with 80.07% (abstract) and 86.2% (abstract, Table V, Section 4.3). This is the result that would make the contribution a hardware demonstration rather than a design study. The manuscript never provides measured chip evidence. Section 4.2 describes implementation and timing but gives no test setup, no list of test images, no raw output pin traces, no confusion matrix, and no description of how output DC levels were converted to labels. Section 4.3 presents only software-computed accuracy for digit combinations (Table VI) and resource estimates. The 3.02 GHz figure comes from Fig. 3(b), a circuit simulation ('Simulation result...'), not from chip measurement; the 6.55 fJ is described in the conclusion as 'estimated switching energy.' The accuracy reporting is internally inconsistent: Section 4.1 and Table VI give 80.07% for 2,3,4 and 86.2% for 0,1,2; Table V and Section 4.3 attribute 86.2% to the 2,3,4 chip. Thus the reader cannot tell which number was observed on hardware, on what test set, or whether any number was observed at all. Because every headline hardware claim depends on this missing measurement, the central empirical result is unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SuperSNN, an end-to-end framework for designing physically realizable superconducting spiking neural networks under MIT LL SFQ5ee constraints. It combines hardware-aware training (pruning, weight quantization, hybrid spike/membrane loss), custom SFQ cells including a high fan-in neuron, a shift-register input scheme, and a locally synchronous/globally synchronous (LAGS) clock distribution. The authors report a complete MNIST network with 96.47% accuracy after quantization/pruning and a fabricated chip for digits 2/3/4 with claimed hardware accuracies of 80.07% (abstract, Table VI) and 86.2% (Abstract, Table V, §4.3), operating at 3.02 GHz, with 5,822 JJs, 2.15 mW static power, and 6.55 fJ per inference. The central claim is that a full SFQ SNN chip was fabricated and successfully classified MNIST subsets.","tokens_in":13592,"tokens_out":3673,"duration_ms":38261,"significance":"If the fabricated-chip results are genuine and reproducible, this would be a notable advance in superconducting neuromorphic hardware: it would demonstrate a complete, physically realized SFQ SNN within a commercial fabrication process, with an extremely high clock rate and low energy per inference. The framework contributions—quantization-aware training, fan-in/fan-out constrained layout, and the LAGS clocking scheme—are valuable and partially verified through layout and circuit simulation. However, the significance is heavily contingent on the chip measurements, which are not reported. The paper also provides concrete resource counts (JJs, area, power) that would be useful to the community if properly backed.","major_comments":[{"comment":"The headline result—'the fabricated SuperSNN chip successfully classified a reduced set of digits with 80.07% accuracy'—is not supported by any measured chip data. Section 4.2 describes implementation and timing but gives no test setup, no description of how output DC levels are converted to labels, no raw output traces, no confusion matrix, and no list of test images. The 3.02 GHz operating point is attributed to Fig. 3(b), whose caption explicitly states 'Simulation result', not chip measurement. The 6.55 fJ per inference is described in the Conclusion as 'estimated switching energy'. Thus the central empirical claim—that the chip physically realized the trained network—is unverified in this manuscript.","section":"Abstract; §4.2; §4.3; Fig. 3(b)"},{"comment":"There is an internal inconsistency about the fabricated chip's accuracy. The abstract states 80.07% for digits 2/3/4 and 'maximum 86.2% for digits 0/1/2', and Table VI reports 86.20% for 0/1/2 and 80.07% for 2/3/4. However, Table V lists 'Accuracy (%)' as 86.2 for the row 'Predictable digits 2, 3, 4', and §4.3 states that the prototype 'focuses on the digit subset 2, 3, and 4, achieving an inference accuracy of 86.2%'. The reader cannot determine which number was observed on the fabricated chip, on which test set, or whether 86.2% refers to the 2/3/4 chip at all. This must be corrected with an unambiguous mapping between digit subsets and accuracies.","section":"Abstract vs Table V and §4.3"},{"comment":"The complete-network accuracy of 96.47% is obtained after a training procedure that selects the model with the highest accuracy on the full test dataset: 'Each stage runs for a predetermined number of epochs, with the model yielding the highest accuracy on the full test dataset being preserved' (stage e, Validation). This is test-set leakage/model selection on the test set, which inflates the reported accuracy and makes the 96.47% figure not an honest estimate of generalization. A held-out validation set must be used for model selection, and the test set should be used only once for final evaluation. This issue affects the correctness of the complete-network claim.","section":"§3.1.1(e); §4.1"},{"comment":"The power and energy figures—2.15 mW static power, 6.55 fJ per inference, and 1.31e-6 nJ—are presented as chip characteristics, but the manuscript does not provide any measurement methodology or instrument details. The conclusion explicitly calls the switching energy 'estimated'. If these are simulation or schematic estimates, they should be labeled as such throughout, and the distinction between simulated and measured quantities must be made explicit in the abstract and results tables. As written, a reader may reasonably but incorrectly infer that these values were measured from the fabricated chip.","section":"§4.3; Conclusion"}],"minor_comments":[{"comment":"The clock scheme is called 'locally synchronous, globally synchronous (LAGS)' in the Abstract but 'locally asynchronous, globally synchronous (LAGS)' in the Introduction's contribution list. Please standardize the terminology and clarify the actual synchronization policy.","section":"Abstract vs Introduction/Contributions"},{"comment":"The conclusion says 'after quantification and pruning' where 'quantization' is meant. Please fix this typo.","section":"Conclusion"},{"comment":"Reference [25] is dated 2020 in the text but the arXiv ID (1812.10354) suggests a 2018 preprint. Please verify the citation details.","section":"Reference [25]"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a potentially important chip, but the absence of any measured chip data is the central problem. The internal accuracy inconsistency (80.07% vs 86.2% for the same digit subset) and the test-set model selection for the 96.47% number are additional load-bearing issues. I recommended major revision because these are fixable in principle: the authors could add a proper measurement section with raw data and test methodology, correct the accuracy inconsistency, and replace test-set selection with a validation set. If the chip data does not exist, the paper should be reframed as a design study and the fabricated-chip claims removed; in that case the current version would be better rejected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Candid take: the paper's own numbers undercut its headline. The abstract says the fabricated chip classified 2/3/4 at 80.07% and \"reaching a maximum of 86.2% for digits 0,1,2\", but Section 4.3 and Table V attribute 86.2% to the 2/3/4 chip. Meanwhile Section 4.1 and Table VI list 80.07% for 2/3/4 and 86.2% for 0/1/2. That inconsistency alone would need a correction. More importantly, there is no measured test evidence anywhere: no test protocol, no raw output traces, no confusion matrix, no measured clock or power. The 3.02 GHz figure comes from a circuit simulation, and the 6.55 fJ is an \"estimated switching energy\" in the conclusion. For a paper whose entire point is that the chip was fabricated and worked, that is a load-bearing gap.\n\nWhat is genuinely new: the end-to-end framework that takes SNN training through pruning and ternary quantization to a full layout that respects SFQ5ee pin/area/fan-in constraints. The shift-register input scheme to handle the 40-pin limit, the standardized eight-input neuron cell (six positive, two negative couplings), and the LAGS clock distribution are all concrete, implementable pieces of engineering. Compared to prior work that stops at single-neuron fabrications or software-only network evaluations, this is a real step toward a complete on-chip inference network. The 96.47% complete-network accuracy after quantization and pruning is a reasonable software result, though the test-set-based model selection in stage (e) is a mild leak, and the network is far from state-of-the-art for MNIST.\n\nThe chip-level contribution is tiny—21 active neurons, 3 classes, 7x7 input—so the significance is subfield-level, not transformative. But for superconducting neuromorphic hardware, a fabricated full network would matter. The problem is that the paper doesn't demonstrate the fabrication worked. The authors may well have the data; it's just not in the manuscript.\n\nBottom line: this is a serious design/engineering effort with a clear thinking trail and no circularity in the derivations. It deserves a serious referee, not a desk rejection—but the referee should be instructed to demand the measured test data, and the authors should fix the accuracy inconsistency before publication. If the measurement evidence doesn't exist, the paper needs to be recast as a design study.","headline":"A tapeout-level design and training framework for a tiny superconducting SNN, but the fabricated-chip accuracy is unverified and internally inconsistent; worth a referee's time only if the authors can produce measured data.","tokens_in":14110,"tokens_out":2555,"would_cite":false,"duration_ms":25663,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Superconducting spiking chip classifies digits at 80% accuracy","keywords":["superconducting spiking neural network","single-flux-quantum circuits","hardware-aware training","weight quantization","network pruning","MNIST classification","Josephson junction neuron","chip design"],"falsifier":"Directly bench-test the fabricated chip: feed the recorded 7×7 MNIST test samples through the seven input pins, capture the three output pins over many cycles, and compare the resulting classifications to the software model. If the chip's accuracy is statistically indistinguishable from chance (33%) for digits 2, 3, and 4, or if no output pulses appear at the claimed clock rate, the central claim fails.","tokens_in":13138,"feed_emoji":"⚡","tokens_out":5321,"duration_ms":53266,"temperature":0.7,"pith_summary":"SuperSNN aims to show that a complete spiking neural network can be designed, trained, and physically fabricated as a single superconducting chip despite severe area, routing, and pin-count limits. The framework couples hardware-aware training—pruning and ternary weight quantization—with a custom high fan-in neuron cell and a locally asynchronous, globally synchronous clock scheme. After quantization and pruning, the full MNIST network reaches 96.47% accuracy. The fabricated chip, constrained to a three-digit subset by 40 pins and a 5-by-5 mm die, classifies digits 2, 3, and 4 at 80.07% accuracy, with a reported maximum of 86.2% on digits 0, 1, and 2. The paper also reports a 3.02 GHz clock, 5,822 Josephson junctions, 2.15 mW static power, and 6.55 fJ per inference, establishing a concrete path from algorithm to taped-out superconducting neuromorphic hardware.","feed_headline":"Superconducting spiking chip classifies digits at 80% accuracy","feed_subtitle":"A full inference network fits the 40-pin, 5 mm square process with 5,822 Josephson junctions and 6.55 fJ per inference.","key_machinery":"The load-bearing piece is the high fan-in neuron cell: eight mutually coupled input branches (six positive, two negative) drive a single Josephson junction through two dendritic bundles, so weighted summation happens by current addition in superconducting loops. Standardizing every neuron to exactly eight inputs keeps the layout regular and reusable. Around that cell, a locally asynchronous, globally synchronous (LAGS) clock distribution coordinates pulse arrival times across passive transmission lines, while a 49-bit shift register reduces the 7×7 input to seven pins at the cost of a several-cycle loading delay.","core_discovery":"The central claim is that physical realizability constraints need not be treated as an afterthought in superconducting neural network design. SuperSNN folds those constraints into the network itself: weights are restricted to +1, 0, and −1, each neuron's fan-in is capped to fit the silicon area, and the resulting discrete connections map directly onto a high fan-in neuron cell in which synaptic currents add magnetically at a single Josephson junction. The framework trains two networks—a large one for the full MNIST benchmark and a small one that obeys the 40-pin, 25-neuron chip limits—and then lays out the small one using standard cells, a shift-register input scheme, and a clock distributio","pith_inferences":["The 7×7 downsampling is likely the dominant accuracy loss; moving to 14×14 inputs or adding a convolutional front-end would probably close much of the gap, but would require more pins and a larger neuron fan-in.","The reported 3.02 GHz clock frequency appears to be a simulation or schematic estimate, because the shift-register path is described as adding 433 ps, which implies a throughput ceiling near 1 GHz for the actual chip.","The 6.55 fJ per-inference figure implies that only a small fraction of the 5,822 Josephson junctions switch during a single classification; counting switching events in the simulation netlist would test this.","The same training and neuron-cell recipe could be run on a non-MNIST dataset, such as Fashion-MNIST or EMNIST letters, to check whether the accuracy holds beyond handwritten digits."],"forward_implications":["A full inference SNN, not just a few neurons, can be placed on a single superconducting chip under real fabrication constraints.","The 40-pin limit can be met by serial shift-register input, at the price of a reduced classification throughput.","The per-inference energy of 6.55 fJ and 2.15 mW static power suggest cryogenic inference far more efficient than CMOS, if the reported figures are reproduced.","The accuracy gap between the full network (96.47%) and the chip (80.07%) comes primarily from 7×7 downsampling and pin limits, not from the neuron cell itself.","The same framework can be retargeted to other digit subsets; digits 0, 1, and 2 yield 86.2% accuracy."],"supporting_citations":[{"why":"Supplies the high fan-in neuron design and the software baseline (96.1% on ten digits) this chip is compared against.","marker":"[26]"},{"why":"Defines the SFQ5ee fabrication process limits, including the 5x5 mm die and 40-pin budget, which set the hardware constraints.","marker":"[27]"},{"why":"Provides the RSFQ logic family and transmission-line primitives that the neuron, splitter trees, and clock distribution are built from.","marker":"[30]"},{"why":"Supplies the spike-count decoding convention used by the complete network to map output spikes to class labels.","marker":"[31]"}],"fun_headline_variants":["SuperSNN fits full MNIST on 5mm chip at 96%","Superconducting SNN chip: 80% accuracy, 6.55 fJ per inference","Hardware-aware SNN hits 96% MNIST under chip constraints","Chip-scale superconducting neural net runs at 3 GHz","Mapping SNNs to real chips: SuperSNN achieves 80% on digits"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The core result depends on the fabricated chip actually performing the trained classification at the reported 80.07% accuracy and 3.02 GHz operation, yet the paper includes no measured test protocol, output waveforms, or per-class spike counts, so those numbers may reflect simulation rather than chip measurements.","fun_headline_variants_meta":{"raw":{"variants":["SuperSNN fits full MNIST on 5mm chip at 96%","Superconducting SNN chip: 80% accuracy, 6.55 fJ per inference","Hardware-aware SNN hits 96% MNIST under chip constraints","Chip-scale superconducting neural net runs at 3 GHz","Mapping SNNs to real chips: SuperSNN achieves 80% on digits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000358,"raw_usage":{"total_tokens":1869,"prompt_tokens":930,"completion_tokens":939,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":850}},"tokens_in":674,"tokens_out":939,"duration_ms":8669,"temperature":1.0,"reasoning_tokens":850,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:23:55.460012+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Directly bench-test the fabricated chip: feed the recorded 7×7 MNIST test samples through the seven input pins, capture the three output pins over many cycles, and compare the resulting classifications to the software model. If the chip's accuracy is statistically indistinguishable from chance (33%) for digits 2, 3, and 4, or if no output pulses appear at the claimed clock rate, the central claim fails.","supporting_citations":[{"cited_title":"Scalable superconductor neuron with ternary synaptic connections for ultra-fast snn hardware,","cited_arxiv_id":null,"evidence_quote":"Supplies the high fan-in neuron design and the software baseline (96.1% on ten digits) this chip is compared against."},{"cited_title":"Advanced fabrication processes for superconducting very large scale integrated circuits,","cited_arxiv_id":null,"evidence_quote":"Defines the SFQ5ee fabrication process limits, including the 5x5 mm die and 40-pin budget, which set the hardware constraints."},{"cited_title":"Rsfq logic/memory family: a new josephson-junction technology for sub-terahertz-clock-frequency digital systems,","cited_arxiv_id":null,"evidence_quote":"Provides the RSFQ logic family and transmission-line primitives that the neuron, splitter trees, and clock distribution are built from."},{"cited_title":"Unsupervised learning of digit recognition using spike-timing-dependent plasticity,","cited_arxiv_id":null,"evidence_quote":"Supplies the spike-count decoding convention used by the complete network to map output spikes to class labels."}],"review_version":1}