{"id":"c3b40720-3339-478e-90fd-d60f2914b6e3","arxiv_id":"2411.11575","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"The claim that SpiNNaker improves GHA classification accuracy is not supported by the paper's own uncontrolled comparison.","lead":"This paper reports running the Generalized Hebbian Algorithm (GHA) on the SpiNNaker neuromorphic platform with MNIST and five small UCI datasets. It claims accuracy improvements over plain Hebbian learning, but the comparison is confounded by different training/test splits and lacks statistical controls.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central claim is undermined by a confounded comparison: the SpiNNaker result uses a 30% training split, the non-SpiNNaker baseline uses 80%, and the manuscript itself says the main change was the split; the supporting tables are also captioned for HA rather than GHA.","rationale":"The reader's rejection is justified. The most load-bearing premise for the central claim is that the higher accuracy in the SpiNNaker condition is caused by the hardware, but the reported comparison uses different training splits: with-SpiNNaker uses 30% training data, while without-SpiNNaker uses 80%. The paper's own text says the main change was the adjustment of the training and test data split, so the manuscript itself supplies an alternative explanation for the improvement. This confound alone is sufficient to invalidate the hardware-advantage claim. In addition, the caption mismatch in Tables IV–VII (labeled HA rather than GHA) and the absence of the table bodies make the algorithm attribution and the numerical evidence unverifiable. No code, hyperparameters, or error bars are provided, and there is no independent support such as machine-checked proofs or a reproducible artifact. The reader's weakest assumption identifies the same confounded comparison, so my assessment agrees with the reader's verdict of REJECT; no change to that verdict is needed.","tokens_in":12072,"tokens_out":6406,"duration_ms":56972,"concrete_test":"Rerun the UCI Wine experiment with identical training splits, e.g., 30/70, both with and without SpiNNaker, using the same GHA implementation, hyperparameters, seeds, and classifier. If the without-SpiNNaker accuracy at 30% training is within statistical error of the with-SpiNNaker accuracy, the hardware-driven improvement claim fails. In the same run, record explicitly whether the 80.56% entry is produced by GHA or by HA, resolving the caption ambiguity in Tables IV–VII.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that GHA on SpiNNaker yields 'significant improvements in classification accuracy' rests on comparing 80.56% accuracy with 2650 J on the Wine dataset 'with 30% train set ... with SpiNNaker' to 72.80% with 7100 J 'with 80% train set ... without SpiNNaker' (Section III.B, paragraph after Table VII). These are different training-set sizes, so the hardware effect is not separated from the data-split effect. The manuscript explicitly states that 'the main change was the adjustment of the training and test data split,' undercutting the hardware attribution. A 30% train run and an 80% train run are not comparable conditions, and post-hoc selection of the best split for each condition makes the accuracy gap uninterpretable as a hardware benefit. Compounding this, Tables IV–VII are captioned 'ON HEBBIAN LEARNING ALGORITHM,' not GHA, even though the text says the results are for the GHA model; the table bodies are absent from the manuscript, so the algorithm that produced the 80.56% figure is not verifiable. No code, hyperparameters, or error bars are provided. Because the two load-bearing conditions differ in the very factor the paper says was changed, the central claim collapses unless the hardware comparison is rerun with identical splits and with correct algorithm attribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an experimental comparison of the Generalized Hebbian Algorithm (GHA) and standard Hebbian learning on the MNIST and UCI Machine Learning datasets, with and without the SpiNNaker neuromorphic platform. The central claim is that GHA on SpiNNaker achieves significant improvements in classification accuracy, with a reported peak of 80.56% on the UCI Wine dataset using a 30% training split, versus 72.80% without SpiNNaker using an 80% training split, along with energy consumption of 2650 J and 7100 J respectively. The authors conclude that GHA outperforms Hebbian learning and that SpiNNaker is a powerful platform for improving classification accuracy.","tokens_in":12353,"tokens_out":4189,"duration_ms":32963,"significance":"If the central claim were supported, demonstrating that running GHA on SpiNNaker improves classification accuracy relative to the same algorithm without the platform would be a notable result for neuromorphic computing. The paper also provides a concise summary of SpiNNaker's architecture and the GHA update rule. However, the experimental design does not support the claim: the hardware and non-hardware conditions differ in training set size, the supporting tables are absent and mislabeled, and no hyperparameters or error bars are provided. As presented, the result is not a reproducible measurement of hardware-driven improvement.","major_comments":[{"comment":"The comparison that grounds the abstract's claim of 'significant improvements' is confounded: the SpiNNaker condition uses a 30% training split and the non-SpiNNaker condition uses an 80% split, and the manuscript itself states that 'the main change was the adjustment of the training and test data split.' Consequently, the difference between 80.56% and 72.80% cannot be attributed to SpiNNaker, and the central claim collapses unless the comparison is rerun with identical splits.","section":"Section III.B, paragraph after Table VII"},{"comment":"The tables are captioned 'ON HEBBIAN LEARNING ALGORITHM' even though the text describes the results as being for the GHA model, and the table bodies are missing entirely. This makes it impossible to verify which algorithm produced the reported accuracy and energy figures, including the headline 80.56% result; the algorithm attribution must be corrected and the tables populated before the results can be assessed.","section":"Tables IV-VII"},{"comment":"The paper reports no hyperparameter values (learning rate, number of epochs, number of output principal components), no preprocessing details for the UCI datasets, and no repeated runs or error bars. Instead, the paper selects the highest accuracy across four training/test splits as the headline result; this post hoc selection is not an independent measurement of GHA or of SpiNNaker and inflates the apparent improvement.","section":"Section III.A and III.B"}],"minor_comments":[{"comment":"The title spells the platform as 'Spinnaker'; it should be 'SpiNNaker' throughout the paper.","section":"Title"},{"comment":"The phrase 'By offers numerous advantages as an extension of the classical Hebbian learning rule' is ungrammatical and should be revised to 'It offers numerous advantages as an extension of the classical Hebbian learning rule.'","section":"Section I, third paragraph"},{"comment":"The sentence 'making it ideal platform for Neuromorphic computing' should read 'making it an ideal platform for neuromorphic computing.'","section":"Section I, fourth paragraph"},{"comment":"The claims that GHA has higher memory usage and training time than HA are unsupported because Table II is empty in the manuscript.","section":"Section III.B, paragraph after Table II"},{"comment":"Reference [10] appears truncated, and the GHA description would be clearer with a formal pseudocode listing or numbered equations for the update rule.","section":"Section II.C"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be an early-stage draft: the experimental section lacks all data tables, the algorithm labels are inconsistent, and the writing contains multiple incomplete sentences. The central comparison, as presented, cannot support the claimed hardware benefit. If the authors can rerun the experiments with matched training splits and provide full tables, a resubmission might be considered, but the current version does not meet the bar for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is an incomplete draft whose central claim—that SpiNNaker improves GHA classification accuracy—is not supported by the experiments as reported. The comparison is confounded, the tables are missing, and the algorithm attribution is sloppy.\n\nThe paper does a decent job explaining GHA and SpiNNaker. The description of Sanger's rule is accurate, and the summary of SpiNNaker's architecture is competent. The authors are also honest that the main change between conditions was the train/test split.\n\nBut the headline result compares 80.56% accuracy with a 30% training split on SpiNNaker against 72.80% with an 80% split without SpiNNaker. Those conditions differ in the very factor the paper says was changed. The paper reports the best of four splits for each condition, so the gap is uninterpretable. There are no error bars, no repeated runs, no hyperparameter values. Tables IV–VII, which are supposed to contain the UCI results, are absent from the manuscript, and their captions say \"HEBBIAN LEARNING ALGORITHM\" rather than GHA. That makes it impossible to verify which algorithm produced the 80.56% figure. The paper also does not cite prior SpiNNaker learning-rule implementations, so the novelty claim is weakly situated.\n\nNovelty is minimal: GHA is a 1989 algorithm, and a small benchmark on five tiny UCI datasets is not a new capability or efficiency result. The paper's own text admits the accuracy improvement came from changing the split, not from the hardware. That undercuts the abstract's claim of significant improvements.\n\nThis deserves a desk reject, not a full peer review. If the authors rerun the experiment with identical splits, correct the table captions, add error bars and code, and compare against appropriate baselines, there might be a modest systems note. As is, I would not send it to referees.","headline":"An incomplete draft whose central hardware-benefit claim is confounded by different train/test splits and missing tables; worth a desk reject, not referee time.","tokens_in":12868,"tokens_out":2001,"would_cite":false,"duration_ms":18022,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Running the Generalized Hebbian Algorithm on the SpiNNaker neuromorphic platform is reported to lift classification accuracy on the UCI Wine dataset to 80.56%, with 2650 J of energy consumption.","keywords":["Generalized Hebbian Algorithm","SpiNNaker","neuromorphic hardware","Hebbian learning","classification accuracy","energy consumption","UCI Machine Learning Repository","MNIST"],"falsifier":"Run GHA on the UCI Wine dataset with the same training/test split (for example 30/70) both on SpiNNaker and on the non-SpiNNaker setup and compare the accuracies and energy readings; if the reported 80.56% versus 72.80% gap vanishes or inverts at matched splits, the hardware-driven improvement is not supported.","tokens_in":11860,"feed_emoji":"🧠","tokens_out":7029,"duration_ms":59231,"temperature":0.7,"pith_summary":"This paper sets out to show that the Generalized Hebbian Algorithm (GHA), a biologically plausible unsupervised learning rule, can be implemented on the SpiNNaker neuromorphic hardware platform and deliver competitive classification accuracy with low energy use. The authors compare GHA with standard Hebbian learning on MNIST and five UCI Machine Learning Repository datasets, measuring error rate, convergence rate, training time, memory usage, accuracy, and energy. They report that GHA outperforms Hebbian learning on error rate, convergence, and accuracy, at the cost of higher training time and memory. Their headline result is 80.56% classification accuracy on the Wine dataset with 2650 J of energy using a 30% training split on SpiNNaker, compared with 72.80% accuracy without SpiNNaker using an 80% split. The sympathetic reading is that biologically inspired learning rules are viable and energy-efficient on parallel neuromorphic hardware.","feed_headline":"GHA on SpiNNaker hits 80.56% on wine dataset","feed_subtitle":"A brain-inspired learning rule beats plain Hebbian learning and uses only 2650 J on neuromorphic hardware.","key_machinery":"The Generalized Hebbian Algorithm (GHA) is the central object: an unsupervised online update rule in which each output unit learns one eigenvector of the input correlation matrix, with outputs ordered by decreasing eigenvalue, extending Hebb's 'fire together, wire together' rule to find principal components from data without first computing the correlation matrix. SpiNNaker is the hardware carrier: a massively parallel neuromorphic machine built from ARM cores, with packet-switched Address Event Representation (AER) multicast communication, designed to simulate spiking neural networks in real time. The paper's argument runs through the interaction of these two: GHA's simple, local, sample-by-sample updates are well matched to SpiNNaker's parallel event-driven architecture. The experiments also rely on adjusting the training/test split (70/30, 50/50, 80/20, 30/70), which the authors identify as the main change that raised GHA's classification accuracy from 30.56% on MNIST to over 80% on the UCI Wine dataset.","core_discovery":"The paper's central claim is that GHA—an online rule that learns the eigenvectors of the input autocorrelation matrix directly from samples—maps cleanly onto SpiNNaker and yields significant improvements in classification accuracy compared with both plain Hebbian learning and the same GHA run without SpiNNaker. On MNIST, GHA shows a lower error rate and a higher average convergence rate than Hebbian learning, but it also uses more memory and takes longer to train. On the UCI datasets, the authors report the highest accuracy of 80.56% for the Wine dataset with a 30% training split on SpiNNaker at 2650 J, and 72.80% without SpiNNaker at 7100 J with an 80% training split. They attribute the improvement to SpiNNaker's neuromorphic, energy-efficient processing and conclude that GHA is well suited to high-performance, rapid-convergence tasks such as image and speech recognition.","pith_inferences":["Because the SpiNNaker and non-SpiNNaker comparisons used different training splits, the accuracy gap cannot be cleanly attributed to the hardware; a matched-split comparison would settle whether SpiNNaker itself improves generalization.","The 80.56% figure is a single small dataset result; a natural extension is to test GHA on larger benchmarks with the same split-adjustment protocol to see whether the accuracy gain generalizes beyond the UCI Wine dataset.","If matched-split comparisons confirm a hardware benefit, it would suggest that event-driven parallel execution adds a regularization-like effect to Hebbian learning, a hypothesis that future work could test directly.","The authors' planned FPGA-based GHA accelerators could reuse the same accuracy and energy metrics to compare neuromorphic versus reconfigurable platforms on equal terms."],"forward_implications":["GHA offers a biologically plausible route to feature extraction that can run on parallel neuromorphic hardware without pre-computing input covariance matrices.","On the reported UCI Wine result, GHA on SpiNNaker reaches 80.56% accuracy with 2650 J, indicating that brain-inspired learning can be energy-efficient on small benchmark tasks.","GHA's higher accuracy and convergence rate come with increased training time and memory usage, so resource management is a direct constraint on its practical use.","Adjusting the training-to-test split is the lever the authors used to improve GHA accuracy from 30.56% on MNIST to above 80% on UCI datasets."],"supporting_citations":[{"why":"Defines the Generalized Hebbian Algorithm and its convergence to eigenvectors of the input autocorrelation matrix; this is the learning rule the paper ports to SpiNNaker.","marker":"[10]"},{"why":"Describes the SpiNNaker system architecture, including ARM cores, memory, and energy-per-connection costs; supplies the hardware model behind the accuracy and energy claims.","marker":"[22]"},{"why":"Presents the SpiNNaker 1-W 18-core system-on-chip, grounding the claim that SpiNNaker is a massively parallel, low-power neuromorphic platform.","marker":"[34]"},{"why":"Source of the Wine, Parkinson's, Heart Disease, Liver Disease, and Breast Cancer datasets used for the classification accuracy comparisons.","marker":"[30]"},{"why":"Survey of neuromorphic computing and neural networks in hardware, framing SpiNNaker's place and the low-power, parallel advantages the paper relies on.","marker":"[1]"},{"why":"Large benchmark suite for machine learning, cited to justify the choice of UCI-ML datasets for empirical comparison.","marker":"[42]"},{"why":"Discusses opportunities for neuromorphic algorithms and applications, supporting the motivation that biologically inspired learning can be effective and energy-efficient.","marker":"[33]"}],"fun_headline_variants":["GHA on SpiNNaker: 80.56% wine accuracy, 2650 J","Brain-inspired GHA beats Hebbian on SpiNNaker","Neuromorphic GHA scores 80.56% on wine dataset","SpiNNaker GHA: energy-efficient learning hits 80.56%","GHA on SpiNNaker: 80.56% accuracy, low energy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that SpiNNaker improves accuracy rests on comparing runs with different training-set sizes, 30% on SpiNNaker versus 80% without, so the hardware effect is never separated from the effect of the data split.","fun_headline_variants_meta":{"raw":{"variants":["GHA on SpiNNaker: 80.56% wine accuracy, 2650 J","Brain-inspired GHA beats Hebbian on SpiNNaker","Neuromorphic GHA scores 80.56% on wine dataset","SpiNNaker GHA: energy-efficient learning hits 80.56%","GHA on SpiNNaker: 80.56% accuracy, low energy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0005,"raw_usage":{"total_tokens":2396,"prompt_tokens":842,"completion_tokens":1554,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":1448}},"tokens_in":458,"tokens_out":1554,"duration_ms":8799,"temperature":1.0,"reasoning_tokens":1448,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:21:00.838561+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run GHA on the UCI Wine dataset with the same training/test split (for example 30/70) both on SpiNNaker and on the non-SpiNNaker setup and compare the accuracies and energy readings; if the reported 80.56% versus 72.80% gap vanishes or inverts at matched splits, the hardware-driven improvement is not supported.","supporting_citations":[{"cited_title":"Hebbian learning Algorithm rule","cited_arxiv_id":null,"evidence_quote":"Defines the Generalized Hebbian Algorithm and its convergence to eigenvectors of the input autocorrelation matrix; this is the learning rule the paper ports to SpiNNaker."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the SpiNNaker system architecture, including ARM cores, memory, and energy-per-connection costs; supplies the hardware model behind the accuracy and energy claims."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Presents the SpiNNaker 1-W 18-core system-on-chip, grounding the claim that SpiNNaker is a massively parallel, low-power neuromorphic platform."},{"cited_title":"A protocol for scalable loop-free multicast routing,","cited_arxiv_id":null,"evidence_quote":"Source of the Wine, Parkinson's, Heart Disease, Liver Disease, and Breast Cancer datasets used for the classification accuracy comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Large benchmark suite for machine learning, cited to justify the choice of UCI-ML datasets for empirical comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Discusses opportunities for neuromorphic algorithms and applications, supporting the motivation that biologically inspired learning can be effective and energy-efficient."}],"review_version":1}