{"id":"0e17afee-2a5d-4e9c-b15a-5a892e93ae98","arxiv_id":"2411.17006","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A PRISMA-style review of 151 SNN-for-CV papers that categorizes datasets, architectures, learning rules, and implementation media, accompanied by a code repository.","lead":"This paper is a systematic review of 151 papers on event-based spiking neural networks (SNNs) for computer vision, organizing them by datasets, architectures, learning rules, and hardware implementations. It also ships a GitHub repository with Python examples for building, training, and simulating SNN models.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's quantitative synthesis is not auditable: Section II.E reports counts that sum to far fewer than 151 (14 implementation-medium, 67 architecture, 40 learning-rule categorizations) and the included-study list is absent, so the claimed codification lacks a verifiable evidentiary base.","rationale":"The reader's weakest assumption was external representativeness of the 151-article sample due to missing search details. My concern is more specific and internal: even the paper's own reported counts do not sum to 151, which suggests the data extraction or reporting is unreliable. This is a concrete, checkable flaw that directly undermines the central 'codifies' claim. The reader noted the implementation-medium inconsistency as one of three weaknesses but did not identify it as the primary load-bearing issue. Because the concern is significant but potentially correctable by releasing the study list and corrected counts, the conditional verdict remains appropriate. If the test reveals that the counts cannot be reconciled, the verdict would need to move toward rejection; if the counts are reconciled and the list is provided, the review could be accepted. The code repository and qualitative narrative are useful, but they do not compensate for an unauditable quantitative base.","tokens_in":44608,"tokens_out":3254,"duration_ms":28651,"concrete_test":"Request a machine-readable supplementary table listing all 151 included papers with per-paper categorical tags for architecture, learning rule, implementation medium, dataset, and evaluation metrics, along with the PRISMA screening log and exact search queries. Then recompute the counts in Section II.E and the percentages in Figures 2 and 3 from this table. If the implementation-medium count is not 151, or if the sum of architecture/learning-rule classifications is not equal to the total number of papers actually classified, the reported distributions are unsupported and the claims must be revised. Additionally, rerun the provided search strings to confirm that the 151 papers are a subset of the 1,169 screened articles.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that analyzing 151 articles lets the review 'codify' the effectiveness of architectures, learning rules, and hardware trade-offs. For that claim to hold, the reported distributions in Figures 2 and 3 and Section II.E must accurately summarize all 151 papers. They do not appear to. Section II.E reports: 'Implementation Mediums: 9 papers reported on the use of simulated environments and 5 papers reported on neuromorphic hardware implementations' (sum 14); architecture counts of 23 SCNN + 12 FCSNN + 12 HSNN + 11 SRNN + 9 SDBN = 67; learning-rule counts of 17 + 10 + 13 = 40. None of these sums to 151. The figures are captioned as distributions of the reviewed papers, so either the text is wrong, the figures cover only a subset, or most papers were unclassifiable. The manuscript does not explain this. Moreover, the full list of included studies, the exact search strings, and the PRISMA screening log are not provided, so the reader cannot verify the sample or reconcile the numbers. If the quantitative synthesis is internally inconsistent, the practitioner framework built on these distributions is not reliably supported. This is a load-bearing weakness because the review's main contribution is precisely the codification of trends from this corpus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a systematic review, following a stated PRISMA-style methodology, of 151 journal and conference articles on event-based spiking neural networks (SNNs) for computer vision, with a focus on object detection. The paper organizes the literature into datasets, architectures, learning rules, implementation mediums, and evaluation metrics, and it proposes a practitioner framework that connects data acquisition, encoding, architecture choice, learning method, implementation, and evaluation. It also provides an open-source repository with Python examples and identifies open challenges for SNN-based computer vision.","tokens_in":44748,"tokens_out":2179,"duration_ms":23326,"significance":"If the quantitative synthesis is reliable, the review fills a real gap: prior surveys treat learning rules, hardware, and applications separately, whereas this paper attempts a unified codification with a practitioner-oriented pipeline and an accompanying code repository. The breadth of coverage—spanning neuron models, encodings, datasets, simulators, neuromorphic chips, learning rules, and architectures—makes it a potentially useful entry point for newcomers. The paper ships concrete artifacts: a public repository, tables comparing frameworks and hardware, and a summary of accuracy numbers drawn from the literature. However, the central value of the review depends on the credibility and auditability of the claimed analysis of 151 articles, and that evidentiary base is currently not transparent or internally consistent.","major_comments":[{"comment":"The quantitative synthesis reported in Section II.E does not reconcile with the stated corpus of 151 papers. The implementation-medium counts are 9 + 5 = 14, the architecture counts are 23 + 12 + 12 + 11 + 9 = 67, and the learning-rule counts are 17 + 10 + 13 = 40. None of these sums equals 151, and the text does not explain whether papers can belong to multiple categories, whether most papers were unclassifiable, or whether the figures describe only a subset. Because the abstract's central claim is that the review 'codifies' trends from 151 articles, these distribution counts are load-bearing. The manuscript must either provide a complete reconciliation (e.g., a full coding table with per-paper categories and percentages, including multi-label counts) or explicitly restrict the claims in the abstract and figures to the subset of papers for which each categorization was possible.","section":"Section II.E"},{"comment":"The PRISMA-based selection process is not auditable as reported. The identification stage lists broad search terms ('SNN applications,' 'SNN learning rules,' etc.) but gives no exact query strings, no database-specific search strings, no screening decision rules beyond broad eligibility bullets, and no log of exclusions. Crucially, the 151 included studies are never listed, so a reader cannot verify that the claimed distributions in Figures 2 and 3, or the narrative conclusions about architecture and learning-rule prevalence, actually follow from the cited corpus. A systematic review should include either a reference list of all included studies (in an appendix or supplementary file) and a PRISMA-style screening table, or the methodology section should be revised to describe the selection process as an illustrative scoping review rather than a fully auditable systematic review.","section":"Section II (Methodology)"},{"comment":"The abstract states that the review 'codifies: 1) the effectiveness of fully connected, convolutional, and recurrent architectures; 2) the performance of direct unsupervised, direct supervised, and indirect learning methods; and 3) the trade-offs in energy consumption, latency, and memory in neuromorphic hardware implementations.' However, the body mostly tabulates reported accuracy numbers and qualitative framework features rather than providing a controlled comparison of effectiveness or performance across architectures, learning rules, or hardware. For example, Table V lists accuracies such as 95% (MNIST, additive STDP), 99.1% (Caltech 101, multiplicative STDP), and 98.89% (MNIST, STBP) from different studies with different datasets, preprocessing, and network sizes; these numbers are not commensurable evidence for 'effectiveness' or 'performance' as codified conclusions. The authors should either add a structured comparative analysis that normalizes or contextualizes these numbers (e.g., by dataset, architecture capacity, and evaluation protocol) or temper the abstract and conclusion claims to describe a taxonomy and reported trends rather than codified effectiveness.","section":"Abstract and Section II.D"},{"comment":"The paper's treatment of hardware and learning rules is largely descriptive and does not substantiate the claimed trade-offs among energy consumption, latency, and memory. Section VI.B discusses TrueNorth, Loihi, BrainScaleS, Tianjic, and SpiNNaker with their nominal specifications, and Section VII reviews learning rules, but the connection between the two—what measurable latency, energy, or memory consequences each learning rule or architecture has on a given hardware platform—is not systematically quantified or compared. Since the abstract explicitly lists these trade-offs as a codified outcome, the review needs either a dedicated comparative synthesis (e.g., a table with per-study energy/latency/memory measurements and their conditions) or a rewritten claim that such trade-offs are surveyed qualitatively without being codified.","section":"Section VI.B and VII"}],"minor_comments":[{"comment":"The sentence 'In addition to practitioners, this framework provides also provides a structured approach for educators...' contains a duplicated 'provides'; it should read 'this framework also provides a structured approach.'","section":"Section III"},{"comment":"Equation (13), describing the output firing rate of the residual membrane potential neuron, is typeset in a way that makes the floor function and the condition 'n ≥ 0' difficult to parse, and the variables n and N are not defined in the surrounding text. Please clarify the notation and the intended domain of the formula.","section":"Section IV.C.1, Eq. (13)"},{"comment":"The section ordering in the introduction is inconsistent with the actual order of sections: the text lists Section VIII before Section VII, and Sections IX and XI are mentioned but not described in sequence. Please align the roadmap with the final section numbering.","section":"Section I"},{"comment":"Table III would benefit from a column or note on whether each simulation framework has been used in the 151 reviewed papers or is included purely as background context; this would help the reader connect the framework descriptions to the quantitative synthesis.","section":"Section VI.A"},{"comment":"The caption of Figure 3 states that the right panel distinguishes 'software tools and hardware-based implementations,' but the text refers to 'simulated environments' and 'neuromorphic hardware implementations'; please align the caption terminology with the text.","section":"Section II.E"},{"comment":"The paper repeatedly refers to its own GitHub repository [71] for tutorials and links. This self-reference is benign, but the main text should make clear that the repository is supplementary material and not one of the 151 analyzed articles, to avoid ambiguity in the corpus description.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The review has a useful scope and a substantial amount of descriptive content, but the central quantitative claim—codifying 151 articles—is currently unsupported by the reported counts and absence of an included-studies list. The issues are fixable within the manuscript's scope: adding a supplementary list of included studies, exact search strings, and a reconciliation table for the category counts would make the synthesis auditable. If the authors cannot provide those items, they should scale back the claims to a scoping review. I do not see evidence of deliberate misrepresentation; the inconsistency appears to stem from incomplete reporting, which is why I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful practitioner-oriented review with a working code repository, but the quantitative synthesis at its center doesn't hold up to arithmetic. Section II.E reports 9 simulated + 5 hardware = 14 implementation-medium papers, 23+12+12+11+9 = 67 architecture categorizations, and 17+10+13 = 40 learning-rule categorizations. None of these sums to 151, and the figures are captioned as distributions of the reviewed papers. Either the text is wrong, the figures cover a subset, or most papers were unclassifiable—the manuscript doesn't say which. Without the list of included studies or exact search strings, a reader cannot reconcile.\n\nCredit where it's due: the paper does something no prior review does in one place. It connects datasets, architectures, learning rules, and hardware into a single applied pipeline, adds a PRISMA flow and a practitioner framework (Figure 5), and ships an open-source repository with runnable examples for encoding, training, and simulation. The neuron-model and coding sections are competent; the STDP variant taxonomy is useful; the hardware summaries are accurate. That is real service, and the code is independently checkable.\n\nSoft spots, in proportion. The 'codifies effectiveness' wording overstates what is, at bottom, a catalog of reported accuracies. There is no quantitative meta-analysis, no error bars, no assessment of whether the reported numbers were reproduced. The internal inconsistency about simulated vs hardware counts is not cosmetic: the central claim of the review is the codification of trends from this corpus, and the corpus distribution is not verifiable. The authors also repeatedly cite their own repository [71] for tutorials; that's benign self-reference and not a problem for the main claims.\n\nWho is this for? A newcomer or a practitioner needing a map of the field and a copy-paste starting point. That reader will get genuine value even in the current form.\n\nRecommendation: a serious editor should send this to peer review—the field needs integrative reviews and the repo is evidence of effort. But acceptance should require the authors to list all 151 included studies, provide exact search strings and screening logs, reconcile the counts in Section II.E, and temper the effectiveness claims. If they do that, this becomes a solid reference.","headline":"Useful practitioner review with a real code base, but the central quantitative synthesis isn't auditable and the counts in Section II.E don't foot to 151.","tokens_in":45364,"tokens_out":1933,"would_cite":true,"duration_ms":16950,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 151 articles on spiking neural networks for computer vision codifies the architectures, learning rules, and hardware trade-offs that a practitioner needs to choose among.","keywords":["spiking neural networks","object detection","event cameras","neuromorphic hardware","STDP","surrogate gradients","ANN-to-SNN conversion","systematic review"],"falsifier":"An independent systematic search over the same databases (IEEE Xplore, Scopus, PubMed, Google Scholar) and years (2000–2023) with explicit query strings that yields a materially different distribution of architectures or learning rules—for example, fewer than 23 spiking-convolutional papers or a different leading training method—would falsify the codification.","tokens_in":44300,"feed_emoji":"⚡","tokens_out":6911,"duration_ms":56579,"temperature":0.7,"pith_summary":"This review of 151 peer-reviewed articles on spiking neural networks (SNNs) for computer vision sets out to organize the field into a small number of comparables: architecture families, learning-rule families, implementation mediums, and evaluation metrics. It claims that, read together, these papers support a codification of which architectural and training choices have been reported to work, and with what trade-offs in energy, latency, and memory on neuromorphic hardware. The paper's practical purpose is to give a newcomer a structured pipeline—from data acquisition and encoding through architecture, learning, implementation, and evaluation—plus an open-source repository of runnable examples, so that design decisions can be grounded in documented evidence rather than trial and error. A sympathetic reader would care because the SNN literature is fragmented across isolated surveys of learning rules, hardware, and applications; this review attempts the integration that makes the field navigable.","feed_headline":"151 SNN papers codify architecture and hardware trade-offs","feed_subtitle":"The review gives newcomers a pipeline to pick datasets, learning rules, and neuromorphic hardware.","key_machinery":"The machinery is the systematic categorization scheme plus the six-stage practitioner pipeline. The scheme classifies every reviewed article along four axes—architecture (FCSNN, HSNN, SCNN, SDBN, SRNN), learning rule (direct unsupervised, direct supervised, indirect), implementation medium (simulation or neuromorphic hardware), and evaluation metric (accuracy, energy, latency, memory)—and the pipeline orders these axes into data type, encoding, architecture, learning, implementation, and evaluation. The scheme carries the argument because all of the paper's distributions, tables, and trade-off statements are derived from it.","core_discovery":"The paper's central claim is that the 151 selected studies, analyzed through a PRISMA-guided systematic review, support a single practitioner framework for event-based SNN object detection. The review codifies three things: the effectiveness of fully connected, hierarchical, convolutional, deep-belief, and recurrent architectures; the performance of direct unsupervised (STDP-family), direct supervised (temporal backpropagation, surrogate gradients), and indirect (ANN-to-SNN conversion) learning methods; and the trade-offs among energy, latency, and memory across simulation and neuromorphic hardware implementation mediums. It reports, for example, that spiking convolutional networks are the most frequently discussed architecture in the surveyed set, that STDP and surrogate-gradient methods dominate the learning side, and that conversion from pre-trained ANNs is a practical but temporally limited shortcut. On that basis, the paper asserts that a practitioner can choose datasets, architectures, learning rules, and hardware by consulting the documented distributions and accuracy tables rather than starting from scratch.","pith_inferences":["If the codification is correct, an implicit implication is that reported accuracy differences across studies may be driven as much by dataset and encoding choices as by architecture or learning rule; the paper does not hold encoding fixed when comparing methods, so an apples-to-apples benchmark varying encoding alone would be a natural test of that implication.","The claim that converted SNNs lack native temporal learning suggests a testable extension: conversion followed by surrogate-gradient fine-tuning on event data should recover some of the temporal performance, a hybrid the review points toward but does not evaluate systematically.","The qualitative trade-off statements could be turned into a quantitative Pareto frontier: mining the 151 papers for raw energy and latency numbers would let practitioners see which hardware platforms dominate which regimes."],"forward_implications":["A newcomer can use the framework and the companion code repository to select a dataset, encoding, architecture, learning rule, and hardware combination that matches documented trade-offs rather than relying on trial and error.","The reported dominance of spiking convolutional networks implies that spatial feature extraction is the current main driver of SNN object-detection performance.","The survey's accuracy tables provide concrete reference points—for example, unsupervised STDP reaching about 95% on MNIST, surrogate-gradient training reaching about 99% on N-MNIST, and ANN-to-SNN conversion exceeding 99% on MNIST—that future work can benchmark against.","The documented energy, latency, and memory trade-offs give practitioners a decision rule for choosing between simulation environments and neuromorphic hardware for a given deployment target.","The challenges the review identifies—training stability, limited hardware accessibility, and the lack of native temporal learning in conversion—define the field's near-term research agenda."],"supporting_citations":[{"why":"Supplies the PRISMA systematic-review methodology that defines the paper's selection of 151 articles.","marker":"[22]"},{"why":"Provides the classic unsupervised STDP FCSNN result on MNIST (95%), a cornerstone of the direct-unsupervised category and a baseline for later comparisons.","marker":"[35]"},{"why":"The ANN-to-SNN conversion framework that anchors the indirect-learning category and the practitioner pipeline's conversion constraints.","marker":"[53]"},{"why":"Establishes the data-based and model-based weight/threshold normalization methods that indirect learning relies on.","marker":"[76]"},{"why":"The deep SCNN with STDP whose Caltech 101 accuracy anchors the unsupervised-SCNN and STDP categories.","marker":"[79]"},{"why":"The snnTorch simulation framework, a load-bearing example in the implementation-medium discussion.","marker":"[21]"},{"why":"Defines the neuromorphic-converted dataset category through N-MNIST and N-Caltech101.","marker":"[94]"},{"why":"The 1 Megapixel Automotive dataset that defines the neuromorphic-captured category and the high-resolution detection benchmark.","marker":"[101]"}],"fun_headline_variants":["151 SNN studies codify object detection trade-offs","Event-based SNN review: architectures, learning, hardware","Systematic review maps SNN object detection choices","Review of 151 SNN papers aids detector design","SNN review: datasets, learning rules, hardware trade-offs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 151 papers that survived screening are a representative and unbiased sample of the whole spiking-neural-network-for-computer-vision literature, so the reported distributions and trade-offs truly describe the field.","fun_headline_variants_meta":{"raw":{"variants":["151 SNN studies codify object detection trade-offs","Event-based SNN review: architectures, learning, hardware","Systematic review maps SNN object detection choices","Review of 151 SNN papers aids detector design","SNN review: datasets, learning rules, hardware trade-offs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000607,"raw_usage":{"total_tokens":2809,"prompt_tokens":905,"completion_tokens":1904,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":1826}},"tokens_in":521,"tokens_out":1904,"duration_ms":12793,"temperature":1.0,"reasoning_tokens":1826,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:36:55.807771+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent systematic search over the same databases (IEEE Xplore, Scopus, PubMed, Google Scholar) and years (2000–2023) with explicit query strings that yields a materially different distribution of architectures or learning rules—for example, fewer than 23 spiking-convolutional papers or a different leading training method—would falsify the codification.","supporting_citations":[],"review_version":1}