{"id":"dcbd248e-0d9e-48c5-93e5-9aa401197476","arxiv_id":"2411.14773","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A spiking neural network with an explicit mode and key theory subsystem learns pitch-class connection patterns resembling the Krumhansl-Schmuckler key profiles and generates four-part music conditioned on the requested mode and key.","lead":"A brain-inspired spiking neural network with an explicit Western mode and key theory module can learn tonality profiles and generate four-part music conditioned on a requested mode and key. A generalist might read this as a test of whether cognitive inductive biases can substitute for large statistical models in music generation, though key evidence is partly statistical.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The KS-similarity claim (Sec. 5.1, cosine 0.92–0.94) is confounded: PSC/PASW (Eq. 6-7, 14-15) reduce to corpus pitch-class co-occurrence statistics, and no histogram baseline is reported.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing concern: PSC/PASW are co-occurrence statistics, so the high cosine similarity to KS profiles may be an artifact of corpus pitch-class frequencies rather than a property of the spiking network or its mode-conditioning mechanism. This is indeed the most serious threat to the paper's headline claim. The concern is concrete: the definitions in Eq. 6-7 and Eq. 14-15 make the reduction transparent, and the paper itself notes dataset-specific deviations (e.g., SHTE subdominant emphasis) that are themselves corpus-statistical effects. The proposed histogram baseline is the minimal control that would settle whether the architecture adds anything beyond counting. Because the paper still presents a working mode-conditioned SNN generator with quantitative evaluation against baselines, the appropriate verdict remains CONDITIONAL pending that control. I agree with the reader's assessment and see no need to change the verdict.","tokens_in":18769,"tokens_out":3009,"duration_ms":31328,"concrete_test":"Compute the mode- and key-conditioned pitch-class histograms of SHTE and Bach (all notes, optionally duration-weighted) and normalize them; calculate cosine similarity to the KS major/minor profiles (and to each of the 24 key profiles for Fig. 6). If these baseline cosine values are within ~0.02 of the reported PSC/PASW values (e.g., 0.92–0.94), the architecture adds no measurable evidence for KS-like internal representation beyond corpus statistics. Ideally, also recompute PSC/PASW after shuffling the mode/key labels or removing the co-firing threshold to isolate the co-occurrence component.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the trained architecture's connections approximate the Krumhansl-Schmuckler profiles rests on PSC and PASW. These are defined in Eq. 14 and 15 from binary synapses (Eq. 6-7) created when mode/key-cluster neurons and pitch neurons co-fire at least five times, and from STDP weights on those synapses. In the training procedure, the mode/key cluster is activated for the entire piece and pitches fire whenever they occur, so PSC(k) is essentially a binarized, thresholded count of how often pitch class k appears in pieces labeled with that mode/key; PASW is a weighted version of the same co-occurrence. The KS profiles themselves are empirically correlated with pitch-class usage in Western tonal music. Therefore the reported cosine similarities (0.92–0.94 for SHTE and Bach) may simply reflect the pitch-class histogram of the training corpus, not any emergent property of the spiking dynamics, the synaptic-creation rule, or the mode-conditioning architecture. The paper never computes a histogram baseline or a 'bag-of-pitches' control. Absent such a control, the claim of a 'connection framework closely similar to the KS model' is not distinguished from a trivial counting statistic, and the neuroscience/psychology interpretation is over-stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multi-region spiking neural network, implemented on BrainCog, for mode- and key-conditioned four-part music generation. A music theory subsystem encodes major/minor modes and 24 keys through hand-wired scale-degree connections, while a sequential memory subsystem encodes pitch and duration. Learning relies on activity-dependent synaptic creation and STDP. The authors define Pitch Synaptic Count (PSC) and Pitch Average Synaptic Weight (PASW) features of the trained connections, compare them against Krumhansl-Schmuckler (KS) key profiles, and evaluate generated pieces against SHTE and Bach datasets and an SCSNN baseline.","tokens_in":19123,"tokens_out":4659,"duration_ms":48642,"significance":"If the KS-similarity claim were established, the model would be an interpretable, brain-inspired generative architecture whose internal connectivity mirrors psychological tonal hierarchies. The paper contributes a new SHTE dataset, releases code on GitHub, and benchmarks against an existing SNN baseline with quantitative KLD/OA metrics. However, the central similarity result is currently not distinguished from trivial pitch-class corpus statistics, and the generation evaluation aggregates over keys without per-key compliance tests or significance statements. The conceptual claims therefore outrun the evidence provided.","major_comments":[{"comment":"PSC and PASW are defined as counts and average weights of synapses created when mode/key-cluster neurons and pitch neurons co-fire at least five times. Because the mode/key clusters are driven by the input note pitches themselves (Eqs. (1)-(2)), a synapse between mode neuron k and pitch-class-k neurons is created essentially whenever pitch class k occurs in a piece labeled with that mode/key. PSC(k) is therefore a thresholded pitch-class co-occurrence count of the training corpus, and PASW(k) is an STDP-weighted version of the same statistic. The reported cosine similarities of 0.92-0.94 to KS profiles may simply reflect the pitch-class distributions of Western tonal music in SHTE and Bach, rather than any emergent property of the spiking dynamics, synaptic-creation rule, or conditioning architecture. The paper does not report a baseline control, for example the cosine similarity between raw (or thresholded) pitch-class histograms of the training corpora per mode and the corresponding KS profiles. Such a control is necessary to support the paper's central claim that the trained connection framework is closely similar to the Krumhansl-Schmuckler model.","section":"Sec. 5.1, Eqs. (14)-(15) together with Eqs. (6)-(7)"},{"comment":"The generation evaluation aggregates 50 generated samples across 'different keys' and reports mean KLD/OA values against SHTE and Bach. These aggregate statistics do not test the conditioning claim: they do not show that a piece generated in a specified key has pitch-class statistics closer to that key than to other keys, and no per-key diatonic pitch rate or per-key KLD is reported. Without per-key compliance tests and associated significance or confidence statements, the claim that the model generates music with the characteristics of the given modes and keys is not established.","section":"Sec. 5.2, Table 3 and Sec. 5.2.2"},{"comment":"Figure 6 reports zero cosine similarities for C# major, F# major, C# minor, Ab minor, A# minor, and D# minor because no training pieces exist in those keys, as confirmed by Table 2. The paper nonetheless claims that the model can generate music in various modes and keys. Section 5.2 does not state which keys were used for the 50 generated samples. If the generated set excludes the unseen keys, the generalization claim is unsupported; if it includes them, per-key results must be reported. Please clarify and, if appropriate, evaluate unseen keys separately.","section":"Sec. 5.1, Fig. 6 and Sec. 5.2"}],"minor_comments":[{"comment":"The caption contains a typo: 'pitch lass' should be 'pitch class'.","section":"Eq. (14) caption"},{"comment":"'Winner-Takes-All priciple' should be 'Winner-Takes-All principle'.","section":"Sec. 3.4"},{"comment":"The word 'Phenorminans' appears in the bullet list; it should likely be 'phenomenon' or similar.","section":"Sec. 5.1"},{"comment":"The acronym is inconsistently written as both 'PSWA' and 'PASW' in the text; please unify.","section":"Sec. 5.1"},{"comment":"The phrase 'oscillatory times' in the description of Eq. (6) is confusing; 'coincidence counts' or 'co-firing counts' would be clearer.","section":"Sec. 3.3.1"},{"comment":"The statement 'Approximately 91.5% of the values are above 0.7' should be recomputed and justified, since the figure includes zero values for absent keys.","section":"Sec. 5.1, Fig. 6"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern is well-founded: the central KS-similarity claim currently lacks a histogram baseline, and PSC/PASW plausibly reduce to thresholded corpus pitch-class statistics. If the authors add a raw-corpus-histogram control and per-key generation tests, the paper may be salvageable; otherwise the main claim collapses to a pitch-counting artifact. I recommend major revision rather than rejection, because the identified issues are addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a plausible, useful piece of work with one major unaddressed confound. The model is a spiking neural network that learns to generate four-part music conditioned on mode and key. What's actually new: the explicit music-theory subsystem (mode cluster and key cluster with hierarchical connections), the new SHTE dataset (193 textbook harmony exercises), and the fact that nobody else, as far as the cited literature goes, has done mode-conditioned SNN generation. The code is public. That is real.\n\nThe paper's central claim is that after training, the synapses between the mode/key clusters and the pitch subnetworks have cosine similarity 0.92–0.94 with the Krumhansl-Schmuckler profiles. The stress-test note is correct: PSC and PASW, as defined in Eq. 14–15, are essentially binarized and weighted counts of co-firing between the always-active mode/key neurons and the pitch neurons. Since those neurons fire whenever a pitch appears, PSC(k) reduces to a thresholded pitch-class histogram of the training corpus. The KS profiles are known to correlate with pitch-class usage. So the high cosine similarity may be nothing more than the model reflecting the corpus statistics. No histogram baseline or bag-of-pitches control is reported. This is not a fatal flaw, but it is a load-bearing gap: without such a control, the neuroscience/psychology interpretation is over-stated.\n\nThe generation evaluation has a similar issue. The aggregated KLD/OA comparisons to the training sets show the output resembles the data, which is nice, but there is no per-key compliance test, no statistical significance, and no listening test. The comparison to the SCSNN baseline is helpful, but again aggregate.\n\nWhat the paper does well: it is transparent about missing keys (the zero values in Fig. 6), it gives the full architecture and parameters, and it is candid about dataset-specific deviations from KS. The SHTE dataset alone is worth having.\n\nIn short: the idea is good, the dataset is reusable, and the architecture is clearly described. The central claim needs a control baseline before publication. With that added, the paper would be a solid contribution to computational music cognition. I would send this to serious review rather than desk-reject it, with the clear expectation of a major revision.","headline":"A mode-conditioned spiking network with a genuinely useful new dataset, but the KS-profile similarity claim needs a histogram baseline before it can be believed.","tokens_in":19629,"tokens_out":2268,"would_cite":true,"duration_ms":22688,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a brain-inspired spiking neural network, trained on four-part harmony and given Western mode/key theory as prior knowledge, develops pitch-connection statistics matching the Krumhansl-Schmuckler key profiles with…","keywords":["spiking neural network","music generation","mode and key perception","Krumhansl-Schmuckler model","STDP","symbolic music","four-part harmony"],"falsifier":"Compute the cosine similarity between the raw pitch-class histogram of the SHTE and Bach corpora and the KS profiles. If those corpus-histogram similarities are close to the reported 0.92–0.94, then the PSC/PASW agreement is inherited from corpus statistics, and a control experiment without spiking dynamics — e.g., counting co-occurrences directly from note tokens — would reproduce the result.","tokens_in":18549,"feed_emoji":"🎵","tokens_out":9540,"duration_ms":79691,"temperature":0.7,"pith_summary":"This paper tries to show that a brain-inspired spiking neural network can internalize Western musical modes and keys, and that the connections it forms between its mode/key neurons and its pitch neurons closely resemble the Krumhansl-Schmuckler (KS) key profiles from music psychology, with cosine similarities between 0.92 and 0.94. If true, the trained network is a mode-conditioned music generator whose internal architecture mirrors human tonal perception, offering a mechanistic account of how key hierarchies might emerge from exposure to music. The authors train the network on four-part harmony exercises and Bach chorales, then measure two connection statistics per pitch class (synaptic count and average synaptic weight) and compare them with the KS profiles. They also show the network can generate four-part pieces in a specified mode and key, with statistics closer to the training corpora than an earlier spiking baseline.","feed_headline":"Spiking network's internal wiring matches human key perception","feed_subtitle":"After chorales, pitch connections hit 0.92–0.94 similarity with psychological key profiles; composes in requested modes.","key_machinery":"The carrying mechanism is a two-subsystem spiking architecture: a music theory subsystem (MTS) with a mode cluster and a 24-group key cluster that encodes the prior knowledge of major/minor modes and all keys, and a sequential memory subsystem (SMS) with pitch and duration subnetworks that encode ordered notes. Learning proceeds by a synaptic creation rule (Eq. 6–7) that forms a connection whenever two neurons co-fire at least five times, followed by STDP weight updates (Eq. 8). The key comparison objects are two statistics computed from the trained MTS-to-pitch connections — Pitch Synaptic Count (PSC), the number of synapses per pitch class, and Pitch Average Synaptic Weight (PASW), the mean weight per pitch class — which are normalized and compared against the Krumhansl-Schmuckler key profiles as a proxy for the psychological tonal hierarchy.","core_discovery":"The central discovery claimed is that a structured spiking network, given Western mode/key theory as prior knowledge and trained on symbolic four-part music via synaptic creation and STDP, develops pitch-class connection statistics whose normalized profiles match the Krumhansl-Schmuckler key profiles. The match is quantified by cosine similarities of 0.93 and 0.92 for the harmony-exercise dataset (major PSC and minor PASW) and 0.94 and 0.94 for the chorale dataset, with about 91.5% of per-key similarities above 0.7. The same architecture, when seeded with a tonic chord and conditioned on a mode and key, generates four-part music whose diatonic pitch rate (0.86) and other statistical features resemble the training sets, and which outperforms the earlier spiking baseline in key adherence and pitch range. The authors interpret the connection-similarity result as evidence that the model's learned representation of tonal importance is consistent with the psychological model of key perception.","pith_inferences":["A reader should treat the KS-similarity number as provisional: because PSC is defined as a binarized co-firing count, it is essentially a pitch-class frequency histogram of the training corpus, and a raw histogram may match the KS profiles just as closely without any spiking dynamics.","A direct control experiment comparing PSC/PASW with the corpus pitch-class histogram and with a non-spiking co-occurrence counter would separate what the spiking machinery contributes from what the music statistics already contain.","The generation quality is only compared against an earlier spiking baseline, not against statistical or deep-learning generators, so the claimed advantage in tonality characteristics and melodic adaptability would be strengthened by testing against a standard n-gram or Transformer baseline.","A behavioral extension would be to ask human listeners to rate the key clarity of generated versus corpus pieces; if listeners cannot distinguish the generated pieces' tonality from the training set, the conditioning claim would be validated perceptually."],"forward_implications":["The trained connection statistics provide a quantitative bridge between neural network learning and the Krumhansl-Schmuckler key-finding algorithm, so the same network can be used to test how different training corpora shift tonal hierarchies.","Because generation is conditioned on mode and key through the MTS prior, the model supplies a concrete route for steering symbolic music generation toward a specified tonality without retraining.","The model's internal representation adapts to dataset-specific harmonic conventions (e.g., higher subdominant values on the harmony-exercise dataset), suggesting that deviations from KS profiles may be usable as fingerprints of a corpus's harmonic style.","The four-part generation results position the model as a brain-inspired baseline whose output statistics resemble the training corpora more closely than an earlier spiking system on key adherence."],"supporting_citations":[{"why":"Defines the Krumhansl-Schmuckler key profiles used as the psychological target for similarity comparison.","marker":"Krumhansl (1990)"},{"why":"Describes and tests the KS key-finding algorithm, grounding the quantitative use of key profiles.","marker":"Temperley (1999)"},{"why":"Supplies the sequential memory subsystem architecture and spike-based learning for musical sequences.","marker":"Liang et al. (2020)"},{"why":"Provides the earlier brain-inspired spiking melody-composition system used as the generation baseline.","marker":"Liang and Zeng (2021)"},{"why":"The neuron model used to simulate dynamics in both subsystems.","marker":"Izhikevich (2003)"},{"why":"The STDP rule used for synaptic weight updates after synaptic creation.","marker":"Bi and Poo (1998)"},{"why":"Provides the Music21 toolkit through which the Bach chorale dataset is loaded.","marker":"Cuthbert and Ariza (2010)"},{"why":"Supplies the evaluation metrics (KLD, overlap area) used for generation assessment.","marker":"Yang and Lerch (2020)"}],"fun_headline_variants":["Spiking net's wiring mirrors human key perception","Brain-inspired AI learns modes, matches psychology model","Network's internal tonality map aligns with human perception","Music-composing spiking net reproduces psychological key profiles","This AI composer internalizes human-like key sense"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the number and average strength of the synapses between mode/key neurons and pitch neurons reveal the model's notion of which pitches matter, rather than simply reflecting how often each pitch appears in the training music.","fun_headline_variants_meta":{"raw":{"variants":["Spiking net's wiring mirrors human key perception","Brain-inspired AI learns modes, matches psychology model","Network's internal tonality map aligns with human perception","Music-composing spiking net reproduces psychological key profiles","This AI composer internalizes human-like key sense"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000461,"raw_usage":{"total_tokens":2350,"prompt_tokens":1029,"completion_tokens":1321,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":645,"completion_tokens_details":{"reasoning_tokens":1247}},"tokens_in":645,"tokens_out":1321,"duration_ms":12768,"temperature":1.0,"reasoning_tokens":1247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:54:56.790841+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the cosine similarity between the raw pitch-class histogram of the SHTE and Bach corpora and the KS profiles. If those corpus-histogram similarities are close to the reported 0.92–0.94, then the PSC/PASW agreement is inherited from corpus statistics, and a control experiment without spiking dynamics — e.g., counting co-occurrences directly from note tokens — would reproduce the result.","supporting_citations":[],"review_version":1}