{"id":"d7122d7c-7ede-49d9-891f-d7b4b76bf427","arxiv_id":"2501.15721","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This paper proposes that the meaning of music emerges from interoceptive predictive coding within a multi-agent symbol emergence system, parallel to language.","lead":"This paper outlines a speculative theory: the meaning of music comes from the way our brain predicts bodily, emotional signals, and those predictions become shared as a social symbol system. The idea connects robotics research on how symbols emerge with neuroscience of emotion and music.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Music's firstness breaks the arbitrariness required for symbol emergence, undermining the Section 4.2 mapping.","rationale":"The reader's conditional verdict is appropriate because the paper is an explicitly speculative perspective with no empirical or formal support for the music-emotion mapping. My concern complements and sharpens the reader's weakest assumption: even if music does not explicitly represent external events, its demonstrable firstness—acknowledged in Section 5—directly undermines the arbitrariness that the symbol emergence framework requires. This is an internal tension, not merely an external disagreement with consensus. The paper could potentially resolve it by modeling music as a partially non-arbitrary sign that both perturbs interoceptive states and is socially coordinated, but it does not do so. I recommend keeping the conditional verdict: the hypothesis is coherent enough to warrant further development, but the unresolved firstness issue must be addressed before the central claim can be accepted even as a programmatic statement. The concrete test above would determine whether the proposed framework can accommodate this feature or whether the mapping requires a fundamental modification.","tokens_in":14058,"tokens_out":5303,"duration_ms":54038,"concrete_test":"Formalize the music-side generative model by writing the listener's interoceptive observations as p(o_A | z_A, w), where w is the musical sequence and z_A the emotional state, instead of the symbol-emergence model's p(o_A | z_A). Then re-run the Metropolis-Hastings naming game inference from [69] on two agents with this modified likelihood. If the resulting posterior over the shared w fails to converge to a common symbol system, or if the derivation of Eq. (10) requires the independence assumption p(o_A | z_A, w) = p(o_A | z_A), then the proposed mapping is not formally valid. Alternatively, an analytical check: determine whether Eq. (10) can be derived when o_m depends directly on w; if not, the analogy's formal backbone fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central analogy in Section 4.2 maps language's perceptual internal representations to music's emotional internal representations, with external sensorimotor interaction replaced by interoceptive interaction. In the symbol emergence model of Section 3 and Figure 2, the shared sign w coordinates agents because both z_A and z_B predict the same external world. For music, however, the proposed z is the listener's emotional state inferred from private interoceptive signals, so there is no common external referent that would allow agents to align their internal representations. More seriously, Section 5 explicitly concedes that music has 'many symbolic aspects of firstness,' meaning the sign itself directly affects visceral senses. If a musical sign directly influences interoceptive observations, the generative process becomes o_m ~ p(o_m | z, w), not merely o_m ~ p(o_m | z). This breaks the head-to-head latent variable structure of Eq. (10), where w is a shared latent variable conditionally independent of observations given z. The paper does not explain how collective predictive coding can produce a shared musical symbol system when the sign is partly non-arbitrary and causally coupled to the private observations it is supposed to signify. Without addressing this, the central claim that music is a socially emerged symbol system grounded in interoceptive predictions is not supported by the framework it invokes.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a speculative parallelism between music and language from the viewpoint of symbol emergence systems modeled with probabilistic generative models (PGMs). It first reviews the author's prior work on symbol emergence in robotics, where a multi-agent symbol system is modeled as decentralized Bayesian inference and described as 'collective predictive coding.' It then extends this framework to music by reinterpreting the latent variables: perceptual internal representations become emotional internal representations, and sensorimotor interaction with the external world becomes interoceptive interaction with the internal environment. The paper hypothesizes that 'the meaning of music' is the change in the listener's mental state, i.e., the inference of emotional internal representations, and that music is a socially emerged symbol system grounded in interoceptive predictions. The paper explicitly states that it does not provide further details or evidence for the correspondence.","tokens_in":14277,"tokens_out":4709,"duration_ms":44057,"significance":"If the hypothesis were substantiated, it would provide a unifying computational framework for music semantics, linking music to emotion, predictive coding, and social symbol emergence, and it would open a concrete research program, such as extending PGM-based language models to music with emotional latent variables. The paper is honest about its speculative status and clearly lays out the key analogies in Figure 2. Its strengths are the explicit computational framing and the connection to a well-developed line of PGM-based symbol emergence research, including machine-checked or reproducible models in prior work. However, as a proposal it does not yet provide a derivation, simulation, or empirical test, and the formal part currently contains an error in Eq. (10) and an unresolved tension with the paper's own admission of music's 'firstness.' The paper's value is as a position statement that may provoke discussion, but its central claim is not yet adequately supported.","major_comments":[{"comment":"Equation (10) as written is formally incorrect: it duplicates p(zA|{oAm}) and omits p(zB|{oBm}). The correct factorization should involve p(w|zA, zB) p(zA|{oAm}) p(zB|{oBm}), or an equivalent symmetric form. This is not a harmless typo, because the equation is presented as the mathematical description of the symbol emergence system, and the reader cannot verify the claimed decentralized Bayesian inference without the correct terms.","section":"Section 4.2, Eq. (10)"},{"comment":"The paper concedes that music has 'many symbolic aspects of firstness' and that musical signs can directly affect visceral senses. In the language model, the shared sign w is conditionally independent of observations o given the internal states z, which is what allows two agents to align their internal representations through a common external world. For music, the sign may directly influence interoceptive observations, yielding o_m ~ p(o_m | z, w) rather than o_m ~ p(o_m | z). This changes the generative structure of Eq. (10) and undermines the proposed mapping. The paper does not explain how a symbol system can still emerge when the sign is partly non-arbitrary and causally coupled to the private observations it is supposed to signify. The author should either extend the model to include firstness and demonstrate that symbol emergence can still occur, or argue explicitly why firstness does not preclude the conventional alignment of musical signs.","section":"Section 5 (and Section 4.2)"},{"comment":"The statement that Eqs. (8) and (9) are 'identical' to Eqs. (1) and (2) is formally true but semantically superficial. In the language case, the latent z is presented as a cause of the utterance and is grounded in external objects and sensorimotor information; in the music case, z is an emotional state inferred from interoceptive signals. The formal similarity of two sequence-generating models is not by itself evidence of parallelism between the underlying cognitive mechanisms; it is exactly the proposed hypothesis that these latent variables play analogous roles. Please clarify that this is an analogy in generative form, not a derivation, and discuss what additional assumptions are needed to make the analogy substantive.","section":"Section 2.2, Eqs. (8)-(9)"}],"minor_comments":[{"comment":"The right panel of Figure 2 uses 'introspective signals' while the text consistently uses 'interoceptive signals'; please unify the terminology.","section":"Figure 2 caption"},{"comment":"The phrase 'the author pointed out' appears in the first-person narrative of a formal paper; please rephrase to avoid a personal voice.","section":"Section 5"},{"comment":"References [41] and [42] are the same Okanoya 2007 paper; remove the duplicate and renumber accordingly.","section":"References"},{"comment":"Equation (5) writes inference as z ~ p(z|{om}); this notation could be confused with sampling from the generative model. Consider using posterior notation such as z* = argmax z p(z|{om}) or a more explicit inference formulation.","section":"Section 2.1, Eq. (5)"},{"comment":"The paper would benefit from a short discussion of testable predictions of the hypothesis, for example, whether a multi-agent PGM with interoceptive observations can actually produce shared musical symbols in simulation.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is a keynote post-proceedings paper, which may explain its speculative and review-style nature. The central new idea (music as a symbol emergence system grounded in interoceptive prediction) is intriguing and could be valuable if properly formalized. The main blockers are the incorrect Eq. (10) and the unresolved firstness objection, both of which can be fixed within the paper's scope by revising the model or the argument. I would not reject, but the current version does not yet support the central claim as stated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a speculative perspective piece, not a result paper, and the author knows it. The only genuinely new content is the mapping in Sec. 4.2 and Fig. 2: replace perceptual internal representations with emotional ones, replace sensorimotor interaction with the external world by interoceptive interaction with the body, and then read symbol emergence in music as collective predictive coding of emotions. Everything else is a competent review of the author's and others' work on probabilistic generative models for language, object/place concepts, and the Metropolis-Hastings naming game. The review part is accurate and useful, and the author is honest that the music hypothesis is a hypothesis, explicitly not backed by evidence in this paper.\n\nWhere the paper is soft:\n\nFirst, the formal parallelism in Eqs. (8) and (9) is real but thin. Those equations say music and language are both sequences generated from a latent cause; that is true of almost any sequential generative model. The content of the parallelism only comes from the mapping to interoceptive emotions, and that mapping is asserted rather than derived.\n\nSecond, Eq. (10) contains an obvious typo: p(zA|{oA_m}) appears twice and p(zB|{oB_m}) is missing. Since this equation is the centerpiece of the claimed correspondence, that is not cosmetic. Section 4.2 also has a garbled sentence about 'the meaning itself is observable,' which I assume should be 'not observable.'\n\nThird, and most important, the stress-test concern is right. Section 5 concedes that music has 'many symbolic aspects of firstness': the sign itself can directly affect visceral sensations, e.g., heartbeat-synchronized sounds. In the symbol emergence PGM, w is a shared latent variable explaining each agent's observations through zA and zB, conditionally independent of observations given z. If a musical sign acts directly on interoceptive observations, the generative model should be o_m ~ p(o_m | z, w), not o_m ~ p(o_m | z). That breaks the head-to-head structure the paper relies on. The paper does not address this, and it undermines the claim that music is a socially emerged symbol system in the same sense as language.\n\nStill, the paper is clearly written, well referenced, and honestly scoped. It is a position piece, and as a position piece it deserves engagement. I would send it to peer review if the venue takes perspectives; I would not let it through as an established account. The author should fix Eq. (10), clarify the 'observable' sentence, and either add a concrete minimal model with testable predictions or explicitly restrict the claim to the thirdness aspects of music.\n\nWho gets value: researchers working on symbol emergence, grounded language, predictive coding, and music semantics. Useful for discussion, not as a result.\n\nRecommendation: referee it, with the expectation of major revision.","headline":"A clear, well-scoped speculative perspective whose central music-emotion analogy is asserted rather than shown; the paper's own admission of musical firstness breaks the proposed PGM mapping.","tokens_in":14779,"tokens_out":2901,"would_cite":false,"duration_ms":28547,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes that the meaning of music is the change a musical passage induces in a listener's inferred emotional state, modeled as collective predictive coding over interoceptive signals.","keywords":["symbol emergence systems","probabilistic generative models","collective predictive coding","interoceptive inference","meaning of music","language and music parallelism","automatic music composition"],"falsifier":"A decisive test would compare listeners' reported or physiological emotional responses to a musical passage while their interoceptive priors are manipulated, for example by making their heartbeats audible or by inducing a reliable visceral state. If musical meaning can be fully predicted from acoustic structure alone, without any measurable dependence on interoceptive state inference, the central claim fails; conversely, if the same acoustic signal gives different inferred emotional meanings solely because of different bodily priors, the claim is supported.","tokens_in":13795,"feed_emoji":"🎵","tokens_out":6215,"duration_ms":54847,"temperature":0.7,"pith_summary":"Languages and music both look like linear sequences of discrete sounds, and the paper argues this surface similarity runs deeper: both can be modeled as symbol systems that emerge when agents predict their own sensory streams. The paper's central hypothesis is that symbol emergence in a society is collective predictive coding, and that the same machinery explains how language and music acquire meaning. For music, the claim is that the meaning of a piece is the change it induces in a listener's inferred internal state, where emotions are themselves predictions about interoceptive (visceral) signals. By replacing external sensorimotor grounding with internal, bodily grounding, the author proposes a direct parallel between language and music as emergent social symbol systems. A sympathetic reader would care because this offers a computational, testable route from emotion research to music semantics.","feed_headline":"Music's meaning is the body-state change it predicts","feed_subtitle":"Music as collective predictive coding over interoceptive signals, linking emotion, language, and composition.","key_machinery":"The load-bearing mechanism is collective predictive coding realized as decentralized Bayesian inference in a probabilistic generative model. In the PGM for symbol emergence, two agents share a latent word $w$ while each maintains private internal representations $z_A$ and $z_B$ generated from their own observations, and communication is an approximate posterior sampling scheme, the Metropolis-Hastings naming game. The extension to music replaces the observation set $\\{o_m\\}$ with interoceptive signals and the internal representation with an emotional state, so that the same head-to-head graphical model describes how musical symbols arise and stabilize. This carries the analogy: what makes a sign meaningful is not a fixed dyadic relation to an object but its role in predicting an agent's own future signals.","core_discovery":"On the paper's own terms, the discovery is a mapping: the equations that describe a robot learning words from multimodal sensorimotor data, $w, \\{o_m\\} \\sim p(w,\\{o_m\\}|z)$ and $y \\sim p(y|w)$, are identical in form to equations for music generation in which $z$ is an emotional state. The author then replaces the perceptual internal representation of symbol emergence systems with an emotional internal representation, and replaces interaction with the external world via sensorimotor signals with interaction with the internal environment via interoception. The resulting proposal is that 'the meaning of music' is a change in the mental state that the listener undergoes, i.e., inference or state updating of internal representations under emotional predictive coding. Music, on this view, is a socially emerged symbol system grounded not in exteroceptive facts but in interoceptive predictions, with individual agents' emotional states coordinated through semiotic communication just as perceptual states are coordinated in language.","pith_inferences":["This suggests a testable asymmetry: musical signs can act directly on visceral sensing, so the sign itself participates in its own grounding, whereas arbitrary linguistic signs must acquire grounding through learned exteroceptive association.","A natural extension is to expressive prosody and performance gesture, which also carry emotional content through interoceptive predictions and could be modeled with the same head-to-head PGM.","If musical meaning is collectively inferred emotional state, then cultural differences in musical taste become differences in shared priors over interoceptive signals, which could be studied by comparing predictive models fitted to listeners from different traditions."],"forward_implications":["If the meaning of music is an update of emotional internal representations, then automatic composition can be reframed as sampling from a posterior distribution over emotional states, extending existing language-model approaches to note sequences.","A musical symbol system, like a language, is not fixed in a score or a composer's intention; it is continuously re-emerged through social coordination of listeners' interoceptive predictions.","The same PGMs used for unsupervised phoneme and word discovery, double articulation analysis, and multimodal object and place concept formation should be transferable to modeling musical structure and its emotional grounding.","Emotion can be brought into computational music research as a latent variable with a well-defined generative semantics rather than a categorical label."],"supporting_citations":[{"why":"establishes symbol emergence as collective predictive coding via the Metropolis-Hastings naming game.","marker":"[69]"},{"why":"supplies the premise that emotions are predictive coding of interoceptive signals.","marker":"[48]"},{"why":"grounds interoceptive prediction in specific brain mechanisms.","marker":"[8]"},{"why":"provides the nonparametric Bayesian double articulation analyzer for unsupervised word and phoneme discovery.","marker":"[63]"},{"why":"defines symbol emergence systems and the semiotic framework for meaning.","marker":"[67]"},{"why":"models symbol emergence as interpersonal multimodal categorization within a PGM.","marker":"[27]"},{"why":"extends the PGM model to multiagent emergent communication.","marker":"[25]"},{"why":"introduces the head-to-head latent word connection used in the central graphical model.","marker":"[24]"},{"why":"supplies the mutual segmentation hypothesis linking songbird vocalization to language acquisition.","marker":"[43]"}],"fun_headline_variants":["Music's meaning is the predicted body-state shift","Music as collective predictive coding of interoception","Music's meaning emerges from emotional predictive coding","Music's meaning: a change in predicted body state"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the assumption that music, unlike language, does not explicitly represent events in the external world, so its meaning can be grounded in interoceptive rather than exteroceptive signals; if music primarily represents external events or is primarily grounded in body movement rather than visceral feeling, the proposed parallel loses its basis.","fun_headline_variants_meta":{"raw":{"variants":["Music's meaning is the predicted body-state shift","Music as collective predictive coding of interoception","Music's meaning emerges from emotional predictive coding","Music's meaning: a change in predicted body state"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000177,"raw_usage":{"total_tokens":1289,"prompt_tokens":934,"completion_tokens":355,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":296}},"tokens_in":550,"tokens_out":355,"duration_ms":3317,"temperature":1.0,"reasoning_tokens":296,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:00:57.619584+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would compare listeners' reported or physiological emotional responses to a musical passage while their interoceptive priors are manipulated, for example by making their heartbeats audible or by inducing a reliable visceral state. If musical meaning can be fully predicted from acoustic structure alone, without any measurable dependence on interoceptive state inference, the central claim fails; conversely, if the same acoustic signal gives different inferred emotional meanings solely because of different bodily priors, the claim is supported.","supporting_citations":[{"cited_title":"Emergent Communication through Metropolis-Hastings Naming Game with Deep Generative Models","cited_arxiv_id":"2205.12392","evidence_quote":"establishes symbol emergence as collective predictive coding via the Metropolis-Hastings naming game."},{"cited_title":"Trends in cognitive sciences 17(11), 565–573 (2013)","cited_arxiv_id":null,"evidence_quote":"supplies the premise that emotions are predictive coding of interoceptive signals."},{"cited_title":"Nature re- views neuroscience 16(7), 419–429 (2015)","cited_arxiv_id":null,"evidence_quote":"grounds interoceptive prediction in specific brain mechanisms."},{"cited_title":"IEEE Transactions on Cogn itive and Develop- mental Systems (2018)","cited_arxiv_id":null,"evidence_quote":"defines symbol emergence systems and the semiotic framework for meaning."},{"cited_title":"Advanced Robotics 36(5-6), 239–260 (2022)","cited_arxiv_id":null,"evidence_quote":"extends the PGM model to multiagent emergent communication."},{"cited_title":"In: IEEE Interna- tional Conference on Development and Learning (ICDL 2022)","cited_arxiv_id":null,"evidence_quote":"introduces the head-to-head latent word connection used in the central graphical model."},{"cited_title":"Psychonomic bulletin & review 24(1), 106–110 (2017)","cited_arxiv_id":null,"evidence_quote":"supplies the mutual segmentation hypothesis linking songbird vocalization to language acquisition."}],"review_version":1}