{"id":"b77ad101-a3ba-4e31-bb1b-c23f691281e5","arxiv_id":"2608.08772","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Emotion neurons in audio-language models are mostly language-specific, but pooling evidence across languages finds a shared set whose causal manipulation transfers to unseen languages.","lead":"This paper probes where emotions are stored inside large audio-language models that understand speech in many languages. It finds that emotion neurons found in one language rarely match those in another, and a fusion method that pools languages finds neurons that give more targeted emotional control, even on languages never seen during identification.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The admitted language/corpus confound is the load-bearing risk: if CR-Fusion's gains come from recording style rather than language, the core 'multilingual' interpretation is unsupported.","rationale":"The confound is the single most load-bearing because it threatens the construct validity of the paper's key term 'multilingual.' Even if all the causal intervention results replicate exactly, they demonstrate control over corpus-style-correlated features, not necessarily over language-backed emotion representations. The paper's own post-hoc correlation strengthens the concern. The lambda tuning concern is real but secondary: the theory predicts small lambda, and the sensitivity analysis shows a plateau, so the specific value is less likely to invert the comparison. The lack of significance tests is also secondary because the pattern is consistent across four models and both intervention types. The confound, however, is structural: no amount of reanalysis of the existing 12 single-corpus languages can separate language from corpus. Only new data collection, as the authors note in the Limitations, can settle it. Therefore the conditional verdict stands: the central claim is plausible and well-supported in its causal mechanics, but the language-level interpretation requires the proposed control.","tokens_in":26773,"tokens_out":5119,"duration_ms":53909,"concrete_test":"Add a second English corpus of a different elicitation style (e.g., an acted corpus such as IEMOCAP alongside naturalistic MSP-Podcast), treat the two English corpora as separate 'languages' in CR-Fusion, and evaluate on the held-out language set. If the within-language cross-corpus fusion reproduces the cross-lingual fusion gain, the effect is corpus-driven; if it fails, the language-level interpretation is supported. Replicate with Mandarin for robustness.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Each language in the study is represented by exactly one corpus, so language identity and recording condition are confounded by construction. The Limitations section explicitly acknowledges this and reports that activation-structure similarity tracks corpus elicitation type at least as strongly as typological relatedness. The central claim—that CR-Fusion isolates a cross-lingually shared emotion component giving superior causal control—therefore rests on an attribution that the evidence cannot yet support: the fusion advantage in zero-shot and low-resource settings could reflect shared recording style, channel, or speaker population rather than shared language. Although §6.3 scrupulously reads the asymmetric transfer as 'non-redundant cross-corpus contributions,' the abstract and conclusion continue to assert language-level transfer. Absent a within-language, cross-corpus control, the study cannot distinguish the language-level hypothesis from a corpus-level alternative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a neuron-level interpretability study of multilingual emotion representation in four large audio-language models (Audio-Flamingo-3, Kimi-Audio, MiniCPM-o-4.5, Qwen2.5-Omni-7B) across 12 languages. The authors define Multilingual Emotion Neurons (MLENs) as units with stable emotional selectivity and aligned causal effects across languages, and propose Consistency-Regularized Fusion (CR-Fusion) to identify them from pooled multilingual activation statistics. Using a margin-based selector (ConAct) on eight identification languages, they report that monolingual emotion-sensitive neuron sets have minimal set overlap, that monolingual identification saturates beyond roughly 50 instances, and that CR-Fusion with a consistency penalty of lambda=0.3 produces stronger deactivation and steering effects than the best monolingual mask in most model-language conditions, with the exception of steering on Qwen2.5-Omni-7B. Leave-one-out ablations show asymmetric transfer patterns, and a variance decomposition attributes 30-59% of emotion-conditioned activation variance to a language-invariant component.","tokens_in":26912,"tokens_out":7745,"duration_ms":81347,"significance":"If the results hold, this would be the first causal, neuron-level account of how large audio-language models encode emotion across languages, with practical implications for training-free cross-lingual affective control. The paper is ambitious in scope (four models, twelve languages, two causal intervention types) and has notable strengths: an explicit estimand for the language-invariant emotion component, a noise-corrected variance decomposition, saturation and sensitivity analyses, and a candid Limitations section that admits the key confound. The central 'multilingual' interpretation, however, is currently undermined by two methodological issues: the one-corpus-per-language confound and the selection of the consistency penalty using the held-out evaluation languages. Both are addressable within a revision, either through additional experiments or through appropriately scoped claims.","major_comments":[{"comment":"The central claim that CR-Fusion identifies cross-lingually shared emotion neurons is confounded by the one-corpus-per-language design: each language in Table 3 is associated with a single dataset, so language identity and recording condition (elicitation style, channel, speaker population) are inseparable by construction. The Limitations section admits this and states that activation-structure similarity tracks corpus elicitation type at least as strongly as typological relatedness. Section 6.3 correctly frames the leave-one-out results as 'non-redundant cross-corpus contributions,' yet the abstract and conclusion re-assert language-level claims ('cross-lingual transfer,' 'shared affective representations that generalize across diverse spoken languages'). Because the fusion advantage and the asymmetric transfer patterns could be driven by shared recording style rather than by language, the defining construct of MLENs as cross-lingual units is not yet supported. The manuscript should either temper the language-level claims to corpus-level claims or add a control language represented by multiple corpora of different elicitation types.","section":"§4 (Table 3), §6.3, Limitations"},{"comment":"The consistency penalty lambda is chosen as the operating point in §6.2 by inspecting Figure 4, which plots average ESS across all 12 evaluation languages, including the four held-out languages. Consequently, the reported zero-shot fusion advantage in Tables 1 and 2 is partly tuned on the test languages, and the confirmation of the §3.4 prediction that 'the optimal penalty is near zero' is circular because the same evaluation data determined the operating point. To support the zero-shot claim, the authors should select lambda using only identification languages (or an inner cross-validation) and report held-out performance at that value, or alternatively present results across a range of lambda and show that the conclusions are robust without selection on the held-out ESS.","section":"§6.2, §3.4, Tables 1-2"},{"comment":"The central comparison between CR-Fusion and the best monolingual mask relies on point estimates with standard deviations across languages but no significance tests or confidence intervals. For example, in Table 1 MiniCPM-o-4.5 deactivation, CR-Fusion (-8.03) differs from the best monolingual mask (Mandarin, -7.64) by only 0.39 pp, with cross-language standard deviations above 4.0; similarly small margins appear in several other rows. The argument in Appendix A.4 that deterministic decoding makes confidence intervals unnecessary conflates seed variance (zero) with sampling variability over evaluation utterances and corpora. The authors should report bootstrap confidence intervals over test utterances or paired significance tests across evaluation languages, or temper the 'outperforms' claims accordingly.","section":"Tables 1 and 2, Appendix A.4"}],"minor_comments":[{"comment":"The claim of 'minimal overlap' would be strengthened by a chance-level baseline: with r=0.5% selection, two random sets of the same size would have an expected Jaccard similarity of roughly r/(2-r) ≈ 0.0025, so the observed JSC values near 0.10 are an order of magnitude above chance; reporting this contrast would make the interpretation more precise.","section":"§5.1, Figure 1"},{"comment":"The notation for the additive decomposition omits the layer and neuron subscripts on the right-hand side (m, u, b, gamma, epsilon), which can confuse the reader; please make the indexing explicit.","section":"§3.4, Eq. (2)"},{"comment":"The leave-one-out heatmaps are averaged over four models; please also provide per-model results or state in the caption that pooling may hide model-specific patterns, since the reader cannot tell whether the asymmetric transfer is driven by one model.","section":"§6.3, Figure 5"},{"comment":"The 'post-hoc analysis' showing that activation-structure similarity tracks corpus elicitation type at least as strongly as typological relatedness is mentioned but no quantitative result is reported; adding a small table or appendix entry would make this claim verifiable.","section":"Limitations"},{"comment":"The definition of E_valid for the global ESS is only mentioned parenthetically; please specify precisely which emotions are excluded per model-language condition when an emotion has zero correctly predicted instances, since this affects the comparability of the aggregated metric.","section":"§4, Appendix A.5"},{"comment":"The phrase 'the first causal, neuron-level account' is a strong novelty claim; consider softening it or explicitly identifying the closest prior work to avoid overclaiming.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"This is a well-executed and unusually honest manuscript, but the load-bearing issues identified by the referee (the language-corpus confound and the selection of lambda on the held-out evaluation set) appear genuine and need to be addressed before the central claims can be accepted. The confound is acknowledged in the Limitations, which is commendable, but the abstract and conclusion still overstate the language-level interpretation. I believe a major revision is appropriate: the authors can either add a multi-corpus control language or rescope the claims to corpus-level effects, and they should fix the lambda selection procedure. The lack of confidence intervals is also a clear requirement for a journal-level publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first causal, neuron-level study of cross-lingual emotion in LALMs, and the central empirical pattern—pooled identification beats any single-corpus mask—holds across four models. The paper deserves a serious referee, but the interpretation needs to be reined in until the language/corpus confound is addressed.\n\nWhat's genuinely new: CR-Fusion is a clean way to pool per-language selector scores with a consistency penalty, and the leave-one-out ablation shows low-resource languages contribute non-redundant evidence. The saturation analysis is useful, and the variance decomposition in §3.4/Appendix B is a nice formal addition—it gives an identifiable estimand and explains why Overlap Fusion underperforms. The causal interventions are real, with deterministic decoding and fixed seeds, and the authors are unusually candid in the Limitations section.\n\nThe soft spots are exactly where the reader flagged. Most importantly, each language is represented by a single corpus, and the paper concedes that activation similarity tracks elicitation style at least as strongly as typological relatedness. That means the fusion gains could be corpus-style transfer rather than language-level sharing. The abstract and conclusion still say \"across languages.\" This is not a minor footnote; it is load-bearing. A within-language, cross-corpus control is needed.\n\nSecond, the operating point λ=0.3 is chosen from the same evaluation ESS used in the main tables. The \"predicted optimal penalty\" in §3.4 is confirmed in §6.2 on the same data, so that specific prediction is circular as evaluated. Fix λ or tune on a held-out subset of languages.\n\nFinally, the main comparisons in Tables 1 and 2 carry no confidence intervals or significance tests. Given deterministic decoding, bootstrap over utterances would give a sense of sampling variability. This is minor compared to the confound, but it would tighten the claims.\n\nThe paper is not definitionally circular—CR-Fusion is a real method and the causal effects are measured. The \"first\" claims are mostly fair but should be scoped against the monolingual work by the same group.\n\nSummary: strong empirical contribution, honest limitations, but the central \"multilingual\" attribution is not yet proven. I'd send it to peer review and ask for a cross-corpus control and a pre-registered penalty. Even if the interpretation ends up being corpus-level rather than language-level, the method and the causal results are worth publishing.","headline":"A real first cut at cross-lingual emotion neurons in LALMs, with a load-bearing confound and a tuned penalty; the causal pattern is consistent, but the 'multilingual' interpretation needs a within-language cross-corpus control before it can be trusted.","tokens_in":27475,"tokens_out":2170,"would_cite":false,"duration_ms":22323,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large audio-language models carry a cross-lingually shared emotion code that pooling languages can isolate, and that single-language neuron sets cannot.","keywords":["multilingual emotion neurons","large audio-language models","causal interpretability","activation steering","cross-lingual transfer","consistency-regularized fusion","speech emotion recognition","neuron-level analysis"],"falsifier":"Take a single language represented by two corpora of different elicitation styles, such as acted versus naturalistic English, and rerun CR-Fusion with one corpus as identification and the other as held-out; if the fusion gain disappears or transfers along elicitation type rather than language, the multilingual-neuron interpretation collapses.","tokens_in":26549,"feed_emoji":"🧠","tokens_out":5326,"duration_ms":50690,"temperature":0.7,"pith_summary":"This paper tries to show that large audio-language models encode emotion partly through neurons shared across languages, and that identifying those neurons by pooling evidence from many languages yields causal control that no single-language procedure can match. The authors define Multilingual Emotion Neurons (MLENs) as units with stable emotional selectivity and aligned causal effects across languages, and propose Consistency-Regularized Fusion (CR-Fusion) to find them. Across four open models and twelve typologically diverse languages, neurons found separately in each language barely overlap, and adding more monolingual identification data saturates quickly. CR-Fusion masks, tested by deactivation and steering, outperform the best monolingual masks on held-out languages in zero-shot and low-resource settings, and leave-one-out analysis shows that low-resource languages contribute non-redundant evidence and are among those that benefit most. If right, this gives a mechanistic handle on cross-lingual emotion generalization and a training-free route to better affective control for languages with scarce data.","feed_headline":"Pooled languages reveal emotion neurons that transfer across languages","feed_subtitle":"Deactivating or steering these shared neurons shifts emotion recognition in languages never used for identification.","key_machinery":"The load-bearing object is the Multilingual Emotion Neuron (MLEN), defined as a neuron with stable emotional selectivity and aligned causal effects across languages, and the identification method is Consistency-Regularized Fusion (CR-Fusion). CR-Fusion takes per-language neuron scores, here the Contrastive Activation Margin between a neuron's best and second-best emotion, normalizes them per language, and ranks neurons by the cross-lingual mean minus λ times the cross-lingual standard deviation, so that λ=0 targets average transfer and larger λ targets quantile or worst-case stability. The paper also decomposes each neuron's activation probability into a baseline, a language effect, an emotion effect, a language-by-emotion interaction, and noise, arguing that fusion recovers the invariant emotion effect, a quantity no single-corpus procedure can isolate. Interventions either deactivate or amplify the selected neurons in the SwiGLU gate outputs to test causal necessity and sufficiency.","core_discovery":"The paper's central claim is that a language-invariant emotion component exists inside modern audio-language models and can be isolated by selecting neurons whose emotion selectivity is consistent across languages rather than maximal in any one language. Monolingual identification produces nearly disjoint neuron sets across languages (Jaccard similarity mostly below 0.10) despite weak-to-moderate rank correlation, and increasing monolingual identification data beyond about fifty instances does not improve causal transfer. CR-Fusion, which penalizes cross-lingual variance of normalized selectivity scores, selects units that, when deactivated or steered, produce emotion-selective accuracy changes that transfer to unseen languages and beat every single-corpus mask in most settings; the exception is steering on the model with the lowest language-invariant share. The paper interprets fusion as estimating a specific identifiable quantity, the invariant emotion effect in a language-by-emotion decomposition of activation probabilities, rather than as consensus filtering, and reports asymmetric leave-one-out contributions, with low-resource languages supplying non-redundant evidence and benefiting most from the resulting transfer.","pith_inferences":["If the corpus confound is resolved, a direct prediction follows: fusing two corpora of the same language with different elicitation styles should reproduce the CR-Fusion transfer gains if the mechanism is language-level, or fail if the gains actually come from recording style.","The variance decomposition suggests a testable extension: applying CR-Fusion to dimensional affect labels such as arousal and valence rather than discrete categories should yield MLENs whose selectivity orders transfer even better, since dimensional axes may align more closely with acoustic universals.","The invariant-share estimates of 0.30–0.59 imply an upper bound on what any pooling procedure can transfer, so the discarded language-specific component, 41–70 percent of variance, is the natural target for lightweight per-language adapters.","One could adversarially test the causal claim by steering MLENs in the opposite direction, suppressing anger while amplifying happiness, and measuring whether cross-lingual confusion patterns shift as predicted; the paper's Emotion Selectivity Score only measures matching interventions."],"forward_implications":["If MLENs exist as described, emotion recognition in LALMs is partly driven by language-agnostic units, so training-free activation steering on fused neuron sets should improve speech emotion recognition on languages never seen during identification.","Because monolingual identification saturates around fifty instances, adding more within-language data will not isolate transferable emotion neurons; cross-lingual pooling is the productive direction.","Low-resource languages such as Amharic, Bengali, and Urdu contribute non-redundant identification evidence, so excluding them from neuron identification does not merely reduce coverage but changes which neurons are selected.","The invariant emotion component accounts for roughly 30–59 percent of emotion-conditioned activation variance across models, leaving a substantial language-specific residue that could support target-aware adaptation.","Per-emotion causal potency, strongest for anger, happiness, and sadness, is a behavioral effect bounded by baseline accuracy and class support, not a sign that those emotions are represented more invariantly; neutral and fear have higher invariant shares.","For emotion-level heterogeneity, per-emotion invariant shares are highest for neutral and fear even though their causal effects are weaker, so representational universality and causal controllability dissociate."],"supporting_citations":[{"why":"Supplies the ConAct margin-based selector used throughout to score each neuron's emotion selectivity per language.","marker":"Zhao et al. 2026a"},{"why":"Establishes the identification-and-control template of deactivating important neurons, which the paper adapts to SwiGLU gate outputs.","marker":"Bau et al. 2019"},{"why":"Contributes the MAD neuron-importance scoring baseline that ConAct is compared against and improves upon.","marker":"Dalvi et al. 2019"},{"why":"One of the four evaluated LALMs, Audio-Flamingo-3, whose activations are logged and causally intervened upon.","marker":"Goel et al. 2025"},{"why":"Provides Qwen2.5-Omni-7B, the model with the lowest language-invariant share and the sole exception where a monolingual mask matches fusion under steering.","marker":"Xu et al. 2025b"},{"why":"Provides MiniCPM-o-4.5, the model with the highest invariant share and the largest CR-Fusion steering gain on held-out languages.","marker":"Yao et al. 2024"},{"why":"Provides Kimi-Audio, one of the four evaluated LALMs used for the cross-lingual neuron analysis.","marker":"KimiTeam 2025"},{"why":"Documents that emotion-correlated neurons identified in one language fail to generalize in multilingual encoders, motivating the need for cross-lingual identification.","marker":"Singh et al. 2026"},{"why":"Represents the contested prior result on cross-lingual activation steering that this paper's causal transfer findings directly address.","marker":"Maraia et al. 2026"}],"fun_headline_variants":["Language-invariant emotion neurons transfer across audio models","Cross-lingual consistency pinpoints emotion neurons that transfer","Shared emotion neurons in audio models transfer across languages","Language-agnostic emotion neurons control affect across languages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Each language in the study is represented by exactly one corpus, so language identity and recording style are confounded by construction, and the cross-lingual consistency credited to languages could actually come from corpus elicitation type; the paper itself flags this.","fun_headline_variants_meta":{"raw":{"variants":["Language-invariant emotion neurons transfer across audio models","Cross-lingual consistency pinpoints emotion neurons that transfer","Shared emotion neurons in audio models transfer across languages","Language-agnostic emotion neurons control affect across languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000668,"raw_usage":{"total_tokens":3072,"prompt_tokens":999,"completion_tokens":2073,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":2011}},"tokens_in":615,"tokens_out":2073,"duration_ms":14818,"temperature":1.0,"reasoning_tokens":2011,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:24:23.784745+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a single language represented by two corpora of different elicitation styles, such as acted versus naturalistic English, and rerun CR-Fusion with one corpus as identification and the other as held-out; if the fusion gain disappears or transfers along elicitation type rather than language, the multilingual-neuron interpretation collapses.","supporting_citations":[],"review_version":1}