Pith. sign in

REVIEW 3 major objections 6 minor 13 references

Multilingual Emotion Neurons in Large Audio-Language Models

T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Large audio-language models carry a cross-lingually shared emotion code that pooling languages can isolate, and that single-language neuron sets cannot.

desk verdict A real first cut at cross-lingual emotion neurons in LALMs, with a load-bearing confound and a tuned penalty; the causal pattern is consistent, but the 'multilingual' interpretation needs a within-language cross-corpus control before it can be trusted. read the letter →

arxiv 2608.08772 v1 pith:IM3UEOF2 submitted 2026-08-09 cs.CL eess.AS

classification cs.CLeess.AS
keywords multilingualemotionneuronslargeaudio-languagemodelscausalinterpretabilityactivationsteeringcross-lingualtransferconsistency-regularizedfusionspeechrecognitionneuron-levelanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that large audio-language models encode emotion partly through neurons shared across languages, and that identifying those neurons by pooling evidence from many languages yields causal control that no single-language procedure can match. The authors define Multilingual Emotion Neurons (MLENs) as units with stable emotional selectivity and aligned causal effects across languages, and propose Consistency-Regularized Fusion (CR-Fusion) to find them. Across four open models and twelve typologically diverse languages, neurons found separately in each language barely overlap, and adding more monolingual identification data saturates quickly. CR-Fusion masks, tested by deactivation and steering, outperform the best monolingual masks on held-out languages in zero-shot and low-resource settings, and leave-one-out analysis shows that low-resource languages contribute non-redundant evidence and are among those that benefit most. If right, this gives a mechanistic handle on cross-lingual emotion generalization and a training-free route to better affective control for languages with scarce data.

What carries the argument

The load-bearing object is the Multilingual Emotion Neuron (MLEN), defined as a neuron with stable emotional selectivity and aligned causal effects across languages, and the identification method is Consistency-Regularized Fusion (CR-Fusion). CR-Fusion takes per-language neuron scores, here the Contrastive Activation Margin between a neuron's best and second-best emotion, normalizes them per language, and ranks neurons by the cross-lingual mean minus λ times the cross-lingual standard deviation, so that λ=0 targets average transfer and larger λ targets quantile or worst-case stability. The paper also decomposes each neuron's activation probability into a baseline, a language effect, an emotion effect, a language-by-emotion interaction, and noise, arguing that fusion recovers the invariant emotion effect, a quantity no single-corpus procedure can isolate. Interventions either deactivate or amplify the selected neurons in the SwiGLU gate outputs to test causal necessity and sufficiency.

What would settle it

Take a single language represented by two corpora of different elicitation styles, such as acted versus naturalistic English, and rerun CR-Fusion with one corpus as identification and the other as held-out; if the fusion gain disappears or transfers along elicitation type rather than language, the multilingual-neuron interpretation collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that a language-invariant emotion component exists inside modern audio-language models and can be isolated by selecting neurons whose emotion selectivity is consistent across languages rather than maximal in any one language. Monolingual identification produces nearly disjoint neuron sets across languages (Jaccard similarity mostly below 0.10) despite weak-to-moderate rank correlation, and increasing monolingual identification data beyond about fifty instances does not improve causal transfer. CR-Fusion, which penalizes cross-lingual variance of normalized selectivity scores, selects units that, when deactivated or steered, produce emotion-selective accuracy changes that transfer to unseen languages and beat every single-corpus mask in most settings; the exception is steering on the model with the lowest language-invariant share. The paper interprets fusion as estimating a specific identifiable quantity, the invariant emotion effect in a language-by-emotion decomposition of activation probabilities, rather than as consensus filtering, and reports asymmetric leave-one-out contributions, with low-resource languages supplying non-redundant evidence and benefiting most from the resulting transfer.

Load-bearing premise

Each language in the study is represented by exactly one corpus, so language identity and recording style are confounded by construction, and the cross-lingual consistency credited to languages could actually come from corpus elicitation type; the paper itself flags this.

Editorial extensions

If this is right

  • If MLENs exist as described, emotion recognition in LALMs is partly driven by language-agnostic units, so training-free activation steering on fused neuron sets should improve speech emotion recognition on languages never seen during identification.
  • Because monolingual identification saturates around fifty instances, adding more within-language data will not isolate transferable emotion neurons; cross-lingual pooling is the productive direction.
  • Low-resource languages such as Amharic, Bengali, and Urdu contribute non-redundant identification evidence, so excluding them from neuron identification does not merely reduce coverage but changes which neurons are selected.
  • The invariant emotion component accounts for roughly 30–59 percent of emotion-conditioned activation variance across models, leaving a substantial language-specific residue that could support target-aware adaptation.
  • Per-emotion causal potency, strongest for anger, happiness, and sadness, is a behavioral effect bounded by baseline accuracy and class support, not a sign that those emotions are represented more invariantly; neutral and fear have higher invariant shares.
  • For emotion-level heterogeneity, per-emotion invariant shares are highest for neutral and fear even though their causal effects are weaker, so representational universality and causal controllability dissociate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the corpus confound is resolved, a direct prediction follows: fusing two corpora of the same language with different elicitation styles should reproduce the CR-Fusion transfer gains if the mechanism is language-level, or fail if the gains actually come from recording style.
  • The variance decomposition suggests a testable extension: applying CR-Fusion to dimensional affect labels such as arousal and valence rather than discrete categories should yield MLENs whose selectivity orders transfer even better, since dimensional axes may align more closely with acoustic universals.
  • The invariant-share estimates of 0.30–0.59 imply an upper bound on what any pooling procedure can transfer, so the discarded language-specific component, 41–70 percent of variance, is the natural target for lightweight per-language adapters.
  • One could adversarially test the causal claim by steering MLENs in the opposite direction, suppressing anger while amplifying happiness, and measuring whether cross-lingual confusion patterns shift as predicted; the paper's Emotion Selectivity Score only measures matching interventions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper presents a neuron-level interpretability study of multilingual emotion representation in four large audio-language models (Audio-Flamingo-3, Kimi-Audio, MiniCPM-o-4.5, Qwen2.5-Omni-7B) across 12 languages. The authors define Multilingual Emotion Neurons (MLENs) as units with stable emotional selectivity and aligned causal effects across languages, and propose Consistency-Regularized Fusion (CR-Fusion) to identify them from pooled multilingual activation statistics. Using a margin-based selector (ConAct) on eight identification languages, they report that monolingual emotion-sensitive neuron sets have minimal set overlap, that monolingual identification saturates beyond roughly 50 instances, and that CR-Fusion with a consistency penalty of lambda=0.3 produces stronger deactivation and steering effects than the best monolingual mask in most model-language conditions, with the exception of steering on Qwen2.5-Omni-7B. Leave-one-out ablations show asymmetric transfer patterns, and a variance decomposition attributes 30-59% of emotion-conditioned activation variance to a language-invariant component.

Significance. If the results hold, this would be the first causal, neuron-level account of how large audio-language models encode emotion across languages, with practical implications for training-free cross-lingual affective control. The paper is ambitious in scope (four models, twelve languages, two causal intervention types) and has notable strengths: an explicit estimand for the language-invariant emotion component, a noise-corrected variance decomposition, saturation and sensitivity analyses, and a candid Limitations section that admits the key confound. The central 'multilingual' interpretation, however, is currently undermined by two methodological issues: the one-corpus-per-language confound and the selection of the consistency penalty using the held-out evaluation languages. Both are addressable within a revision, either through additional experiments or through appropriately scoped claims.

major comments (3)
  1. [§4 (Table 3), §6.3, Limitations] The central claim that CR-Fusion identifies cross-lingually shared emotion neurons is confounded by the one-corpus-per-language design: each language in Table 3 is associated with a single dataset, so language identity and recording condition (elicitation style, channel, speaker population) are inseparable by construction. The Limitations section admits this and states that activation-structure similarity tracks corpus elicitation type at least as strongly as typological relatedness. Section 6.3 correctly frames the leave-one-out results as 'non-redundant cross-corpus contributions,' yet the abstract and conclusion re-assert language-level claims ('cross-lingual transfer,' 'shared affective representations that generalize across diverse spoken languages'). Because the fusion advantage and the asymmetric transfer patterns could be driven by shared recording style rather than by language, the defining construct of MLENs as cross-lingual units is not yet supported. The manuscript should either temper the language-level claims to corpus-level claims or add a control language represented by multiple corpora of different elicitation types.
  2. [§6.2, §3.4, Tables 1-2] The consistency penalty lambda is chosen as the operating point in §6.2 by inspecting Figure 4, which plots average ESS across all 12 evaluation languages, including the four held-out languages. Consequently, the reported zero-shot fusion advantage in Tables 1 and 2 is partly tuned on the test languages, and the confirmation of the §3.4 prediction that 'the optimal penalty is near zero' is circular because the same evaluation data determined the operating point. To support the zero-shot claim, the authors should select lambda using only identification languages (or an inner cross-validation) and report held-out performance at that value, or alternatively present results across a range of lambda and show that the conclusions are robust without selection on the held-out ESS.
  3. [Tables 1 and 2, Appendix A.4] The central comparison between CR-Fusion and the best monolingual mask relies on point estimates with standard deviations across languages but no significance tests or confidence intervals. For example, in Table 1 MiniCPM-o-4.5 deactivation, CR-Fusion (-8.03) differs from the best monolingual mask (Mandarin, -7.64) by only 0.39 pp, with cross-language standard deviations above 4.0; similarly small margins appear in several other rows. The argument in Appendix A.4 that deterministic decoding makes confidence intervals unnecessary conflates seed variance (zero) with sampling variability over evaluation utterances and corpora. The authors should report bootstrap confidence intervals over test utterances or paired significance tests across evaluation languages, or temper the 'outperforms' claims accordingly.
minor comments (6)
  1. [§5.1, Figure 1] The claim of 'minimal overlap' would be strengthened by a chance-level baseline: with r=0.5% selection, two random sets of the same size would have an expected Jaccard similarity of roughly r/(2-r) ≈ 0.0025, so the observed JSC values near 0.10 are an order of magnitude above chance; reporting this contrast would make the interpretation more precise.
  2. [§3.4, Eq. (2)] The notation for the additive decomposition omits the layer and neuron subscripts on the right-hand side (m, u, b, gamma, epsilon), which can confuse the reader; please make the indexing explicit.
  3. [§6.3, Figure 5] The leave-one-out heatmaps are averaged over four models; please also provide per-model results or state in the caption that pooling may hide model-specific patterns, since the reader cannot tell whether the asymmetric transfer is driven by one model.
  4. [Limitations] The 'post-hoc analysis' showing that activation-structure similarity tracks corpus elicitation type at least as strongly as typological relatedness is mentioned but no quantitative result is reported; adding a small table or appendix entry would make this claim verifiable.
  5. [§4, Appendix A.5] The definition of E_valid for the global ESS is only mentioned parenthetically; please specify precisely which emotions are excluded per model-language condition when an emotion has zero correctly predicted instances, since this affects the comparability of the aggregated metric.
  6. [Abstract] The phrase 'the first causal, neuron-level account' is a strong novelty claim; consider softening it or explicitly identifying the closest prior work to avoid overclaiming.

Circularity Check

2 steps flagged · score 6.0 of 10

Selection of λ on the evaluation languages makes the 'predicted optimal penalty' and zero-shot gains partly fitted rather than predicted.

  1. fitted input called prediction [Section 6.2 (Figure 4) and Section 3.4, 'The Predicted Optimal Penalty']
    "We therefore adopt λ=0.3 as the operating point, where the penalty is active but weak. This profile favors both predictions of §3.4: the optimal penalty is near zero under the expected-transfer objective..."

    Section 3.4 derives that the expected-transfer objective implies an optimal penalty near zero. Section 6.2 selects the operating point from the sensitivity curve of average ESS over the same 12 evaluation languages and then reports that the curve 'favors both predictions.' The confirmation is thus a restatement of the selection: the data used to confirm the predicted optimum is the same data used to pick λ, so the 'prediction' is not tested out-of-sample.

  2. fitted input called prediction [Section 4 (Held-out Set) and Section 6.2 (Figure 4)]
    "The Held-out Set (Lheld) contains languages reserved strictly for zero-shot evaluation: French (CaFE (Gournay et al., 2018)), German (EmoDB (Burkhardt et al., 2025)), Persian (ShEMO (Mohamad Nezami et al., 2019)) and Russian (RESD (Amentes et al., 2023)). ... Average ESS (±1 standard error of the mean across 12 languages, shown as the shaded band) under (a) deactivation and (b) steering for MiniCPM."

    The four languages 'reserved strictly for zero-shot evaluation' are included in the 12-language average used in Figure 4 to choose λ=0.3. The zero-shot results in Tables 1-2 for French, German, Persian, and Russian are therefore not independent of hyperparameter selection; the operating point was fitted to maximize average ESS on these very languages, so the reported held-out advantage of CR-Fusion is partly forced by construction rather than predicted.

full rationale

The paper's central derivation—CR-Fusion pooling evidence across languages versus monolingual identification—is not definitionally circular: it is tested against monolingual baselines on held-out languages, across four models, and the empirical claims (minimal JSC overlap, saturation of monolingual evidence) are independent observations. The self-citations (ConAct in Zhao et al. 2026a; monolingual neuron work in Zhao et al. 2026b,c) supply tools and background, not the load-bearing evidence for the cross-lingual fusion claim. However, two linked steps do reduce to fitting. Section 3.4 predicts the optimal consistency penalty is near zero; Section 6.2 confirms this on the same average-ESS curve used to adopt λ=0.3. Moreover, the 'strictly reserved' held-out set is among the 12 languages over which that average is computed, so the zero-shot advantage of CR-Fusion is partly a selected-operating-point artifact. Separately, the admitted one-corpus-per-language confound is a serious attribution risk for the language-level interpretation, but it is a validity threat, not a circularity, and does not by itself raise the score. Weighing the fitted-prediction circularity against the otherwise independent benchmarking, score 6.

Assumptions & free parameters 4 free parameters · 5 assumptions · 1 invented entities

The central claim rests on the CR-Fusion selection procedure with a hand-tuned penalty lambda, on an assumed additive decomposition of activation probabilities, and on the assumption that gate activations are the causally relevant substrate. The most serious structural issue is the single-corpus-per-language confound, which is acknowledged in the Limitations. No new physical entities are introduced; MLENs are an operational construct.

free parameters (4)
  • lambda (CR-Fusion consistency penalty) = 0.3
    Selected as the operating point from sensitivity analysis on the evaluation ESS (Section 6.2). The predicted 'optimal' lambda is near zero, and 0.3 is justified as 'active but weak', but it is tuned on the same evaluation data used to report the main fusion gains.
  • selection fraction r = 0.5%
    Fixed fraction of top-ranked neurons selected per emotion; chosen by hand for all models and languages, with no sensitivity analysis provided for r.
  • identification budget c = 50 instances per emotion
    Chosen based on the saturation analysis in Section 5.2 showing plateaus beyond about 50 instances; this is a hand-picked budget, though the plateau result provides some support.
  • steering gain alpha = 0.5
    Default amplification factor for steering interventions, used throughout without per-model tuning; it is a fixed intervention setting rather than a fitted parameter.
assumptions (5)
  • domain assumption Selector scores decompose as theta(e) + delta(l,e) + xi(l,e) with E[delta]=0 and Gaussian noise, and languages are exchangeable for the expected-transfer objective.
    Used in Section 3.4 to derive the claim that the optimal consistency penalty is near zero; exchangeability is an idealization that is not tested directly.
  • domain assumption Activation probabilities satisfy the additive decomposition P = m + u(l) + b(e) + gamma(l,e) + epsilon (Equation 2), with no higher-order interactions.
    Basis of the variance decomposition in Section 3.4 and Appendix B; the least-squares fit assumes this model, and the invariant shares are interpreted under it.
  • domain assumption SwiGLU gate activations in decoder MLP modules are a causally relevant substrate for emotion recognition, and masking or amplifying them does not confound the ESS metric.
    Interventions in Section 3.5 operate on post-activation gate outputs; the causal interpretation assumes these units are the right level of abstraction.
  • domain assumption Restricting activation logging to correctly predicted utterances yields cleaner emotion-conditioned statistics and does not bias neuron selection.
    Section 3.1 states this choice; selection bias is possible because correct predictions are correlated with emotion class support and baseline accuracy.
  • domain assumption The five discrete emotion categories are commensurable across the 12 languages and corpora.
    Pooling evidence across languages requires that the same emotion labels refer to aligned constructs; the paper notes cultural variation in emotion semantics but does not model it.
invented entities (1)
  • Multilingual Emotion Neurons (MLENs)
    purpose: A set of neurons selected by CR-Fusion that exhibit stable emotional selectivity and aligned causal effects across languages; used as the target of causal interventions.
    Defined operationally by the paper's own identification procedure; the only validation is in-paper causal experiments on held-out languages, which share the same evaluation pipeline used to set lambda, so there is no fully external benchmark.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multilingual Emotion Neurons in Large Audio-Language Models." pith.science (2026). https://pith.science/paper/IM3UEOF2

@misc{pith2026260808772,
  author       = {Pith},
  title        = {Pith review of: Multilingual Emotion Neurons in Large Audio-Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IM3UEOF2}},
  note         = {Machine review of arXiv:2608.08772}
}
read the original abstract

Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion through language-specific correlations or language-agnostic representations. We present the first neuron-level interpretability study of this question. We define Multilingual Emotion Neurons (MLENs) as functional units exhibiting stable emotional selectivity and aligned causal effects across languages, and introduce Consistency-Regularized Fusion (CR-Fusion) to identify them. Across four modern LALMs and 12 typologically diverse languages, emotion-sensitive neurons identified independently per language show minimal overlap, and additional monolingual identification data saturates quickly without isolating more transferable units, motivating identification from pooled cross-lingual evidence. Causal interventions demonstrate that MLENs identified by CR-Fusion provide more precise and transferable affective control than monolingual neuron sets in both zero-shot and low-resource settings. Leave-one-out ablations further reveal asymmetric transfer: individual identification languages, including low-resource ones, contribute non-redundant evidence, while several low-resource languages benefit most from the resulting cross-lingual transfer. Together, our findings provide the first causal, neuron-level account of how LALMs encode emotion across languages, and establish multilingual neuron identification as an effective mechanism for understanding cross-lingual affective behavior.

Figures

Figures reproduced from arXiv: 2608.08772 by the authors.

Figure 1
Figure 1. Cross-lingual agreement of monolingually [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Saturation of monolingual evidence: UAR (%) [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. ESS distributions under (a) deactivation and [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Leave-one-out ∆ESS heatmaps averaged over four models: effect of excluding each identification language (columns) on each evaluation language (rows). are removed. The fused mask thus aggregates par￾tially complementary evidence, with low-resource corpora supplying non-…
Figure 6
Figure 6. Figure 6: Per-emotion ESS magnitudes under (a) deac [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 8
Figure 8. Figure 8: Saturation of monolingual evidence: UAR (%) [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Sensitivity to consistency penalty λ under deactivation (left) and steering (α=0.5, right), comple￾menting [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: ESS distributions across monolingual and [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 5 canonical work pages

  1. [3]

    Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang, Yuxiang Lin, Zheng Lian, Xiaojiang Peng, and Alexander G Hauptmann

    Emozionalmente: A crowdsourced corpus of simulated emotional speech in italian.IEEE Trans- actions on Audio, Speech and Language Processing, 33:1142–1155. Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang, Yuxiang Lin, Zheng Lian, Xiaojiang Peng, and Alexander G Hauptmann. 2024. Emotion-LLaMA: Multimodal emotion recognition and reasoning with instruction t...

  2. [7]

    Preprint, arXiv:2509.24322

    Multimodal large language models meet multi- modal emotion recognition and reasoning: A survey. Preprint, arXiv:2509.24322. Anant Singh and Akshat Gupta. 2023. Decoding emo- tions: A comprehensive multilingual study of speech models for speech emotion recognition.Preprint, arXiv:2308.08713. Pranaydeep Singh, Orphee De Clercq, and Els Lefever

  3. [9]

    Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A Saurous

    Do llms "feel"? emotion circuits discovery and control.Preprint, arXiv:2510.11328. Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A Saurous. 2018. Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis. InInternational confer- ence on machine lea...

  4. [11]

    Audiolens: A closer look at auditory attribute perception of large audio-language models.Preprint, arXiv:2506.05140. Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, and 4 others. 2024...

  5. [12]

    Xiutian Zhao, Rochelle Choenni, Rohit Saxena, and Ivan Titov

    Cross corpus multi-lingual speech emotion recognition using ensemble learning.Complex & Intelligent Systems, 7(4):1845–1854. Xiutian Zhao, Rochelle Choenni, Rohit Saxena, and Ivan Titov. 2026a. Finding culture-sensitive neurons in vision-language models. InProceedings of the 19th Conference of the European Chapter of the As- sociation for Computational Li...

  6. [13]

    Share” is the invariant fraction σ2 b /(σ2 b +σ 2 γ), raw and noise-corrected. “JSC

    with a 20-token generation limit and apply lightweight post-processing to extract the option letter from model outputs. Deterministic decod- ing ensures that model outputs are as reproducible as possible given fixed inputs and model weights. Consequently, repeated runs on identical data pro- duce identical results, and variance in reported met- rics refle...

  7. [430]

    Kristen A Lindquist, Jennifer K MacCormack, and Holly Shablack

    IEEE. Kristen A Lindquist, Jennifer K MacCormack, and Holly Shablack. 2015. The role of language in emo- tion: Predictions from psychological constructionism. Frontiers in psychology, 6:121301. Rui Liu, Berrak Sisman, Guanglai Gao, and Haizhou Li

  8. [627]

    Tung-Yu Wu, Yu-Xiang Lin, and Tsui-Wei Weng

    IEEE. Tung-Yu Wu, Yu-Xiang Lin, and Tsui-Wei Weng. 2024. And: audio network dissection for interpreting deep acoustic models. InProceedings of the 41st Interna- tional Conference on Machine Learning, ICML’24. JMLR.org. Tianxin Xie, Shan Yang, Chenxing Li, Dong Yu, and Li Liu. 2025. Emosteer-tts: Fine-grained and training-free emotion-controllable text-to-...

Show all 13 references
  1. [2018]

    emotion neurons

    A canadian french emotional speech dataset. InProceedings of the 9th ACM Multimedia Systems Conference, MMSys ’18, page 399–402, New York, NY , USA. Association for Computing Machinery. Wes Gurnee, Theo Horsley, Zifan Carl Guo, Tara Rezaei Kheirkhah, Qinyi Sun, Will Hathaway, ...

  2. [2021]

    Jaime Lorenzo-Trueba, Gustav Eje Henter, Shinji Takaki, Junichi Yamagishi, Yosuke Morino, and Yuta Ochiai

    Expressive tts training with frame and style re- construction loss.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:1806–1818. Jaime Lorenzo-Trueba, Gustav Eje Henter, Shinji Takaki, Junichi Yamagishi, Yosuke Morino, and Yuta Ochiai. 2018. Investigating diff...

  3. [2023]

    Mhamed amine Soumiaa

    Resd (revision 75ed61a). Mhamed amine Soumiaa. 2024. Moroccan dialect emo- tion recognition dataset. Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019. Iden- tifying and controlling important neurons in neural machine translation. ...

  4. [2025]

    Badriyya B

    Improving audio explanations using audio language models.IEEE Signal Processing Letters, 32:741–745. Badriyya B. Al-onazi, Muhammad Asif Nauman, Rashid Jahangir, Muhmmad Mohsin Malik, Eman H. Alkhammash, and Ahmed M. Elshewey. 2022. Transformer-based multilingual speech emotio...

  5. [2026]

    In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Lin- guistics (Volume 2: Short Papers), pages 154–159, Rabat, Morocco

    Lost in activations: A neuron-level analysis of encoders for cross-lingual emotion detection. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Lin- guistics (Volume 2: Short Papers), pages 154–159, Rabat, Morocco. Association f...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.