REVIEW 5 major objections 5 minor 46 references
"How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases
T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A user study finds no statistically significant age or language differences in how a pre-trained speech-emotion model reads deliberately acted emotions, with a persistent weakness on high-arousal emotions like happy and angry.
desk verdict Inventive interactive SER evaluation with rich qualitative data, but the headline null result is not statistically grounded because the t-tests treat repeated utterances as independent. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a target-hitting task in valence-arousal space: a participant vocally acts an emotion to hit a fixed coordinate, while a backend runs a pre-trained Wav2Vec 2.0-based SER model on the utterance and logs its predicted valence-arousal coordinate. The paper defines expressive accuracy as the Euclidean distance between prediction and target, so the model's coordinate system is the measuring instrument for every age and language comparison. The task design, inspired by prior targeting and elicitation paradigms, converts emotional expression into a goal-directed, quantifiable act and produces a real-time log of human intent versus machine interpretation.
What would settle it
Re-score the same 384 logged utterances with human listeners who rate valence and arousal, then test for age and language effects; if human raters find differences the model missed, the claimed robustness is an artifact of model insensitivity. A simpler check is to rerun the t-tests after removing all trials where the system recorded silence or hesitation, since the model defaulted those to sad.
Extended reading notes
Core claim
Contrary to the authors' assumptions, neither age nor language produced a statistically significant difference in how accurately the model read deliberately expressed emotions. Expressive accuracy was measured as the Euclidean distance between the model's predicted valence-arousal coordinates and the fixed target coordinates for each of four emotions; across 384 utterances from 24 native Danish speakers, independent-samples t-tests found no group differences in either the English or the Danish condition. The data do show a pattern the authors emphasize: high-arousal emotions, happy especially, had the largest distances for both groups, and the model often read laughter during angry attempts as happy. The paper frames the contribution as a shift from system-centered accuracy to human-centered expressive accuracy, exposing the real-time gap between what a speaker intends and what the model interprets.
Load-bearing premise
The load-bearing premise is that the pre-trained model's valence-arousal readings are a valid and sensitive measure of how accurately a person expressed an emotion; if the model is blind to age or language differences, or if it collapses high-arousal and silent input into biased categories like sad, the null t-tests describe the model, not the speakers.
Editorial extensions
If this is right
- If the null result holds, an English-trained SER model can interpret deliberately acted basic emotions from teenage and older-adult native Danish speakers as accurately as from English utterances, at least for the four studied emotions.
- The systematic difficulty with happy and angry points to a concrete target for model improvement: high-arousal emotions need recalibration or additional training data rather than age- or language-specific fixes.
- Logging the real-time gap between intended and predicted valence-arousal coordinates gives designers a direct way to audit when an SER system misaligns with a user's emotional intent, which can be used to evaluate inclusiveness before deployment.
- Users' tendency to describe emotions as blended states (calm/happy, angry/sad) challenges discrete-label SER benchmarks and supports the paper's argument for models that represent emotional complexity rather than fixed categories.
Reading between the lines
- Because expressive accuracy is measured entirely inside the model's own coordinate system, a plausible reading is that the null result says more about the model's insensitivity than about human equality; a direct test would be to have human listeners rate the same utterances on valence and arousal and compare the two instruments.
- The same targeting-game protocol could be run with a different pre-trained SER model; if that model shows age or language effects the first model misses, the robustness claim is model-specific rather than a general property of speech emotion AI.
- The paper's suggestion to test typologically distant languages such as Greenlandic is a natural extension: positive results would support universal acoustic encoding of emotion, while negative results would locate the boundary of cross-linguistic transfer.
- The silent-input-defaults-to-sad behavior is a concrete artifact that future audits could exploit as a targeted probe for systematic bias in SER systems, since hesitation and silence are common in real interactions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a user study (N=24: 12 teenagers and 12 adults aged 55+) in which participants deliberately acted four emotions (happy, sad, angry, calm) in both Danish and English while a pre-trained SER model (Wagner et al.) predicted valence-arousal coordinates in real time. The authors quantify expressive accuracy as the Euclidean VA distance between the model's predictions and the target coordinates, and they report independent-samples t-tests comparing age groups within each language condition. The central claim is that no statistically significant age or language differences were found (all p > .05), suggesting that, under structured elicitation, both age groups are comparably effective at deliberately conveying emotions as interpreted by this model. Qualitative interview findings about emotion expression, trust in SER, and the perceived best emotions to convey are also reported, together with a discussion of limitations including the model's silence-to-Sad default and difficulties with high-arousal emotions.
Significance. The paper's main strength is its human-centered experimental paradigm: a real-time VA-targeting game with full logging of model predictions, a thoughtful qualitative analysis, and an unusually explicit limitations section. If the reported null results survive correct statistical analysis, they would be a useful empirical counterpoint to assumptions about age and language bias in SER and would support the paper's call for more inclusive, human-centered speech emotion AI. The qualitative findings, especially participants' descriptions of blended emotions and the perceived divergence between self-report and model output, are valuable contributions in their own right. However, as presented, the statistical analysis does not currently support the central null claim, and the absence of any reported test for the language comparison leaves part of the headline result unsupported.
major comments (5)
- [§4.1, §4.3] The statistical unit is inconsistent with the reported degrees of freedom. Section 4.1 states that the analysis partitions the data by language, yielding 192 utterances per language condition (96 per age group), and implies that independent-samples t-tests were run on these utterances. Yet the only t statistic with degrees of freedom, t(21.74) = -1.49 for sad in Danish (Section 4.3), has df ≈ 22, which corresponds to a test on 12 observations per group (i.e., per-participant means) rather than the 24 raw utterances per age group per emotion-language cell. If the two within-participant repetitions were treated as independent observations, the effective sample size is inflated, the observations are non-independent, and the p-values are anti-conservative. If the repetitions were averaged to participant-level means, that aggregation is nowhere stated and the description in Section 4.1 is misleading. Please clarify which unit was used, and re-run the analyses with a mixed-effects model or with explicitly aggregated per-participant means.
- [§4.2, §4.3, §5] The Discussion's central claim that 'no statistically significant differences were found between the age groups or across language conditions (all p > .05)' is not supported by any reported test of the language condition. Sections 4.2 and 4.3 only compare age groups within English and within Danish; no English-versus-Danish comparison (e.g., a paired t-test on participant-level means or a language factor in a mixed model) is reported. Either add the missing language comparison or revise the Conclusion and Discussion to claim only that no age differences were found within each language.
- [§4.1, §6] The VA-distance metric is computed entirely from the Wagner et al. model's predicted coordinates, and the study provides no human-listener baseline, chance level, or any other validation that these distances are sensitive to meaningful differences in expressive accuracy. Without such a baseline, the null t-tests cannot distinguish 'no difference in expressive ability' from 'this model's predictions are insensitive to age or language differences.' This concern is reinforced by the acknowledged silence-to-Sad default and the model's difficulties with high-arousal emotions (Section 6). Please either validate the metric with human ratings or a reference corpus, or rephrase the conclusions to state that no differences were found in the model's VA-distance outputs, rather than in expressive accuracy per se.
- [§3.1, §3.3] The relationship between the pilot study and the final study participants is not specified. Section 3.1 describes a pilot with 12 participants from the 55+ age group, and Section 3.3 reports that the final study had 12 Adults 55+ and 12 Teenagers. If the same 12 adults completed both the pilot and the final study, their prior exposure to the interface and task would confound the age comparison (adults would have had extra practice; teenagers would not). If the pilot participants were a separate group, this should be stated explicitly. Please clarify this methodological detail.
- [§6] The silence-to-Sad default is acknowledged as a technical limitation, but the quantitative analysis in Section 4 does not report how many trials involved silence or hesitation, nor any sensitivity analysis excluding these trials. If silent trials were included in the VA-distance means, they could artifactually lower the sad distances and raise the happy/angry distances, directly affecting the reported means and the null results. Please report the number of affected trials and re-run the analysis without them, or with silence as a covariate.
minor comments (5)
- [§4.3] The heading of Section 4.3 duplicates Section 4.2 ('English Condition: Age Group Differences in Expressive Accuracy') even though the section content concerns Danish utterances; the heading should be corrected.
- [Figures 9 and 11] The captions for Figures 9 and 11 appear to be duplicated; each figure caption says the figure displays mean VA distances 'comparing Teenagers and Adults in English (left panel) and Danish (right panel)', but each figure should describe only its own language condition.
- [§2.1] There is a missing citation placeholder in the sentence introducing the Circumplex Model of Affect: 'Russell and Barrett offers ... [? ]'.
- [§4.3, §7] There are several typos: 'Teenagrs' in Section 4.3, 'in-debt' in the Conclusion (should be 'in-depth'), 'targetted' in the Conclusion, and 'dependable variable' (should be 'dependent variable').
- [§3.2] The text 'This resulted in 16 total utterances per participant (four emotions× four repetitions)' is imprecise; the structure is four emotions × two languages × two repetitions, and stating it this way would avoid confusion.
Circularity Check
No significant circularity; the study is an empirical measurement of SER model outputs, with only minor non-load-bearing self-citations for the prototype architecture.
full rationale
This is an empirical HCI user study rather than a derivation chain. The dependent variable, VA distance, is explicitly defined as the Euclidean distance between the SER model's predicted valence-arousal coordinates and fixed target coordinates, so every age and language comparison is a comparison of model outputs. That is the paper's stated object of study ('how a pre-trained SER model from Wagner et al. [39] interprets those vocal expressions'), not a hidden equivalence: the authors fit no parameter to the outcome and then relabel it as a prediction. The only self-citations are [6,7] for the UI/backend architecture ('To this end, we build on the architecture of previous work [7]'), and those are implementation reuse, not load-bearing evidence for the central null result. The Wagner et al. model [39] is an external, fixed instrument; whether it is valid or unbiased is a measurement-correctness concern, not circularity. Section 6's admission that silent input defaults to 'Sad' is a stated instrument limitation and does not reduce any result to its own inputs. The non-independence of within-participant replicates in the t-tests is a genuine statistical-validity problem, but it is not a circularity: the reported p-values may be invalid without the claim being equivalent to the inputs by construction. Verdict: no significant circularity; only minor, non-load-bearing self-citation, so score 1.
Assumptions & free parameters
free parameters (1)
- Target VA coordinates for the four emotion targets =
Not reported in the paper
assumptions (2)
- domain assumption Russell's Circumplex Model of Affect is a valid 2D representation of the four target emotions, so Euclidean distance to a fixed target coordinate is a meaningful measure of expressive accuracy.
- domain assumption The pre-trained Wagner et al. SER model gives unbiased and sufficiently sensitive VA predictions for Danish and English, for teenagers and adults over 55, so that null t-test results can be interpreted as age and language robustness.
Cite this review
Pith. "Pith review of "How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases." pith.science (2026). https://pith.science/paper/ABKP4WRW
@misc{pith2026250712580,
author = {Pith},
title = {Pith review of: "How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases},
year = {2026},
howpublished = {\url{https://pith.science/paper/ABKP4WRW}},
note = {Machine review of arXiv:2507.12580}
}
read the original abstract
This study explores how age and language shape the deliberate vocal expression of emotion, addressing underexplored user groups, Teenagers (N = 12) and Adults 55+ (N = 12), within speech emotion recognition (SER). While most SER systems are trained on spontaneous, monolingual English data, our research evaluates how such models interpret intentionally performed emotional speech across age groups and languages (Danish and English). To support this, we developed a novel experimental paradigm combining a custom user interface with a backend for real-time SER prediction and data logging. Participants were prompted to hit visual targets in valence-arousal space by deliberately expressing four emotion targets. While limitations include some reliance on self-managed voice recordings and inconsistent task execution, the results suggest contrary to expectations, no significant differences between language or age groups, and a degree of cross-linguistic and age robustness in model interpretation. Though some limitations in high-arousal emotion recognition were evident. Our qualitative findings highlight the need to move beyond system-centered accuracy metrics and embrace more inclusive, human-centered SER models. By framing emotional expression as a goal-directed act and logging the real-time gap between human intent and machine interpretation, we expose the risks of affective misalignment.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Communication Trends in the Post-literacy Era: Multilingualism, Multimodality, Multiculturalism
Sofia Abramova, Natalia Antonova, and Anna Gurarii. 2020. Speech behavior and multimodality in online communication among teenagers. In IV International Scientific Conference “Communication Trends in the Post-literacy Era: Multilingualism, Multimodality, Multiculturalism”.—Ekaterinburg,
work page 2020
-
[2]
Adebanji Adeleye, Samaneh Madanian, and Olayinka Adeleye. 2024. Emotion Variation Detection in Discrete English Speech: A Wavelet Transform Use Case in Mental Health Monitoring. In Proceedings of the 2024 Australasian Computer Science Week (Sydney, NSW, Australia) (ACSW ’24). Association for Computing Machinery, New York, NY, USA, 115–119. https://doi.org...
-
[3]
Shahin Amiriparian, Artem Sokolov, Ilhan Aslan, Lukas Christ, Maurice Gerczuk, Tobias Hübner, Dmitry Lamanov, Manuel Milling, Sandra Ottl, Ilya Poduremennykh, Evgeniy Shuranov, and Björn W. Schuller. 2021. On the Impact of Word Error Rate on Acoustic-Linguistic Speech Emotion Recognition: An Update for the Deep Learning Era. arXiv:2104.10121 [cs.SD] https...
work page Pith review arXiv 2021
-
[4]
How to Explore Biases in Speech Emotion AI with Users?
Pengcheng An, Chaoyu Zhang, Haichen Gao, Ziqi Zhou, Yage Xiao, and Jian Zhao. 2025. AniBalloons: Animated chat balloons as affective augmentation for social messaging and chatbot interaction. International Journal of Human-Computer Studies 194 (2025), 103365. https://doi.org/10. 1016/j.ijhcs.2024.103365 "How to Explore Biases in Speech Emotion AI with Users?" 19
arXiv 2025
-
[5]
Pengcheng An, Jiawen Stefanie Zhu, Zibo Zhang, Yifei Yin, Qingyuan Ma, Che Yan, Linghao Du, and Jian Zhao. 2024. EmoWear: Exploring Emotional Teasers for Voice Message Interaction on Smartwatches. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY,...
arXiv 2024
-
[6]
Ilhan Aslan, Carla F. Griggio, Henning Pohl, Timothy Merritt, and Niels van Berkel. 2025. Speejis: Enhancing User Experience of Mobile Voice Messaging with Automatic Visual Speech Emotion Cues. arXiv:2502.05296 [cs.HC] https://arxiv.org/abs/2502.05296
work page Pith review arXiv 2025
-
[7]
Ilhan Aslan, Timothy Merritt, Stine S. Johansen, and Niels van Berkel. 2025. Speech Command + Speech Emotion: Exploring Emotional Speech Commands as a Compound and Playful Modality. arXiv:2504.08440 [cs.HC] https://arxiv.org/abs/2504.08440
work page Pith review arXiv 2025
-
[8]
Ilhan Aslan, Tabea Schmidt, Jens Woehrle, Lukas Vogel, and Elisabeth André. 2018. Pen + Mid-Air Gestures: Eliciting Contextual Gestures. In Proceedings of the 20th ACM International Conference on Multimodal Interaction (Boulder, CO, USA)(ICMI ’18). Association for Computing Machinery, New York, NY, USA, 135–144. https://doi.org/10.1145/3242969.3242979
arXiv 2018
Show all 46 references
-
[9]
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. arXiv:2006.11477 [cs.CL] https://arxiv.org/abs/2006.11477
2020 arXiv
-
[10]
Patricia E. G. Bestelmeyer, Sonja A. Kotz, and Pascal Belin. 2017. Effects of emotional valence and arousal on the voice perception network. Social Cognitive and Affective Neuroscience 12, 8 (04 2017), 1351–1358. https://doi.org/10.1093/scan/nsx059 arXiv:https://academic.oup.c...
2017 doi
-
[11]
Samuel Cahyawijaya, Holy Lovenia, Willy Chung, Rita Frieske, Zihan Liu, and Pascale Fung. 2023. Cross-Lingual Cross-Age Group Adaptation for Low-Resource Elderly Speech Emotion Recognition. arXiv:2306.14517 [cs.CL] https://arxiv.org/abs/2306.14517
2023 arXiv
-
[12]
Fabio Catania. 2023. Speech emotion recognition in italian using wav2vec 2.Authorea Preprints (2023). https://doi.org/10.36227/techrxiv.22821992.v1
2023 doi
-
[13]
Charlene H Chu, Rune Nyrup, Kathleen Leslie, Jiamin Shi, Andria Bianchi, Alexandra Lyn, Molly McNicholl, Shehroz Khan, Samira Rahimi, and Amanda Grenier. 2022. Digital Ageism: Challenges and Opportunities in Artificial Intelligence for Older Adults. The Gerontologist 62, 7 (01...
2022 doi
-
[14]
Douglas-Cowie, Suzie Savvidou, E
Roddy Cowie, E. Douglas-Cowie, Suzie Savvidou, E. McMahon, M. Sawey, and M. Schröder. 2000. ’FEELTRACE’: An instrument for recording perceived emotion in real time. Proceedings of the ISCA Workshop on Speech and Emotion
2000
-
[15]
Zachary Dair, Ryan Donovan, and Ruairi O’Reilly. 2022. Linguistic and Gender Variation in Speech Emotion Recognition using Spectral Features. arXiv:2112.09596 [cs.SD] https://arxiv.org/abs/2112.09596
2022 arXiv
-
[16]
Jarod Duret, Mickael Rouvier, and Yannick Estève. 2024. MSP-Podcast SER Challenge 2024: L’antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition. arXiv:2407.05746 [cs.AI] https://arxiv.org/abs/2407.05746
2024 arXiv
-
[17]
Paul Ekman. 1992. Are there basic emotions? (1992). https://doi.org/10.1037/0033-295X.99.3.550
1992 doi
-
[18]
Paul Ekman. 1992. Facial Expressions of Emotion: New Findings, New Questions. Psychological Science 3, 1 (1992), 34–38. http://www.jstor.org/ stable/40062750
1992
-
[19]
Google Translate. 2024. Google Translate. https://support.google.com/translate/answer/15139004?hl=en&utm_source=chatgpt.com. Accessed: 2025-06-27
2024
-
[20]
Cristina Gorrostieta, Reza Lotfian, Kye Taylor, Richard Brutti, and John Kane. 2019. Gender De-Biasing in Speech Emotion Recognition.. In Interspeech. 2823–2827. https://www.isca-archive.org/interspeech_2019/gorrostieta19_interspeech.pdf
2019
-
[21]
Igor Grossmann, Alex C Huynh, and Phoebe C Ellsworth. 2016. Emotional complexity: Clarifying definitions and cultural correlates. Journal of personality and social psychology 111, 6 (2016), 895. https://doi.org/10.1037/pspp0000084
2016 doi
-
[22]
Lucía Gómez-Zaragozá, Rocío del Amor, María José Castro-Bleda, Valery Naranjo, Mariano Alcañiz Raya, and Javier Marín-Morales. 2024. EMOVOME: A Dataset for Emotion Recognition in Spontaneous Real-Life Speech. arXiv:2403.02167 [eess.AS] https://arxiv.org/abs/2403.02167
2024 arXiv
-
[23]
Lasse Hansen, Yan-Ping Zhang, Detlef Wolf, Konstantinos Sechidis, Nicolai Ladegaard, and Riccardo Fusaroli. 2022. A generalizable speech emotion recognition model reveals depression and remission. Acta Psychiatrica Scandinavica 145, 2 (2022), 186–199
2022
-
[24]
Alexandra Israelsson, Anja Seiger, and Petri Laukka. 2023. Blended emotions can be accurately recognized from dynamic facial and vocal expressions. Journal of Nonverbal Behavior 47, 3 (2023), 267–284. https://doi.org/10.1007/s10919-023-00426-9
2023 doi
-
[25]
Jesin James, Felix Marattukalam, Kaustubh Shukla, Enuri Kolugala, Sunny Choi, and Owen Eng. [n. d.]. Emotiongui: A Web-Based Tool for Visualisation and Annotation of Emotions for Speech and Video Signals. A vailable at SSRN 5063641([n. d.]). http://dx.doi.org/10.2139/ssrn.5063641
-
[26]
Ruhul Amin Khalil, Edward Jones, Mohammad Inayatullah Babar, Tariqullah Jan, Mohammad Haseeb Zafar, and Thamer Alhussain. 2019. Speech Emotion Recognition Using Deep Learning Techniques: A Review.IEEE Access 7 (2019), 117327–117345. https://doi.org/10.1109/ACCESS.2019.2936124
2019
-
[27]
Michael W Kraus. 2017. Voice-only communication enhances empathic accuracy. American Psychologist 72, 7 (2017), 644. https://doi.org/10.1037/ amp0000147
2017
-
[28]
Gaku Kutsuzawa, Hiroyuki Umemura, Koichiro Eto, and Yoshiyuki Kobayashi. 2022. Classification of 74 facial emoji’s emotional states on the valence-arousal axes. Scientific Reports 12, 1 (2022), 398. https://doi.org/10.1038/s41598-021-04357-7
2022 doi
-
[29]
Yi-Cheng Lin, Haibin Wu, Huang-Cheng Chou, Chi-Chun Lee, and Hung yi Lee. 2024. Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition. In Interspeech 2024. 4633–4637. https://doi.org/10.21437/Interspeech.2024-1073
2024 doi
-
[30]
Sandra Luo. 2024. Navigating the Diverse Challenges of Speech Emotion Recognition: A Deep Learning Perspective. In Proceedings of the 27th International Academic Mindtrek Conference (Tampere, Finland) (Mindtrek ’24). Association for Computing Machinery, New York, NY, USA, 133–...
2024
-
[31]
Yong Ma, Yuchong Zhang, Miroslav Bachinski, and Morten Fjeld. 2024. Emotion-Aware Voice Assistants: Design, Implementation, and Preliminary Insights. In Proceedings of the Eleventh International Symposium of Chinese CHI (Denpasar, Bali, Indonesia) (CHCHI ’23). Association for ...
2024
-
[32]
I Scott MacKenzie. 1992. Fitts’ law as a research and design tool in human-computer interaction. Human-computer interaction 7, 1 (1992), 91–139. https://doi.org/10.1207/s15327051hci0701_3
1992 doi
-
[33]
Brent Daniel Mittelstadt, Patrick Allo, Mariarosaria Taddeo, Sandra Wachter, and Luciano Floridi. 2016. The ethics of algorithms: Mapping the debate. Big Data & Society 3, 2 (2016), 2053951716679679. https://doi.org/10.1177/2053951716679679 arXiv:https://doi.org/10.1177/205395...
2016 doi
-
[34]
Michael Neumann and N goc Thang Vu. 2018. CRoss-lingual and Multilingual Speech Emotion Recognition on English and French. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 5769–5773. https://doi.org/10.1109/ICASSP.2018.8462162
2018
-
[35]
Bernstein, Robin N
Joon Sung Park, Michael S. Bernstein, Robin N. Brewer, Ece Kamar, and Meredith Ringel Morris. 2021. Understanding the Representation and Representativeness of Age in AI Data Sets. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (A...
2021
-
[36]
Marc D Pell, Laura Monetta, Silke Paulmann, and Sonja A Kotz. 2009. Recognizing emotions in a foreign language. Journal of Nonverbal Behavior 33, 2 (2009), 107–120. https://doi.org/10.1007/s10919-008-0065-7
2009 doi
-
[37]
Davit Rizhinashvili, Abdallah Hussein Sham, and Gholamreza Anbarjafari. 2024. Enhanced speech emotion recognition using averaged valence arousal dominance mapping and deep neural networks. Signal, Image and Video Processing 18, 10 (2024), 7445–7454. https://doi.org/10.1007/s11...
2024 doi
-
[38]
James A Russell. 1980. A circumplex model of affect. Journal of personality and social psychology 39, 6 (1980), 1161. https://psycnet.apa.org/fulltext/ 1981-25062-001.pdf
1980
-
[39]
Schuller
Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben, and Björn W. Schuller
-
[40]
Yingzhi Wang, Abdelmoumene Boumadane, and Abdelwahab Heba. 2022. A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding. arXiv:2111.02735 [cs.CL] https://arxiv.org/abs/2111.02735
2022 arXiv
-
[41]
Yiming Wang, Yi Yang, and Jiahong Yuan. 2025. Normalization through Fine-tuning: Understanding Wav2vec 2.0 Embeddings for Phonetic Analysis. arXiv:2503.04814 [cs.CL] https://arxiv.org/abs/2503.04814
2025 arXiv
-
[42]
Taiba Majid Wani, Teddy Surya Gunawan, Syed Asif Ahmad Qadri, Mira Kartiwi, and Eliathamby Ambikairajah. 2021. A Comprehensive Review of Speech Emotion Recognition Systems. IEEE Access 9 (2021), 47795–47814. https://doi.org/10.1109/ACCESS.2021.3068045
2021
-
[43]
Wobbrock, Meredith Ringel Morris, and Andrew D
Jacob O. Wobbrock, Meredith Ringel Morris, and Andrew D. Wilson. 2009. User-defined gestures for surface computing. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Boston, MA, USA) (CHI ’09). Association for Computing Machinery, New York, NY, USA,...
2009
-
[44]
Robert Wolfe, Aayushi Dangol, Bill Howe, and Alexis Hiniker. 2024. Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study. arXiv:2408.01961 [cs.CY] https://arxiv.org/abs/2408.01961
2024 arXiv
-
[2020]
https://doi.org/10.18502/kss.v4i2.6312
Knowledge E, 75–83. https://doi.org/10.18502/kss.v4i2.6312
-
[2023]
IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10745–10759
Dawn of the Transformer Era in Speech Emotion Recognition: Closing the Valence Gap. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10745–10759. https://doi.org/10.1109/TPAMI.2023.3263585
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.