Pith. sign in

REVIEW 5 major objections 5 minor 46 references

"How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases

T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A user study finds no statistically significant age or language differences in how a pre-trained speech-emotion model reads deliberately acted emotions, with a persistent weakness on high-arousal emotions like happy and angry.

desk verdict Inventive interactive SER evaluation with rich qualitative data, but the headline null result is not statistically grounded because the t-tests treat repeated utterances as independent. read the letter →

arxiv 2507.12580 v1 pith:ABKP4WRW submitted 2025-07-16 cs.HC

classification cs.HC
keywords speechemotionrecognitionaffectivecomputingagebiaslanguagevalence-arousalspaceexpressionhuman-centeredAIuserstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether speech emotion recognition (SER) systems, which are mostly trained on spontaneous English speech, are biased against users who are not young-adult native English speakers. To test this, the authors built a voice-controlled emotion-targeting game in which 12 teenagers and 12 adults aged 55+ deliberately expressed happy, sad, angry, and calm in both Danish and English while a pre-trained SER model scored their distance to target coordinates in valence-arousal space. The central finding is a null result: no statistically significant difference in expressive accuracy by age group or language ($p > .05$), suggesting the model interprets deliberately acted emotions comparably across these groups. The paper also documents a consistent weakness on high-arousal emotions, a mismatch between users' self-perception and model output, and a qualitative call for human-centered SER design that accommodates blended, context-dependent emotions.

What carries the argument

The central mechanism is a target-hitting task in valence-arousal space: a participant vocally acts an emotion to hit a fixed coordinate, while a backend runs a pre-trained Wav2Vec 2.0-based SER model on the utterance and logs its predicted valence-arousal coordinate. The paper defines expressive accuracy as the Euclidean distance between prediction and target, so the model's coordinate system is the measuring instrument for every age and language comparison. The task design, inspired by prior targeting and elicitation paradigms, converts emotional expression into a goal-directed, quantifiable act and produces a real-time log of human intent versus machine interpretation.

What would settle it

Re-score the same 384 logged utterances with human listeners who rate valence and arousal, then test for age and language effects; if human raters find differences the model missed, the claimed robustness is an artifact of model insensitivity. A simpler check is to rerun the t-tests after removing all trials where the system recorded silence or hesitation, since the model defaulted those to sad.

Watch

Extended reading notes

Core claim

Contrary to the authors' assumptions, neither age nor language produced a statistically significant difference in how accurately the model read deliberately expressed emotions. Expressive accuracy was measured as the Euclidean distance between the model's predicted valence-arousal coordinates and the fixed target coordinates for each of four emotions; across 384 utterances from 24 native Danish speakers, independent-samples t-tests found no group differences in either the English or the Danish condition. The data do show a pattern the authors emphasize: high-arousal emotions, happy especially, had the largest distances for both groups, and the model often read laughter during angry attempts as happy. The paper frames the contribution as a shift from system-centered accuracy to human-centered expressive accuracy, exposing the real-time gap between what a speaker intends and what the model interprets.

Load-bearing premise

The load-bearing premise is that the pre-trained model's valence-arousal readings are a valid and sensitive measure of how accurately a person expressed an emotion; if the model is blind to age or language differences, or if it collapses high-arousal and silent input into biased categories like sad, the null t-tests describe the model, not the speakers.

Editorial extensions

If this is right

  • If the null result holds, an English-trained SER model can interpret deliberately acted basic emotions from teenage and older-adult native Danish speakers as accurately as from English utterances, at least for the four studied emotions.
  • The systematic difficulty with happy and angry points to a concrete target for model improvement: high-arousal emotions need recalibration or additional training data rather than age- or language-specific fixes.
  • Logging the real-time gap between intended and predicted valence-arousal coordinates gives designers a direct way to audit when an SER system misaligns with a user's emotional intent, which can be used to evaluate inclusiveness before deployment.
  • Users' tendency to describe emotions as blended states (calm/happy, angry/sad) challenges discrete-label SER benchmarks and supports the paper's argument for models that represent emotional complexity rather than fixed categories.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because expressive accuracy is measured entirely inside the model's own coordinate system, a plausible reading is that the null result says more about the model's insensitivity than about human equality; a direct test would be to have human listeners rate the same utterances on valence and arousal and compare the two instruments.
  • The same targeting-game protocol could be run with a different pre-trained SER model; if that model shows age or language effects the first model misses, the robustness claim is model-specific rather than a general property of speech emotion AI.
  • The paper's suggestion to test typologically distant languages such as Greenlandic is a natural extension: positive results would support universal acoustic encoding of emotion, while negative results would locate the boundary of cross-linguistic transfer.
  • The silent-input-defaults-to-sad behavior is a concrete artifact that future audits could exploit as a targeted probe for systematic bias in SER systems, since hesitation and silence are common in real interactions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper presents a user study (N=24: 12 teenagers and 12 adults aged 55+) in which participants deliberately acted four emotions (happy, sad, angry, calm) in both Danish and English while a pre-trained SER model (Wagner et al.) predicted valence-arousal coordinates in real time. The authors quantify expressive accuracy as the Euclidean VA distance between the model's predictions and the target coordinates, and they report independent-samples t-tests comparing age groups within each language condition. The central claim is that no statistically significant age or language differences were found (all p > .05), suggesting that, under structured elicitation, both age groups are comparably effective at deliberately conveying emotions as interpreted by this model. Qualitative interview findings about emotion expression, trust in SER, and the perceived best emotions to convey are also reported, together with a discussion of limitations including the model's silence-to-Sad default and difficulties with high-arousal emotions.

Significance. The paper's main strength is its human-centered experimental paradigm: a real-time VA-targeting game with full logging of model predictions, a thoughtful qualitative analysis, and an unusually explicit limitations section. If the reported null results survive correct statistical analysis, they would be a useful empirical counterpoint to assumptions about age and language bias in SER and would support the paper's call for more inclusive, human-centered speech emotion AI. The qualitative findings, especially participants' descriptions of blended emotions and the perceived divergence between self-report and model output, are valuable contributions in their own right. However, as presented, the statistical analysis does not currently support the central null claim, and the absence of any reported test for the language comparison leaves part of the headline result unsupported.

major comments (5)
  1. [§4.1, §4.3] The statistical unit is inconsistent with the reported degrees of freedom. Section 4.1 states that the analysis partitions the data by language, yielding 192 utterances per language condition (96 per age group), and implies that independent-samples t-tests were run on these utterances. Yet the only t statistic with degrees of freedom, t(21.74) = -1.49 for sad in Danish (Section 4.3), has df ≈ 22, which corresponds to a test on 12 observations per group (i.e., per-participant means) rather than the 24 raw utterances per age group per emotion-language cell. If the two within-participant repetitions were treated as independent observations, the effective sample size is inflated, the observations are non-independent, and the p-values are anti-conservative. If the repetitions were averaged to participant-level means, that aggregation is nowhere stated and the description in Section 4.1 is misleading. Please clarify which unit was used, and re-run the analyses with a mixed-effects model or with explicitly aggregated per-participant means.
  2. [§4.2, §4.3, §5] The Discussion's central claim that 'no statistically significant differences were found between the age groups or across language conditions (all p > .05)' is not supported by any reported test of the language condition. Sections 4.2 and 4.3 only compare age groups within English and within Danish; no English-versus-Danish comparison (e.g., a paired t-test on participant-level means or a language factor in a mixed model) is reported. Either add the missing language comparison or revise the Conclusion and Discussion to claim only that no age differences were found within each language.
  3. [§4.1, §6] The VA-distance metric is computed entirely from the Wagner et al. model's predicted coordinates, and the study provides no human-listener baseline, chance level, or any other validation that these distances are sensitive to meaningful differences in expressive accuracy. Without such a baseline, the null t-tests cannot distinguish 'no difference in expressive ability' from 'this model's predictions are insensitive to age or language differences.' This concern is reinforced by the acknowledged silence-to-Sad default and the model's difficulties with high-arousal emotions (Section 6). Please either validate the metric with human ratings or a reference corpus, or rephrase the conclusions to state that no differences were found in the model's VA-distance outputs, rather than in expressive accuracy per se.
  4. [§3.1, §3.3] The relationship between the pilot study and the final study participants is not specified. Section 3.1 describes a pilot with 12 participants from the 55+ age group, and Section 3.3 reports that the final study had 12 Adults 55+ and 12 Teenagers. If the same 12 adults completed both the pilot and the final study, their prior exposure to the interface and task would confound the age comparison (adults would have had extra practice; teenagers would not). If the pilot participants were a separate group, this should be stated explicitly. Please clarify this methodological detail.
  5. [§6] The silence-to-Sad default is acknowledged as a technical limitation, but the quantitative analysis in Section 4 does not report how many trials involved silence or hesitation, nor any sensitivity analysis excluding these trials. If silent trials were included in the VA-distance means, they could artifactually lower the sad distances and raise the happy/angry distances, directly affecting the reported means and the null results. Please report the number of affected trials and re-run the analysis without them, or with silence as a covariate.
minor comments (5)
  1. [§4.3] The heading of Section 4.3 duplicates Section 4.2 ('English Condition: Age Group Differences in Expressive Accuracy') even though the section content concerns Danish utterances; the heading should be corrected.
  2. [Figures 9 and 11] The captions for Figures 9 and 11 appear to be duplicated; each figure caption says the figure displays mean VA distances 'comparing Teenagers and Adults in English (left panel) and Danish (right panel)', but each figure should describe only its own language condition.
  3. [§2.1] There is a missing citation placeholder in the sentence introducing the Circumplex Model of Affect: 'Russell and Barrett offers ... [? ]'.
  4. [§4.3, §7] There are several typos: 'Teenagrs' in Section 4.3, 'in-debt' in the Conclusion (should be 'in-depth'), 'targetted' in the Conclusion, and 'dependable variable' (should be 'dependent variable').
  5. [§3.2] The text 'This resulted in 16 total utterances per participant (four emotions× four repetitions)' is imprecise; the structure is four emotions × two languages × two repetitions, and stating it this way would avoid confusion.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; the study is an empirical measurement of SER model outputs, with only minor non-load-bearing self-citations for the prototype architecture.

full rationale

This is an empirical HCI user study rather than a derivation chain. The dependent variable, VA distance, is explicitly defined as the Euclidean distance between the SER model's predicted valence-arousal coordinates and fixed target coordinates, so every age and language comparison is a comparison of model outputs. That is the paper's stated object of study ('how a pre-trained SER model from Wagner et al. [39] interprets those vocal expressions'), not a hidden equivalence: the authors fit no parameter to the outcome and then relabel it as a prediction. The only self-citations are [6,7] for the UI/backend architecture ('To this end, we build on the architecture of previous work [7]'), and those are implementation reuse, not load-bearing evidence for the central null result. The Wagner et al. model [39] is an external, fixed instrument; whether it is valid or unbiased is a measurement-correctness concern, not circularity. Section 6's admission that silent input defaults to 'Sad' is a stated instrument limitation and does not reduce any result to its own inputs. The non-independence of within-participant replicates in the t-tests is a genuine statistical-validity problem, but it is not a circularity: the reported p-values may be invalid without the claim being equivalent to the inputs by construction. Verdict: no significant circularity; only minor, non-load-bearing self-citation, so score 1.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

The paper makes no derivation and estimates no fitted parameters. The central claim rests on hand-chosen target coordinates and on the pre-trained SER model as a valid measurement instrument, both treated as given. No new entities, forces, or dimensions are introduced.

free parameters (1)
  • Target VA coordinates for the four emotion targets = Not reported in the paper
    Happy, sad, angry, and calm are mapped to specific valence-arousal target coordinates (Sections 3.2 and 4.1). The VA distance and all t-tests depend on these hand-chosen values, which are not listed or justified in the manuscript.
assumptions (2)
  • domain assumption Russell's Circumplex Model of Affect is a valid 2D representation of the four target emotions, so Euclidean distance to a fixed target coordinate is a meaningful measure of expressive accuracy.
    Invoked in Sections 2.1, 3.2, and 4.1. If the dimensional mapping or the chosen target coordinates are wrong, every accuracy score shifts.
  • domain assumption The pre-trained Wagner et al. SER model gives unbiased and sufficiently sensitive VA predictions for Danish and English, for teenagers and adults over 55, so that null t-test results can be interpreted as age and language robustness.
    The model is the measurement instrument throughout Sections 3.4 and 4.1. The paper's own limitation section shows the model is not neutral, because silence defaults to the "Sad" quadrant.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases." pith.science (2026). https://pith.science/paper/ABKP4WRW

@misc{pith2026250712580,
  author       = {Pith},
  title        = {Pith review of: "How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ABKP4WRW}},
  note         = {Machine review of arXiv:2507.12580}
}
read the original abstract

This study explores how age and language shape the deliberate vocal expression of emotion, addressing underexplored user groups, Teenagers (N = 12) and Adults 55+ (N = 12), within speech emotion recognition (SER). While most SER systems are trained on spontaneous, monolingual English data, our research evaluates how such models interpret intentionally performed emotional speech across age groups and languages (Danish and English). To support this, we developed a novel experimental paradigm combining a custom user interface with a backend for real-time SER prediction and data logging. Participants were prompted to hit visual targets in valence-arousal space by deliberately expressing four emotion targets. While limitations include some reliance on self-managed voice recordings and inconsistent task execution, the results suggest contrary to expectations, no significant differences between language or age groups, and a degree of cross-linguistic and age robustness in model interpretation. Though some limitations in high-arousal emotion recognition were evident. Our qualitative findings highlight the need to move beyond system-centered accuracy metrics and embrace more inclusive, human-centered SER models. By framing emotional expression as a goal-directed act and logging the real-time gap between human intent and machine interpretation, we expose the risks of affective misalignment.

Figures

Figures reproduced from arXiv: 2507.12580 by the authors.

Figure 1
Figure 1. The sentences that were uttered in each emotional quadrant of our model in the pilot study. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The sentences that were uttered in each emotional quadrant of our model. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. (1) Participants indicate their emotional state. (2) Participants start recording the model. (3) SER results are mapped to a [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Participants’ current emotional state before our user study. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Screenshots displaying the user interface used in our SER-based emotion targeting game, for participants to interact with. (a) [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Predicted vs. Target Emotion Coordinates: Adults. [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Predicted vs. Target Emotion Coordinates: Teenagers. [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Overview of Mean V A Distances by Age Group for English Utterances.(Note: Lower VA distances indicate greater expressive [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Mean V A Distances by Emotion, Age Group and Language(English). This figure displays mean V A distances for the four [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]
Figure 10
Figure 10. Figure 10: Mean V A distances by Age Group for Danish utterances. (Note: Lower VA distances indicate greater expressive accuracy. The [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 11
Figure 11. Figure 11: Mean V A Distances by Emotion, Age Group and Language(Danish). This figure displays mean V A distances for the four [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

46 extracted references · 29 canonical work pages

  1. [1]

    Communication Trends in the Post-literacy Era: Multilingualism, Multimodality, Multiculturalism

    Sofia Abramova, Natalia Antonova, and Anna Gurarii. 2020. Speech behavior and multimodality in online communication among teenagers. In IV International Scientific Conference “Communication Trends in the Post-literacy Era: Multilingualism, Multimodality, Multiculturalism”.—Ekaterinburg,

  2. [2]

    Adebanji Adeleye, Samaneh Madanian, and Olayinka Adeleye. 2024. Emotion Variation Detection in Discrete English Speech: A Wavelet Transform Use Case in Mental Health Monitoring. In Proceedings of the 2024 Australasian Computer Science Week (Sydney, NSW, Australia) (ACSW ’24). Association for Computing Machinery, New York, NY, USA, 115–119. https://doi.org...

  3. [3]

    On the Impact of Word Error Rate on Acoustic-Linguistic Speech Emotion Recognition: An Update for the Deep Learning Era

    Shahin Amiriparian, Artem Sokolov, Ilhan Aslan, Lukas Christ, Maurice Gerczuk, Tobias Hübner, Dmitry Lamanov, Manuel Milling, Sandra Ottl, Ilya Poduremennykh, Evgeniy Shuranov, and Björn W. Schuller. 2021. On the Impact of Word Error Rate on Acoustic-Linguistic Speech Emotion Recognition: An Update for the Deep Learning Era. arXiv:2104.10121 [cs.SD] https...

  4. [4]

    How to Explore Biases in Speech Emotion AI with Users?

    Pengcheng An, Chaoyu Zhang, Haichen Gao, Ziqi Zhou, Yage Xiao, and Jian Zhao. 2025. AniBalloons: Animated chat balloons as affective augmentation for social messaging and chatbot interaction. International Journal of Human-Computer Studies 194 (2025), 103365. https://doi.org/10. 1016/j.ijhcs.2024.103365 "How to Explore Biases in Speech Emotion AI with Users?" 19

  5. [5]

    Pengcheng An, Jiawen Stefanie Zhu, Zibo Zhang, Yifei Yin, Qingyuan Ma, Che Yan, Linghao Du, and Jian Zhao. 2024. EmoWear: Exploring Emotional Teasers for Voice Message Interaction on Smartwatches. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY,...

  6. [6]

    Speejis: Enhancing User Experience of Mobile Voice Messaging with Automatic Visual Speech Emotion Cues

    Ilhan Aslan, Carla F. Griggio, Henning Pohl, Timothy Merritt, and Niels van Berkel. 2025. Speejis: Enhancing User Experience of Mobile Voice Messaging with Automatic Visual Speech Emotion Cues. arXiv:2502.05296 [cs.HC] https://arxiv.org/abs/2502.05296

  7. [7]

    Speech Command + Speech Emotion: Exploring Emotional Speech Commands as a Compound and Playful Modality

    Ilhan Aslan, Timothy Merritt, Stine S. Johansen, and Niels van Berkel. 2025. Speech Command + Speech Emotion: Exploring Emotional Speech Commands as a Compound and Playful Modality. arXiv:2504.08440 [cs.HC] https://arxiv.org/abs/2504.08440

  8. [8]

    Ilhan Aslan, Tabea Schmidt, Jens Woehrle, Lukas Vogel, and Elisabeth André. 2018. Pen + Mid-Air Gestures: Eliciting Contextual Gestures. In Proceedings of the 20th ACM International Conference on Multimodal Interaction (Boulder, CO, USA)(ICMI ’18). Association for Computing Machinery, New York, NY, USA, 135–144. https://doi.org/10.1145/3242969.3242979

Show all 46 references
  1. [9]

    Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations. arXiv:2006.11477 [cs.CL] https://arxiv.org/abs/2006.11477

  2. [10]

    Patricia E. G. Bestelmeyer, Sonja A. Kotz, and Pascal Belin. 2017. Effects of emotional valence and arousal on the voice perception network. Social Cognitive and Affective Neuroscience 12, 8 (04 2017), 1351–1358. https://doi.org/10.1093/scan/nsx059 arXiv:https://academic.oup.c...

  3. [11]

    Samuel Cahyawijaya, Holy Lovenia, Willy Chung, Rita Frieske, Zihan Liu, and Pascale Fung. 2023. Cross-Lingual Cross-Age Group Adaptation for Low-Resource Elderly Speech Emotion Recognition. arXiv:2306.14517 [cs.CL] https://arxiv.org/abs/2306.14517

  4. [12]

    Fabio Catania. 2023. Speech emotion recognition in italian using wav2vec 2.Authorea Preprints (2023). https://doi.org/10.36227/techrxiv.22821992.v1

  5. [13]

    Charlene H Chu, Rune Nyrup, Kathleen Leslie, Jiamin Shi, Andria Bianchi, Alexandra Lyn, Molly McNicholl, Shehroz Khan, Samira Rahimi, and Amanda Grenier. 2022. Digital Ageism: Challenges and Opportunities in Artificial Intelligence for Older Adults. The Gerontologist 62, 7 (01...

  6. [14]

    Douglas-Cowie, Suzie Savvidou, E

    Roddy Cowie, E. Douglas-Cowie, Suzie Savvidou, E. McMahon, M. Sawey, and M. Schröder. 2000. ’FEELTRACE’: An instrument for recording perceived emotion in real time. Proceedings of the ISCA Workshop on Speech and Emotion

  7. [15]

    Zachary Dair, Ryan Donovan, and Ruairi O’Reilly. 2022. Linguistic and Gender Variation in Speech Emotion Recognition using Spectral Features. arXiv:2112.09596 [cs.SD] https://arxiv.org/abs/2112.09596

  8. [16]

    Jarod Duret, Mickael Rouvier, and Yannick Estève. 2024. MSP-Podcast SER Challenge 2024: L’antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition. arXiv:2407.05746 [cs.AI] https://arxiv.org/abs/2407.05746

  9. [17]

    Paul Ekman. 1992. Are there basic emotions? (1992). https://doi.org/10.1037/0033-295X.99.3.550

  10. [18]

    Paul Ekman. 1992. Facial Expressions of Emotion: New Findings, New Questions. Psychological Science 3, 1 (1992), 34–38. http://www.jstor.org/ stable/40062750

  11. [19]

    Google Translate. 2024. Google Translate. https://support.google.com/translate/answer/15139004?hl=en&utm_source=chatgpt.com. Accessed: 2025-06-27

  12. [20]

    Cristina Gorrostieta, Reza Lotfian, Kye Taylor, Richard Brutti, and John Kane. 2019. Gender De-Biasing in Speech Emotion Recognition.. In Interspeech. 2823–2827. https://www.isca-archive.org/interspeech_2019/gorrostieta19_interspeech.pdf

  13. [21]

    Igor Grossmann, Alex C Huynh, and Phoebe C Ellsworth. 2016. Emotional complexity: Clarifying definitions and cultural correlates. Journal of personality and social psychology 111, 6 (2016), 895. https://doi.org/10.1037/pspp0000084

  14. [22]

    Lucía Gómez-Zaragozá, Rocío del Amor, María José Castro-Bleda, Valery Naranjo, Mariano Alcañiz Raya, and Javier Marín-Morales. 2024. EMOVOME: A Dataset for Emotion Recognition in Spontaneous Real-Life Speech. arXiv:2403.02167 [eess.AS] https://arxiv.org/abs/2403.02167

  15. [23]

    Lasse Hansen, Yan-Ping Zhang, Detlef Wolf, Konstantinos Sechidis, Nicolai Ladegaard, and Riccardo Fusaroli. 2022. A generalizable speech emotion recognition model reveals depression and remission. Acta Psychiatrica Scandinavica 145, 2 (2022), 186–199

  16. [24]

    Alexandra Israelsson, Anja Seiger, and Petri Laukka. 2023. Blended emotions can be accurately recognized from dynamic facial and vocal expressions. Journal of Nonverbal Behavior 47, 3 (2023), 267–284. https://doi.org/10.1007/s10919-023-00426-9

  17. [25]

    Jesin James, Felix Marattukalam, Kaustubh Shukla, Enuri Kolugala, Sunny Choi, and Owen Eng. [n. d.]. Emotiongui: A Web-Based Tool for Visualisation and Annotation of Emotions for Speech and Video Signals. A vailable at SSRN 5063641([n. d.]). http://dx.doi.org/10.2139/ssrn.5063641

  18. [26]

    Ruhul Amin Khalil, Edward Jones, Mohammad Inayatullah Babar, Tariqullah Jan, Mohammad Haseeb Zafar, and Thamer Alhussain. 2019. Speech Emotion Recognition Using Deep Learning Techniques: A Review.IEEE Access 7 (2019), 117327–117345. https://doi.org/10.1109/ACCESS.2019.2936124

  19. [27]

    Michael W Kraus. 2017. Voice-only communication enhances empathic accuracy. American Psychologist 72, 7 (2017), 644. https://doi.org/10.1037/ amp0000147

  20. [28]

    Gaku Kutsuzawa, Hiroyuki Umemura, Koichiro Eto, and Yoshiyuki Kobayashi. 2022. Classification of 74 facial emoji’s emotional states on the valence-arousal axes. Scientific Reports 12, 1 (2022), 398. https://doi.org/10.1038/s41598-021-04357-7

  21. [29]

    Yi-Cheng Lin, Haibin Wu, Huang-Cheng Chou, Chi-Chun Lee, and Hung yi Lee. 2024. Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition. In Interspeech 2024. 4633–4637. https://doi.org/10.21437/Interspeech.2024-1073

  22. [30]

    Sandra Luo. 2024. Navigating the Diverse Challenges of Speech Emotion Recognition: A Deep Learning Perspective. In Proceedings of the 27th International Academic Mindtrek Conference (Tampere, Finland) (Mindtrek ’24). Association for Computing Machinery, New York, NY, USA, 133–...

  23. [31]

    Yong Ma, Yuchong Zhang, Miroslav Bachinski, and Morten Fjeld. 2024. Emotion-Aware Voice Assistants: Design, Implementation, and Preliminary Insights. In Proceedings of the Eleventh International Symposium of Chinese CHI (Denpasar, Bali, Indonesia) (CHCHI ’23). Association for ...

  24. [32]

    I Scott MacKenzie. 1992. Fitts’ law as a research and design tool in human-computer interaction. Human-computer interaction 7, 1 (1992), 91–139. https://doi.org/10.1207/s15327051hci0701_3

  25. [33]

    Brent Daniel Mittelstadt, Patrick Allo, Mariarosaria Taddeo, Sandra Wachter, and Luciano Floridi. 2016. The ethics of algorithms: Mapping the debate. Big Data & Society 3, 2 (2016), 2053951716679679. https://doi.org/10.1177/2053951716679679 arXiv:https://doi.org/10.1177/205395...

  26. [34]

    Michael Neumann and N goc Thang Vu. 2018. CRoss-lingual and Multilingual Speech Emotion Recognition on English and French. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 5769–5773. https://doi.org/10.1109/ICASSP.2018.8462162

  27. [35]

    Bernstein, Robin N

    Joon Sung Park, Michael S. Bernstein, Robin N. Brewer, Ece Kamar, and Meredith Ringel Morris. 2021. Understanding the Representation and Representativeness of Age in AI Data Sets. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (A...

  28. [36]

    Marc D Pell, Laura Monetta, Silke Paulmann, and Sonja A Kotz. 2009. Recognizing emotions in a foreign language. Journal of Nonverbal Behavior 33, 2 (2009), 107–120. https://doi.org/10.1007/s10919-008-0065-7

  29. [37]

    Davit Rizhinashvili, Abdallah Hussein Sham, and Gholamreza Anbarjafari. 2024. Enhanced speech emotion recognition using averaged valence arousal dominance mapping and deep neural networks. Signal, Image and Video Processing 18, 10 (2024), 7445–7454. https://doi.org/10.1007/s11...

  30. [38]

    James A Russell. 1980. A circumplex model of affect. Journal of personality and social psychology 39, 6 (1980), 1161. https://psycnet.apa.org/fulltext/ 1981-25062-001.pdf

  31. [39]

    Schuller

    Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben, and Björn W. Schuller

  32. [40]

    Yingzhi Wang, Abdelmoumene Boumadane, and Abdelwahab Heba. 2022. A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding. arXiv:2111.02735 [cs.CL] https://arxiv.org/abs/2111.02735

  33. [41]

    Yiming Wang, Yi Yang, and Jiahong Yuan. 2025. Normalization through Fine-tuning: Understanding Wav2vec 2.0 Embeddings for Phonetic Analysis. arXiv:2503.04814 [cs.CL] https://arxiv.org/abs/2503.04814

  34. [42]

    Taiba Majid Wani, Teddy Surya Gunawan, Syed Asif Ahmad Qadri, Mira Kartiwi, and Eliathamby Ambikairajah. 2021. A Comprehensive Review of Speech Emotion Recognition Systems. IEEE Access 9 (2021), 47795–47814. https://doi.org/10.1109/ACCESS.2021.3068045

  35. [43]

    Wobbrock, Meredith Ringel Morris, and Andrew D

    Jacob O. Wobbrock, Meredith Ringel Morris, and Andrew D. Wilson. 2009. User-defined gestures for surface computing. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (Boston, MA, USA) (CHI ’09). Association for Computing Machinery, New York, NY, USA,...

  36. [44]

    Robert Wolfe, Aayushi Dangol, Bill Howe, and Alexis Hiniker. 2024. Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study. arXiv:2408.01961 [cs.CY] https://arxiv.org/abs/2408.01961

  37. [2020]

    https://doi.org/10.18502/kss.v4i2.6312

    Knowledge E, 75–83. https://doi.org/10.18502/kss.v4i2.6312

  38. [2023]

    IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10745–10759

    Dawn of the Transformer Era in Speech Emotion Recognition: Closing the Valence Gap. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10745–10759. https://doi.org/10.1109/TPAMI.2023.3263585

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.