Pith. sign in

SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The interpretation of human voices holds importance across various applications. This study ventures into predicting age, gender, and emotion from vocal cues, a field with vast applications. Voice analysis tech advancements span domains, from improving customer interactions to enhancing healthcare and retail experiences. Discerning emotions aids mental health, while age and gender detection are vital in various contexts. Exploring deep learning models for these predictions involves comparing single, multi-output, and sequential models highlighted in this paper. Sourcing suitable data posed challenges, resulting in the amalgamation of the CREMA-D and EMO-DB datasets. Prior work showed promise in individual predictions, but limited research considered all three variables simultaneously. This paper identifies flaws in an individual model approach and advocates for our novel multi-output learning architecture Speech-based Emotion Gender and Age Analysis (SEGAA) model. The experiments suggest that Multi-output models perform comparably to individual models, efficiently capturing the intricate relationships between variables and speech inputs, all while achieving improved runtime.

citation-role summary

method 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

roles

method 1

polarities

use method 1

representative citing papers

CoLMbo: Speaker Language Model for Descriptive Profiling

cs.CL · 2025-06-11 · reject · novelty 5.0

CoLMbo pairs a fixed speaker encoder with a small language model to write descriptive profiles from voice, reporting high zero-shot accuracy for age, gender, ethnicity, and dialect.

citing papers explorer

Showing 1 of 1 citing paper.

  • CoLMbo: Speaker Language Model for Descriptive Profiling cs.CL · 2025-06-11 · reject · none · ref 2 · internal anchor

    CoLMbo pairs a fixed speaker encoder with a small language model to write descriptive profiles from voice, reporting high zero-shot accuracy for age, gender, ethnicity, and dialect.