Pith. sign in

Speaker Embedding Extraction with Phonetic Information

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Speaker embeddings achieve promising results on many speaker verification tasks. Phonetic information, as an important component of speech, is rarely considered in the extraction of speaker embeddings. In this paper, we introduce phonetic information to the speaker embedding extraction based on the x-vector architecture. Two methods using phonetic vectors and multi-task learning are proposed. On the Fisher dataset, our best system outperforms the original x-vector approach by 20% in EER, and by 15%, 15% in minDCF08 and minDCF10, respectively. Experiments conducted on NIST SRE10 further demonstrate the effectiveness of the proposed methods.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

unclear 1

representative citing papers

CoLMbo: Speaker Language Model for Descriptive Profiling

cs.CL · 2025-06-11 · reject · novelty 5.0

CoLMbo pairs a fixed speaker encoder with a small language model to write descriptive profiles from voice, reporting high zero-shot accuracy for age, gender, ethnicity, and dialect.

citing papers explorer

Showing 1 of 1 citing paper.

  • CoLMbo: Speaker Language Model for Descriptive Profiling cs.CL · 2025-06-11 · reject · none · ref 18 · internal anchor

    CoLMbo pairs a fixed speaker encoder with a small language model to write descriptive profiles from voice, reporting high zero-shot accuracy for age, gender, ethnicity, and dialect.