Pith. sign in

REVIEW 1 cited by

Shouted Speech Compensation for Speaker Verification Robust to Vocal Effort Conditions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.02487 v1 pith:3Z6BSOHU submitted 2020-08-06 eess.AS cs.HCcs.LGcs.SD

classification eess.AScs.HCcs.LGcs.SD
keywords shoutedcompensationspeechspeakerverificationconditionseffortnormal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The performance of speaker verification systems degrades when vocal effort conditions between enrollment and test (e.g., shouted vs. normal speech) are different. This is a potential situation in non-cooperative speaker verification tasks. In this paper, we present a study on different methods for linear compensation of embeddings making use of Gaussian mixture models to cluster shouted and normal speech domains. These compensation techniques are borrowed from the area of robustness for automatic speech recognition and, in this work, we apply them to compensate the mismatch between shouted and normal conditions in speaker verification. Before compensation, shouted condition is automatically detected by means of logistic regression. The process is computationally light and it is performed in the back-end of an x-vector system. Experimental results show that applying the proposed approach in the presence of vocal effort mismatch yields up to 13.8% equal error rate relative improvement with respect to a system that applies neither shouted speech detection nor compensation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

    eess.AS 2026-08 conditional novelty 4.0 of 10

    Because voices vary with mood, health, age, speaking style, and recording conditions, the voiceprint metaphor is scientifically misleading and voice evidence should be expressed as calibrated degrees of support.

Pith tools