A benchmark of five frozen speech encoders across 22 Indic languages finds that out-of-domain synthetic speech recall is predicted by centroid proximity to unseen TTS systems, rising from 7% to 51% with a four-system training pool.
Modeling Prosody for Language Identification on Read and Spontaneous Speech , booktitle =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages
A benchmark of five frozen speech encoders across 22 Indic languages finds that out-of-domain synthetic speech recall is predicted by centroid proximity to unseen TTS systems, rising from 7% to 51% with a four-system training pool.