SQ-LLM, trained on the new SpeechEval dataset of 32k multilingual clips with 128k annotations, enables LLMs to perform interpretable multi-task speech quality evaluation including assessment, comparison, improvement suggestions, and deepfake detection.
InICASSP 2021-2021 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), pages 6369–6373
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.SD 3verdicts
UNVERDICTED 3representative citing papers
Pre-training in TTS boosts naturalness during phoneme addition but provides no efficiency gain for learning new phonemes compared to training from scratch.
The study compares speaker embeddings from more than 40 speech foundation models with human subjective similarity scores and identifies model factors that better align with human perception.
citing papers explorer
-
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
SQ-LLM, trained on the new SpeechEval dataset of 32k multilingual clips with 128k annotations, enables LLMs to perform interpretable multi-task speech quality evaluation including assessment, comparison, improvement suggestions, and deepfake detection.
-
Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis
Pre-training in TTS boosts naturalness during phoneme addition but provides no efficiency gain for learning new phonemes compared to training from scratch.
-
Do speech foundation models perceive speaker similarity as humans do?
The study compares speaker embeddings from more than 40 speech foundation models with human subjective similarity scores and identifies model factors that better align with human perception.