REVIEW 8 cited by
Speech Resynthesis from Discrete Disentangled Self-Supervised Representations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker identity. This allows to synthesize speech in a controllable manner. We analyze various state-of-the-art, self-supervised representation learning methods and shed light on the advantages of each method while considering reconstruction quality and disentanglement properties. Specifically, we evaluate the F0 reconstruction, speaker identification performance (for both resynthesis and voice conversion), recordings' intelligibility, and overall quality using subjective human evaluation. Lastly, we demonstrate how these representations can be used for an ultra-lightweight speech codec. Using the obtained representations, we can get to a rate of 365 bits per second while providing better speech quality than the baseline methods. Audio samples can be found under the following link: speechbot.github.io/resynthesis.
Forward citations
Cited by 8 Pith papers
-
Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
By appending optimized token sequences to harmful speech, the authors achieve up to 89% attack success rate on SpeechGPT across six forbidden categories.
-
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
A text-like 'unit language' mined from discrete speech units via n-gram modeling, plus task-prompt multi-task training, improves textless speech-to-speech translation to near text-supervised performance.
-
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
PFlow-VC performs expressive voice conversion by conditioning a flow-matching Mel-spectrogram decoder on discrete speaker-normalized pitch tokens and a target speaker prompt, improving emotion style transfer.
-
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis
Continuous autoregressive text-to-speech with a Gaussian-mixture codec matches or beats a discrete-codec VALL-E baseline with a fraction of the language model parameters.
-
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
The paper presents a unit-based direct speech-to-speech translation system and a paired English-Spanish movie dataset, claiming better preservation of paralinguistic information while maintaining translation quality.
-
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...
-
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
A rectified-flow voice conversion model with speaker feature fusion achieves zero-shot conversion in one sampling step with quality close to 30-step diffusion baselines.
-
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
A cascaded speech-to-text translation model that feeds five aligned ASR candidates and self-supervised speech units to a translation model matches end-to-end performance on GigaST, with an English-to-Chinese BLEU of 38.1.
Discussion (0). Continue with ORCID to comment.