Pith. sign in

REVIEW 4 cited by

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.00355 v3 pith:SXARY6L5 submitted 2021-04-01 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords speechrepresentationsresynthesisqualityself-superviseddiscretedisentangledmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker identity. This allows to synthesize speech in a controllable manner. We analyze various state-of-the-art, self-supervised representation learning methods and shed light on the advantages of each method while considering reconstruction quality and disentanglement properties. Specifically, we evaluate the F0 reconstruction, speaker identification performance (for both resynthesis and voice conversion), recordings' intelligibility, and overall quality using subjective human evaluation. Lastly, we demonstrate how these representations can be used for an ultra-lightweight speech codec. Using the obtained representations, we can get to a rate of 365 bits per second while providing better speech quality than the baseline methods. Audio samples can be found under the following link: speechbot.github.io/resynthesis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

    cs.CL 2025-05 conditional novelty 6.0 of 10

    By appending optimized token sequences to harmful speech, the authors achieve up to 89% attack success rate on SpeechGPT across six forbidden categories.

  2. Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A text-like 'unit language' mined from discrete speech units via n-gram modeling, plus task-prompt multi-task training, improves textless speech-to-speech translation to near text-supervised performance.

  3. Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...

  4. ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A rectified-flow voice conversion model with speaker feature fusion achieves zero-shot conversion in one sampling step with quality close to 30-step diffusion baselines.

Pith tools