Pith. sign in

REVIEW 8 cited by

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.00355 v3 pith:SXARY6L5 submitted 2021-04-01 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords speechrepresentationsresynthesisqualityself-superviseddiscretedisentangledmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker identity. This allows to synthesize speech in a controllable manner. We analyze various state-of-the-art, self-supervised representation learning methods and shed light on the advantages of each method while considering reconstruction quality and disentanglement properties. Specifically, we evaluate the F0 reconstruction, speaker identification performance (for both resynthesis and voice conversion), recordings' intelligibility, and overall quality using subjective human evaluation. Lastly, we demonstrate how these representations can be used for an ultra-lightweight speech codec. Using the obtained representations, we can get to a rate of 365 bits per second while providing better speech quality than the baseline methods. Audio samples can be found under the following link: speechbot.github.io/resynthesis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

    cs.CL 2025-05 conditional novelty 6.0 of 10

    By appending optimized token sequences to harmful speech, the authors achieve up to 89% attack success rate on SpeechGPT across six forbidden categories.

  2. Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A text-like 'unit language' mined from discrete speech units via n-gram modeling, plus task-prompt multi-task training, improves textless speech-to-speech translation to near text-supervised performance.

  3. Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

    cs.SD 2025-02 conditional novelty 6.0 of 10

    PFlow-VC performs expressive voice conversion by conditioning a flow-matching Mel-spectrogram decoder on discrete speaker-normalized pitch tokens and a target speaker prompt, improving emotion style transfer.

  4. Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Continuous autoregressive text-to-speech with a Gaussian-mixture codec matches or beats a discrete-codec VALL-E baseline with a fraction of the language model parameters.

  5. A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    The paper presents a unit-based direct speech-to-speech translation system and a paired English-Spanish movie dataset, claiming better preservation of paralinguistic information while maintaining translation quality.

  6. Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...

  7. ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A rectified-flow voice conversion model with speaker feature fusion achieves zero-shot conversion in one sampling step with quality close to 30-step diffusion baselines.

  8. When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A cascaded speech-to-text translation model that feeds five aligned ASR candidates and self-supervised speech units to a translation model matches end-to-end performance on GigaST, with an English-to-Chinese BLEU of 38.1.

Pith tools