Pith. sign in

REVIEW 30 cited by

Zero-shot Voice Conversion with Diffusion Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.09943 v1 pith:2Q6L6V7L submitted 2024-11-15 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords timbreconversionvoicespeechzero-shotseed-vctrainingdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre leakage, insufficient timbre representation, and mismatches between training and inference tasks. We propose Seed-VC, a novel framework that addresses these issues by introducing an external timbre shifter during training to perturb the source speech timbre, mitigating leakage and aligning training with inference. Additionally, we employ a diffusion transformer that leverages the entire reference speech context, capturing fine-grained timbre features through in-context learning. Experiments demonstrate that Seed-VC outperforms strong baselines like OpenVoice and CosyVoice, achieving higher speaker similarity and lower word error rates in zero-shot voice conversion tasks. We further extend our approach to zero-shot singing voice conversion by incorporating fundamental frequency (F0) conditioning, resulting in comparative performance to current state-of-the-art methods. Our findings highlight the effectiveness of Seed-VC in overcoming core challenges, paving the way for more accurate and versatile voice conversion systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 30 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data

    cs.SD 2026-06 unverdicted novelty 7.0 of 10

    KNN retrieval on WavLM features constructs synthetic-to-real pairs for training a zero-shot voice conversion model that generalizes across languages when trained exclusively on English.

  2. SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

    eess.AS 2026-06 unverdicted novelty 7.0 of 10

    SpeechEditBench provides seven atomic editing tasks, compositional multi-operation instructions, and an anchor-based protocol yielding target success, preservation success, and joint success metrics; evaluations show ...

  3. Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    Poly-SVC converts singing voices from polyphonic recordings while keeping melody, lyrics, and harmonies by combining CQT-based pitch extraction with a conditional flow matching diffusion decoder.

  4. X-VC: Zero-shot Streaming Voice Conversion in Codec Space

    eess.AS 2026-04 unverdicted novelty 7.0 of 10

    X-VC achieves zero-shot streaming voice conversion via one-step codec-space conversion with dual-conditioning acoustic converter and role-assignment training on generated paired data.

  5. From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

    cs.HC 2026-03 unverdicted novelty 7.0 of 10

    Voice conversion in interactive studies boosts user trust in SpeechLLM responses while automated metrics detect accent-by-gender disparities in alignment and verbosity.

  6. A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

    eess.AS 2026-07 accept novelty 6.0 of 10

    Fixed-threshold voice-clone attribution is unreliable on professional voice actors: a re-ranking-resistant misidentification floor persists, and generic encoders falsely implicate enrolled actors for around half of no...

  7. A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

    eess.AS 2026-07 conditional novelty 6.0 of 10

    On 1,168 professional voice actors, a misidentification floor in speaker embeddings survives calibration, normalization, and discriminative re-ranking, and the same floor makes fixed-threshold voice-clone attribution ...

  8. VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Pretrained spoofing detectors reach at best 28.98% EER on a new English-Spanish benchmark of 10 LLM-era TTS/VC systems under 10 post-processing conditions, with most near chance.

  9. GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

    cs.LG 2026-07 accept novelty 6.0 of 10

    GRAFT splices a short spoken word sample into a neural codec TTS prompt and uses voice-conversion training so the model copies that pronunciation into any target voice, cutting target-word phoneme error 22-39%.

  10. TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Introduces DyadEE dataset and TRACE window-level framework using sequences of acoustic embeddings for emotional entrainment detection, reporting 97.01% accuracy when context and relationship information are included.

  11. TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

    cs.CL 2026-06 conditional novelty 6.0 of 10

    TRACE detects synthetic emotional entrainment disruption in dyadic speech at up to 93.47% accuracy when conditioned on relationship, using windowed emotion-Whisper sequences on the new DyadEE dataset.

  12. ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    ProsoCodec models prosody as a conditional residual in a speech codec via text and speaker prefix conditioning, yielding improved prosody preservation and less timbre leakage in voice conversion experiments.

  13. Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    A MoE-enhanced model with conditional distillation reduces speech-NVV EER from 38.93% to 22.66% and speech EER from 13.17% to 9.24% across 10 NVV types.

  14. RTCFake: Speech Deepfake Detection in Real-Time Communication

    cs.SD 2026-04 unverdicted novelty 6.0 of 10

    RTCFake is the first large-scale dataset of real-time communication speech deepfakes paired with offline versions, paired with a phoneme-guided consistency learning method that improves cross-platform and noise-robust...

  15. How Far Are Video Models from True Multimodal Reasoning?

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Current video models succeed on basic understanding but achieve under 25% success on logically grounded generation and near 0% on interactive generation, exposing gaps in multimodal reasoning.

  16. Entropy-based Coarse and Compressed Semantic Speech Representation Learning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.

  17. AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    AugCodec disentangles speech into semantic, speaker, and prosody tokens via tailored data augmentations, achieving 12.5 Hz operation with three streams and outperforming prior codecs on LibriSpeech reconstruction and ...

  18. Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    Zero-VC applies speaker anonymization as a perturbation to achieve strictly causal zero-lookahead streaming voice conversion by balancing timbre leakage against prosodic utility.

  19. Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    VibE-SVC2 extends prior singing voice conversion work with new modules for independent pitch-style and timbre-style control, claiming better performance and finer controllability than existing methods.

  20. From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    KNN retrieval over WavLM representations creates synthetic source-target pairs from non-parallel data for supervised voice conversion training with a speaker loss, achieving strong results on multilingual test sets de...

  21. Universal Speech Content Factorization

    eess.AS 2026-03 conditional novelty 5.0 of 10

    A universal least-squares speech-to-content map plus few-second speaker transforms yields open-set, low-rank, timbre-suppressed features competitive for zero-shot VC and TTS.

  22. QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis

    cs.LG 2026-01 reject novelty 5.0 of 10

    Diffusion-generated video/audio samples weighted by a learned quality scorer are claimed to improve multimodal sentiment analysis on CH-SIMS, CMU-MOSI, and MUStARD.

  23. REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers

    eess.AS 2025-08 unverdicted novelty 5.0 of 10

    REF-VC is a zero-shot voice conversion system that random-erases redundant parts of speech-embedding features to stay robust to noise, and uses shortcut-distilled flow matching to convert speech in only four steps.

  24. Kimi-Audio Technical Report

    eess.AS 2025-04 unverdicted novelty 5.0 of 10

    Kimi-Audio is an open-source audio foundation model that achieves state-of-the-art results on speech recognition, audio understanding, question answering, and conversation after pre-training on more than 13 million ho...

  25. Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR

    eess.AS 2026-07 unverdicted novelty 4.0 of 10

    Three data-centric strategies are studied to improve rare non-verbal vocalization recognition in ASR while preserving lexical accuracy.

  26. Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis

    cs.SD 2026-07 unverdicted novelty 4.0 of 10

    Unified guidance framework for Flow Matching speech synthesis achieves nearly 3x faster inference and improved speaker similarity by combining heterogeneous data augmentation with intrinsic model guidance to eliminate...

  27. MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion

    eess.AS 2026-06 unverdicted novelty 4.0 of 10

    MeanVC 2 introduces future-receptive chunking and a universal timbre token encoder to achieve lower-latency and more robust streaming zero-shot voice conversion than the original MeanVC.

  28. Semantic-Aware Ship Detection with Vision-Language Integration

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Abstract claims a VLM-plus-adaptive-window framework and a new semantic ship dataset, but the manuscript body is a different voice-timbre paper.

  29. AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

    cs.SD 2026-04 unverdicted novelty 3.0 of 10

    AT-ADD introduces standardized tracks and datasets for evaluating audio deepfake detectors on speech under real-world conditions and on diverse unknown audio types to promote generalization beyond speech-centric methods.

  30. Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects

    cs.HC 2025-10 unverdicted novelty 2.0 of 10

    A holistic survey of affective computing for intelligent agents covering emotion understanding via multimodal data, affective cognition, emotional expression synthesis, key challenges, and future directions emphasizin...

Pith tools