Seed-TTS models produce speech matching human naturalness and speaker similarity, with added controllability via self-distillation and reinforcement learning.
V oiceShop: A unified speech-to-speech framework for identity- preserving zero-shot voice editing
3 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
Encoder-decoder manifold alignment framework to achieve exact idempotency in generative models.
RIVET enforces an idempotency objective during training of voice attribute editing models to improve robustness to noisy labels, outperforming standard training on controlled noise and the GLOBE dataset.
citing papers explorer
-
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Seed-TTS models produce speech matching human naturalness and speaker similarity, with added controllability via self-distillation and reinforcement learning.
-
Encoder-Decoder Manifold Alignment for Idempotent Generation
Encoder-decoder manifold alignment framework to achieve exact idempotency in generative models.
-
RIVET: Robust Idempotent Voice Attribute Editing
RIVET enforces an idempotency objective during training of voice attribute editing models to improve robustness to noisy labels, outperforming standard training on controlled noise and the GLOBE dataset.