REVIEW 30 cited by
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Text-to-Speech (TTS) systems face ongoing challenges in processing complex linguistic features, handling polyphonic expressions, and producing natural-sounding multilingual speech - capabilities that are crucial for future AI applications. In this paper, we present Fish-Speech, a novel framework that implements a serial fast-slow Dual Autoregressive (Dual-AR) architecture to enhance the stability of Grouped Finite Scalar Vector Quantization (GFSQ) in sequence generation tasks. This architecture improves codebook processing efficiency while maintaining high-fidelity outputs, making it particularly effective for AI interactions and voice cloning. Fish-Speech leverages Large Language Models (LLMs) for linguistic feature extraction, eliminating the need for traditional grapheme-to-phoneme (G2P) conversion and thereby streamlining the synthesis pipeline and enhancing multilingual support. Additionally, we developed FF-GAN through GFSQ to achieve superior compression ratios and near 100\% codebook utilization. Our approach addresses key limitations of current TTS systems while providing a foundation for more sophisticated, context-aware speech synthesis. Experimental results show that Fish-Speech significantly outperforms baseline models in handling complex linguistic scenarios and voice cloning tasks, demonstrating its potential to advance TTS technology in AI applications. The implementation is open source at \href{https://github.com/fishaudio/fish-speech}{https://github.com/fishaudio/fish-speech}.
Forward citations
Cited by 30 Pith papers
-
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
SemBridge supervises autoregressive states with discrete semantic tokens during training, improving content fidelity of continuous-latent speech generation without changing inference.
-
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
LRLspoof corpus and threshold-transfer evaluation demonstrate that spoof detection performance varies markedly across languages, identifying language as an independent domain shift factor.
-
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
ArVoice is a new 83.5-hour, 11-voice Modern Standard Arabic speech corpus with diacritized transcripts for multi-speaker TTS and voice conversion.
-
Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption
Energy of text-to-video diffusion models is predicted from architectural first principles and observable generation parameters with under 3% MAPE, without needing weights or model size.
-
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.
-
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Block discrete diffusion over X-Codec2 tokens yields competitive zero-shot TTS with 0.6B parameters, 20K training hours, and a 0.15 real-time factor.
-
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.
-
HISPASpoof: A New Dataset For Spanish Speech Forensics
HISPASpoof is a new large Spanish synthetic-speech dataset for detection and attribution, with evidence that English-trained detectors fail on Spanish and Spanish training helps.
-
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
A unified audio-video generator built by stitching pre-trained video and music diffusion models, trained on 7,600 hours of data, with a new evaluation benchmark.
-
MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
MoTAS combines TTS speech augmentation with MoE-guided feature selection to reach 85.71% accuracy on ADReSSo, the highest among the baselines listed.
-
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.
-
Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
By fine-tuning an LLM with audio and motion tokens, MECo generates co-speech gestures that follow a user-provided motion example, and reports state-of-the-art FGD and diversity on BEAT2 and ZeroEGGS.
-
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
A transcript-prompted Whisper model with dictionary-based decoding automatically produces phonemic and prosodic annotations for Japanese audio-transcript pairs, improving Japanese TTS naturalness.
-
Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
XS-CoT trains speech LLMs to answer non-core language questions by generating an English reasoning chain before the final target-language answer, and a semi-implicit variant compresses that chain to cut latency.
-
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
GenSE enhances speech by first denoising semantic tokens with a language model and then generating acoustic tokens from a single-quantizer codec, reporting higher DNSMOS, speaker similarity, and lower WER than prior systems.
-
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
TouchTTS reports a 51.6% data retention rate using a noise-robust tokenizer and two-ASR cross-validation, and a Qwen-backbone flow model that unifies streaming and non-streaming synthesis while matching CosyVoice on PER.
-
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs
Faster IndexTTS-2 compiles all components of IndexTTS-2 into TensorRT engines, cutting end-to-end latency by up to 3.6x while adding streaming and batching.
-
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.
-
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
A domain tag marking synthetic training audio plus 3× oversampling of real audio lets a TTS model absorb large synthetic augmentation without losing speaker similarity.
-
DarkStream: real-time speech anonymization with low latency
DarkStream combines a causal content encoder with limited lookahead, k-means quantization, and a GAN-based pseudo-speaker embedding to anonymize speech in real time with near-chance speaker-verification error rates.
-
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
A speech-deepfake dataset for ten public figures built with transcription-based segmentation reports high synthetic naturalness (NISQA 3.69) and a human misclassification rate of 61.9%.
-
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
The authors propose DistilCodec, a 32,768-code single-codebook audio codec, and UniTTS, a Qwen2.5-7B TTS model trained with audio, text, and cross-modal autoregressive tasks on interleaved prompts.
-
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
FlexSpeech is a zero-shot TTS system that predicts phoneme durations autoregressively, renders speech with flow matching, and applies direct preference optimization to durations for fast style transfer.
-
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
A series of Qwen2.5-based speech chatbots (0.5B to 14B) that stream speech output with a CosyVoice2-style autoregressive decoder and outperform prior SpeechLMs using only 200K synthetic multi-turn dialogues.
-
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
IndexTTS combines character and pinyin modeling to make Chinese polyphone pronunciation controllable, and reports improved zero-shot voice cloning and naturalness over open-source TTS baselines.
-
BoSS: Beyond-Semantic Speech
Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.
-
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
A speculative decoding variant with a tolerance-based acceptance rule speeds up CosyVoice 2 inference by 1.4x while keeping subjective quality comparable.
-
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
Muyan-TTS, a 3B-parameter LLM-based TTS model trained on 100,000+ hours of podcast audio, produces competitive zero-shot speech and runs at 0.33 seconds of inference per second of speech.
-
Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
Fine-tuned open LLMs can imitate individual writing styles from small samples, evade detection tools, and are not yet addressed by current safeguards or law.
-
WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.
Discussion (0). Continue with ORCID to comment.