Pith. sign in

REVIEW 30 cited by

Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.01156 v2 pith:DRLOER4X submitted 2024-11-02 cs.SD eess.AS

classification cs.SDeess.AS
keywords fish-speechlinguisticmodelsmultilingualsynthesisapplicationsarchitecturecloning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Text-to-Speech (TTS) systems face ongoing challenges in processing complex linguistic features, handling polyphonic expressions, and producing natural-sounding multilingual speech - capabilities that are crucial for future AI applications. In this paper, we present Fish-Speech, a novel framework that implements a serial fast-slow Dual Autoregressive (Dual-AR) architecture to enhance the stability of Grouped Finite Scalar Vector Quantization (GFSQ) in sequence generation tasks. This architecture improves codebook processing efficiency while maintaining high-fidelity outputs, making it particularly effective for AI interactions and voice cloning. Fish-Speech leverages Large Language Models (LLMs) for linguistic feature extraction, eliminating the need for traditional grapheme-to-phoneme (G2P) conversion and thereby streamlining the synthesis pipeline and enhancing multilingual support. Additionally, we developed FF-GAN through GFSQ to achieve superior compression ratios and near 100\% codebook utilization. Our approach addresses key limitations of current TTS systems while providing a foundation for more sophisticated, context-aware speech synthesis. Experimental results show that Fish-Speech significantly outperforms baseline models in handling complex linguistic scenarios and voice cloning tasks, demonstrating its potential to advance TTS technology in AI applications. The implementation is open source at \href{https://github.com/fishaudio/fish-speech}{https://github.com/fishaudio/fish-speech}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 30 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

    eess.AS 2026-08 conditional novelty 7.0 of 10

    SemBridge supervises autoregressive states with discrete semantic tokens during training, improving content fidelity of continuous-latent speech generation without changing inference.

  2. When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus

    cs.SD 2026-03 unverdicted novelty 7.0 of 10

    LRLspoof corpus and threshold-transfer evaluation demonstrate that spoof detection performance varies markedly across languages, identifying language as an independent domain shift factor.

  3. ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

    cs.CL 2025-05 conditional novelty 7.0 of 10

    ArVoice is a new 83.5-hour, 11-voice Modern Standard Arabic speech corpus with diacritized transcripts for multi-speaker TTS and voice conversion.

  4. Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption

    cs.MM 2026-07 conditional novelty 6.5 of 10

    Energy of text-to-video diffusion models is predicted from architectural first principles and observable generation parameters with under 3% MAPE, without needing weights or model size.

  5. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  6. DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

    cs.CL 2026-06 conditional novelty 6.0 of 10

    Block discrete diffusion over X-Codec2 tokens yields competitive zero-shot TTS with 0.6B parameters, 20K training hours, and a 0.15 real-time factor.

  7. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0 of 10

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  8. HISPASpoof: A New Dataset For Spanish Speech Forensics

    cs.LG 2025-09 conditional novelty 6.0 of 10

    HISPASpoof is a new large Spanish synthetic-speech dataset for detection and attribution, with evidence that English-trained detectors fail on Spanish and Spanish training helps.

  9. UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A unified audio-video generator built by stitching pre-trained video and music diffusion models, trained on 7,600 hours of data, with a new evaluation benchmark.

  10. MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening

    cs.SD 2025-08 conditional novelty 6.0 of 10

    MoTAS combines TTS speech augmentation with MoE-guided feature selection to reach 85.71% accuracy on ADReSSo, the highest among the baselines listed.

  11. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  12. Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    By fine-tuning an LLM with audio and motion tokens, MECo generates co-speech gestures that follow a user-provided motion example, and reports state-of-the-art FGD and diversity on BEAT2 and ZeroEGGS.

  13. Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A transcript-prompted Whisper model with dictionary-based decoding automatically produces phonemic and prosodic annotations for Japanese audio-transcript pairs, improving Japanese TTS naturalness.

  14. Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning

    cs.SD 2025-04 conditional novelty 6.0 of 10

    XS-CoT trains speech LLMs to answer non-core language questions by generating an English reasoning chain before the final target-language answer, and a semi-implicit variant compresses that chain to cut latency.

  15. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

    eess.AS 2025-02 conditional novelty 6.0 of 10

    GenSE enhances speech by first denoising semantic tokens with a language model and then generating acoustic tokens from a single-quantizer codec, reporting higher DNSMOS, speaker similarity, and lower WER than prior systems.

  16. TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

    cs.SD 2024-12 conditional novelty 6.0 of 10

    TouchTTS reports a 51.6% data retention rate using a noise-robust tokenizer and two-ASR cross-validation, and a Qwen-backbone flow model that unifies streaming and non-streaming synthesis while matching CosyVoice on PER.

  17. Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Faster IndexTTS-2 compiles all components of IndexTTS-2 into TensorRT engines, cutting end-to-end latency by up to 3.6x while adding streaming and batching.

  18. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  19. ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

    cs.SD 2026-03 conditional novelty 5.0 of 10

    A domain tag marking synthetic training audio plus 3× oversampling of real audio lets a TTS model absorb large synthetic augmentation without losing speaker similarity.

  20. DarkStream: real-time speech anonymization with low latency

    eess.AS 2025-09 conditional novelty 5.0 of 10

    DarkStream combines a causal content encoder with limited lookahead, k-means quantization, and a GAN-based pseudo-speaker embedding to anonymize speech in real time with near-chance speaker-verification error rates.

  21. Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A speech-deepfake dataset for ten public figures built with transcription-based segmentation reports high synthetic naturalness (NISQA 3.69) and a human misclassification rate of 61.9%.

  22. UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

    cs.SD 2025-05 conditional novelty 5.0 of 10

    The authors propose DistilCodec, a 32,768-code single-codebook audio codec, and UniTTS, a Qwen2.5-7B TTS model trained with audio, text, and cross-modal autoregressive tasks on interleaved prompts.

  23. FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech

    eess.AS 2025-05 conditional novelty 5.0 of 10

    FlexSpeech is a zero-shot TTS system that predicts phoneme durations autoregressively, renders speech with flow matching, and applies direct preference optimization to durations for fast style transfer.

  24. LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A series of Qwen2.5-based speech chatbots (0.5B to 14B) that stream speech output with a CosyVoice2-style autoregressive decoder and outperform prior SpeechLMs using only 200K synthetic multi-turn dialogues.

  25. IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

    cs.SD 2025-02 conditional novelty 5.0 of 10

    IndexTTS combines character and pinyin modeling to make Chinese polyphone pronunciation controllable, and reports improved zero-shot voice cloning and naturalness over open-source TTS baselines.

  26. BoSS: Beyond-Semantic Speech

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.

  27. Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A speculative decoding variant with a tolerance-based acceptance rule speeds up CosyVoice 2 inference by 1.4x while keeping subjective quality comparable.

  28. Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget

    cs.SD 2025-04 conditional novelty 4.0 of 10

    Muyan-TTS, a 3B-parameter LLM-based TTS model trained on 100,000+ hours of podcast audio, produces competitive zero-shot speech and runs at 0.33 seconds of inference per second of speech.

  29. Position: It's Time to Act on the Risk of Efficient Personalized Text Generation

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Fine-tuned open LLMs can imitate individual writing styles from small samples, evade detection tools, and are not yet addressed by current safeguards or law.

  30. WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration

    cs.SD 2025-08 conditional novelty 3.0 of 10

    WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.

Pith tools