Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2411.01156.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T06:05:11.377969Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:39:38.706029Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 14708426-a94d-4b89-b445-2ee36eab8058 · inbound
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c711ee6-7141-46e4-a003-909bbdfd5f41 · inbound
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56fb5224-3e15-411f-9401-d2b9b29b2629 · inbound
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cfcd3fe-219d-4502-a5fa-f0bcea9aa9f3 · inbound
Position: It's Time to Act on the Risk of Efficient Personalized Text Generation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 861424a8-255a-49e2-a223-0064ca68b613 · inbound
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c0d5022-b7ef-4d7b-bd24-85b85705c8cd · inbound
Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6851a7f-c87b-480b-ab04-3ceaf0e81ce5 · inbound
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ed7c07-7ba5-4896-856a-ab21e0eb11b6 · inbound
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8383ff8-ddd2-4fd7-8b0c-546f08809e16 · inbound
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bafbe8f-77ba-41e9-bc48-584777613c6e · inbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f620d47a-d187-42f4-93e0-6e13081bcf3c · inbound
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f8a2ca-2413-402c-a719-cf2fe2236001 · inbound
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5a0f09-b588-4d3b-be6d-6f1fa0693359 · inbound
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5804bb3-fddd-4e20-a57e-7810b4c9c2af · inbound
BoSS: Beyond-Semantic Speech Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f21bb22d-6379-4d43-9121-f091d7b085d5 · inbound
Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80da568d-b212-4b0d-b9dc-962443201971 · inbound
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cada4e6f-d952-465e-8a7c-276f224b2f62 · inbound
MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7c9d51-f42b-4e92-8c76-cb2f56566c67 · inbound
WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b11f19d-93af-49a8-a8ea-d4fcb7aed126 · inbound
DarkStream: real-time speech anonymization with low latency Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38140346-be0d-4da8-8620-30bce8c6ea37 · inbound
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d3cb53e-2dbb-4ef6-8259-6ad3e65b75d7 · inbound
HISPASpoof: A New Dataset For Spanish Speech Forensics Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a836ee88-0739-48bc-ad6b-83f6c2c22224 · inbound
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aa1182b9-20f7-4b6a-a44f-c777dfcd9268 · inbound
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d7f06d45-5136-43bd-a656-a0dbfa48de7d · inbound
Qwen3-TTS Technical Report Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8f943109-eb0b-4c6d-ab35-2049e1e068d4 · inbound
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0fb93cd-1be2-4fb5-b53c-ca71ed9541f1 · inbound
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 074b5629-107c-4113-aa8c-b87eb816ccde · inbound
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca7150ad-4923-4f01-b7b0-f6e190f72502 · inbound
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf9ea63e-d386-4fff-98af-102418eae34c · inbound
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9bf6fc20-e0d0-4bbf-980b-5129030961ee · inbound
Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 47cbbba5-e1a2-460e-b72a-137d87a9c5b3 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f348cb99-400d-4a35-b18c-dc165d39c950 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0349b6f-db83-4856-a82b-da36a5808452 · inbound
NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bd7d2443-a996-4c83-a19f-54f5998998e0 · inbound
RTCFake: Speech Deepfake Detection in Real-Time Communication Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e93d7b81-8cf5-4449-ade3-78e53c56bfe5 · inbound
V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c12b94b0-311e-4bcf-8a7e-1a18febcd28d · inbound
Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8f7d2581-11fe-4319-930e-df8a9269f01e · inbound
Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 45966ad9-4f26-4fb6-92e8-e5cb3328bc24 · inbound
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5772ce05-c8ce-4ede-9068-4f767d606373 · inbound
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 456d7a1a-912c-428f-a5a1-474ab12e7993 · inbound
FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8e9be2d8-7655-49ef-870b-50849499dc90 · inbound
An Evaluation Framework for Text-to-Speech Voice Reconstruction Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf62a5da-8fb4-4b81-92b8-9318cda5d810 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0efa5e6a-eb1c-4f4e-9f7b-b7be0635e344 · inbound
Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96e8430a-ea47-465d-8339-2ee17845339c · inbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6310a5c4-2395-4954-bb9d-84ea5c281897 · inbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fda3f1f-e7e8-4272-9da1-7fd420f65fcc · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3559cc-902c-4a5e-8aa7-a43b9f23cc53 · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f07b94-66e8-42f0-970a-a83355966612 · inbound
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.