Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2409.00750.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:03.299001Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T12:19:49.609343Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 086fc22e-91a8-4ecc-9e4f-0f74191a7385 · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2e907032-4eca-4045-a25a-8266e682f8a0 · inbound
WavChat: A Survey of Spoken Dialogue Models MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 217
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac6d73c-4f62-49b0-a77c-5a5d04a92fbf · inbound
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca3a5fb0-7de2-4f11-abd9-268cc17b4ef2 · inbound
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67747de-bf49-4491-84c0-2700e7c1cf2e · inbound
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 829c2aa3-ad90-4160-9ef9-f513bb06313b · inbound
Overview of the Amphion Toolkit (v0.2) MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 119
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 637cb4a9-52c4-4d75-abe9-81f326f42521 · inbound
Emotional Face-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d8a91e-e789-4371-b7d5-103e43f7811c · inbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80fd235d-9dec-42a2-8da9-4a87396ebd46 · inbound
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da8156d4-ecc8-44b2-a160-ff66a880bfa4 · inbound
SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c26f3345-058c-47a6-927e-1fc6efdd23e9 · inbound
Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4648c6c9-9105-46a3-9050-85ee6376b56d · inbound
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e71f19e-d50f-4ae4-8075-535d96840a6a · inbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 323ba844-defb-424d-9b4d-76a2e892d01f · inbound
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e70ec91-873c-42c2-894f-28e9d834de26 · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 03333264-800a-4980-89ee-16fd2c942e25 · inbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba44ab7-e2ce-4b97-b982-f296b1bfdde9 · inbound
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88db68ac-66a9-450a-8a21-beb069a6fa7f · inbound
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee274426-26c7-41a7-83ea-6e304359fe61 · inbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22e0c373-417c-4506-8d24-986b3f37aa1c · inbound
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 663c1c04-c3e4-47e7-9303-e95d307ed5f0 · inbound
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04e4ef9c-3e63-4826-b1d8-4ba7a760e217 · inbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824d47ee-60fe-4859-b56e-c26291a5af5e · inbound
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa92c4b-ae16-48b0-9694-c7636e770379 · inbound
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a32f064e-3d61-4988-8f59-e4bee014f2fa · inbound
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3cbe4af-0b4a-4535-b907-5547b8ff3d60 · inbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63183282-1e75-44c9-94dd-9a9f6e44dfa2 · inbound
Step-Audio 2 Technical Report MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 752c8ba0-a959-4247-8d1e-c3737aee4cf3 · inbound
Adaptive Duration Model for Text Speech Alignment MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd815554-0fb7-4877-a0b9-f2406c475e7b · inbound
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ccd477-b2ae-4e42-ac81-b7a3e019362b · inbound
Entropy-based Coarse and Compressed Semantic Speech Representation Learning MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d8e8d91-3e33-49eb-8da4-a442370d71d5 · inbound
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 751bb11e-20cb-4fe1-a0c7-cc73a5d917cc · inbound
Audio Deepfake Verification MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93e93911-012d-41c8-95b2-a035cc226d4f · inbound
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f634f89e-bfc2-4b66-8aa0-d610f9a94a59 · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45072d0a-4ed1-4e45-8486-d449f290cad2 · inbound
TokenChain: A Discrete Speech Chain via Semantic Token Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8f2ccafa-0bd7-47fd-bd41-f2e79f8858f2 · inbound
Qwen3-TTS Technical Report MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation be168de6-39a7-449c-9b49-70883c5b8409 · inbound
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b7b74c95-bbe2-4c41-8ceb-22ebd5ff1b46 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fae34acb-80fa-4cd5-a6f1-5152c13aca10 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a5a0c1-14fd-4266-8ec7-75aa23856e8f · inbound
MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cd698a02-fa82-4e54-85bc-f48c8cc23020 · inbound
Hierarchical Codec Diffusion for Video-to-Speech Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9548e9dd-0c86-4d8a-8d3e-e1dd3b2ecd07 · inbound
Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2587db95-d30e-4031-941a-ee853a6e8982 · inbound
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a8e0ff9b-a39d-4744-837e-4ce17ed55243 · inbound
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b8583be2-aef0-4be5-98d7-8b637f38461b · inbound
Taming Audio VAEs via Target-KL Regularization MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d84b2da4-4d8a-4534-98c4-d35e80872289 · inbound
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 878f64a7-6f99-48d7-b9e5-f790dcad259a · inbound
SegTune: Structured and Fine-Grained Control for Song Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2ee16ea4-22a0-4f0d-8c81-d09cd0a93592 · inbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 52ae7175-3f4a-40bf-9a0c-b0535720a546 · inbound
TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 372151ba-879e-49d0-939c-093574a86af6 · inbound
FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3adf4026-52b5-4038-87fa-1fee4aa4c4b5 · inbound
End-to-End Training for Discrete Token LLM based TTS System MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 026bcbb9-2253-4d3a-bbbf-3a0715c81c4c · inbound
An Evaluation Framework for Text-to-Speech Voice Reconstruction MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d10e45ad-c73b-40fe-90c6-c06deea7f08e · inbound
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c04579a6-0ded-4ff9-907c-34885ad06637 · inbound
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 044e4142-2a3a-43d1-8098-8415a551a242 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 202
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 418d0736-613b-4ee4-98d4-f3e77d558404 · inbound
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 60ab2a50-c696-4b5d-905b-9e54efd22029 · inbound
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed90ffa-6d2f-40e0-b4aa-a8a78347f78c · inbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4423d7a2-cb63-4f14-a32e-0d9655df8656 · inbound
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 158
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc0ab8f-49f9-4847-84ec-fe1532ef48a4 · inbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2eecb1a-ea0b-4bf8-93c3-ff37f5c41f39 · inbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae40bb28-6d73-43a9-a72c-fa9cccee1085 · inbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.