Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T02:47:38.049516Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2608.00011.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T02:47:38.049516Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T02:47:34.168653Z
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation be70cc1e-e6cd-41a6-8a24-7a9dfbca4a4b · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Autoregressive codec language models [1, 2, 3] achieve high-quality zero-shot synthesis but require 60K–250K hours of data and generate tokens sequentially, incurring high latency
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 518e8a65-a969-4a42-b294-fea1704f6e7a · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1be558b-a6d2-4965-b088-020f8ef2de2f · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Background Neural Audio Codecs.Neural audio codecs compress contin- uous audio waveforms into discrete token sequences through learned quantization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d69a8f-de82-416a-b0f5-5d4524f7f3f9 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f71871-7bb4-43a3-b3a9-093f60121af6 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21ccdb6-ac3f-49c0-9972-a7fc04b454ef · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis All research contributions, including the methodology, experimen- tal design, results, and scientific claims, are the authors’ own
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beceb961-0f28-4d3f-ad6c-29600f2490ce · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee3d8f0-ae27-4ffb-a62f-2f41c2b2fb4c · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb01b07b-0bd0-496f-804b-1a1b2e42f3fb · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aeebb4c-2152-4128-a65a-1a9aeef37e13 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis V oice- box: Text-guided multilingual universal speech generation at scale,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 599d5572-5398-415a-9896-66d7c5a80a71 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fdb55c4-8aee-452c-98d0-69c556dc568e · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33813783-458b-47a1-b02a-7b41d02c4a8f · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdc947d-a354-4ade-8063-f082075dfdf3 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Simple and effective masked diffusion language models,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcecf7e3-3f49-49a4-9347-f8da11b38c47 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Large Language Diffusion Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e660b47-e65a-4800-a25c-6dc16e0e653a · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Block diffusion: Interpolating between autoregressive and diffu- sion language models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1dbf8d9-fd18-4917-9f24-fb8e92986e97 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c258d91d-e565-4bd9-b6d9-10a4bda65405 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9adef03b-cd60-4f95-91c2-458aa7e115f0 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis DiTAR: Diffusion transformer autoregressive modeling for speech generation,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f74d0e87-b1ef-4f30-bd3f-791758b35ef9 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aca05a97-dab2-477e-bf51-19acde3e2781 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis SoundStorm: Efficient Parallel Audio Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1562ad06-8805-48b2-a8ad-4d7b16698ae6 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d4e5d1a-cca7-4f84-a52a-71383fc85b7c · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis High Fidelity Neural Audio Compression
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e731f3ab-d26c-4e85-bbb0-177d8a75886d · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Qwen2 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 415c7e6d-50c9-4f9b-adaa-8187600a2456 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis RoFormer: Enhanced transformer with rotary position embedding,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552f8c7f-7728-4854-8caa-0392f27ae99d · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655011b0-91a7-4bfb-a0d0-9390d8bd3fda · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Decoupled weight decay regulariza- tion,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85777bd1-ba21-46c9-a4cf-ad1ff0df5a93 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Robust speech recognition via large-scale weak su- pervision,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c40261d-3e0b-4992-a38f-9dd41b552a80 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis WavLM: Large-scale self- supervised pre-training for full stack speech processing,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 193c228c-c04a-4052-9eea-674c8ac8d391 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis CodecMOS-Accent: A MOS benchmark of resynthesized and TTS speech from neural codecs across English accents,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28686fc6-7cc7-4693-887f-911c13a40eba · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 973c327d-e420-4af0-a9b0-eea58a97c956 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Qwen2.5-Omni Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed90ffa-6d2f-40e0-b4aa-a8a78347f78c · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8f2a09-c0b0-41bc-9593-8ebc6ba8cb21 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6310a5c4-2395-4954-bb9d-84ea5c281897 · outbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 518e8a65-a969-4a42-b294-fea1704f6e7a · inbound
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.