Pith. sign in

Paper Citation Record · LEDGER

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2608.00011.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00011 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T02:47:38.049516Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T02:47:34.168653Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be70cc1e-e6cd-41a6-8a24-7a9dfbca4a4b · outbound

This paper cites Autoregressive codec language models [1, 2, 3] achieve high-quality zero-shot synthesis but require 60K–250K hours of data and generate tokens sequentially, incurring high latency.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Autoregressive codec language models [1, 2, 3] achieve high-quality zero-shot synthesis but require 60K–250K hours of data and generate tokens sequentially, incurring high latency

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:34.082711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:34.082711Z digest=sha256:bb96cb9fa6734802ade3cdd26ace919cd2065b5086ea011f02bb35d51cfa595a

Observation 518e8a65-a969-4a42-b294-fea1704f6e7a · outbound

This paper cites DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:34.168653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:34.168653Z digest=sha256:3aeb4e8345e97716a19883efd26b52c54b00aa07fd5c3cdaed4e08977be638d1

Observation a1be558b-a6d2-4965-b088-020f8ef2de2f · outbound

This paper cites Background Neural Audio Codecs.Neural audio codecs compress contin- uous audio waveforms into discrete token sequences through learned quantization.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Background Neural Audio Codecs.Neural audio codecs compress contin- uous audio waveforms into discrete token sequences through learned quantization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:34.330382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:34.330382Z digest=sha256:7131fd499355648b48acbd7ae18a2c9292cf5b4cc68f96884ad632972f9a1b27

Observation 74d69a8f-de82-416a-b0f5-5d4524f7f3f9 · outbound

This paper cites an unresolved cited work.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:34.475935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:34.475935Z digest=sha256:1d88aa84eb7eef6823812d4990e26db2089259f4acf47f2d03d58220c44cb5db

Observation 77f71871-7bb4-43a3-b3a9-093f60121af6 · outbound

This paper cites an unresolved cited work.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:34.651465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:34.651465Z digest=sha256:4c7968a40996c0ea1a5fe768e302acdb30896c9b4df03c5d5be5fd568222d608

Observation c21ccdb6-ac3f-49c0-9972-a7fc04b454ef · outbound

This paper cites All research contributions, including the methodology, experimen- tal design, results, and scientific claims, are the authors’ own.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis All research contributions, including the methodology, experimen- tal design, results, and scientific claims, are the authors’ own

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:34.844134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:34.844134Z digest=sha256:6f913cd424ad2e77337c2f07f383f1cf20a0d94cc83409f90b4f6b494da46bbc

Observation beceb961-0f28-4d3f-ad6c-29600f2490ce · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:34.991882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:34.991882Z digest=sha256:3c5b9655c3a674dd99d505af89cb8cae9b8b15c2d25ea9b1ce12b55e53b82585

Observation 3ee3d8f0-ae27-4ffb-a62f-2f41c2b2fb4c · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:35.142024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:35.142024Z digest=sha256:72f0456e2815038a2182c15ce51e398a8e914e25949608e4006c6786d3052de2

Observation eb01b07b-0bd0-496f-804b-1a1b2e42f3fb · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:35.291836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:35.291836Z digest=sha256:1d7b9efbbcefb4a139260cc7a7cfa14a403b4a63ddcadecf608384e8b85ebb1e

Observation 4aeebb4c-2152-4128-a65a-1a9aeef37e13 · outbound

This paper cites V oice- box: Text-guided multilingual universal speech generation at scale,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis V oice- box: Text-guided multilingual universal speech generation at scale,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:35.434923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:35.434923Z digest=sha256:372c5dc1311c8350a332cd13a322bc2a01942af09c29e91bc00c778a1a9a17fa

Observation 599d5572-5398-415a-9896-66d7c5a80a71 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:35.599102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:35.599102Z digest=sha256:6af75ef7f4a01df33cea85573dc265775f97c96878fcf14747989b57e5970c4b

Observation 3fdb55c4-8aee-452c-98d0-69c556dc568e · outbound

This paper cites NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:35.745081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:35.745081Z digest=sha256:953c95949ab720bd1ab67e7a02cdc04dcde301732887639c79dd4d40f7905614

Observation 33813783-458b-47a1-b02a-7b41d02c4a8f · outbound

This paper cites StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:35.918355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:35.918355Z digest=sha256:4a3fe81fe4f2f22a820bfb9afb8678b8928460fb23c78095101777f5285de164

Observation 1fdc947d-a354-4ade-8063-f082075dfdf3 · outbound

This paper cites Simple and effective masked diffusion language models,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Simple and effective masked diffusion language models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.067536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.067536Z digest=sha256:10ed1b07d0c8309bd64a10b4c52792e70ca2667985d8c2c3131f36f3ca0577a2

Observation dcecf7e3-3f49-49a4-9347-f8da11b38c47 · outbound

This paper cites Large Language Diffusion Models.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Large Language Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.152258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.152258Z digest=sha256:c9373d08a15b2b52df7e2a611b8387347ddfefa3bcb0920fccce19db9edd5a69

Observation 3e660b47-e65a-4800-a25c-6dc16e0e653a · outbound

This paper cites Block diffusion: Interpolating between autoregressive and diffu- sion language models,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Block diffusion: Interpolating between autoregressive and diffu- sion language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.286198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.286198Z digest=sha256:a9b8d8ea31a142b97532f51a8c433ccaf0bb9ecafdf173a86471761a1b2d4699

Observation c1dbf8d9-fd18-4917-9f24-fb8e92986e97 · outbound

This paper cites Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.419427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.419427Z digest=sha256:847d3775ca5f20b3e5cc6dad1979c4751dcfc54da4d6fe5238889a22755f9045

Observation c258d91d-e565-4bd9-b6d9-10a4bda65405 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.533889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.533889Z digest=sha256:12722a5dd52b5f366661723ce92980803005ee3bbdada5ce168ed9a89aeb2e88

Observation 9adef03b-cd60-4f95-91c2-458aa7e115f0 · outbound

This paper cites DiTAR: Diffusion transformer autoregressive modeling for speech generation,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis DiTAR: Diffusion transformer autoregressive modeling for speech generation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.620536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.620536Z digest=sha256:f8c0c3d2c74a9cfa1118afa76e7dab3a06ade2920c0ceacb95751b890dc3111c

Observation f74d0e87-b1ef-4f30-bd3f-791758b35ef9 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.673778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.673778Z digest=sha256:9f6658ca49bd02777116d461e6c16ddd0c709ef71b3f55a1eb890821aa1ceebf

Observation aca05a97-dab2-477e-bf51-19acde3e2781 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis SoundStorm: Efficient Parallel Audio Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.770711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.770711Z digest=sha256:b57434558359196a93504f848ef73c2faffcdce900c18707aa3bb74b340308e9

Observation 1562ad06-8805-48b2-a8ad-4d7b16698ae6 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.903349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.903349Z digest=sha256:01884f322b63ac97f9d4bd36d4d4412d942f9bba9016bb8e15d20e24c47b1714

Observation 0d4e5d1a-cca7-4f84-a52a-71383fc85b7c · outbound

This paper cites High Fidelity Neural Audio Compression.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis High Fidelity Neural Audio Compression

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.015401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.015401Z digest=sha256:46c55b414ddb42beffbe6d60150999260ff4c6b46c2e5eee836f48f8f909bc2d

Observation e731f3ab-d26c-4e85-bbb0-177d8a75886d · outbound

This paper cites Qwen2 Technical Report.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Qwen2 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.124600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.124600Z digest=sha256:68691e424bb0bcad41924d15ea943326b29bdb9b20225c543ff94c88beaa3419

Observation 415c7e6d-50c9-4f9b-adaa-8187600a2456 · outbound

This paper cites RoFormer: Enhanced transformer with rotary position embedding,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis RoFormer: Enhanced transformer with rotary position embedding,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.218585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.218585Z digest=sha256:4889a54c4ed54c5d6fcd58b446043726b5b48aff6307e7ec47c81f51de8fe004

Observation 552f8c7f-7728-4854-8caa-0392f27ae99d · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.291105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.291105Z digest=sha256:2bf9c2d9523105fb416f765eec2d77ebd0a32bc880956c4adc890a99d2818c90

Observation 655011b0-91a7-4bfb-a0d0-9390d8bd3fda · outbound

This paper cites Decoupled weight decay regulariza- tion,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Decoupled weight decay regulariza- tion,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.395690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.395690Z digest=sha256:d6988e6bd0c33f0f16c028a3f95990e810bd4de0bf40f5aeb61356fcb0683c0d

Observation 85777bd1-ba21-46c9-a4cf-ad1ff0df5a93 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Robust speech recognition via large-scale weak su- pervision,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.470170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.470170Z digest=sha256:46db096b7f29e78c57efc00f92460860f02f9867d75e4e2d94c568907eb747ac

Observation 8c40261d-3e0b-4992-a38f-9dd41b552a80 · outbound

This paper cites WavLM: Large-scale self- supervised pre-training for full stack speech processing,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis WavLM: Large-scale self- supervised pre-training for full stack speech processing,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.581829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.581829Z digest=sha256:b6860c023241cf95c3003d847f4b429d02c70ce3e943fe951f5cafb03231fd92

Observation 193c228c-c04a-4052-9eea-674c8ac8d391 · outbound

This paper cites CodecMOS-Accent: A MOS benchmark of resynthesized and TTS speech from neural codecs across English accents,.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis CodecMOS-Accent: A MOS benchmark of resynthesized and TTS speech from neural codecs across English accents,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.653775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.653775Z digest=sha256:90415ad5329c8b98bad4756b691823f48ce0041e2e8ba6eb8fca13f200522606

Observation 28686fc6-7cc7-4693-887f-911c13a40eba · outbound

This paper cites IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.726590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.726590Z digest=sha256:17d982b330631d319148b9f60092870727b230895d3826f80e07acce0f9c6062

Observation 973c327d-e420-4af0-a9b0-eea58a97c956 · outbound

This paper cites Qwen2.5-Omni Technical Report.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Qwen2.5-Omni Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.800808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.800808Z digest=sha256:d3f6eb2508a5262a559fde531a91e9501600d928f3d3a8254faad07e1e89c176

Observation 0ed90ffa-6d2f-40e0-b4aa-a8a78347f78c · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.876645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.876645Z digest=sha256:4d5dab32dcfd98c7623faea60e02aaf5f85b49d072cf7f404765d4f096ef7e58

Observation 2c8f2a09-c0b0-41bc-9593-8ebc6ba8cb21 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.938205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.938205Z digest=sha256:881ff63d6262728e52728112d66be2cb0dec723ba7e124279f545c92783e193c

Observation 6310a5c4-2395-4954-bb9d-84ea5c281897 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:38.049516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:38.049516Z digest=sha256:599f6c9ea8c96ab126dd532f548513ea35012ff4ef8afc857a76b977673b7698

Pith citing papers

Observation 518e8a65-a969-4a42-b294-fea1704f6e7a · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:34.168653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:34.168653Z digest=sha256:3aeb4e8345e97716a19883efd26b52c54b00aa07fd5c3cdaed4e08977be638d1