Pith. sign in

Paper Citation Record · LEDGER

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.22362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22362 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:12:00.395703Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:11:56.263021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:12:00.605940Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · outbound

This paper cites DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:12:00.657415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:56.263021Z digest=sha256:4c0b47b443d4227695770936453755596f879204149085f758534547f765c3d8

Observation 8df8c223-4cfd-4c6b-82d9-88e276c36f6b · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.884534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:56.378331Z digest=sha256:06b607b0906210615dc09c2732d5da8dbc91b3d5e21c381473afbe98a429dcc3

Observation 61ed7bf7-03e8-4cb2-92e4-ad7026f60aaf · outbound

This paper cites SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.545419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:56.883055Z digest=sha256:ed734350a008f61d8a7ff95d2eb58c4d63f1701e2f72b6c22f2fcfb13527e307

Observation 58d2b895-c15f-4d5c-979d-86e1565358a0 · outbound

This paper cites There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.432611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:57.068976Z digest=sha256:b7c121c9a1d573b5594b6f726255a2f6c9146f380475fe6cfb011da78ffda2a2

Observation 697cd690-a421-48a1-b57a-f10af5163824 · outbound

This paper cites High Fidelity Neural Audio Compression.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:57.819506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:57.819506Z digest=sha256:df1e5740ef7793a01a05a2e76ab2d4b2bf7403e1771c97d3aa775927bbf0aca2

Observation 470f8e5c-fd4b-44f3-a4ed-73491a7869a3 · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.785845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:56.556586Z digest=sha256:28af7fdaf4783444fd4e91f1fd396d4bc3fdf4c9b197481d07f08097dd2aa67b

Observation cb89ca61-b5f7-4f40-bbcf-3a46db86612f · outbound

This paper cites HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.326298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:57.215975Z digest=sha256:b4c8584c0fd7860c3ed04f3051ada9e05f0e557c445fdba0e55840b54f1f491c

Observation 7f71171f-daf1-4a7f-be44-40d65d974f96 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.192372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:57.401222Z digest=sha256:0584907827d8bb2cabf6e3e195a20249e54142d5938fc709bc372fb39d9fd310

Observation d5a9493a-cb89-4734-9dbf-e5e0eeb6fddf · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:57.528724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:57.528724Z digest=sha256:694277fe116dc56012f39d40904b10ae4fb7b2925aa3d1da8ff29df5475faa1d

Observation fd77f7af-be89-4f03-bdd7-70e59f5c7aab · outbound

This paper cites w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.084487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:57.697694Z digest=sha256:0ec866a490e49819b143b65fbedb3769430a1bee818c7066be5136f9777d42d8

Observation 28cb9e28-d715-4793-8603-bff255369fec · outbound

This paper cites PolyV oice: Language Models for Speech to Speech Translation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding PolyV oice: Language Models for Speech to Speech Translation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.407814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:58.332796Z digest=sha256:5ea79356c4b2422c3b61190d907a1da78fb5a900b0f90f33d4b781fd45b00fa2

Observation 2ae3cbb9-8dce-49cb-8963-43e3a5d44fe8 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Soundstream: An end-to-end neural audio codec,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.950584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:57.887238Z digest=sha256:0d9f8842b2825a9e7a5a597eace972622e2a70e725a0c464fd5614eb4d739129

Observation 5616e741-3b16-4a35-b52a-46ffd11ba287 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding AudioLM: A Language Modeling Approach to Audio Generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.753500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:57.968067Z digest=sha256:9bb2e0adfb7b40a6af1e6ca784ec339d18105729d318f30d7cf475a959d7928a

Observation f371b50b-589e-4282-903e-299acaccadd5 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.039438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.039438Z digest=sha256:3cb9293eaa5c1c4b2344699419793e1358950ad208ff9918f3ec898860a76b4d

Observation f6314508-c797-4b3c-b0c6-6201c39e2af5 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.112443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.112443Z digest=sha256:ffbea8e8084729cea85381e5198f2e8a0470fd6b84d6620120b00c9150d5698a

Observation 53225d5e-261b-47b4-b570-6f689c6a2011 · outbound

This paper cites TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.590748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:58.204048Z digest=sha256:85ec639ce0e51d05c1f276576494efefa0ec7a7e7fa8a81ba23546fc450bdedf

Observation 741d6f2e-ded0-471f-8981-7d2a8d44a1cd · outbound

This paper cites Denoising Diffusion Probabilistic Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Denoising Diffusion Probabilistic Models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.658852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:58.800630Z digest=sha256:e568bab60bb545c3587eb1f6d30827393f6f044ac4d7c60b0acada28ba7cce12

Observation 541cf954-a92b-44d0-ac2b-02138a4b836a · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.428052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.428052Z digest=sha256:f616a5e48a78a51b1bac1a11d5caf7db6dfe34536a0394e21527db643f889e47

Observation 346f092b-0acc-4437-bcc2-9fd0777a651b · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Moshi: a speech-text foundation model for real-time dialogue

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.489098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.489098Z digest=sha256:037b95538d3213cd611d3b6ad0fb9c1a19cfc1f5500f22a3a0bce6e8a92f187c

Observation ba766575-ff7d-4c67-8a2f-1b753275e311 · outbound

This paper cites Autoregressive Image Generation Using Residual Quantization,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Autoregressive Image Generation Using Residual Quantization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.193477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:58.579941Z digest=sha256:9aa39d0ed500777d9850c7f4a8ea9804b902e09a9511b13a2ae145ed28ec3876

Observation 0503de73-1854-4765-b019-aa31e7982554 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.004761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:58.652586Z digest=sha256:0c8e0bf5fda79d8c2287b2e6c6d783cc64d29aecd0246c2c4247ec0ea722314b

Observation 2eee440f-7147-4fdf-b57b-56e8a377ca85 · outbound

This paper cites Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.825848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:58.731676Z digest=sha256:32b3f01ae0b1d809fe354fe37576d655ee57da8ca3e01a5c6e84153b10747efd

Observation b518cc89-7d79-492f-b39e-9406b9a98ced · outbound

This paper cites Auto-Encoding Variational Bayes,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Auto-Encoding Variational Bayes,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.760597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.303877Z digest=sha256:f39bed94871966919999ffdb1efccc164e028a62a66c9144f775d371690dd4e7

Observation 29dbd204-1178-4dc2-aa33-9055e61bcc79 · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermody- namics,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Deep Unsupervised Learning using Nonequilibrium Thermody- namics,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.492841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:58.879689Z digest=sha256:bfeb1a158527a5de1d59417614454d5cb61f969a335e7d0480ae049d96f89650

Observation 415f794e-88e2-4cbb-9868-005b8a65a511 · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.654095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:56.737595Z digest=sha256:7fed717bcb9e1a572684b81de71d0f17e144b0984f9bbe9374a35ee303dc7712

Observation 75e26e44-6ab8-478a-9520-2502c3b003aa · outbound

This paper cites Multi- step Distillation of Diffusion Models via Moment Matching,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Multi- step Distillation of Diffusion Models via Moment Matching,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.345517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:58.962209Z digest=sha256:a6cbfd38f30d925a62b57048df298e1b14d32fc3cd59533670c63eabb544153c

Observation 09f42aa1-4ebc-4c2d-a76a-6ec7a3930a25 · outbound

This paper cites SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.154376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.044535Z digest=sha256:910718ed41b5b9e292d659df1c4e9247353682bc07a5fe03189304359ede281b

Observation 1a30a0b0-92f1-4a23-8d32-43abdcd71bd9 · outbound

This paper cites Simple and Controllable Music Gen- eration,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Simple and Controllable Music Gen- eration,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.959308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.138687Z digest=sha256:a4c5b5e9eab012c5078f861ff339b3c6b070fed1094cd37fb9133434a5aa6465

Observation d6b0c266-afa8-44f9-8036-d6e1b059ab5c · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SoundStorm: Efficient Parallel Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:59.228454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:59.228454Z digest=sha256:38017d6d1db5e29313a7b666434458d2a32bb096905b9ec4e809167a43052322

Observation 57ce0ac9-6869-4c4a-ad8d-8a7f7314c980 · outbound

This paper cites FiLM: Visual Reasoning with a General Conditioning Layer,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding FiLM: Visual Reasoning with a General Conditioning Layer,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.583977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.401405Z digest=sha256:280ba979e94ed62cebf6a460218f0d9c1dbc435f54062cd91a963079613152d4

Observation 5ff8a0bc-abcd-4e0b-bcd1-eb5534083d2f · outbound

This paper cites High-Resolution Image Synthesis With Latent Diffusion Mod- els,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Resolution Image Synthesis With Latent Diffusion Mod- els,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.376140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.467113Z digest=sha256:95bd4957c43fa0e4b3bb57a742adcf5d06f72d268dfed684bf646ed2c311d301

Observation 77eab515-24e4-4ea3-b1e9-cd5e767fa542 · outbound

This paper cites Improved Denoising Diffusion Probabilistic Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Denoising Diffusion Probabilistic Models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.220833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.562721Z digest=sha256:5e1ed9b5806264c059917ad7e5d38db5d62666e82959245c876d09a8c50ebd96

Observation 67e37efe-fc40-4e9b-992e-6b321fc40388 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Progressive Distillation for Fast Sampling of Diffusion Models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.030022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.627762Z digest=sha256:154e0e1a79ad3111d252cf3533e273be97639607dc88847e1d8869f8355a5e6d

Observation 43f4d6bd-f5e9-4c83-a00c-da998cabf0cc · outbound

This paper cites Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.835199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.723665Z digest=sha256:3410ce23374e7c06846c6d74cd66a9811c1187e8b8c4869758d2264286899a45

Observation 98e74125-b74a-4b75-8179-05ebc6103f4f · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding WaveNet: A Generative Model for Raw Audio

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:59.809510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:59.809510Z digest=sha256:13bc5bec36d465b030a394babbfcab6564c4206da8027a35d91f96ba7a185fd4

Observation ee08011d-db9e-4b47-af8a-a1ced36515b6 · outbound

This paper cites Improved Distribution Matching Distillation for Fast Image Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Distribution Matching Distillation for Fast Image Synthesis,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.628591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.912833Z digest=sha256:8b31d7bada1b5244c9c793b5ec6449b90f16893851f49a9cfe58fd8daf28e968

Observation 98d1f1e6-947e-45f9-8bd9-2e14b9a64f87 · outbound

This paper cites Adam: A Method for Stochastic Opti- mization,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Adam: A Method for Stochastic Opti- mization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.443509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:59.998448Z digest=sha256:47be918e86501c4268aafe63d497efb1c89c6ac2e999862d545bf62d1f79fc44

Observation 660927e8-3baa-4905-bf20-2df453f72676 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.166687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:12:00.112002Z digest=sha256:34b1c45511b0959c0ad6b237c8e969af136638996ab4def26e21ee8972575051

Observation bc7abb8f-32a5-404e-a102-33693344d727 · outbound

This paper cites DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.013883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:12:00.221210Z digest=sha256:680ed3c0bad2d3fb879ed6400702c0efcbdd7aff14b7cd7b8db692bc8162dd71

Observation 941ab650-d1d2-4991-8545-9bba2435104b · outbound

This paper cites Method for the subjective assessment of intermediate quality level of audio systems,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Method for the subjective assessment of intermediate quality level of audio systems,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:00.848285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:12:00.325633Z digest=sha256:f21b924fbe0e25645052403fd406cbaa25ce0a2c2d4b9392ebdc9a27de0471ab

Observation 5ef30e46-e58e-47f3-97e6-20c11c41a23b · outbound

This paper cites From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:12:00.395703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:12:00.395703Z digest=sha256:a8fada90d46a065975b32ec7d27bc40277d41ea24df70f90abe372e0abf11bad

Pith citing papers

Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · inbound

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding cites this paper.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:12:00.657415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:11:56.263021Z digest=sha256:4c0b47b443d4227695770936453755596f879204149085f758534547f765c3d8