Pith. sign in

Paper Citation Record · LEDGER

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

As of 22 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.22362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22362 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:12:00.395703Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:11:56.263021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:12:00.605940Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · outbound

This paper cites DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:12:00.657415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:56.263021Z digest=sha256:4c97933db9d7a7b7fa0c0598a4eebc4eeb5d45edc40c38e9b9ab0b74f1368e54

Observation 8df8c223-4cfd-4c6b-82d9-88e276c36f6b · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.884534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:56.378331Z digest=sha256:22bb2369b2e5a8282fba6fe0d36abfbee891a108a6d168cba2dccef643da9176

Observation 61ed7bf7-03e8-4cb2-92e4-ad7026f60aaf · outbound

This paper cites SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.545419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:56.883055Z digest=sha256:b8b1cce35a73ac545107cbc96d8c6c9dc40f266d0936bedbbbd57d59adb0d29a

Observation 58d2b895-c15f-4d5c-979d-86e1565358a0 · outbound

This paper cites There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.432611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:57.068976Z digest=sha256:5757d58cc95a28fa7f45411ea6bfe16d4623811edf1100e7607757d8635180d1

Observation 697cd690-a421-48a1-b57a-f10af5163824 · outbound

This paper cites High Fidelity Neural Audio Compression.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:57.819506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:57.819506Z digest=sha256:e67c1823de25070896ccf9893342bd592bb0176267214bcdef1b394f19c332e8

Observation 470f8e5c-fd4b-44f3-a4ed-73491a7869a3 · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.785845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:56.556586Z digest=sha256:ba07b414bfc73cc9585251ac3c14d816869c21a85e271c8d5ee53683d6e7f82e

Observation cb89ca61-b5f7-4f40-bbcf-3a46db86612f · outbound

This paper cites HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.326298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:57.215975Z digest=sha256:ad7a4876ee492ec80e678eb4d19992a90e3690a5a8f76365abae27f003f3d849

Observation 7f71171f-daf1-4a7f-be44-40d65d974f96 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.192372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:57.401222Z digest=sha256:942ec6f28b51bd966b6b1cd6928817ff3983f7fd7a140cfc91a7b1876b37d41e

Observation d5a9493a-cb89-4734-9dbf-e5e0eeb6fddf · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:57.528724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:57.528724Z digest=sha256:1b5e3d80c232ed006670ff0de88eed4ebe8369cfbfdb34061e524c6e9e4c3331

Observation fd77f7af-be89-4f03-bdd7-70e59f5c7aab · outbound

This paper cites w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.084487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:57.697694Z digest=sha256:35a7ec92b1c892a316de046d061d2a2ea964d2905cc348d32b499b297c2a6ac0

Observation 28cb9e28-d715-4793-8603-bff255369fec · outbound

This paper cites PolyV oice: Language Models for Speech to Speech Translation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding PolyV oice: Language Models for Speech to Speech Translation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.407814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:58.332796Z digest=sha256:106a4c1d437d37e269c20e291dc0e284642f7f5c2825d829d9d59bc4d36ff1fa

Observation 2ae3cbb9-8dce-49cb-8963-43e3a5d44fe8 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Soundstream: An end-to-end neural audio codec,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.950584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:57.887238Z digest=sha256:f87682e483c8a89d89a0cd91456c869a9b235de7dc5a979fd8420f32b1fd0e20

Observation 5616e741-3b16-4a35-b52a-46ffd11ba287 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding AudioLM: A Language Modeling Approach to Audio Generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.753500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:57.968067Z digest=sha256:6e358564521def89dacda0434b6f6a0c9520152f88732ef5ede7c7fc1a9326a7

Observation f371b50b-589e-4282-903e-299acaccadd5 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.039438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.039438Z digest=sha256:4e849db6a1db5b0d40036741821713224fe692f7aeed8eb8f36b3169021e514d

Observation f6314508-c797-4b3c-b0c6-6201c39e2af5 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.112443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.112443Z digest=sha256:8dbf8342afcfecc7353a4792d3f3b86c451ba9a22b2dc3f8a24ea0d1edb65244

Observation 53225d5e-261b-47b4-b570-6f689c6a2011 · outbound

This paper cites TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.590748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:58.204048Z digest=sha256:1555ec984df9bf70475cbd947db0389945ce3c38f45b549e4f3dd473eda1207d

Observation 741d6f2e-ded0-471f-8981-7d2a8d44a1cd · outbound

This paper cites Denoising Diffusion Probabilistic Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Denoising Diffusion Probabilistic Models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.658852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:58.800630Z digest=sha256:df06f4ae4a46a95c9ffcca49d937da9d683a711ebd325590752162d64928258c

Observation 541cf954-a92b-44d0-ac2b-02138a4b836a · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.428052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.428052Z digest=sha256:086cf666e05013c72780f55a5ace347a8f526c2329eb20adbf0f24e1cd1225e1

Observation 346f092b-0acc-4437-bcc2-9fd0777a651b · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Moshi: a speech-text foundation model for real-time dialogue

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.489098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.489098Z digest=sha256:cfce16bbfadb6a612d588658c3c256e001feee8004f77911ae647f8a20a50c19

Observation ba766575-ff7d-4c67-8a2f-1b753275e311 · outbound

This paper cites Autoregressive Image Generation Using Residual Quantization,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Autoregressive Image Generation Using Residual Quantization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.193477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:58.579941Z digest=sha256:aaae3e250dc7a94bbc54a96a326556387a5a6cb6c192d6e31475d3d4b5b6c76b

Observation 0503de73-1854-4765-b019-aa31e7982554 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.004761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:58.652586Z digest=sha256:8d0aab547ab870743c6f671a5ec378c31d9eb57107c16f9cc25496e68f384704

Observation 2eee440f-7147-4fdf-b57b-56e8a377ca85 · outbound

This paper cites Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.825848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:58.731676Z digest=sha256:138a95af4fba19070d31251178d73d80220fcf6495ffd4fec6cf161ed29d3a54

Observation b518cc89-7d79-492f-b39e-9406b9a98ced · outbound

This paper cites Auto-Encoding Variational Bayes,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Auto-Encoding Variational Bayes,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.760597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.303877Z digest=sha256:63ac260a06c75e02b6eb13d1b603f22d204a90c81225ac00e9cb8e675d32215b

Observation 29dbd204-1178-4dc2-aa33-9055e61bcc79 · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermody- namics,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Deep Unsupervised Learning using Nonequilibrium Thermody- namics,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.492841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:58.879689Z digest=sha256:d5993bf70364e565b8f0ec17bfae70283d1ae9655c36f48e886d9db72b3b715f

Observation 415f794e-88e2-4cbb-9868-005b8a65a511 · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.654095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:56.737595Z digest=sha256:b79324d9407946cbcf04009b04b9aa75c2c0bdac81372b3894e4316cb5100cda

Observation 75e26e44-6ab8-478a-9520-2502c3b003aa · outbound

This paper cites Multi- step Distillation of Diffusion Models via Moment Matching,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Multi- step Distillation of Diffusion Models via Moment Matching,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.345517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:58.962209Z digest=sha256:1b8112471b238507a4e9dcef32015ce21d1375cf2d51d658331b00739c0cd625

Observation 09f42aa1-4ebc-4c2d-a76a-6ec7a3930a25 · outbound

This paper cites SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.154376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.044535Z digest=sha256:b949c42ce374d3b3a3d1c1e036ab7b9fb2bf2b2423f7d6839817842403ae80bb

Observation 1a30a0b0-92f1-4a23-8d32-43abdcd71bd9 · outbound

This paper cites Simple and Controllable Music Gen- eration,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Simple and Controllable Music Gen- eration,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.959308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.138687Z digest=sha256:0138a2ee462ca13d72177cc83dc55a037cea47721c8907771044aed858a176e4

Observation d6b0c266-afa8-44f9-8036-d6e1b059ab5c · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SoundStorm: Efficient Parallel Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:59.228454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:59.228454Z digest=sha256:f80d259e45b6c570ed8ef7ab1ba4aa8b88a3b78178c60dbc34b5f0cfb45a16f3

Observation 57ce0ac9-6869-4c4a-ad8d-8a7f7314c980 · outbound

This paper cites FiLM: Visual Reasoning with a General Conditioning Layer,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding FiLM: Visual Reasoning with a General Conditioning Layer,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.583977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.401405Z digest=sha256:8aeaf03b65b9c58e18a31515a5908a0bab83840eb23c0166a104c9136d17ec08

Observation 5ff8a0bc-abcd-4e0b-bcd1-eb5534083d2f · outbound

This paper cites High-Resolution Image Synthesis With Latent Diffusion Mod- els,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Resolution Image Synthesis With Latent Diffusion Mod- els,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.376140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.467113Z digest=sha256:7e24a4d86937bef80755d448c506108c519aa5c2f1478df94d1fe251085b3446

Observation 77eab515-24e4-4ea3-b1e9-cd5e767fa542 · outbound

This paper cites Improved Denoising Diffusion Probabilistic Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Denoising Diffusion Probabilistic Models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.220833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.562721Z digest=sha256:acafa46e543f137505e0e48cb5367776902c1b56d2dfb44a748f0d4b7d4786a6

Observation 67e37efe-fc40-4e9b-992e-6b321fc40388 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Progressive Distillation for Fast Sampling of Diffusion Models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.030022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.627762Z digest=sha256:cd422e18c086d659b4cc714fc8e7daa6b14a2a1be9a86c5ca9773456623bd002

Observation 43f4d6bd-f5e9-4c83-a00c-da998cabf0cc · outbound

This paper cites Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.835199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.723665Z digest=sha256:dfabf4eb924ea30a49099f51bf9e763a9ab17b76224b9658152fc99b953f5bb6

Observation 98e74125-b74a-4b75-8179-05ebc6103f4f · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding WaveNet: A Generative Model for Raw Audio

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:59.809510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:59.809510Z digest=sha256:2b129cbe46713d5aa98377d01773b349eed0bf3bf80b9bf210d60271e0d56ae6

Observation ee08011d-db9e-4b47-af8a-a1ced36515b6 · outbound

This paper cites Improved Distribution Matching Distillation for Fast Image Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Distribution Matching Distillation for Fast Image Synthesis,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.628591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.912833Z digest=sha256:af4755b8f9c4f9f12404f3c9a49b246299059f19344ccbc705ae4e3217f5b40a

Observation 98d1f1e6-947e-45f9-8bd9-2e14b9a64f87 · outbound

This paper cites Adam: A Method for Stochastic Opti- mization,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Adam: A Method for Stochastic Opti- mization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.443509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:59.998448Z digest=sha256:8c175c5dd55f48ca7b489b81740e3fcf4a5109404015edb3d24e8c6b333897c4

Observation 660927e8-3baa-4905-bf20-2df453f72676 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.166687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:12:00.112002Z digest=sha256:1ebe8f1013bf95d50174f9e5858f08a79ee559751d5b3b84d3ae79027f35e8b5

Observation bc7abb8f-32a5-404e-a102-33693344d727 · outbound

This paper cites DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.013883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:12:00.221210Z digest=sha256:313c6aa46f9890b574fab7397e61a7069e4caf6f914e6a4a3774fd4fa9209c3a

Observation 941ab650-d1d2-4991-8545-9bba2435104b · outbound

This paper cites Method for the subjective assessment of intermediate quality level of audio systems,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Method for the subjective assessment of intermediate quality level of audio systems,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:00.848285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:12:00.325633Z digest=sha256:6ecc9a33ef86cb76cc3411deeb9706f22a033a74060edae80730acd6e6e90b18

Observation 5ef30e46-e58e-47f3-97e6-20c11c41a23b · outbound

This paper cites From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:12:00.395703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:12:00.395703Z digest=sha256:47e23e62d19df8fd210a6b5ac46328c56e0f341a7582ca6d0b1433a1a7616bd6

Pith citing papers

Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · inbound

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding cites this paper.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:12:00.657415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:11:56.263021Z digest=sha256:4c97933db9d7a7b7fa0c0598a4eebc4eeb5d45edc40c38e9b9ab0b74f1368e54