Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:12:00.395703Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.22362.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:12:00.395703Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:11:56.263021Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T22:12:00.605940Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8df8c223-4cfd-4c6b-82d9-88e276c36f6b · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61ed7bf7-03e8-4cb2-92e4-ad7026f60aaf · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58d2b895-c15f-4d5c-979d-86e1565358a0 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 697cd690-a421-48a1-b57a-f10af5163824 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High Fidelity Neural Audio Compression
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 470f8e5c-fd4b-44f3-a4ed-73491a7869a3 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb89ca61-b5f7-4f40-bbcf-3a46db86612f · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f71171f-daf1-4a7f-be44-40d65d974f96 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5a9493a-cb89-4734-9dbf-e5e0eeb6fddf · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Wavlm: Large-scale self-supervised pre-training for full stack speech processing,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd77f7af-be89-4f03-bdd7-70e59f5c7aab · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28cb9e28-d715-4793-8603-bff255369fec · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding PolyV oice: Language Models for Speech to Speech Translation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ae3cbb9-8dce-49cb-8963-43e3a5d44fe8 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Soundstream: An end-to-end neural audio codec,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5616e741-3b16-4a35-b52a-46ffd11ba287 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding AudioLM: A Language Modeling Approach to Audio Generation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f371b50b-589e-4282-903e-299acaccadd5 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6314508-c797-4b3c-b0c6-6201c39e2af5 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53225d5e-261b-47b4-b570-6f689c6a2011 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 741d6f2e-ded0-471f-8981-7d2a8d44a1cd · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Denoising Diffusion Probabilistic Models,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 541cf954-a92b-44d0-ac2b-02138a4b836a · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Fidelity Simultaneous Speech-To-Speech Translation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 346f092b-0acc-4437-bcc2-9fd0777a651b · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Moshi: a speech-text foundation model for real-time dialogue
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba766575-ff7d-4c67-8a2f-1b753275e311 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Autoregressive Image Generation Using Residual Quantization,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0503de73-1854-4765-b019-aa31e7982554 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2eee440f-7147-4fdf-b57b-56e8a377ca85 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b518cc89-7d79-492f-b39e-9406b9a98ced · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Auto-Encoding Variational Bayes,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29dbd204-1178-4dc2-aa33-9055e61bcc79 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Deep Unsupervised Learning using Nonequilibrium Thermody- namics,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 415f794e-88e2-4cbb-9868-005b8a65a511 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75e26e44-6ab8-478a-9520-2502c3b003aa · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Multi- step Distillation of Diffusion Models via Moment Matching,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09f42aa1-4ebc-4c2d-a76a-6ec7a3930a25 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a30a0b0-92f1-4a23-8d32-43abdcd71bd9 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Simple and Controllable Music Gen- eration,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6b0c266-afa8-44f9-8036-d6e1b059ab5c · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SoundStorm: Efficient Parallel Audio Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ce0ac9-6869-4c4a-ad8d-8a7f7314c980 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding FiLM: Visual Reasoning with a General Conditioning Layer,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ff8a0bc-abcd-4e0b-bcd1-eb5534083d2f · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Resolution Image Synthesis With Latent Diffusion Mod- els,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77eab515-24e4-4ea3-b1e9-cd5e767fa542 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Denoising Diffusion Probabilistic Models,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67e37efe-fc40-4e9b-992e-6b321fc40388 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Progressive Distillation for Fast Sampling of Diffusion Models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43f4d6bd-f5e9-4c83-a00c-da998cabf0cc · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98e74125-b74a-4b75-8179-05ebc6103f4f · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding WaveNet: A Generative Model for Raw Audio
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee08011d-db9e-4b47-af8a-a1ced36515b6 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Distribution Matching Distillation for Fast Image Synthesis,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98d1f1e6-947e-45f9-8bd9-2e14b9a64f87 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Adam: A Method for Stochastic Opti- mization,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 660927e8-3baa-4905-bf20-2df453f72676 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc7abb8f-32a5-404e-a102-33693344d727 · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 941ab650-d1d2-4991-8545-9bba2435104b · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Method for the subjective assessment of intermediate quality level of audio systems,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ef30e46-e58e-47f3-97e6-20c11c41a23b · outbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · inbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.