Pith. sign in

Paper Citation Record · LEDGER

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

As of 6 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2606.04418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.04418 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T05:20:49.952030Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact31
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35573f12-6d52-4096-8d69-969cca7b64d3 · outbound

This paper cites Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-28T05:21:39.869212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:8dd15fd35be7820f8329c2d2c75064e6fb08d8c9b12a87cc7042967fdb405b1e

Observation 9c2578c1-928e-422a-bd70-3173d739cb81 · outbound

This paper cites URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T05:21:39.847082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:8fc6487e9e02d97d83a082804a6e6f06f180fba8d44de13c0cbee88e7f0d952f

Observation bd04bbec-8e24-4d2c-818b-167f0737bb4b · outbound

This paper cites Attentive Statistics Pooling for Deep Speaker Embedding , booktitle =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Attentive Statistics Pooling for Deep Speaker Embedding , booktitle =

Reference 3

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.863130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:59a519a5e2aa8333d4a13c9f44cbae3c1cf1f2f4841affa3a7a470ab1a4e0cf5

Observation 7c97e162-ed0c-4bea-8229-8a427656cdbd · outbound

This paper cites In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV).

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T05:21:39.830247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:bba4731435ca04b4100a1001298edce00108cd7bb5a85c2546f214fc868a6215

Observation 4e1b6965-9f57-4cd4-8bff-78b4d542c552 · outbound

This paper cites author Zhou, A.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding author Zhou, A

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T05:21:39.843477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:8afffeef6a13dba5d747f3df6d6a562ed8b31b93567baf29a0ebecdb5c678027

Observation 4fd98e9a-bfcf-4e56-bc5f-8fea8212fb17 · outbound

This paper cites HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.776420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:56d6162766b9deeedc1e43817f59f53cb1080c2c9a4482cfd8a5c423bf9abc20

Observation 97e38ae8-bcfa-4288-9f88-27e4c026daf9 · outbound

This paper cites A ConvNet for the 2020s.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding A ConvNet for the 2020s

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T05:21:39.839310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:64774645a0f87ceb5609bf7fe8d567ec56a8e4de47efbffa386d7772d7487527

Observation 3174fc14-969f-4e16-b1f8-dda8ee890e65 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-28T05:21:39.804685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:c51a4023df39d837d80ef1c162e11d417b7c33bc6aec3925cd94ea4f9b7c65bb

Observation 95eee689-4932-446a-956a-1495335b02e5 · outbound

This paper cites ECAPA- TDNN: Emphasized Channel Attention, Propagation and Ag- gregation in TDNN Based Speaker Verification.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding ECAPA- TDNN: Emphasized Channel Attention, Propagation and Ag- gregation in TDNN Based Speaker Verification

Reference 9

Resolution
metadata mismatch
doi, observed 2026-06-28T05:21:39.835969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:559f1cb909fdbed6eeb7bc9efdc0f6081a47cb63c8bbbac44c387a2ea676bd25

Observation 9b1025b7-e108-4949-883a-aab276e2864f · outbound

This paper cites & Khudanpur, S.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding & Khudanpur, S

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.816980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:94e5c036fe5ea54af8e096d745b8cd130fe68896f02744fd184ba444431e9e51

Observation dc34fffa-5700-4974-819e-68397651b6f7 · outbound

This paper cites neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.791932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:116b38aadbdf3b621e8c1a08c4487c66b349dcf0ee9063479e5e48c5f09accbe

Observation 74dbc207-449a-42af-a47c-bf2e2be74fd4 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding HuBERT: Self-supervised speech representation learning by masked prediction of hidden units

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.790085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:b69e98475f848b58e7e15e7eb6e3df7ec061882c33a8e044fa52829a96463e4d

Observation 7872efa4-430f-4b19-83f5-e6d925f66af7 · outbound

This paper cites wav2vec 2.0:.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding wav2vec 2.0:

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:e2512631c798bbd7799be6f9013c6073a685687f36fe3828ff4e0a818915ab00

Observation d6c85265-6a8c-4b63-b637-cff10b36e87d · outbound

This paper cites Neural Discrete Representation Learning , booktitle =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Neural Discrete Representation Learning , booktitle =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:d9ecb5185dc3ffbc9d0fb6207a9c84af6517ccb579f6e978cd00004c37b1d0ca

Observation 641224f9-d03d-4917-b45f-8828ef06a1c4 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding The Twelfth International Conference on Learning Representations,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:195636a05b297a99ac8844740458d1437f9c87bb97dcc5d44b84fd7f6fcdba10

Observation 433ec5b2-1545-488d-9fa3-90d9737a9233 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding The Thirteenth International Conference on Learning Representations,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:fe3d071ce482bf9cf0136fda177537b870d1cef9cf7146e5159d55016362c602

Observation 5990753a-b156-4d5e-902b-383be41dde78 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis , booktitle =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis , booktitle =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:860bdf2a9c57cba2be715952a84070aac0494a9a4e20946b04060adf4c45b67a

Observation 61b3cb67-83fd-4d57-ad9e-76af0bc9e056 · outbound

This paper cites High Fidelity Neural Audio Compression , journal =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding High Fidelity Neural Audio Compression , journal =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:dab5dcd10a0917cb53fc721cdb7b0bfd506cf607060059770b775193d9db196b

Observation 46c7a0be-3b9b-4bfc-b1bd-6d1aa08827bc · outbound

This paper cites High-Fidelity Audio Compression with Improved.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding High-Fidelity Audio Compression with Improved

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:e9b4ef6ee42881bc6b618a5f9459473882029e8153f37c312a8ff4c98c3d4a9c

Observation 56ef43fb-16ff-497b-941b-427fb1ab3497 · outbound

This paper cites Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation , booktitle =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation , booktitle =

Reference 20

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.855452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:85fa6039369e54fb987dac3260a4aa0cc2d8cd2b5cf12c08a62adc9c4f81571f

Observation 86d262d4-58ba-4b0e-b2f4-ed7c3f780bd6 · outbound

This paper cites Focalcodec: Low-bitrate speech coding via focal modulation networks.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Focalcodec: Low-bitrate speech coding via focal modulation networks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.860903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:c771453051b4d8fbd890f6371815f9a2fb15aa1a5f532fe51b8be43383aa4aed

Observation 5224681c-b3d5-4408-82f3-594a74bd512d · outbound

This paper cites VibeVoice Technical Report.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding VibeVoice Technical Report

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.852485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:346efbbffdc6eb179d97830fd465a3d0d933e07fdc93d5ced8a0377e91cc5605

Observation dca2dab1-3332-4b84-8c17-c065c96ac84c · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding The Twelfth International Conference on Learning Representations,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:665b710836c5ee5762097a340da3d4ee361b1c30c6d499d00dfcb14ec826c46c

Observation 534e3033-9c5b-40ec-87db-62011f4ff70d · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.866030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:407f81e78d426abfe838125f5d9923b27c4ca39d0e52f04113222f732f6f0912

Observation 2e723673-d047-4408-a420-3eaab5c0fd44 · outbound

This paper cites UnivNet:.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding UnivNet:

Reference 25

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.824795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:7f1c7960a1b3c01969dfdf872bba2f57181afab2cc315a8ffe7088aa061e96b2

Observation 99ccfa1c-cb5c-4245-9d41-c41843f7f2ef · outbound

This paper cites BigVGAN:.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding BigVGAN:

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:6d4541d2598179d45d666d6b3c0f57dc422b06e924f8c4b156eff645a2dd3393

Observation 5dabbddb-8d36-4852-8cc6-b8948af088d5 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Moshi: a speech-text foundation model for real-time dialogue

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T05:21:39.857848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:b4774ebd04b4c096cd3f1777c137f21f30a42dafd05df8ffbc45a5110cd8b97f

Observation 8c2e8bca-077e-4e20-bef8-d007808562f9 · outbound

This paper cites Qwen3-TTS Technical Report.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Qwen3-TTS Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-28T05:21:39.772589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:3934b24f6b4ca3824e7144d48c57e833a9917043ab1991ea153038b9122ce226

Observation c3485699-94bc-4197-95ec-da97668ac4a9 · outbound

This paper cites arXiv preprint arXiv:2509.14128 , year =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding arXiv preprint arXiv:2509.14128 , year =

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T05:21:39.807860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:dcae9a8052a7c6c26dc358b7735fc15bb9fca4b6bf40b2e4cc868e493a9b5af0

Observation 13dade32-c32f-43ef-8740-80457b84f0be · outbound

This paper cites In: Interspeech 2022.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding In: Interspeech 2022

Reference 30

Resolution
metadata mismatch
doi, observed 2026-06-28T05:21:39.862974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:663a3cd39701a23509deb71cc9583c1e9ea99cbeb4f21e23f147d0892993bf57

Observation f2cbe538-2a52-4994-8ce9-0a854cd1a3f7 · outbound

This paper cites In: ICASSP 2023 - 2023 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding In: ICASSP 2023 - 2023 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T05:21:39.835550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:ccc263b9986d46cc7e4b8130eb2c749e9aae20ca68a899710f68ee8f0796d134

Observation 70c309cd-fd0e-4d55-90b0-8a1de86ab0f5 · outbound

This paper cites Reshape Dimensions Network for Speaker Recognition.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Reshape Dimensions Network for Speaker Recognition

Reference 32

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.849229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:0455a85ce5e43dd51eaaeaf0adb34c0bed596da95c895343672e02a905a343d2

Observation 611a5e2d-f0e4-437a-bfb9-744acbbd9f58 · outbound

This paper cites C Users J.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding C Users J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:e653965b18d23584862b09a769eacfd1f455a63475796400127f70b8bc5e9633

Observation 47b49f36-2feb-4c52-873b-3736084ae197 · outbound

This paper cites Gray , title=.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Gray , title=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:613fcda41e944975bebe88d714bf16978c23788859a0cd2065cc1e2f927f94df

Observation 286f5cff-c65f-4896-a739-080252b1c75e · outbound

This paper cites Sensors , volume =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Sensors , volume =

Reference 35

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.792454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:134ef63f7d19611904ecc03c9b8b5dd5aedc4816f7c0bca6a7afd1679466933f

Observation 394623c8-11e0-43c3-89bd-80e664631a86 · outbound

This paper cites Neural computation , volume=.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Neural computation , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:7f849f36ea23199789cff298426e690da83de6772a7c2daed9f48c1230f68b87

Observation 142cc661-e6ce-4873-b5e0-316f87edd103 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 37

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.860420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:8fcf98a1ed2aae34195aa458aff90c157d083874a03168d5853398cfc6efdf9b

Observation e6fdd493-b935-49df-8143-53f7fd9b854d · outbound

This paper cites Emilia: A large-scale, extensive, multilin- gual, and diverse dataset for speech generation.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Emilia: A large-scale, extensive, multilin- gual, and diverse dataset for speech generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.814083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:0f35ee27361ae1311119ed1ed0572a9b3fbb429b3467272db7c64ba2716ecdf0

Observation b0620059-3141-41b9-8a55-c5098ddaf857 · outbound

This paper cites GigaSpeech: An Evolving, Multi-Domain.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding GigaSpeech: An Evolving, Multi-Domain

Reference 39

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.827386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:3f306f9ce17608cfdef4b4fc7bc1e69d236cf2bb9575effa198dc558a1e5f150

Observation 43d54c4f-67a3-4593-9612-a003da37fcd7 · outbound

This paper cites 2023 , booktitle =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding 2023 , booktitle =

Reference 40

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.830094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:ec8480f53cb2d96d9e57c7b56ee67db61b8bda269ef8f4a6aaea0db9083e05de

Observation baa303eb-3ffa-40bf-9cbc-c5ba33564e56 · outbound

This paper cites FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T05:21:39.844230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:f0b822518ad368d995e2ab38cd7fd76889fcd3a190da19e7674e4154ef52564a

Observation 4025d02d-9805-4f45-bf49-324464bc4ed5 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models , booktitle =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models , booktitle =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:f25032212700cdea79c6b3a2d0221ef7fb3d771538b1e8c3deaee668179c4a23

Observation 005ced62-6479-4db8-884e-cd8d05a0f791 · outbound

This paper cites URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T05:21:39.838702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:93b20e93407059d2fd0b8c47b7bef4e550b4e291650efdffb945b28b642b7d11

Observation e4465350-649e-44f5-b8ea-3416ec4e05fb · outbound

This paper cites CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92).

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92)

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:35b7cc32a83e5eacef37a58558237aa20a5a52de2b6f824fcf6b86c1edeeec80

Observation ab74e4b5-93f3-4ec9-87cf-55ea663dbaa3 · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 45

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.845584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:b61e7fee60cb1faa8a65f52cf3dc61531ef9c64bed6a22c53a0368ca8abdf99f

Observation 69c5bc55-6f47-4062-b342-88c466535288 · outbound

This paper cites Simple and Controllable Music Generation , booktitle =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Simple and Controllable Music Generation , booktitle =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:9fbf6d63854dc82c197750e45c8deb3b12ed1b52c6c8662b6af3c3b93e288232

Observation 56cca99c-ac62-4935-844d-3307dc3e73b2 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-28T05:21:39.872068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:5c15445cad73df925172b5544d2385b4ec0a935559bdc68166b8feff8effb7fd

Observation 29d8c45d-976f-4774-9f78-fcce42560d97 · outbound

This paper cites Sidon: Fast and robust open-source multilingual speech restoration for large-scale dataset cleansing.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Sidon: Fast and robust open-source multilingual speech restoration for large-scale dataset cleansing

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.851036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:8692e07df923116afad8bdf9bd47b759dd56aeb7762aa2e834757e855e646e9f

Observation 72c82a01-9826-47e1-8ef5-7b649dee5955 · outbound

This paper cites Pyroomacoustics: A python package for audio room simulation and array processing algorithms,.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Pyroomacoustics: A python package for audio room simulation and array processing algorithms,

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.819911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:1c36cc4b3ed0a93b8604c76ce82264fae27d48502daa8903b0d87c2259d6225e

Observation a24b9e12-0c1b-4104-ab1b-984c3ed6cd86 · outbound

This paper cites Gemmeke, Daniel P.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Gemmeke, Daniel P

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.806809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:2eb7d2c8004883bfbaad10866d62a1cdad61446a5e303cf7b576063402137f8f

Observation 635c1b9b-3348-4b19-9b35-908219e1aec9 · outbound

This paper cites Ericsson, L., Espinosa, M., Yang, C., Antoniou, A., Storkey, A., Cohen, S.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Ericsson, L., Espinosa, M., Yang, C., Antoniou, A., Storkey, A., Cohen, S

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.798843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:6e74d09153b0944ca02e5b833321b909e485479e8258c0d6f24fa796499a4faf

Observation 590af918-c75e-4743-b6c7-2417f413d59b · outbound

This paper cites WHAM!: Extending speech separation to noisy environments.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding WHAM!: Extending speech separation to noisy environments

Reference 52

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.803724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:8cbd3a995c3f541168ad6d2e3f263d93a208e1a8effa3a9b96939ebc6ebbc55d

Observation 69a15ad5-1818-4586-ad04-70fbff2f53cc · outbound

This paper cites AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:56:51.761438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:0c7c264614c745333ef33ea8460cc4113b25ac7f18ba357d08af9685daaec445

Observation 831b6115-b449-40ae-8567-677fbb583546 · outbound

This paper cites VoxCeleb: A Large-Scale Speaker Identification Dataset.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding VoxCeleb: A Large-Scale Speaker Identification Dataset

Reference 54

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.853383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:924074e7dac7b5f79bfab582680713cb26a82232fe53d5c1925c9fd0f57fe0f7

Observation 0209404c-e9bb-4f67-b002-953f162dc551 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding The Twelfth International Conference on Learning Representations,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:d3d7fbccfb455d975b1332370049e10c7f9827fc8af94efe3eb53d024ad04bc6

Observation 46ebadd5-dedc-4f62-b567-e0c5fccbeb1e · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-28T05:21:39.848241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:0f855271def9e27efa9268ebbd08a573e22ae0db28fb137ebb61925fbf3dbf32

Observation ef916669-df6c-4c35-a6fc-11524f13bb13 · outbound

This paper cites Parker and Anton Smirnov and Jordi Pons and CJ Carr and Zack Zukowski and Zach Evans and Xubo Liu , title =.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Parker and Anton Smirnov and Jordi Pons and CJ Carr and Zack Zukowski and Zach Evans and Xubo Liu , title =

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T05:20:49.952030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:adf5fe24adcff8e8a6603af8c77c5213e367766f3968567850e3f89ff2237dda

Observation 93bdf356-00d5-4e9f-ac77-e60a68e8d3cd · outbound

This paper cites DualCodec:.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding DualCodec:

Reference 58

Resolution
verified exact
doi, observed 2026-06-28T05:21:39.854490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:ca7bd24be9f2f43ea293f954e897be1099df487bf65bea358ce06653370d5c9a

Observation cb6cebef-6221-4a61-bf5e-efd76448a699 · outbound

This paper cites PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.788836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:d0c71faf0ff36d8c5405ba525d2d016ba72c4f49902db82565dfa851da3e9543

Pith citing papers

No inbound Pith citation observations are available.