Pith. sign in

Paper Citation Record · LEDGER

Factorized RVQ-GAN For Disentangled Speech Tokenization

As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2506.15456.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15456 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:39:29.000932Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:39:28.911720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T19:39:29.056834Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17acfd32-4fc5-4f28-8271-d3a621c456be · outbound

This paper cites hj/iAePKhU/wEzYUdFoprleUHDo=.

Factorized RVQ-GAN For Disentangled Speech Tokenization hj/iAePKhU/wEzYUdFoprleUHDo=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.309084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.907303Z digest=sha256:5084109f461a1def2cfd824544a7affeb8a67010b872f693f7533d9f5a6247af

Observation 823089f8-397a-4d80-bc32-da4d5adb1c7f · outbound

This paper cites Factorized RVQ-GAN For Disentangled Speech Tokenization.

Factorized RVQ-GAN For Disentangled Speech Tokenization Factorized RVQ-GAN For Disentangled Speech Tokenization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:39:29.060642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.911720Z digest=sha256:a614a13f2b159fdc2df6db8de0f67abf76baf065d0a882245f28dddf7c0081e4

Observation bb2ef3d0-6154-4f9b-a418-cb28c846f132 · outbound

This paper cites These objectives are described in detail in [3].

Factorized RVQ-GAN For Disentangled Speech Tokenization These objectives are described in detail in [3]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.298909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.915197Z digest=sha256:5c7ffb015e56f8c23e8bf23e03468e77228cfb8b4e6f7facf1f3649569ff7e31

Observation 31996334-126a-4d2c-b4bb-afacd1786516 · outbound

This paper cites For evaluation, we use forced-aligned test sets from Lib- riSpeech and Multilingual LibriSpeech (MLS) [16].

Factorized RVQ-GAN For Disentangled Speech Tokenization For evaluation, we use forced-aligned test sets from Lib- riSpeech and Multilingual LibriSpeech (MLS) [16]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.288543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.919214Z digest=sha256:9b47162c8730d399c29bcf5498df179a7ca32c481f69e2fc2afa6cb4ab5a9f9e

Observation 00a7ef10-7418-4ed1-a9c0-6f8c5c70c131 · outbound

This paper cites Like Encodec, ST updates its codebooks via exponential moving average (EMA) and period- ically re-initializes them to maximize utilization.

Factorized RVQ-GAN For Disentangled Speech Tokenization Like Encodec, ST updates its codebooks via exponential moving average (EMA) and period- ically re-initializes them to maximize utilization

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.278477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.922726Z digest=sha256:060934d6fbdaf75581370601c9baf6a32995c12294fbd03180c0e9ed6e6f6ef6

Observation 1ee2ae7a-eb97-48a4-8209-433f13449f20 · outbound

This paper cites an unresolved cited work.

Factorized RVQ-GAN For Disentangled Speech Tokenization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:39:29.268132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.926087Z digest=sha256:bc474822e11cac63fd94133a93a1bc73655f9f9f5189da7bb91f690b03e74098

Observation 53396364-a75e-4635-a4c2-25d8c352b34f · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

Factorized RVQ-GAN For Disentangled Speech Tokenization SoundStream: An end-to-end neural audio codec,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:39:28.929045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:39:28.929045Z digest=sha256:5f73a136636536f6faf43e11abe2ae27fd9b94ac03b63cc7264ba54a8967e949

Observation 6bc4daf8-cc26-4a67-8223-e769b22e9268 · outbound

This paper cites High fidelity neural audio compression,.

Factorized RVQ-GAN For Disentangled Speech Tokenization High fidelity neural audio compression,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:39:28.932410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:39:28.932410Z digest=sha256:978c8a05643b7a7c6afb39225db24df6fdffb1ca6613ba97c434350cd77896a1

Observation 264a4ccb-060b-49f4-9e57-3b2c51b88bd8 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN,.

Factorized RVQ-GAN For Disentangled Speech Tokenization High-fidelity audio compression with improved RVQGAN,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.245002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.936622Z digest=sha256:a6208dce59f2228b27ec29b74b9515a8598adda3a28010ffc1ce10e4e18aa72e

Observation 652621fa-0f61-442d-b083-36592251ef83 · outbound

This paper cites ESPnet-Codec: Com- prehensive training and evaluation of neural codecs for audio, mu- sic, and speech,.

Factorized RVQ-GAN For Disentangled Speech Tokenization ESPnet-Codec: Com- prehensive training and evaluation of neural codecs for audio, mu- sic, and speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.235174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.940045Z digest=sha256:66e129c8106eeda4e52c534ad302c9257d99c312736d7a2247a88b743067e310

Observation 45e8e3c6-cfb3-4c09-9b28-a6c3bcdb3bb6 · outbound

This paper cites AudioLM: a language modeling approach to audio gener- ation,.

Factorized RVQ-GAN For Disentangled Speech Tokenization AudioLM: a language modeling approach to audio gener- ation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.225363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.943495Z digest=sha256:0a7bdba010714ba7871cf5dbae4ce743457eda62d7eaad7b199cac7f6eab7d64

Observation 8508c9a3-4bb6-4af4-9a49-a9725cf96e99 · outbound

This paper cites Direct speech-to-speech translation with discrete units,.

Factorized RVQ-GAN For Disentangled Speech Tokenization Direct speech-to-speech translation with discrete units,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.214728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.947499Z digest=sha256:42e677a99462d4f3d259d07b61ae9a78fad46d52b8441e51d0c0dcde4458b054

Observation d0b72eda-3e21-4b8b-b60e-862d2cb91473 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

Factorized RVQ-GAN For Disentangled Speech Tokenization Neural codec language models are zero-shot text to speech synthesizers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.204716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.950681Z digest=sha256:b78f16e3675481a167fff1967901bf84ab3592a8ca4ca8fefaa5d7d1344c745c

Observation b0d91740-e2c2-4c09-92f9-8f502c7fdc80 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

Factorized RVQ-GAN For Disentangled Speech Tokenization HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.194460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.954394Z digest=sha256:8afaee534e6857f267ed2956e800604049a472a26fe6d23b2a510b4d469186a3

Observation 9573b92a-6efc-4b33-bfd7-86b9ae562824 · outbound

This paper cites Exploring speech recognition, translation, and understanding with discrete speech units: A com- parative study,.

Factorized RVQ-GAN For Disentangled Speech Tokenization Exploring speech recognition, translation, and understanding with discrete speech units: A com- parative study,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.182883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.957711Z digest=sha256:af62a4359c3699d5b6a78dbf7a5e2c552bdd312aef1379e728245194c3b4ed05

Observation cd9498f8-6315-40f6-8fc0-77a6bc9a23cf · outbound

This paper cites On generative spoken language modeling from raw audio,.

Factorized RVQ-GAN For Disentangled Speech Tokenization On generative spoken language modeling from raw audio,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.172175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.960865Z digest=sha256:e4135de6642dc77042c682eef9c5dafaedf8265e6ba01ebb4d457b1903958e2e

Observation 60d3c19e-cfbb-4a10-a5fd-0fb54671810f · outbound

This paper cites SpeechTok- enizer: Unified speech tokenizer for speech large language mod- els,.

Factorized RVQ-GAN For Disentangled Speech Tokenization SpeechTok- enizer: Unified speech tokenizer for speech large language mod- els,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.161492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.964465Z digest=sha256:ae33133f88650892dd4fe53ce9a0fe1596018e0877b488d45bed863d4b9db532

Observation a98dcb61-4645-47d9-9cdb-163c652fd600 · outbound

This paper cites Language-agnostic BERT sentence embedding,.

Factorized RVQ-GAN For Disentangled Speech Tokenization Language-agnostic BERT sentence embedding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.151363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.967595Z digest=sha256:c507fdc46346742edb3d9f8a6a25b596f207cd6b6d0b3348ae7be70962067078

Observation a9aec702-8656-4623-a236-301a20c4a995 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,.

Factorized RVQ-GAN For Disentangled Speech Tokenization wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.141514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.971282Z digest=sha256:20bb3960385c98da3758ddfd7333b1a4ef515587b4e03f3704d76be9492f3b53

Observation 08353250-3506-4b29-a868-7d5fe339a811 · outbound

This paper cites Lib- rispeech: An asr corpus based on public domain audio books,.

Factorized RVQ-GAN For Disentangled Speech Tokenization Lib- rispeech: An asr corpus based on public domain audio books,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:39:28.974547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:39:28.974547Z digest=sha256:d6deee2ccb5ea07057de5187aee9fd7df0f16528b7bdeae6c1aa68fcd65c6231

Observation e40a318f-ce4a-447b-b217-951b2ebc052f · outbound

This paper cites V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,.

Factorized RVQ-GAN For Disentangled Speech Tokenization V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.124787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.977657Z digest=sha256:29b299ea8f608045321a119380ef0a14ff55dd40b53a12a74bb9ae1d66f9d1e1

Observation 39bb37a1-b3a3-4fae-9eae-c7edf779f8fd · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research,.

Factorized RVQ-GAN For Disentangled Speech Tokenization MLS: A large-scale multilingual dataset for speech research,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.113947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.981225Z digest=sha256:57aa396d51a3b4e169df31caa3d0cc4782a4e4397ef99f08112ebac075b6b583

Observation 3ddbe758-26de-41da-a50c-1093baa7f9e3 · outbound

This paper cites Montreal Forced Aligner: Trainable text-speech align- ment using Kaldi,.

Factorized RVQ-GAN For Disentangled Speech Tokenization Montreal Forced Aligner: Trainable text-speech align- ment using Kaldi,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.102343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.984436Z digest=sha256:af00ebb017e66fbd069f14dacf30c63174e0d49fd63202f7f56d37b2d6a25784

Observation 86d4e17d-e9f0-42a3-92c4-72ce2bb145dc · outbound

This paper cites mHuBERT-147: A Compact Multilingual HuBERT Model.

Factorized RVQ-GAN For Disentangled Speech Tokenization mHuBERT-147: A Compact Multilingual HuBERT Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:39:28.987701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:39:28.987701Z digest=sha256:d5a564a572ab7b07252cd9d207f0d0e3339c275dedd93707b957452c68c3a48e

Observation 8d28b839-019e-4593-8f9b-71982ed02adb · outbound

This paper cites SAMU-XLSR: Semantically-aligned multimodal utterance-level cross-lingual speech representation,.

Factorized RVQ-GAN For Disentangled Speech Tokenization SAMU-XLSR: Semantically-aligned multimodal utterance-level cross-lingual speech representation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.091922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.991280Z digest=sha256:46c9427fc6301be7e2b6191ab2252bc765775602364dda6c87e54580df42d8ef

Observation 0d202429-7ec0-4460-a16d-f29f6638cf41 · outbound

This paper cites Abx-discriminability measures and applications,.

Factorized RVQ-GAN For Disentangled Speech Tokenization Abx-discriminability measures and applications,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.081374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.994420Z digest=sha256:ae06557c76e3f87d61076898a73fc632e2a46970827172b7d8a1dedc66d4e291

Observation f3fd29c5-089d-4b1a-80f5-045b128a1dad · outbound

This paper cites The Zero Resource Speech Challenge 2021: Spoken language modelling.

Factorized RVQ-GAN For Disentangled Speech Tokenization The Zero Resource Speech Challenge 2021: Spoken language modelling

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:39:29.036877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.997541Z digest=sha256:321e17b8e122f61b49c397731deeeaf5604ef93af198cd90b559ebcb629457c3

Observation 19e21dc0-11c0-4e8d-a43c-a9fe40114560 · outbound

This paper cites Learning hierarchical dis- crete linguistic units from visually-grounded speech,.

Factorized RVQ-GAN For Disentangled Speech Tokenization Learning hierarchical dis- crete linguistic units from visually-grounded speech,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:39:29.071255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:29.000932Z digest=sha256:0d1c70876efea06989316dae4fd56a2ca445c97d70a60f50acf37f4e21b0624d

Pith citing papers

Observation 823089f8-397a-4d80-bc32-da4d5adb1c7f · inbound

Factorized RVQ-GAN For Disentangled Speech Tokenization cites this paper.

Factorized RVQ-GAN For Disentangled Speech Tokenization Factorized RVQ-GAN For Disentangled Speech Tokenization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:39:29.060642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T19:39:28.911720Z digest=sha256:a614a13f2b159fdc2df6db8de0f67abf76baf065d0a882245f28dddf7c0081e4