Pith. sign in

Paper Citation Record · LEDGER

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2607.19033.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19033 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.656380Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.420218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:15:23.475461Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact3
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2af3fbb-4cc0-42f9-ace0-8327fc1aeac3 · outbound

This paper cites Content is What Remains: Invariant Speech Tokenization from Parallel Utterances.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.420218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.420218Z digest=sha256:15cb0ebf46563294fae97d674110707d544a7b9c9ca3b882bfafde15aaeb4e76

Observation 36122b67-fcf0-4c05-b518-26339874cf68 · outbound

This paper cites an unresolved cited work.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:37:37.773073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.426600Z digest=sha256:210c8fd333d4889903a68658579a28fb786b98a3f77c639e97b6d9d8824a3969

Observation 314f73bb-bf00-4cba-84c2-caa1eccad9cc · outbound

This paper cites Further we evaluate the downstream compres- sion benefit unlocked by the above.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Further we evaluate the downstream compres- sion benefit unlocked by the above

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T15:37:37.757774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.431802Z digest=sha256:9242e2bf142d5cf7c90d6956d3ba425b752b83e52efeb227858134ebedb7d4e0

Observation e619767f-ffcb-42b7-92b3-ccf4b6312d00 · outbound

This paper cites an unresolved cited work.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:37:37.742427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.436685Z digest=sha256:5ca18c171f31e85c2b4085e79af94f5e8c4fd2633d83a286f6b1c78f1c698a1e

Observation f6d5b190-beab-4257-8444-22f4875b90b3 · outbound

This paper cites The experiments were run manually and results were manually verified.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The experiments were run manually and results were manually verified

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.726711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.441464Z digest=sha256:8392252976e4c0df158650320b76e4dde46ed10623a1036c875cc463a9c16762

Observation de499ba7-9b34-418c-9973-674a7badd1cc · outbound

This paper cites AudioLM: a language modeling approach to audio generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances AudioLM: a language modeling approach to audio generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.710583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.446270Z digest=sha256:0b3d080ea102af732c0156a2c1407c8e458e87f347015ce5530fc20f39ae5a00

Observation fc188f6a-54ef-425d-9b17-b5e03b654f03 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.696744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.450990Z digest=sha256:a3dafd8ef33b729dab25537bb5af842efceab8d3cf76b7fccfc18561d6f942a2

Observation 33894dc6-3747-4fd1-9516-1c6b297ca72e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Moshi: a speech-text foundation model for real-time dialogue

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.456476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.456476Z digest=sha256:dbe53dafdd2e143ddbb42066c1f73ad47ed6d849b63637f9cdcb16c78be7f176

Observation 63a33f02-57d5-4c38-ba68-5d94e182fff3 · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.461335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.461335Z digest=sha256:441430effeed4aa176c8256da66b71fbea06f52a23e8bfc97a5c9e0d0ed901dc

Observation 8ee4ab36-a958-40c1-a6f9-d0f6bd0ea221 · outbound

This paper cites SpeechTok- enizer: Unified speech tokenizer for speech language models,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances SpeechTok- enizer: Unified speech tokenizer for speech language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.681063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.466266Z digest=sha256:b6b4d434dc6fc56e1430a68db4ae293385e3f4bf724c24c1a0e2abe63ba38833

Observation ef9bd97c-7591-4b86-a813-f90d20517ba1 · outbound

This paper cites AudioDec: An open-source streaming high-fidelity neural audio codec,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances AudioDec: An open-source streaming high-fidelity neural audio codec,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.664857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.470649Z digest=sha256:55a9fda17d980b07f19196d5ac23e27acc201a7cf68d7d5f4af451893996615f

Observation 1cbd49f7-3425-4cf6-9c55-a96ff4db9109 · outbound

This paper cites Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.474582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.474582Z digest=sha256:09446f2e9439fdb6296aa995eebbfb6070377b47a9a2435ec887b01813723f85

Observation 6e300fcc-ab13-44cf-9cad-b4d7168681a6 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.478614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.478614Z digest=sha256:b2ebc186b7b9d31069e4b6833ac6eff9f800e6559781f1a1b5dbc11137fdb046

Observation 7799387e-c3bd-4bd1-9ae7-8f47588a0dcd · outbound

This paper cites ContentVec: An improved self-supervised speech representation by disentangling speakers,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances ContentVec: An improved self-supervised speech representation by disentangling speakers,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.486606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.486606Z digest=sha256:0790596b076efef25a7a51dfcb5fe5ae224edf14b200719d4fb8bf5530cffcc6

Observation faba2e41-4dfc-4b35-a2dd-931154e42af9 · outbound

This paper cites Estimating the completeness of discrete speech units,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Estimating the completeness of discrete speech units,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.620815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.490350Z digest=sha256:f8725424bcec7c1ef56765dca0027799513e21d293a61348a22124e8cb1f5874

Observation 1aff2f3f-d144-43b1-9ae7-26d7c0194291 · outbound

This paper cites Augmentation invariant discrete representation for generative spoken language modeling,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Augmentation invariant discrete representation for generative spoken language modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.606133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.494023Z digest=sha256:cc22cd470a05bb811fa5b49c4f580ac0fa39e326a28a6ea737686cd32a5fbd12

Observation 88d4a449-ca1f-4533-bce0-993379af8834 · outbound

This paper cites StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:37:37.080968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.531979Z digest=sha256:1799dfaeab14f6701cd4be0f279dc9f1534ed9b80572b31be8f6989906852d9a

Observation 788381a9-7117-4f0a-8075-db053f96d105 · outbound

This paper cites STAB: Speech Tokenizer Assessment Benchmark.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances STAB: Speech Tokenizer Assessment Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.502980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.502980Z digest=sha256:986d8fcc82ed1b8f6358783d78169d3394bd323a40fb94856e246023a66e018b

Observation 7ac573bf-c9f7-4578-9aff-11cb22ddf585 · outbound

This paper cites Dc- spin: A speaker-invariant speech tokenizer for spoken language models,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Dc- spin: A speaker-invariant speech tokenizer for spoken language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.577027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.507781Z digest=sha256:91795ca51b8aa5769be9c6be1e4d27766a7178b3b909397c8f6bd2363bae58e5

Observation bac9b4da-0f77-450a-a7af-a908838c7e47 · outbound

This paper cites Rethinking discrete speech representation tokens for accent generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Rethinking discrete speech representation tokens for accent generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.563488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.512769Z digest=sha256:45690c2fd2de42940d9ee842ba23440432c3921a5e0d5714457f69cb07118a82

Observation a0a54f54-523e-406b-90c0-9dee637fb615 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Robust Speech Recognition via Large-Scale Weak Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.555809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.555809Z digest=sha256:59042eaf4a685d6e331576367ff40ea1f022e960e43d04de42553f3deb78bacc

Observation 86d3fbae-e589-4fea-8288-7939cd8d6701 · outbound

This paper cites Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.522433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.522433Z digest=sha256:cec494d18ca9333a31de439c94a5f28a8b5018052cfbc1bac65850699490952e

Observation 1f13b3ec-36bc-437f-a1e3-c157d76e754b · outbound

This paper cites NAST: Noise Aware Speech Tokenization for Speech Language Models.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances NAST: Noise Aware Speech Tokenization for Speech Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.527205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.527205Z digest=sha256:510388f2021fa850c7d8d76885e0a81d1eedc5e91c55259770c8eeb73dbf6f2c

Observation 3c456030-231e-48df-b613-ed6a1f945cb7 · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.506664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.573839Z digest=sha256:c2bb8c5beb9a57acadf424c976149488dfa954f5e69fd3fe079d2edfd91121a8

Observation 82920521-3379-497c-9f60-eed4e20f0fba · outbound

This paper cites Dualcodec: A low-frame-rate, semantically-enhanced neural audio codec for speech generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Dualcodec: A low-frame-rate, semantically-enhanced neural audio codec for speech generation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.537592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.537592Z digest=sha256:dc050de4c13475f7403e64b5da862f1c76d22391c4023be1d2da59c412fbd517

Observation ecb525e6-c774-4269-aad9-03ca55301624 · outbound

This paper cites XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.541957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.541957Z digest=sha256:48f60ff925d91f6564b94a1416399fb93d4df89d0207d621a3f5c020639d75dc

Observation 53c8d88c-fbf8-44e6-be24-86d970fdba9e · outbound

This paper cites Sac: Neural speech codec with semantic-acoustic dual-stream quantization,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Sac: Neural speech codec with semantic-acoustic dual-stream quantization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.549691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.546380Z digest=sha256:9562d2f8361ca02baa7d05e01257e18af90ee71869f644d8f33b31d259d6d776

Observation da88886d-b705-43c7-ab98-b0ecb3e04f3d · outbound

This paper cites The chains corpus: Characterizing individual speakers,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The chains corpus: Characterizing individual speakers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.444811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.591693Z digest=sha256:c8c5829e4c2f920230ce254cb9425b6bc255804d5b22aea813e9e4b640818d1f

Observation 8fd06ec6-4823-49f3-a1ab-76f6db781d33 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.596071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.596071Z digest=sha256:55d21be3b335f106ba202a1e982ab0a922715e4fa9872095eeb308c1382e94f0

Observation 0513336b-4b7e-4198-a1d4-69a8f94e2445 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.536085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.560362Z digest=sha256:82fba012e9d69b5ae486e2d361c98a265b16de06bdd82199257d79c0f552fce7

Observation 23478bcb-0e67-4a27-b311-d4e27796c0a7 · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.413903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.604243Z digest=sha256:4d8b5218751684afe53ce77c158336be7725a12d2eb3084c5bc827c14fb96829

Observation 36c02c41-c9f7-4d4e-91b8-9ff7888675b9 · outbound

This paper cites Kokoro-82m,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Kokoro-82m,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.521724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.569307Z digest=sha256:249843f7148fbbc4c1781110e381523a82e19c4bf91c36158c435fa6790a395b

Observation 249fdd17-c16b-4c20-801b-f1fc4f7f76cc · outbound

This paper cites Darpa timit acoustic phonetic con- tinuous speech corpus cdrom,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Darpa timit acoustic phonetic con- tinuous speech corpus cdrom,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.612873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.612873Z digest=sha256:86ca0c2f10bff168b380b2ee51bb68ea695cb115c6689e8ae6b956cf4c0098d9

Observation 202ec8d5-3725-495b-b14e-ff0d587c5921 · outbound

This paper cites Learning sound event classifiers from web audio with noisy labels,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Learning sound event classifiers from web audio with noisy labels,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.490573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.578405Z digest=sha256:61c585f8f445f47d285738a90b805112d57eda1f835269424874b63ade20a23b

Observation 5d8b0ac3-d261-4242-b2a6-0d84029a0fd8 · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using kaldi,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Montreal forced aligner: Trainable text-speech align- ment using kaldi,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.473348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.582751Z digest=sha256:f160343d867f627aabb79f099d828bb8106369c94928411107905b21f90fafd0

Observation b4983d30-63a2-4b2c-9469-83fc06b98bbb · outbound

This paper cites The cmu arctic speech databases,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The cmu arctic speech databases,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.459318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.587397Z digest=sha256:8861038d31bf4f29316a175affcc9383df69ec63830b171f8ef06e522acd1c6f

Observation d5f5f318-0d7e-400b-a3b6-fd88c3fa8697 · outbound

This paper cites Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.362800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.629857Z digest=sha256:7119ba801966f9100d8158203fc5dd1c365b6c136c9c04d18ffe9f70294c46e3

Observation 522fa691-bc23-42c5-a351-dc5f53aab19e · outbound

This paper cites Soft-dtw: a differentiable loss func- tion for time-series,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Soft-dtw: a differentiable loss func- tion for time-series,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.345344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.634292Z digest=sha256:5a9cbc442c15b32f0f4db805d9ef35de7428edfc268c296c59ce6b0b42c42831

Observation 5574bffc-00ab-4c3f-a271-54df18374705 · outbound

This paper cites Open-source multi-speaker corpora of the English accents in the British isles,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Open-source multi-speaker corpora of the English accents in the British isles,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.429609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.600245Z digest=sha256:ca068db64a1f839ab1709476082a04a15c69d9a946cbbe65db3e3a2002a42993

Observation 82939289-67ef-4a9f-9fdf-fbf74550678d · outbound

This paper cites Mean teachers are better role models: Weight-averaged consistency targets improve semi- supervised deep learning results,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Mean teachers are better role models: Weight-averaged consistency targets improve semi- supervised deep learning results,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.312918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.642925Z digest=sha256:93da8fe9781fe711b323d2d9aa709afe44c843997948b216caef035b7d2edec9

Observation 329955bc-1630-41c4-b757-18b1d5e1a461 · outbound

This paper cites Learning Disentangled Speech Representations.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Learning Disentangled Speech Representations

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:37:36.892990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.608319Z digest=sha256:5ae9ecca7ee42acc2bc626dbcc64986c06bca465bcaedaaefa1b3ff9dd39ec42

Observation 05f03d3a-5db6-4f46-992a-1b274dd7c4fd · outbound

This paper cites Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.297591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.651768Z digest=sha256:48abe3a3d45efb188918fe104ab0d1598d344294f6e4e14a756b49079977645b

Observation d4590b98-b95a-4637-87c8-9ba486d5eae0 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Lib- rispeech: an asr corpus based on public domain audio books,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.617067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.617067Z digest=sha256:b287fd966012cec053572cf6eab3004af86032b9e6a8d0d033f3ac2bce6013cd

Observation 7e20d7c1-efd6-4333-8d97-b160364a448f · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.378582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.621152Z digest=sha256:b0560fcc1c05baace4fdf40859239615094e93a2fcf8cdc8d5be2efcb3b6956a

Observation eaf1e97c-3c83-40af-8879-2ae53a5176df · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Libri-light: A benchmark for asr with limited or no supervision,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.625423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.625423Z digest=sha256:83815cff0b39ce9b57ff9ca14de264bcf5a5192a1949d58ecd31260d828b257b

Observation 598fdfe8-c382-4126-aeac-c811a3627b47 · outbound

This paper cites Textless speech-to-speech translation on real data,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Textless speech-to-speech translation on real data,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.329727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.638543Z digest=sha256:235c3e83a2da46660fd797ca140a393368b48ad3b53700dc89baaeb7f4a5268e

Observation 99d84a0e-d1f6-4749-aed6-5d6fce1da5a9 · outbound

This paper cites Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.647187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.647187Z digest=sha256:758ed64f68486cb4229c659db375a81cc1149a53830d06f90376b901494778f4

Observation c145d71b-3f7a-490b-9298-274e2256af8f · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances X-vectors: Robust dnn embeddings for speaker recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.281487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.656380Z digest=sha256:327a7d2260eed2f282907a69aa206053af2b96d137f65f18304ef5b4d293d7bf

Observation f0fcf3c6-851d-4378-8a71-0801f38f29b2 · outbound

This paper cites Available: https://aclanthology.org/2023.iwslt-1.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://aclanthology.org/2023.iwslt-1

Reference 477

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.591389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.498516Z digest=sha256:cc8f43fe4db81dd3caaea86ab9be06407e51c951d751706f9830ce2c99dccf5b

Observation 623f462a-0fe4-4be0-9f4e-bf8f3c8b2a43 · outbound

This paper cites W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.564661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.564661Z digest=sha256:00f5d4cb4f3f9f00ec19757973efabcf6b214c0060d0c9b153060810a5502ca3

Observation bd2e5f66-ea7e-40f2-b825-38d5ceec6c26 · outbound

This paper cites Available: https://arxiv.org/abs/2510.16841.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://arxiv.org/abs/2510.16841

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.551536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.551536Z digest=sha256:e4a40214db813a2a097636f1bd857e2ef922f649794e1c2779dacfd58bc951e4

Observation b7c1552a-d0ca-4218-83b3-dbcf5865334b · outbound

This paper cites Available: https://arxiv.org/abs/2601.19786.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://arxiv.org/abs/2601.19786

Reference 2026

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:37:37.184946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T15:37:36.517410Z digest=sha256:9ad249524e051e12787b419c2a02bd20f1ed56483088c8547947f5f5203ec0de

Pith citing papers

Observation b2af3fbb-4cc0-42f9-ace0-8327fc1aeac3 · inbound

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances cites this paper.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.420218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.420218Z digest=sha256:15cb0ebf46563294fae97d674110707d544a7b9c9ca3b882bfafde15aaeb4e76

Observation 8f600229-0b91-43ba-9765-7dbfcefa954f · inbound

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure cites this paper.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:15:23.478725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T00:15:23.039683Z digest=sha256:a9808e5b40baa4908d8df38c16e1462274849576e9f2e45f077556cb9fcee630