Pith. sign in

Paper Citation Record · LEDGER

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

As of 16 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2607.19033.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19033 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.656380Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.420218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:15:23.475461Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact3
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2af3fbb-4cc0-42f9-ace0-8327fc1aeac3 · outbound

This paper cites Content is What Remains: Invariant Speech Tokenization from Parallel Utterances.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.420218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.420218Z digest=sha256:877cd3ca45e75bfc52031f0f613dff2875d2e1ffbec1d16067039dfa28b52d9e

Observation 36122b67-fcf0-4c05-b518-26339874cf68 · outbound

This paper cites an unresolved cited work.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:37:37.773073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.426600Z digest=sha256:83f81cb0a9c1c1b493612eedb74e8e3dd4ddeb90bdf79c916e2b9cf3149dcb0c

Observation 314f73bb-bf00-4cba-84c2-caa1eccad9cc · outbound

This paper cites Further we evaluate the downstream compres- sion benefit unlocked by the above.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Further we evaluate the downstream compres- sion benefit unlocked by the above

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T15:37:37.757774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.431802Z digest=sha256:b153d7bcc4ecb4c67c83b81f43c7a3adab7d237e17484d514dcbc0e2c6af9b40

Observation e619767f-ffcb-42b7-92b3-ccf4b6312d00 · outbound

This paper cites an unresolved cited work.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:37:37.742427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.436685Z digest=sha256:70ce9ea321528f268d5bea149d5c1e2fa4a4b4c06915b74cc1cbfeb8f8d39e79

Observation f6d5b190-beab-4257-8444-22f4875b90b3 · outbound

This paper cites The experiments were run manually and results were manually verified.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The experiments were run manually and results were manually verified

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.726711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.441464Z digest=sha256:2a1d5b23c47f1d68f1aab103da6bf6f5570dd52dcce6f8e3256fd9ae8fd4fe0b

Observation de499ba7-9b34-418c-9973-674a7badd1cc · outbound

This paper cites AudioLM: a language modeling approach to audio generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances AudioLM: a language modeling approach to audio generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.710583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.446270Z digest=sha256:76ce5ae10a4f6bb7b552bdd090a526a21363ad05c2b6aeccf5c17547cdfc19dc

Observation fc188f6a-54ef-425d-9b17-b5e03b654f03 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.696744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.450990Z digest=sha256:c9fef91410b073b4202a114870ea8110230b437bc8e38e973ad854a0c354b465

Observation 33894dc6-3747-4fd1-9516-1c6b297ca72e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Moshi: a speech-text foundation model for real-time dialogue

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.456476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.456476Z digest=sha256:1df8fe081a5f9a7c3a359118e6b0a0a1d55eaa0f92da7fd2061d18ec1ee8808b

Observation 63a33f02-57d5-4c38-ba68-5d94e182fff3 · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.461335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.461335Z digest=sha256:f591ce968c5b495c566dd5f6dc0d9c401663344bf64ea8fc8a388f67aeea7850

Observation 8ee4ab36-a958-40c1-a6f9-d0f6bd0ea221 · outbound

This paper cites SpeechTok- enizer: Unified speech tokenizer for speech language models,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances SpeechTok- enizer: Unified speech tokenizer for speech language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.681063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.466266Z digest=sha256:cd83dd3f7f5a115aa0e2cbbd1dc86f8fd3371ec09b40217220c8b183ef140035

Observation ef9bd97c-7591-4b86-a813-f90d20517ba1 · outbound

This paper cites AudioDec: An open-source streaming high-fidelity neural audio codec,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances AudioDec: An open-source streaming high-fidelity neural audio codec,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.664857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.470649Z digest=sha256:9ed8d182006856e4e9f6905c3eec8f128d50408c9869da45072fcfbdb1758cba

Observation 1cbd49f7-3425-4cf6-9c55-a96ff4db9109 · outbound

This paper cites Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.474582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.474582Z digest=sha256:641e33114a2a3173428a9c11365656f297f68233238393c1f2cf6d6ff6b17340

Observation 6e300fcc-ab13-44cf-9cad-b4d7168681a6 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.478614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.478614Z digest=sha256:ccd7be3233ed55490b797aa4923ecb21341d12e4598627c1ba8155db901fc106

Observation 7799387e-c3bd-4bd1-9ae7-8f47588a0dcd · outbound

This paper cites ContentVec: An improved self-supervised speech representation by disentangling speakers,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances ContentVec: An improved self-supervised speech representation by disentangling speakers,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.486606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.486606Z digest=sha256:4854367f3ff49a6e12c443974efc62d3dd19f0c10de933457b3eb0913988690d

Observation faba2e41-4dfc-4b35-a2dd-931154e42af9 · outbound

This paper cites Estimating the completeness of discrete speech units,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Estimating the completeness of discrete speech units,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.620815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.490350Z digest=sha256:e9b7c0db4737dd14a3e8da1ca31de81668a131cb03c344f983cf66f59fa2692f

Observation 1aff2f3f-d144-43b1-9ae7-26d7c0194291 · outbound

This paper cites Augmentation invariant discrete representation for generative spoken language modeling,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Augmentation invariant discrete representation for generative spoken language modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.606133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.494023Z digest=sha256:5a57e909651f7995301650cf4fd3d558b8f3bf2ff5888d8a548f3e6e72ce1d57

Observation 88d4a449-ca1f-4533-bce0-993379af8834 · outbound

This paper cites StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:37:37.080968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.531979Z digest=sha256:ef551004dfe7198544e1af59a375dd599a4c8e2ee571c11b95a08ba287e40c59

Observation 788381a9-7117-4f0a-8075-db053f96d105 · outbound

This paper cites STAB: Speech Tokenizer Assessment Benchmark.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances STAB: Speech Tokenizer Assessment Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.502980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.502980Z digest=sha256:ba9e56e7b7b040ecc8158a3c156696347a5afc4de194d0ca1c6ebcf15d802aba

Observation 7ac573bf-c9f7-4578-9aff-11cb22ddf585 · outbound

This paper cites Dc- spin: A speaker-invariant speech tokenizer for spoken language models,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Dc- spin: A speaker-invariant speech tokenizer for spoken language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.577027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.507781Z digest=sha256:8f1cced65f1af97d59bfd737d8871a837c9b846878e4b4737edf7e011466eff5

Observation bac9b4da-0f77-450a-a7af-a908838c7e47 · outbound

This paper cites Rethinking discrete speech representation tokens for accent generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Rethinking discrete speech representation tokens for accent generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.563488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.512769Z digest=sha256:f1e0ab624d2628ece6b3b09e8a8350b0f784531bcf4800bbcceeb347214bf136

Observation a0a54f54-523e-406b-90c0-9dee637fb615 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Robust Speech Recognition via Large-Scale Weak Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.555809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.555809Z digest=sha256:2be5b4fbba40e4e6c66589cc3d7d42bf2b45b10e72b53223a6d79a88a6c53324

Observation 86d3fbae-e589-4fea-8288-7939cd8d6701 · outbound

This paper cites Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.522433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.522433Z digest=sha256:00185b9da883568a73e347e241ca0705f9eaa842e4a591767ab70b932cc48612

Observation 1f13b3ec-36bc-437f-a1e3-c157d76e754b · outbound

This paper cites NAST: Noise Aware Speech Tokenization for Speech Language Models.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances NAST: Noise Aware Speech Tokenization for Speech Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.527205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.527205Z digest=sha256:8430089a9079e39fd2caec5793ed157510ac9e76d9b77135cc6b420eecd4ec14

Observation 3c456030-231e-48df-b613-ed6a1f945cb7 · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.506664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.573839Z digest=sha256:0e9b0627fe7af48e53a558e93c369e029bcc5cccf0a31d31cad60bc5807c3ef0

Observation 82920521-3379-497c-9f60-eed4e20f0fba · outbound

This paper cites Dualcodec: A low-frame-rate, semantically-enhanced neural audio codec for speech generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Dualcodec: A low-frame-rate, semantically-enhanced neural audio codec for speech generation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.537592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.537592Z digest=sha256:9b086d655762a0c8e50bddc33a70a74c082e808aacd6beac89ae804568af744f

Observation ecb525e6-c774-4269-aad9-03ca55301624 · outbound

This paper cites XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.541957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.541957Z digest=sha256:5a5bfaa2f7d53dd352d8cdcfb44edc6e0d98a93ce1063ddcd2cd063c09c0a93d

Observation 53c8d88c-fbf8-44e6-be24-86d970fdba9e · outbound

This paper cites Sac: Neural speech codec with semantic-acoustic dual-stream quantization,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Sac: Neural speech codec with semantic-acoustic dual-stream quantization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.549691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.546380Z digest=sha256:8a47f1651ff2b8853a0bd0581bb8fa2278f5fdd9f48f483e369117818f96422d

Observation da88886d-b705-43c7-ab98-b0ecb3e04f3d · outbound

This paper cites The chains corpus: Characterizing individual speakers,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The chains corpus: Characterizing individual speakers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.444811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.591693Z digest=sha256:4eaa5be1f04003df0274b557e588e040af1b98a22ee6e4e8d39f520a8ed5234b

Observation 8fd06ec6-4823-49f3-a1ab-76f6db781d33 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.596071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.596071Z digest=sha256:2474857608e6ee58251d405bdc938f681c20dcfed63d73418da74c57b5c81340

Observation 0513336b-4b7e-4198-a1d4-69a8f94e2445 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.536085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.560362Z digest=sha256:defe4d582f58f523eafd178f7f435d3b2d202f8e1bb40c7d96ff57730387e421

Observation 23478bcb-0e67-4a27-b311-d4e27796c0a7 · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.413903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.604243Z digest=sha256:d99e6f7d973b897690d20b5d3f43441d16379d4402e49702589d4ee1463f4add

Observation 36c02c41-c9f7-4d4e-91b8-9ff7888675b9 · outbound

This paper cites Kokoro-82m,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Kokoro-82m,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.521724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.569307Z digest=sha256:e0f4abb514225c849d34d1de1632f75c00a3c058201fcca8843ff4581f58fe10

Observation 249fdd17-c16b-4c20-801b-f1fc4f7f76cc · outbound

This paper cites Darpa timit acoustic phonetic con- tinuous speech corpus cdrom,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Darpa timit acoustic phonetic con- tinuous speech corpus cdrom,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.612873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.612873Z digest=sha256:d2ebc424cf5e184382268c37ace9e82bf1e40316edbb5b53eb8b2d493c2b9c2d

Observation 202ec8d5-3725-495b-b14e-ff0d587c5921 · outbound

This paper cites Learning sound event classifiers from web audio with noisy labels,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Learning sound event classifiers from web audio with noisy labels,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.490573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.578405Z digest=sha256:57fac6eb01fc75a0f9b14aa11c7515268ffc613aeb82cbf288a4c409fbd45ca9

Observation 5d8b0ac3-d261-4242-b2a6-0d84029a0fd8 · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using kaldi,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Montreal forced aligner: Trainable text-speech align- ment using kaldi,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.473348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.582751Z digest=sha256:a5da363ed23e50c285b38fc24849e3fe9b81511582c51e40e12c52cce4f34561

Observation b4983d30-63a2-4b2c-9469-83fc06b98bbb · outbound

This paper cites The cmu arctic speech databases,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The cmu arctic speech databases,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.459318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.587397Z digest=sha256:512de2ce144740b96591d0d30ea9a6ce8f306ac53240743f8362c997a8c6896b

Observation d5f5f318-0d7e-400b-a3b6-fd88c3fa8697 · outbound

This paper cites Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.362800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.629857Z digest=sha256:9133cb8e80065e390a0ed45f15ce40f459fe8120854e345c34e99bb718adc9fe

Observation 522fa691-bc23-42c5-a351-dc5f53aab19e · outbound

This paper cites Soft-dtw: a differentiable loss func- tion for time-series,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Soft-dtw: a differentiable loss func- tion for time-series,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.345344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.634292Z digest=sha256:32ad36a84e4f5a948e411599c3b3e8135bfcc81ad076d02a6b3651f5a582ac47

Observation 5574bffc-00ab-4c3f-a271-54df18374705 · outbound

This paper cites Open-source multi-speaker corpora of the English accents in the British isles,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Open-source multi-speaker corpora of the English accents in the British isles,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.429609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.600245Z digest=sha256:4950ce41cd691dfff21a3a14afabcf7fed71a8bb1997e50cdf9795730fd3716f

Observation 82939289-67ef-4a9f-9fdf-fbf74550678d · outbound

This paper cites Mean teachers are better role models: Weight-averaged consistency targets improve semi- supervised deep learning results,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Mean teachers are better role models: Weight-averaged consistency targets improve semi- supervised deep learning results,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.312918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.642925Z digest=sha256:7db4f9d363d752ccb364909b94d671c27c8d4165408ff3af234a099b68cf03c3

Observation 329955bc-1630-41c4-b757-18b1d5e1a461 · outbound

This paper cites Learning Disentangled Speech Representations.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Learning Disentangled Speech Representations

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:37:36.892990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.608319Z digest=sha256:490f6d3f0597083995d54945fb8eee6c254fa4053193895f9679eb4b45bbe93a

Observation 05f03d3a-5db6-4f46-992a-1b274dd7c4fd · outbound

This paper cites Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.297591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.651768Z digest=sha256:80fd7980f6cec826cc2c7b04b819bd03291b1fdae12f49c4560910ca1069b0ec

Observation d4590b98-b95a-4637-87c8-9ba486d5eae0 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Lib- rispeech: an asr corpus based on public domain audio books,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.617067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.617067Z digest=sha256:dc199be1c28837b1f58d34a7185a83ced9b0ab1cf9614bcac881ec7f36ed85f2

Observation 7e20d7c1-efd6-4333-8d97-b160364a448f · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.378582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.621152Z digest=sha256:bb8234a0bbcf7c4b798d72a529dde5e158870fc91ce143ac01e589208e011398

Observation eaf1e97c-3c83-40af-8879-2ae53a5176df · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Libri-light: A benchmark for asr with limited or no supervision,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.625423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.625423Z digest=sha256:07909a017012f8240f3f50a0094a11c5065eec5bc87e577307b591669487ef40

Observation 598fdfe8-c382-4126-aeac-c811a3627b47 · outbound

This paper cites Textless speech-to-speech translation on real data,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Textless speech-to-speech translation on real data,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.329727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.638543Z digest=sha256:6ee7ea9c476163e4a1ecdd0c8655d5e180ae19655ab08834fa5ff87a497266b0

Observation 99d84a0e-d1f6-4749-aed6-5d6fce1da5a9 · outbound

This paper cites Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.647187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.647187Z digest=sha256:755e147e6d3cb1e3e73c8205a747acc9bfa3284f9593dcf62bb3eb767ee17798

Observation c145d71b-3f7a-490b-9298-274e2256af8f · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances X-vectors: Robust dnn embeddings for speaker recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.281487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.656380Z digest=sha256:7ff1c03ce95bc91be4bec2d8f722ee81cb2bf5963c0d7aa447d7b5d6dfae1af9

Observation f0fcf3c6-851d-4378-8a71-0801f38f29b2 · outbound

This paper cites Available: https://aclanthology.org/2023.iwslt-1.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://aclanthology.org/2023.iwslt-1

Reference 477

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.591389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.498516Z digest=sha256:bd889e629696522875c6f7abd4f8e51f7f2d1a41663894ad8f4f9befb7bb20f8

Observation 623f462a-0fe4-4be0-9f4e-bf8f3c8b2a43 · outbound

This paper cites W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.564661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.564661Z digest=sha256:4215e2a5c64f341aa41d6304427fe0ccf1c36f92d780b81a4fec98d84345b765

Observation bd2e5f66-ea7e-40f2-b825-38d5ceec6c26 · outbound

This paper cites Available: https://arxiv.org/abs/2510.16841.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://arxiv.org/abs/2510.16841

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.551536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.551536Z digest=sha256:53c84ab50b1dd8f79f16637304ea2e0a37e7998ff5a6f10bc4be761b4ce8af22

Observation b7c1552a-d0ca-4218-83b3-dbcf5865334b · outbound

This paper cites Available: https://arxiv.org/abs/2601.19786.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://arxiv.org/abs/2601.19786

Reference 2026

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:37:37.184946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T15:37:36.517410Z digest=sha256:cfa6f38c0138b227f5983097fd4d7b0c3638c3252718dc465d6d82b57374a555

Pith citing papers

Observation b2af3fbb-4cc0-42f9-ace0-8327fc1aeac3 · inbound

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances cites this paper.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.420218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.420218Z digest=sha256:877cd3ca45e75bfc52031f0f613dff2875d2e1ffbec1d16067039dfa28b52d9e

Observation 8f600229-0b91-43ba-9765-7dbfcefa954f · inbound

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure cites this paper.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:15:23.478725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:23.039683Z digest=sha256:251feba3accb78a546c41000490413b1b2778eeaa9d79ebdc9c1fddbb4626e3c