Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.656380Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2607.19033.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.656380Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.420218Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-12T00:15:23.475461Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b2af3fbb-4cc0-42f9-ace0-8327fc1aeac3 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36122b67-fcf0-4c05-b518-26339874cf68 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 314f73bb-bf00-4cba-84c2-caa1eccad9cc · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Further we evaluate the downstream compres- sion benefit unlocked by the above
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e619767f-ffcb-42b7-92b3-ccf4b6312d00 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f6d5b190-beab-4257-8444-22f4875b90b3 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The experiments were run manually and results were manually verified
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation de499ba7-9b34-418c-9973-674a7badd1cc · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances AudioLM: a language modeling approach to audio generation,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fc188f6a-54ef-425d-9b17-b5e03b654f03 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 33894dc6-3747-4fd1-9516-1c6b297ca72e · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Moshi: a speech-text foundation model for real-time dialogue
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a33f02-57d5-4c38-ba68-5d94e182fff3 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee4ab36-a958-40c1-a6f9-d0f6bd0ea221 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances SpeechTok- enizer: Unified speech tokenizer for speech language models,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ef9bd97c-7591-4b86-a813-f90d20517ba1 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances AudioDec: An open-source streaming high-fidelity neural audio codec,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1cbd49f7-3425-4cf6-9c55-a96ff4db9109 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e300fcc-ab13-44cf-9cad-b4d7168681a6 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Wavlm: Large-scale self-supervised pre-training for full stack speech processing,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7799387e-c3bd-4bd1-9ae7-8f47588a0dcd · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances ContentVec: An improved self-supervised speech representation by disentangling speakers,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faba2e41-4dfc-4b35-a2dd-931154e42af9 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Estimating the completeness of discrete speech units,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1aff2f3f-d144-43b1-9ae7-26d7c0194291 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Augmentation invariant discrete representation for generative spoken language modeling,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 88d4a449-ca1f-4533-bce0-993379af8834 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 788381a9-7117-4f0a-8075-db053f96d105 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances STAB: Speech Tokenizer Assessment Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac573bf-c9f7-4578-9aff-11cb22ddf585 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Dc- spin: A speaker-invariant speech tokenizer for spoken language models,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bac9b4da-0f77-450a-a7af-a908838c7e47 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Rethinking discrete speech representation tokens for accent generation,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a0a54f54-523e-406b-90c0-9dee637fb615 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Robust Speech Recognition via Large-Scale Weak Supervision
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86d3fbae-e589-4fea-8288-7939cd8d6701 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f13b3ec-36bc-437f-a1e3-c157d76e754b · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances NAST: Noise Aware Speech Tokenization for Speech Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c456030-231e-48df-b613-ed6a1f945cb7 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 82920521-3379-497c-9f60-eed4e20f0fba · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Dualcodec: A low-frame-rate, semantically-enhanced neural audio codec for speech generation,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb525e6-c774-4269-aad9-03ca55301624 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c8d88c-fbf8-44e6-be24-86d970fdba9e · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Sac: Neural speech codec with semantic-acoustic dual-stream quantization,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation da88886d-b705-43c7-ab98-b0ecb3e04f3d · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The chains corpus: Characterizing individual speakers,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8fd06ec6-4823-49f3-a1ab-76f6db781d33 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0513336b-4b7e-4198-a1d4-69a8f94e2445 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 23478bcb-0e67-4a27-b311-d4e27796c0a7 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 36c02c41-c9f7-4d4e-91b8-9ff7888675b9 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Kokoro-82m,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 249fdd17-c16b-4c20-801b-f1fc4f7f76cc · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Darpa timit acoustic phonetic con- tinuous speech corpus cdrom,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202ec8d5-3725-495b-b14e-ff0d587c5921 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Learning sound event classifiers from web audio with noisy labels,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5d8b0ac3-d261-4242-b2a6-0d84029a0fd8 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Montreal forced aligner: Trainable text-speech align- ment using kaldi,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b4983d30-63a2-4b2c-9469-83fc06b98bbb · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The cmu arctic speech databases,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d5f5f318-0d7e-400b-a3b6-fd88c3fa8697 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 522fa691-bc23-42c5-a351-dc5f53aab19e · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Soft-dtw: a differentiable loss func- tion for time-series,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5574bffc-00ab-4c3f-a271-54df18374705 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Open-source multi-speaker corpora of the English accents in the British isles,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 82939289-67ef-4a9f-9fdf-fbf74550678d · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Mean teachers are better role models: Weight-averaged consistency targets improve semi- supervised deep learning results,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 329955bc-1630-41c4-b757-18b1d5e1a461 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Learning Disentangled Speech Representations
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 05f03d3a-5db6-4f46-992a-1b274dd7c4fd · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d4590b98-b95a-4637-87c8-9ba486d5eae0 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Lib- rispeech: an asr corpus based on public domain audio books,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e20d7c1-efd6-4333-8d97-b160364a448f · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eaf1e97c-3c83-40af-8879-2ae53a5176df · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Libri-light: A benchmark for asr with limited or no supervision,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598fdfe8-c382-4126-aeac-c811a3627b47 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Textless speech-to-speech translation on real data,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 99d84a0e-d1f6-4749-aed6-5d6fce1da5a9 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c145d71b-3f7a-490b-9298-274e2256af8f · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances X-vectors: Robust dnn embeddings for speaker recognition,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f0fcf3c6-851d-4378-8a71-0801f38f29b2 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://aclanthology.org/2023.iwslt-1
Reference 477
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 623f462a-0fe4-4be0-9f4e-bf8f3c8b2a43 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd2e5f66-ea7e-40f2-b825-38d5ceec6c26 · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://arxiv.org/abs/2510.16841
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7c1552a-d0ca-4218-83b3-dbcf5865334b · outbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://arxiv.org/abs/2601.19786
Reference 2026
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b2af3fbb-4cc0-42f9-ace0-8327fc1aeac3 · inbound
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f600229-0b91-43ba-9765-7dbfcefa954f · inbound
ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.