Pith. sign in

Paper Citation Record · LEDGER

Representing Speech Through Autoregressive Prediction of Cochlear Tokens

As of 23 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2508.11598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11598 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:58.755219Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:58.525812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T19:54:59.403728Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact4
  • verified fuzzy35
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c87c66bf-5d6a-4851-91b1-d73aff9c5ccf · outbound

This paper cites Representing Speech Through Autoregressive Prediction of Cochlear Tokens.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representing Speech Through Autoregressive Prediction of Cochlear Tokens

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T19:54:59.463829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.525812Z digest=sha256:3bc222763039ef6d89387b42165b152d48b9dbbae466ca84ecf44096628197ce

Observation e70517cc-c426-43e4-b519-7ec9eba9316d · outbound

This paper cites water” and “river.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens water” and “river

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.623324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.530279Z digest=sha256:12e0370201541939da9edce0b523da37489893c37fa11cafbe675836585f0a0c

Observation 1dbc760f-0472-4349-bb5c-34037b0f7723 · outbound

This paper cites er” was often confused with “r.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens er” was often confused with “r

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T19:55:03.613327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.534959Z digest=sha256:a6fbbf04ac49d081d42daf90546a4678dc704d850af05ef8a8bc04d8a7c4d077

Observation 8418ae8c-cc49-47e6-8870-8f8d089400a8 · outbound

This paper cites Transformation Imitation.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Transformation Imitation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.603935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.538930Z digest=sha256:ac1e9a3dd0a134e71aa8880a51e40bec7302ccec43d095e6fa815134b18b5016

Observation 87b8727c-cc16-4faa-a61c-702bb755647c · outbound

This paper cites acknowledges support from The K.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens acknowledges support from The K

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.594277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.542617Z digest=sha256:a663da95771fc47c2c0a375c97a2b33478991c58f299ee97a651ae3b96f0a6a9

Observation d35525ca-2a30-43d0-9329-ae571f573311 · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens SUPERB: Speech processing Universal PERformance Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.546041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.546041Z digest=sha256:9ebece583c4bd5e6d1f1e16320815e8f8e27c0957ae088b1df4177285d2a6825

Observation 7257bd88-eb8c-483a-84c1-1faded21bf78 · outbound

This paper cites Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.341956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.550290Z digest=sha256:9ac7ce704923ef5fb4b68061467fe42c89704e3b679fca37d6090c842f037fae

Observation ba6f1da7-8ec5-4af4-a722-1c40c19dfc80 · outbound

This paper cites A Review of Deep Learning Techniques for Speech Processing.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens A Review of Deep Learning Techniques for Speech Processing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.553897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.553897Z digest=sha256:5832ecbcf9f734537195a90cc01d2505fb352076fa361c66f2dc5b617db399e3

Observation 86bff55e-0e93-4447-aa6c-b3938d7f5878 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Soundstream: An end-to-end neural audio codec,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.558742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.558742Z digest=sha256:c8dbab6eb4a38b0bee3ad0f5939b54a3efaa3a8801808e7f4b13eb10a1b4cc71

Observation 587e6140-14f7-4998-a625-0c5aa2904316 · outbound

This paper cites High Fidelity Neural Audio Compression.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens High Fidelity Neural Audio Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.561899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.561899Z digest=sha256:5d49b5619623ba24a14a7de45f0fbb0dfff71655221ddec3057ebd7066a4c4cf

Observation f225f353-cfd7-4e63-b6bc-70213e002bb5 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.565677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.565677Z digest=sha256:f13e809273bdd51832f688da9747a1847e8e8ee221c89b8e1661220f119fa9fa

Observation 4075519f-d176-4b8a-bce3-882d355bb389 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens High-fidelity audio compression with improved rvqgan,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.577698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.569242Z digest=sha256:d647fb4a9a01af460da7c03a92c362edbab4f5363fc780ab3743d64ea96f4491

Observation b43b3d0a-1d76-4908-9c67-cf1874151514 · outbound

This paper cites Language-Codec: Bridging Discrete Codec Representations and Speech Language Models.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language-Codec: Bridging Discrete Codec Representations and Speech Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.572361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.572361Z digest=sha256:fabb01c8627e73f87c8266c0a2ae10580caa6bf03f71db2a4a5d2a9cc9daea50

Observation 3f67f68b-b935-425c-86b6-f24317671d3e · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.566990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.575845Z digest=sha256:d40bb00ae680dfd29e4855d61adbbfe0f4db86eff939162febc44c050aac50e3

Observation 61224e10-6883-4e89-a553-dd944a8b051c · outbound

This paper cites CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.579073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.579073Z digest=sha256:ec154f35ffde02817ed8c78afb7c5b619824bdaa78239df81f0d6c71af8d92a7

Observation 1886f8f4-88fc-4818-9a5b-9a573165cfbb · outbound

This paper cites Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.582582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.582582Z digest=sha256:7aa5f868d2a4cd5f5feb4c056694cd8491ee3894e4a962ad61e81e3b940528e6

Observation 7f87b3e5-0f62-484c-a2bf-6ee9e158b49b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.586186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.586186Z digest=sha256:9763636faeacaaefeabc47143043cfd25e7c5f573beef550e83bfdaf743e5a2e

Observation e7744f64-4ab0-4588-a3cc-55690244c93e · outbound

This paper cites Hierarchical organization of hu- man auditory cortex: evidence from acoustic invariance in the re- sponse to intelligible speech,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Hierarchical organization of hu- man auditory cortex: evidence from acoustic invariance in the re- sponse to intelligible speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.462945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.589508Z digest=sha256:b3751d7e18f9c1f0fdab4521f5f8a3a8016efa19657dfa81f0d91230b9555d92

Observation 5d4cc089-2728-4852-b9a9-0427f40f18a8 · outbound

This paper cites Intonational speech prosody encoding in the human auditory cortex,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Intonational speech prosody encoding in the human auditory cortex,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.286871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.592847Z digest=sha256:4ce1f89fbd179e24cd3209a23eda8e5e8cc9f555731183dffc450d1d970dcc03

Observation 8464c6ed-1479-49e3-b937-6b7d393d173a · outbound

This paper cites Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.595978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.595978Z digest=sha256:8559b9c63386f264027567d063cc3b0b0c85e676f76a53b7ff9b72de12b576f8

Observation 9516fb69-478a-4798-bfa4-95e9d6cda6bf · outbound

This paper cites W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.599394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.599394Z digest=sha256:706c4046887eb8adc27402dcd5e4f1b36c6e63e79293bd375e7363af98c54abb

Observation 43f76c7d-ebb6-4dc9-9aa6-e9fbefa6023f · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.602756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.602756Z digest=sha256:77cc35a02f5819f9c211afa4829df40448a61dc9a8870d54a3520ec6e6e3befd

Observation ed71dc35-6842-4fef-9545-5fdea44e45e5 · outbound

This paper cites Blind phoneme segmentation with temporal prediction errors.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Blind phoneme segmentation with temporal prediction errors

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.172535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.605940Z digest=sha256:73d348afb91811f184470691b141b9689f816d74cc145f976052d7193ac6ca95

Observation b177be0e-b3b8-4219-b791-7006268441cc · outbound

This paper cites An Unsupervised Autoregressive Model for Speech Representation Learning.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens An Unsupervised Autoregressive Model for Speech Representation Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.609844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.609844Z digest=sha256:6bd96bdf818e6a79bdb3f10f25223e5ee10d1ba25601383933feb4b6c0785ccc

Observation 73d39a96-2ea6-41e1-881d-0a986ab7417b · outbound

This paper cites Acquiring language from speech by learning to remember and predict,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Acquiring language from speech by learning to remember and predict,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.113543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.613387Z digest=sha256:d2e5f3b1959f1e253f05109f2b03c056bed6a33d2d45b805209312edf65af92d

Observation 04458adf-6e74-495f-ae62-26106e634722 · outbound

This paper cites Vector-Quantized Autoregressive Predictive Coding.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Vector-Quantized Autoregressive Predictive Coding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.616637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.616637Z digest=sha256:b4f2cc9f8b82a6ac7cf8c67567f5399e6532882900a29b3faeef5921ae04a051

Observation 4125ea9c-5623-4cdf-97da-333987beef75 · outbound

This paper cites Audio albert: A lite bert for self-supervised learning of audio representation,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Audio albert: A lite bert for self-supervised learning of audio representation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.985774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.621129Z digest=sha256:e1e7d752fa37490ebaecf20306318f7ea8d6841bd43ec05f077f590ad63ca66c

Observation 6774d475-e580-4647-9742-715858ad99a6 · outbound

This paper cites On generative spoken language modeling from raw au- dio,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens On generative spoken language modeling from raw au- dio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.859177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.624665Z digest=sha256:38ebf7e54112df732c46221985c3b12c131cd7e58a1e8f26a2b80efcfdd4d391

Observation c35d794d-6ba2-43a0-a662-565838525c7c · outbound

This paper cites Audiolm: a language modeling approach to audio gener- ation,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Audiolm: a language modeling approach to audio gener- ation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.628627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.628627Z digest=sha256:3764cde1f2d3271dcb17de0c48ee3164b7b8f6c27c08c5b69b5190ec31171d95

Observation 4a55b06c-d015-495f-a306-39f267076dbf · outbound

This paper cites Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.113274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.632070Z digest=sha256:d31a8000573a46c83efaccf80ab7c9bc0557cf284461f2097d83954ff147a4f5

Observation 4eb29282-33de-4d8d-b2ac-40aa568d4df9 · outbound

This paper cites BERT: Pre- training of Deep Bidirectional Transformers for Language Under- standing,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens BERT: Pre- training of Deep Bidirectional Transformers for Language Under- standing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.733522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.635519Z digest=sha256:2fcc40ce1ca0cb4882ce42fc3d99a23a41ebf21bab4f4e0bbcdefbc7dd42c262

Observation 54f7547d-408e-47fd-b58f-3c2612a5e04e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.639389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.639389Z digest=sha256:3a1c0cf555f3e34df69d1bd362ec48c34db7ca9671867cc77d9ff16b473dc130

Observation 05670626-1df8-47d1-a877-7bccc75d583d · outbound

This paper cites Wav2vec 2.0: a framework for self-supervised learning of speech representa- tions,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Wav2vec 2.0: a framework for self-supervised learning of speech representa- tions,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.595018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.642906Z digest=sha256:d5e569e8fc94f2ad75e1fc1deaf385886def678bc4830c9b9cfff3e7a29a271d

Observation d0ea69db-a483-47df-a0f9-960dc3ab132b · outbound

This paper cites Contrastive learning of general-purpose audio representations,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Contrastive learning of general-purpose audio representations,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.473345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.647061Z digest=sha256:f793ed4b4c825f732033f9861b75e343aecd737fed0f50f3823087bdd407687b

Observation c036ce17-2c66-4ac3-a5f0-bec0c25cf112 · outbound

This paper cites Unsupervised speech segmentation and variable rate repre- sentation learning using segmental contrastive predictive coding,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Unsupervised speech segmentation and variable rate repre- sentation learning using segmental contrastive predictive coding,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.333859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.650700Z digest=sha256:fa037848ab2ab2b784ede752ceb155654c066ea8d130abe7ce25461f03c260fe

Observation d8f38926-fe18-4bef-b40f-325195474c9d · outbound

This paper cites Self-normalization and noise- robustness in early auditory representations,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Self-normalization and noise- robustness in early auditory representations,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.231236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.653998Z digest=sha256:62e9c179423605d7e2887840402a0996d4ae2230bc15140ecab9a5534a73ee05

Observation a5178f83-656d-412d-a88b-be14c90d6f35 · outbound

This paper cites an unresolved cited work.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.657328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.657328Z digest=sha256:42d6bea77d7b017c29046fb7d5a94edb50babf4777724b8e471952a3cfbd4eea

Observation 764197d2-608a-47ea-a155-1fae0d4a3019 · outbound

This paper cites Multiresolution spectrotem- poral analysis of complex sounds,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Multiresolution spectrotem- poral analysis of complex sounds,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.100079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.660417Z digest=sha256:5e93605edfe620409137c2edb8ddbfe951f2b04c8619da1cc271d184bf6eddb7

Observation 7dfd53ae-2f9a-4016-9a9a-a5fa14fbd665 · outbound

This paper cites Model metamers reveal divergent invariances between biological and ar- tificial neural networks,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Model metamers reveal divergent invariances between biological and ar- tificial neural networks,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.976683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.663600Z digest=sha256:dccc75614c1ec418646c58ddac07e8e200bb74efd6ec49be24ae4d85f4640843

Observation 6a985c03-b404-43b8-903e-93fc23763d3c · outbound

This paper cites Derivation of auditory filter shapes from notched-noise data,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Derivation of auditory filter shapes from notched-noise data,

Reference 40

Resolution
verified exact
raw_fallback, observed 2026-08-05T19:54:58.998024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.666884Z digest=sha256:6ec0d4929bd52d2293154ac7b9940088b66b05865d71e945bb40c9878933eb7c

Observation 71303311-fc97-4f31-aa77-67021373af4d · outbound

This paper cites Sound texture perception via statistics of the auditory periphery: evidence from sound syn- thesis,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Sound texture perception via statistics of the auditory periphery: evidence from sound syn- thesis,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.847895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.670026Z digest=sha256:862aa340752bc7d98c4e0004f8bba0a5a1330bb7c9412d38b0d85a5aa5a91eba

Observation c0098db9-3dc5-4551-b07b-e52f6010aaf4 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.673271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.673271Z digest=sha256:0af75676c931db33e7561cc2f871493c5735dcf83f7b9552243b180602e1b0d6

Observation 2b59ce30-34e4-413c-8261-07bcdeafbaef · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Lib- rispeech: an asr corpus based on public domain audio books,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.677108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.677108Z digest=sha256:228abad42108f2036d64b64a9318538739601701b6871996397fc9cc37b8bf5f

Observation 653e79ba-084b-461e-be72-dbbf8515a5e7 · outbound

This paper cites Improving language understanding by generative pre-training,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Improving language understanding by generative pre-training,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.721880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.680937Z digest=sha256:2f42c2eb09378aa5c283935d2278129a8844a0753114de00ae4b10cb989df209

Observation f5858c51-8ec3-42f6-899c-28e3d76aa871 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Gaussian Error Linear Units (GELUs)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.684181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.684181Z digest=sha256:1311235334e26ea10ab5e9cf88f16291110a0f2735d26cb3c1d53bb7fa68cd79

Observation 1e70758c-5b26-475c-ab84-35265a244a88 · outbound

This paper cites Root Mean Square Layer Normalization.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Root Mean Square Layer Normalization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.688240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.688240Z digest=sha256:0dc83bafcfc02cb28b38bbf0dfb51c9cb501ac5a6afdd9f05dc73b94b2ffb6ee

Observation 3d165a1c-5b8b-4e06-a373-0f6b2f87e3a8 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Libri-light: A benchmark for asr with limited or no supervision,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.611021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.691750Z digest=sha256:55397f0a038a4cc459b3666bbb1f03277d65d39e5de2358a42e0b99d78627cba

Observation e2c0c792-4af0-45a9-a891-87aa239c8340 · outbound

This paper cites Timit acoustic phonetic continuous speech cor- pus,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Timit acoustic phonetic continuous speech cor- pus,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.458253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.695205Z digest=sha256:8d2065c98f3ccc4d049e25197df8c482a6e3a560e04cea5954d9fdffe6405f0c

Observation 98d1da4f-be71-4a68-852d-3bc94bbb395d · outbound

This paper cites Speaker-independent phone recogni- tion using hidden markov models,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speaker-independent phone recogni- tion using hidden markov models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.279439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.699082Z digest=sha256:9a5597db0f60f00cb57b495108dee0a1fed32532924cb63c69f63314bf7d917a

Observation eb7042ec-1ac7-468a-b583-3db817e346b2 · outbound

This paper cites Scikit-learn: Machine Learning in Python,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Scikit-learn: Machine Learning in Python,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.140934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.703299Z digest=sha256:73b0340f9fcf0fd96528b2a293a4209c679761a5ea49770fa2e62eeef214b10c

Observation a2d0db29-082a-4ab4-8bf5-da843b7777f7 · outbound

This paper cites The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.706814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.706814Z digest=sha256:b67002b9b95b50d095dd169fc82ea212e4baf715d274f19631438df165b468e7

Observation 38b04a89-0a8f-4ecd-b6ce-a817b2096755 · outbound

This paper cites Over-reliance on english hinders cognitive science,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Over-reliance on english hinders cognitive science,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.002185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.710255Z digest=sha256:d2068b0ec5375abbce66ac500fc8c1e79dcd6bdae488f5cc956f34bb7c7b507b

Observation 099f2fd0-5e9a-45d3-8cef-f1ff9189d8b0 · outbound

This paper cites To- wards inclusive automatic speech recognition,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens To- wards inclusive automatic speech recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.891669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.713597Z digest=sha256:3b2cd21700ba013acb7656517f6d0adcf066d57e5cf1272d8e769a885d3fdcf5

Observation 3aaeb462-9d71-4e5c-ba65-82b91e492f2a · outbound

This paper cites Say- cam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Say- cam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.778519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.716758Z digest=sha256:a8db915ff3c0720bc094be03bb152e3ed0962c41b66b43d47f6af4173537249a

Observation 071b6908-1925-49bf-b58e-757b99b7b746 · outbound

This paper cites Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.720381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.720381Z digest=sha256:f53e5158cc2cd953efb81d96f3ee9cdfd8226d0168dd90f2843c595afe24f46c

Observation 19b671da-5252-4497-9d4b-a0caffead11c · outbound

This paper cites A task-optimized neural network repli- cates human auditory behavior, predicts brain responses, and re- veals a cortical processing hierarchy,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens A task-optimized neural network repli- cates human auditory behavior, predicts brain responses, and re- veals a cortical processing hierarchy,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.644771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.723735Z digest=sha256:ef72a4e78d644db7bd7db4003c005ee135f5694ea3be0062bc110e96d0fdcb66

Observation d7f67bdb-df30-48c7-bf43-d124bd131507 · outbound

This paper cites Toward a realistic model of speech processing in the brain with self-supervised learning,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Toward a realistic model of speech processing in the brain with self-supervised learning,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.557865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.727007Z digest=sha256:f3241bcc1041b3cd7ee4666fca0c1e95f9124e75336ffbc8c7000fd9c985ad69

Observation 02a3ce6b-c9d5-4b31-b65d-c42660187c23 · outbound

This paper cites Dissecting neural computations of the human auditory pathway using deep neural networks for speech,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Dissecting neural computations of the human auditory pathway using deep neural networks for speech,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.429799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.730190Z digest=sha256:de3d2064ac6c3d303756d430a9ec662baccd3d27e8b9f60810e7354e7ccf7789

Observation aee78f11-cfe0-443e-8813-2da3653d8bce · outbound

This paper cites Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.297487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.733488Z digest=sha256:a34f3734fc6e69b87c7e8732a10075e71e642ab337a97d31d4dfcdc82e36e43b

Observation 91f4b4e3-66e6-4218-ac8c-0418f1508ccb · outbound

This paper cites Speech taskonomy: Which speech tasks are the most predictive of fmri brain activity?.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speech taskonomy: Which speech tasks are the most predictive of fmri brain activity?

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.192596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.736777Z digest=sha256:46af83001b2b0d46a3dde3fc292fa9fba5230cc726ab074dee46579facfadbc1

Observation 3aff2a49-132d-4df5-803b-27c0a1a5079a · outbound

This paper cites Language in brains, minds, and machines,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language in brains, minds, and machines,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.058048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.741001Z digest=sha256:61314c557aaaee54dcb3b961a9f6986dd7130c93755f8a86c67a89f09e7530c2

Observation 0264a1d7-45ec-4d1d-a2b8-29e6e9f2387a · outbound

This paper cites Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.744160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.744160Z digest=sha256:11b49fbdcc92564cd2f36bc5aeeb810b100d2430f0be536e8be912fc7cf7ce74

Observation c3849522-5e98-42e5-b262-13da0475a8e1 · outbound

This paper cites An algorithm for the machine calculation of complex fourier series,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens An algorithm for the machine calculation of complex fourier series,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.936362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.747668Z digest=sha256:8d2c580480f2af6994727c5a6faef141ef1fec73d912df28d65ce83a980d05eb

Observation 13352801-f597-4a72-8ac0-fc4301aedecb · outbound

This paper cites Lib- rispeech: An ASR corpus based on public domain audio books,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Lib- rispeech: An ASR corpus based on public domain audio books,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.764253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.751677Z digest=sha256:8845cee97bb4ab9202bbf9832deb28596a5816339290441e7e1d770f7f4a5de6

Observation ec17cb52-6a2a-46f9-a014-6dc96e011ce3 · outbound

This paper cites codebook usage.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens codebook usage

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.574844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.755219Z digest=sha256:a9872ad6278ea620c9b6ca7045b75138edd5bff7602494286258cd2d5c341a6c

Pith citing papers

Observation c87c66bf-5d6a-4851-91b1-d73aff9c5ccf · inbound

Representing Speech Through Autoregressive Prediction of Cochlear Tokens cites this paper.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representing Speech Through Autoregressive Prediction of Cochlear Tokens

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T19:54:59.463829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:54:58.525812Z digest=sha256:3bc222763039ef6d89387b42165b152d48b9dbbae466ca84ecf44096628197ce