Pith. sign in

Paper Citation Record · LEDGER

Optimized Self-supervised Training with BEST-RQ for Speech Recognition

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2501.16131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16131 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:47:17.319518Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a96fabe4-bc40-43d8-9635-47cb881872ad · outbound

This paper cites Dimensionality reduction by learning an invariant mapping,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Dimensionality reduction by learning an invariant mapping,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.603753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.243630Z digest=sha256:eed8a54505a8638633fea6232382678c3d1348ae11d86025fdf2c275f8072be0

Observation aff1bf93-d8ed-4166-871d-c23091c0d72a · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.586921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.249219Z digest=sha256:3fab9109cf67a3abb58c6eb921467d58ba31a6df8772a350082d76934b75420f

Observation d789b403-8e51-4569-943e-ce6f0987d45d · outbound

This paper cites wav2vec 2.0: a framework for self-supervised learning of speech representations,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition wav2vec 2.0: a framework for self-supervised learning of speech representations,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.570600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.254391Z digest=sha256:272a5021f64660e10baf7b7c8320da16fa3f0c5518f69cb5582e9e15e55a6eb4

Observation 725ebc9e-3e5d-46e7-97b4-1e3ff03a5a1f · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.553824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.259748Z digest=sha256:165e9fdeffa33b3bb43224c0a9e09316946516257f2507fa94fa60a83f5ad02d

Observation 85051f18-8a61-470f-99b9-b8b9af738568 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:17.265186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:17.265186Z digest=sha256:51091e1efe8d22f2b204f686c4223f865dbfab77179c7e74333d8120edf112be

Observation 60a3f2e3-d435-4d38-9d9e-a1da5fe54704 · outbound

This paper cites w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.526007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.270491Z digest=sha256:4e15cd92cffdf683cd2c5a4dbdd98cc86bfb27e4880cf48e7f8c405200be9894

Observation 97264ef8-6afa-4330-865e-c6b785768671 · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recognition,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Self-supervised learning with random-projection quantizer for speech recognition,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.509306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.275913Z digest=sha256:cde618d31b6b4f187633c11e6b3226cc995885dfa8591de0b42a1450137e0be4

Observation 0ad1cff6-a138-4fcb-88c5-b416468683e7 · outbound

This paper cites Accented speech recognition with accent-specific codebooks,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Accented speech recognition with accent-specific codebooks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.491993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.280931Z digest=sha256:6eed74196ef908b5349935431ce73636a0b8a4e7699327d039774bfacb0cbc93

Observation 3e719735-0d6e-485e-b694-54f129509594 · outbound

This paper cites Conformer: Convolution- augmented transformer for speech recognition,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Conformer: Convolution- augmented transformer for speech recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.473170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.285620Z digest=sha256:a01f63c847f90671208345bb5e24441c12f26a6a61a93a1a8387a9904b5b78d5

Observation 5f973d7c-9d6b-406d-b372-1ed5cbc720f6 · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:17.290335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:17.290335Z digest=sha256:d4d6007cc9d9d5cf7e94263c7b574f5bccba61af67bc9c961d541d5aa5bf0346

Observation 8e98a78f-249c-413a-800e-98f62d2a6e4c · outbound

This paper cites Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.455374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.295339Z digest=sha256:589093494a8b4646c49cd3039ad798a8117197fea0a00df10a4c0b2f065c6595

Observation 521f2622-3fb5-4a68-b9e1-21d54b2c21e4 · outbound

This paper cites Music type classification by spectral contrast feature,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Music type classification by spectral contrast feature,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.438158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.300055Z digest=sha256:b65c25b17f91fbac45f3a7a2d5f51d3f6b36ebd53196e9e84908d44ddb54afb7

Observation 77dc82c6-5b12-4629-a2ac-e051cacc9bcc · outbound

This paper cites Open imple- mentation and study of best-rq for speech processing,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Open imple- mentation and study of best-rq for speech processing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.421102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.304882Z digest=sha256:647927b77fbea5adc37b87c425c766059b04dc15f3934eeff2b400c6d691ad77

Observation a5002357-695d-46a3-b381-d84be581b932 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition SpeechBrain: A General-Purpose Speech Toolkit

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:17.309498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:17.309498Z digest=sha256:46580da119b04f94596088914446cfff541ba7adf0bab456a8b4878a0891091c

Observation d0119ed0-8355-4fca-9a92-4cc59d0ee36b · outbound

This paper cites Open-source conversational ai with speechbrain 1.0,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Open-source conversational ai with speechbrain 1.0,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.403654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.314656Z digest=sha256:213e861e58fffc266f6cff0f2d87fbef1a08925f633b305eb1da1cc10ab7790e

Observation 59e97340-7e2e-4222-97d9-8a31ca8cd48a · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Librispeech: an asr corpus based on public domain audio books,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:17.319518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:17.319518Z digest=sha256:99e17e4e306ee436102f0132b611d7466c3b0b21446194be7220586e607901db

Pith citing papers

No inbound Pith citation observations are available.