Pith. sign in

Paper Citation Record · LEDGER

Optimized Self-supervised Training with BEST-RQ for Speech Recognition

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2501.16131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16131 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:47:17.319518Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a96fabe4-bc40-43d8-9635-47cb881872ad · outbound

This paper cites Dimensionality reduction by learning an invariant mapping,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Dimensionality reduction by learning an invariant mapping,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.603753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.243630Z digest=sha256:6273d059923b91fdf7d2e7e9aa7e465d4d3b2876005ca5e54df88b98db2c0f84

Observation aff1bf93-d8ed-4166-871d-c23091c0d72a · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.586921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.249219Z digest=sha256:8982c42a2a29e49e53187ab8ef517db41f9d2d0bd0b0e089915997728e0ff6d2

Observation d789b403-8e51-4569-943e-ce6f0987d45d · outbound

This paper cites wav2vec 2.0: a framework for self-supervised learning of speech representations,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition wav2vec 2.0: a framework for self-supervised learning of speech representations,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.570600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.254391Z digest=sha256:34fff82f6e989bf06d652d4a1a4eba524521da2628d337fd72c0da169cd87440

Observation 725ebc9e-3e5d-46e7-97b4-1e3ff03a5a1f · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.553824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.259748Z digest=sha256:88e750492562cd33338028ef8dd462c500f35c3230743c0e54d66ffa21ab1caa

Observation 85051f18-8a61-470f-99b9-b8b9af738568 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:17.265186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:17.265186Z digest=sha256:164af41500029307f237038f5302607697e014e47348c299029447c6d887751a

Observation 60a3f2e3-d435-4d38-9d9e-a1da5fe54704 · outbound

This paper cites w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.526007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.270491Z digest=sha256:a183d478589d5ac5a09d279c208ae9cbf55e4c01f7a7eae2c4cfc33ad3e40747

Observation 97264ef8-6afa-4330-865e-c6b785768671 · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recognition,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Self-supervised learning with random-projection quantizer for speech recognition,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.509306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.275913Z digest=sha256:c0964da96d954d3b4ead017d9dfaed3084328cfc46a11f49d5a826411077bf94

Observation 0ad1cff6-a138-4fcb-88c5-b416468683e7 · outbound

This paper cites Accented speech recognition with accent-specific codebooks,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Accented speech recognition with accent-specific codebooks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.491993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.280931Z digest=sha256:c9bfcd5937a2e10c2e5265852a7913dab185541ced4c1637c984b58ca1685609

Observation 3e719735-0d6e-485e-b694-54f129509594 · outbound

This paper cites Conformer: Convolution- augmented transformer for speech recognition,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Conformer: Convolution- augmented transformer for speech recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.473170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.285620Z digest=sha256:78f7c58745d8822ad17b80f14bd52669cc484f45d4db06699e92e2fe3ce61951

Observation 5f973d7c-9d6b-406d-b372-1ed5cbc720f6 · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:17.290335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:17.290335Z digest=sha256:2fae655f6eda3df429ae409ba0ca658f5bbd22a1004b2e432355c5769dd948d8

Observation 8e98a78f-249c-413a-800e-98f62d2a6e4c · outbound

This paper cites Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.455374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.295339Z digest=sha256:57ee8c84fbc09ee6f0d91e2846689e7e62cafb802126b1b2894b805c8d6d3f6b

Observation 521f2622-3fb5-4a68-b9e1-21d54b2c21e4 · outbound

This paper cites Music type classification by spectral contrast feature,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Music type classification by spectral contrast feature,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.438158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.300055Z digest=sha256:550898bf8d8ef55125d0fcd819a3969d829656c88765510301dfa23463e8dbbc

Observation 77dc82c6-5b12-4629-a2ac-e051cacc9bcc · outbound

This paper cites Open imple- mentation and study of best-rq for speech processing,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Open imple- mentation and study of best-rq for speech processing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.421102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.304882Z digest=sha256:8b2fe9eda16c8ce83f7ca1e3ca1ba9e58ccfb68bffb0b8eb898be7310be35065

Observation a5002357-695d-46a3-b381-d84be581b932 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition SpeechBrain: A General-Purpose Speech Toolkit

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:17.309498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:17.309498Z digest=sha256:83077aab72b65d76302542d461b6ba18d068192db7849e0973b37eae2c85dc7d

Observation d0119ed0-8355-4fca-9a92-4cc59d0ee36b · outbound

This paper cites Open-source conversational ai with speechbrain 1.0,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Open-source conversational ai with speechbrain 1.0,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:17.403654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T13:47:17.314656Z digest=sha256:ebff71b3468d1fc7a3937063b80d63b8c419908398844311a77eb14e28a48a05

Observation 59e97340-7e2e-4222-97d9-8a31ca8cd48a · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Optimized Self-supervised Training with BEST-RQ for Speech Recognition Librispeech: an asr corpus based on public domain audio books,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:17.319518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:17.319518Z digest=sha256:2f6b8b93b858cecbffeef32316f0298cdca6fb344c8de6d612dd80491ec328c2

Pith citing papers

No inbound Pith citation observations are available.