Pith. sign in

Paper Citation Record · LEDGER

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models

As of 23 July 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2601.20896.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.20896 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T10:17:15.726726Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T10:17:15.726726Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-16T10:17:43.812845Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact3
  • verified fuzzy26
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8dd1f4eb-ff4b-4e8a-9c44-27699075cf2b · outbound

This paper cites an unresolved cited work.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Unresolved cited work

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.858980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:9af1faf6900ecf72e16bce939435912557f18db5241511b8b636c392916d1f82

Observation 2ecd42d5-b945-4a81-a1a5-93ed26280da3 · outbound

This paper cites A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:17:43.815322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:137c4cf9995c449a19c852236f8131b5af95904e76106e53fd0265509cb01e8b

Observation 59b4bed5-22f0-4029-90df-6fdcba6b9643 · outbound

This paper cites an unresolved cited work.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Unresolved cited work

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.851295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:dc17f4b01bf521e3c3e742b95869385795baefef09968b6b3e8a95f570aff3f1

Observation 7fa37eab-eb8c-4f4c-bd66-2c7e0043fecd · outbound

This paper cites Overall, diversity-based sampling methods did not yield significant improvements over the either baseline.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Overall, diversity-based sampling methods did not yield significant improvements over the either baseline

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.854081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:a96fb35252e90fdc0133293d1c84cc052ef66339141610d7c302db2a43cf7bf8

Observation cb7b2272-0f4a-4878-a417-c1c7bf424a46 · outbound

This paper cites In contrast, selecting subsets of longer utterances consistently leads to lower WER, despite these subsets being most out-of-distribution relative to the fine-tuning data.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models In contrast, selecting subsets of longer utterances consistently leads to lower WER, despite these subsets being most out-of-distribution relative to the fine-tuning data

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.861737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:bf57b9fa80acb6ffcced5a1d96320ecb2c8f68ce0c207aae736afed9f479deee

Observation 6f124903-2124-4815-b654-12dd078e9fa9 · outbound

This paper cites Self-supervised speech representation learning: A review.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Self-supervised speech representation learning: A review

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.864013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:c6e27d3e628cb0a0220e1cb1a04ffdc5246ff72ca9ca5c009d8b73b24587c350

Observation 034ede8d-2b78-41b5-94d1-3dec20fc43eb · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech rep- resentations.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models wav2vec 2.0: A framework for self-supervised learning of speech rep- resentations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.866460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:ab20e3fa29cbd5e03f05975d990667b746c95ccb44ecc8a8440c46bb401acb0a

Observation 1be6b491-bd44-4a34-8ff6-57f0a6be1b88 · outbound

This paper cites An analysis of linear complexity attention substitutes with best-rq.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models An analysis of linear complexity attention substitutes with best-rq

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.839269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:69bb262df24dbc9fe28fa2e7c04a376bde310251edb03b889a7be1b1ca4966ea

Observation 6a5507cf-2025-455c-a365-538d836cb8bd · outbound

This paper cites Efficient self- supervised learning with contextualized target representations for vision, speech and language.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Efficient self- supervised learning with contextualized target representations for vision, speech and language

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.841728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:9f6abb553607d00b9159bd922847291de134f7ac47a6e82ff675bfe407646b7d

Observation 36728b3d-6414-4758-bbb2-51a71583861c · outbound

This paper cites Reducing barriers to self-supervised learning: Hubert pre-training with academic compute.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Reducing barriers to self-supervised learning: Hubert pre-training with academic compute

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.844246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:cda680885a36f11494b26c8df411d38f7e76683dee2c9cfd1f84ef73922440cf

Observation b62627ae-d07a-497b-b75b-73ab810d44e3 · outbound

This paper cites Self- supervised learning with random-projection quantizer for speech recognition.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Self- supervised learning with random-projection quantizer for speech recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.831812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:82d02756320e90e90991f98ee416a0588b88ee9c612a43c124cd1950a81096d1

Observation a132405f-3758-4806-ba5f-47ad5002eb97 · outbound

This paper cites Towards Early Prediction of Self-Supervised Speech Model Performance.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Towards Early Prediction of Self-Supervised Speech Model Performance

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.826751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:c16b59b0f521f820a079ddb21e3297055a46a339abfaba16a30331a04ecfbace

Observation d90a938c-6437-4c1b-a526-e633096ec97a · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:17:43.819953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:d10a05753e3697fd1dfc9c5168bec0dbef0bf15a9fcc860ac36cd69a191331c9

Observation 30cd90f5-566e-4505-b608-52d7a26cdb5b · outbound

This paper cites Towards robust speech representation learning for thousands of languages.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Towards robust speech representation learning for thousands of languages

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.829390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:5a8e23ef78a9645475faced53bfe38b61f8f70188f802735ec2571d647cb7590

Observation bb6a0c6d-dc06-4f59-abbc-30006d3c9a6c · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Robust speech recognition via large-scale weak supervision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.834276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:916b2210fc62989befd7b73923c5f117da9e0e5b967fe6db5d937d4564f90744

Observation 7492fb47-cb9f-4e65-8bfc-bf4e3f4dc26f · outbound

This paper cites DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.836825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:1003afb6d5c504fb90715e4829aad305c52e0a90cade97c04a6ad45a88640224

Observation fdad770d-cf9b-4ac2-bbe2-7fe2a15c9ba7 · outbound

This paper cites Towards automatic assessment of self- supervised speech models using rank.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Towards automatic assessment of self- supervised speech models using rank

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.846624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:53c07b6c68250de7cea53787e97fcddde54b8479c83d65808c160a9c507facbb

Observation 500e606f-fc2b-4b35-8325-e91136ae9664 · outbound

This paper cites Active learning methods for low resource end-to-end speech recognition.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Active learning methods for low resource end-to-end speech recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.818856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:1a15f5068a2e8d2b14301b79078dd2a6a60b8bc856c82b44a9a347ce16316df2

Observation 04b2d758-95b7-45a8-b244-d15d62d81fba · outbound

This paper cites Unsupervised Data Selection via Discrete Speech Representation for ASR.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Unsupervised Data Selection via Discrete Speech Representation for ASR

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.808787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:f62a6778d4373aa7500092b4df16b11e49fa2ed0f846972eb2519564d98f339f

Observation 4a2ac93c-3555-4897-a6e4-64ce729ceefd · outbound

This paper cites More speaking or more speakers?.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models More speaking or more speakers?

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.816238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:6a625703f657e06fb046bbb756a6bf942e8e6349f8ee212273cfd891640120f2

Observation 473584dc-71bd-490e-a3f2-1ce318dd311f · outbound

This paper cites Loqua- cious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Loqua- cious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.849134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:3d73821888730d3d6d3c7135bee25839a5324c7e4ea6bf1da748732a33ec9457

Observation 1f5a4e05-1ee3-416b-b6ed-0c375b82a48e · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Wespeaker: A research and production oriented speaker embedding learning toolkit

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.805834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:1fdbf005f8586b7aa41321b2a5f8696e7390851d350b08e3cb84b14916fe55e4

Observation 95466d8a-6a56-4bc7-958d-a9df6d68f7c2 · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.821732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:4ed1ad144926f9efaa61587f5495fee012ae248a642ba9878e7225daf008506f

Observation 0fcea6a5-18d7-4fd9-aabb-eee5abc71e08 · outbound

This paper cites Sense models: an open source solution for multilingual and multimodal semantic-based tasks.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Sense models: an open source solution for multilingual and multimodal semantic-based tasks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.814010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:a82d256dcf390481a40f9aff56276862c06d50c7b05031f8f0342a98a5241f69

Observation 85bef381-aa8d-44fb-a9ee-29f0b1f343ce · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Roformer: Enhanced transformer with rotary position embedding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.811451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:d8ce417eb9deda9e29f9efc69206742b5fba55eff8e5043f02ed676d32160951

Observation 2fbbae7a-86e3-44b0-9df5-a25b23675d03 · outbound

This paper cites Open implementation and study of best-rq for speech processing.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Open implementation and study of best-rq for speech processing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.824214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:25fd10ad0c1b632e0811c31b892f2b8f29bf3e641f62a1b01440ba2b55d0fd0b

Observation efe285bf-bfab-4787-a307-57fb14c1594f · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models SpeechBrain: A General-Purpose Speech Toolkit

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:43.810929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:e88ee46a207a0a5a42709ee7bee3bcd62f9348aa0c4f51a2ea8c9a2397ed1121

Observation c15b48dd-76b3-4ddd-be3e-19c2916165fc · outbound

This paper cites Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:43.805945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:7e1474aa69ca364f0c8e0126aff2a97b571ae067eb6bb7d06ea887cde85c51e1

Observation 43cae7dc-0ecc-48e5-9261-171322402e78 · outbound

This paper cites Dynamic chunk convolution for unified streaming and non- streaming conformer asr.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Dynamic chunk convolution for unified streaming and non- streaming conformer asr

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.803014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:1d9f8e56e4d511fb351f5763fe9e4f207b47a3b4b131619260763feda2732f18

Observation 2e9732f7-7360-4271-89c5-6ccdddab4694 · outbound

This paper cites In-domain ssl pre-training and streaming asr.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models In-domain ssl pre-training and streaming asr

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:17:44.856493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:c7cb4ec86ba4f60129cbd6f38b27f9912bd3ff6d378142ff0801803dabe3fa6e

Pith citing papers

Observation 2ecd42d5-b945-4a81-a1a5-93ed26280da3 · inbound

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models cites this paper.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:17:43.815322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:137c4cf9995c449a19c852236f8131b5af95904e76106e53fd0265509cb01e8b

Observation 2b7fc040-a4e5-4f6f-b50d-1700828db6d5 · inbound

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models cites this paper.

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:31:09.929396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T11:43:47.693392Z digest=sha256:3a2657fedc9bd9fd562cfc000f6f24c53113fa0009058820af1da20bafc5e723