Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:21.174309Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2506.00338.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:21.174309Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:17.177212Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:12:21.729943Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation da98eed8-8609-4fd1-87f4-9a9d2c0507e2 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fecf34e-4ab5-4486-8dbe-b28d4b53b12e · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning YODAS data cleaning The raw YODAS data has not undergone a rigorous cleaning process and may contain annotation errors [24]
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62ffcb80-8de8-42e8-be83-0757cdfc274f · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning transcription
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9ce723c-d9be-4b7c-b87b-6e61c7bcaf36 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning We reveal that large-scale web-crawled data contains incorrect lan- guage labels and audio-text misalignments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc45377e-b74f-4575-826c-4b2c2bab28f7 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f81a6e80-6bf0-4300-ad75-9e2560fae4d9 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Robust speech recognition via large-scale weak supervision,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a49829e8-8114-4c2a-b598-09edd3f5f677 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c293df4-4f20-4953-b75d-8f948b50f06e · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Scaling speech technology to 1,000+ languages,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfbe3c66-a058-4b60-9191-a84e9b7324c4 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Less is more: Accurate speech recognition & translation without web- scale data,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 538be67f-567f-4d89-9806-f9d4c5892108 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Reproducing Whisper-Style Training Using an Open-Source Toolkit and Pub- licly Available Data,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cef51312-f6bb-4e3d-9d3e-b61f35ee3ab5 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning ESPnet: End- to-End Speech Processing Toolkit,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42b5f007-b1a7-4cbb-b9b5-7db73be9266f · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Conformer: Convolution-augmented Transformer for Speech Recognition,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efcccd6c-06d7-4d53-b4eb-9218918c66d4 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8abf4c3-f8bf-4c5c-b159-f1501bae89d4 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Zipformer: A faster and better encoder for automatic speech recognition,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb75094c-5bfc-457f-8ba9-344d1fe37679 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Atten- tion is all you need,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d6c92d4-8967-4486-bcf5-337b9e8732a6 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Squeezeformer: An efficient transformer for automatic speech recognition,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc6f5a37-a3b3-41f0-8d97-7f35e083183d · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Fast conformer with linearly scalable attention for efficient speech recognition,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86c09436-cba7-4031-b368-fd2da1233b7f · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Sum- maryMixing: A linear-complexity alternative to self-attention for speech recognition and understanding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21568b9a-d2a0-425e-819e-d272968b80f9 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v3.1: Bet- ter and faster open whisper-style speech models based on E- Branchformer,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 837310f8-d203-4448-84fb-80b4cf0d328f · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning E-Branchformer: Branch- former with enhanced merging for speech recognition,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfcef2ea-205f-4f2c-abe1-7026a12b9920 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Transla- tion, and Understanding Tasks,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ca5f411-4012-44a5-9995-68ca9fbad818 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56d343c9-0199-4c6e-9ab5-09ac9178e1b8 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea1b3402-bfe7-4ff4-b7a7-7812ed9f1bed · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unsupervised data selection via discrete speech representation for ASR,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d971e49d-8a2c-4745-bc46-4fcaf0b7582d · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unsupervised data selec- tion for speech recognition with contrastive loss ratios,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 272d04d1-4869-4c60-b6c8-f692c1539935 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Spgispeech: 5, 000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c835fcb-4178-443b-bed0-c9e0bb01c651 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Gigaspeech: An evolv- ing, multi-domain ASR corpus with 10, 000 hours of transcribed audio,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ec2465d-b686-4696-a457-74c0a22ec0aa · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 475da79a-c3f1-4fcf-b4df-06d6e81bd146 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning YODAS: Youtube-Oriented Dataset for Audio and Speech,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13297370-5350-435d-8613-107046f5c0d7 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning On the effects of het- erogeneous data sources on speech-to-text foundation models,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77bf07e5-9d02-4ed0-a257-5038fcbd6f38 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4575edf5-26aa-40ee-8c78-c593b0634b56 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8bfbcdb-e401-4f64-8bbb-f01fba88ba5a · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 405fdaac-7213-4d65-b898-1e5404abd152 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61a362a-029f-4875-a106-8c302d761e8b · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Founda- tion Model Training on EU Languages,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 037e1b2f-f68f-4680-9328-3b76ff9c7ab4 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning CTC-Segmentation of Large Corpora for German End-to-End Speech Recognition,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff89a244-e405-41d6-a38f-b4a7e9909921 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Bag of Tricks for Efficient Text Classification
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2207002f-9773-4956-b514-6e8c6148b10c · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FastText.zip: Compressing text classification models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1278efc-5e62-4b40-95e9-1fc712b23bd1 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning SpeechBrain: A General-Purpose Speech Toolkit
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39e7fa83-9333-474a-b8a6-85f1307ad3fb · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Common Voice: A Massively-Multilingual Speech Corpus
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b5be670-959e-4d36-82bc-9f0fba6c716b · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Pytorch: An imperative style, high- performance deep learning library,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e7d6a8a-75ae-48f8-ab33-a35fe2579cab · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FlashAttention-2: Faster Attention with Better Paral- lelism and Work Partitioning,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59c95f76-2645-43e3-8463-5e7bbf4706eb · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Decoupled weight decay regular- ization,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ab61f24-357c-4dfc-915a-db9a72d7dfc7 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning CoV oST 2 and Massively Multilingual Speech Translation,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 820967f7-1ead-4b55-b5d4-46b2b45fe252 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FLEURS: Few-Shot Learning Evaluation of Universal Representations of Speech,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d196b4a7-095d-443c-b774-7b1f45f97f0c · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MLS: A Large-Scale Multilingual Dataset for Speech Research
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48a6f4de-4db2-4861-8697-d222a52175a2 · outbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Srivastav, S
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da98eed8-8609-4fd1-87f4-9a9d2c0507e2 · inbound
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.