Pith. sign in

Paper Citation Record · LEDGER

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2506.00338.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00338 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:21.174309Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:17.177212Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:12:21.729943Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved10
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation da98eed8-8609-4fd1-87f4-9a9d2c0507e2 · outbound

This paper cites OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-08-07T12:12:21.828937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.177212Z digest=sha256:fbfa1fcf9af3df16d3e7d694827df3cf02069e84a200591af958514af1b4365f

Observation 8fecf34e-4ab5-4486-8dbe-b28d4b53b12e · outbound

This paper cites YODAS data cleaning The raw YODAS data has not undergone a rigorous cleaning process and may contain annotation errors [24].

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning YODAS data cleaning The raw YODAS data has not undergone a rigorous cleaning process and may contain annotation errors [24]

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:12:27.493239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.271724Z digest=sha256:d9e672330e9441cccb6a9ab2fe93eb325cf20efa13281e8f7ee11983351c84c5

Observation 62ffcb80-8de8-42e8-be83-0757cdfc274f · outbound

This paper cites transcription.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning transcription

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:12:21.617665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.393419Z digest=sha256:40da33b0a001a55734dca60e425a886c8d77f29f46368b0fde14f16ed3db1b61

Observation e9ce723c-d9be-4b7c-b87b-6e61c7bcaf36 · outbound

This paper cites We reveal that large-scale web-crawled data contains incorrect lan- guage labels and audio-text misalignments.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning We reveal that large-scale web-crawled data contains incorrect lan- guage labels and audio-text misalignments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:27.321548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.504498Z digest=sha256:9acb1981002c8b33922ed45beb72c9a3148bfe9647e2a3260c5848af0a2b72e4

Observation fc45377e-b74f-4575-826c-4b2c2bab28f7 · outbound

This paper cites an unresolved cited work.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:12:27.141555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.586769Z digest=sha256:5e370e7cfb6dcd4c80a2585ac6bf0cbee49f0c73b8fcdfe601d6fe342f793cd9

Observation f81a6e80-6bf0-4300-ad75-9e2560fae4d9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Robust speech recognition via large-scale weak supervision,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.954012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.675610Z digest=sha256:d7f23decdecbf56a05fcfc24a4346267d1849648d746483b4dca5990deb6a52d

Observation a49829e8-8114-4c2a-b598-09edd3f5f677 · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:17.790407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:17.790407Z digest=sha256:ef59a67abe5e2adb1dc9256e71b40ac25bb325c2379d5484350fa0228538259d

Observation 1c293df4-4f20-4953-b75d-8f948b50f06e · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Scaling speech technology to 1,000+ languages,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.785010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.904554Z digest=sha256:d54aad250523ee8c75b7c4c0012ecfaf45c29c05cd7033b39a57ec3b5c79ca79

Observation cfbe3c66-a058-4b60-9191-a84e9b7324c4 · outbound

This paper cites Less is more: Accurate speech recognition & translation without web- scale data,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Less is more: Accurate speech recognition & translation without web- scale data,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.633111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.996141Z digest=sha256:5aabd489288e3a1e0717c54fa7187eb95aac082ea1630498b6906170c5647a75

Observation 538be67f-567f-4d89-9806-f9d4c5892108 · outbound

This paper cites Reproducing Whisper-Style Training Using an Open-Source Toolkit and Pub- licly Available Data,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Reproducing Whisper-Style Training Using an Open-Source Toolkit and Pub- licly Available Data,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.501683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.091981Z digest=sha256:adf07c4abf8d3efb6dcbfa72f411e527ce6fde655230f1358fd32964604b43f2

Observation cef51312-f6bb-4e3d-9d3e-b61f35ee3ab5 · outbound

This paper cites ESPnet: End- to-End Speech Processing Toolkit,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning ESPnet: End- to-End Speech Processing Toolkit,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.377366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.171426Z digest=sha256:81e7870bc2ad4c4d1c3d5cac6a4cce178aaabe291d3fdaec175945d0b3c760dc

Observation 42b5f007-b1a7-4cbb-b9b5-7db73be9266f · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Conformer: Convolution-augmented Transformer for Speech Recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.201127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.265022Z digest=sha256:e985ba5f2af6400da3dc4a3994756950b832c5941ada3d716f67e196eeaff233

Observation efcccd6c-06d7-4d53-b4eb-9218918c66d4 · outbound

This paper cites Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.994805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.358460Z digest=sha256:d0e0573ff8dfbc77f26c51dd540f0a5a450b63cbc6302b8347226784da8c781f

Observation d8abf4c3-f8bf-4c5c-b159-f1501bae89d4 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Zipformer: A faster and better encoder for automatic speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.745680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.435408Z digest=sha256:d612bf47b62e053b16f05d661b21d1ba5cd43351d72494d1138d8676f9ac04d1

Observation cb75094c-5bfc-457f-8ba9-344d1fe37679 · outbound

This paper cites Atten- tion is all you need,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Atten- tion is all you need,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.603809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.504149Z digest=sha256:9c1924a9c058c1d18d352936a52ddd9b0c151be6e08c133befa743c5d7e4dcff

Observation 9d6c92d4-8967-4486-bcf5-337b9e8732a6 · outbound

This paper cites Squeezeformer: An efficient transformer for automatic speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Squeezeformer: An efficient transformer for automatic speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.467750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.599148Z digest=sha256:691a605a4db254d8d8c3547f0eb929e73efb03a30a0397285c58466bf9720200

Observation fc6f5a37-a3b3-41f0-8d97-7f35e083183d · outbound

This paper cites Fast conformer with linearly scalable attention for efficient speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Fast conformer with linearly scalable attention for efficient speech recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.321074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.726297Z digest=sha256:777004d5c165279664afed2407b9c19bb71a65ca521793ac721fdb5821daafaf

Observation 86c09436-cba7-4031-b368-fd2da1233b7f · outbound

This paper cites Sum- maryMixing: A linear-complexity alternative to self-attention for speech recognition and understanding,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Sum- maryMixing: A linear-complexity alternative to self-attention for speech recognition and understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.178642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.814410Z digest=sha256:697284307e6cc5d374bf72650837a48351478c21fdb39a12dcc4a13eacdbbf95

Observation 21568b9a-d2a0-425e-819e-d272968b80f9 · outbound

This paper cites OWSM v3.1: Bet- ter and faster open whisper-style speech models based on E- Branchformer,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v3.1: Bet- ter and faster open whisper-style speech models based on E- Branchformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.047851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.852042Z digest=sha256:2ec0276461d0237e7ec67ea2a2ecb3d2e149491f49f39f700949b2fe5634fa8c

Observation 837310f8-d203-4448-84fb-80b4cf0d328f · outbound

This paper cites E-Branchformer: Branch- former with enhanced merging for speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning E-Branchformer: Branch- former with enhanced merging for speech recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.905555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.855538Z digest=sha256:689a6cb15bde2fd97102940baa32aa8a008c2bf9eaa47ffb74aac0f206d3461b

Observation bfcef2ea-205f-4f2c-abe1-7026a12b9920 · outbound

This paper cites A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Transla- tion, and Understanding Tasks,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Transla- tion, and Understanding Tasks,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.764470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.888232Z digest=sha256:79b79bb9ecc95c30375e0e21b0353ea8b50966749ab26c3dc4d2b74325b75060

Observation 8ca5f411-4012-44a5-9995-68ca9fbad818 · outbound

This paper cites OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.649967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:18.949572Z digest=sha256:494ab3001cbead3b0435ef0024b8d4cb3a70bed55dd277759487b760fa101299

Observation 56d343c9-0199-4c6e-9ab5-09ac9178e1b8 · outbound

This paper cites Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.544340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.036480Z digest=sha256:7c6361bdcb311b46534398363367f001b84c75d822575e74fa19bb7b20148d1a

Observation ea1b3402-bfe7-4ff4-b7a7-7812ed9f1bed · outbound

This paper cites Unsupervised data selection via discrete speech representation for ASR,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unsupervised data selection via discrete speech representation for ASR,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.424537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.139139Z digest=sha256:fd20304465e9c5e5f0ceabce91a830f2a95fe5650ec2b4af072dec418a879ae9

Observation d971e49d-8a2c-4745-bc46-4fcaf0b7582d · outbound

This paper cites Unsupervised data selec- tion for speech recognition with contrastive loss ratios,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unsupervised data selec- tion for speech recognition with contrastive loss ratios,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.308824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.209648Z digest=sha256:fff11e735ee59d1b8de445774fcda47c753419bea864c99ce2bd7644d3083fb9

Observation 272d04d1-4869-4c60-b6c8-f692c1539935 · outbound

This paper cites Spgispeech: 5, 000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Spgispeech: 5, 000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.147242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.301270Z digest=sha256:3b49b2d4198a3be03cf4160408c366fff94f0df4360fc779f1ea0d02f5507d12

Observation 2c835fcb-4178-443b-bed0-c9e0bb01c651 · outbound

This paper cites Gigaspeech: An evolv- ing, multi-domain ASR corpus with 10, 000 hours of transcribed audio,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Gigaspeech: An evolv- ing, multi-domain ASR corpus with 10, 000 hours of transcribed audio,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.980936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.428013Z digest=sha256:7a157dadfa574c3e9d75218d07af0e2d6d605d94641bf340a7e0b8c7094a8f0e

Observation 9ec2465d-b686-4696-a457-74c0a22ec0aa · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:19.512688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:19.512688Z digest=sha256:051925a6d6ef5f3458fb53f94c300af718f121ac07b2d801465c22624621a628

Observation 475da79a-c3f1-4fcf-b4df-06d6e81bd146 · outbound

This paper cites YODAS: Youtube-Oriented Dataset for Audio and Speech,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning YODAS: Youtube-Oriented Dataset for Audio and Speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.795234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.597084Z digest=sha256:385c3792692ef1a0a61baa8bd4afa1d452f50d19b3f6d86aa0d4ce51ecc36eb4

Observation 13297370-5350-435d-8613-107046f5c0d7 · outbound

This paper cites On the effects of het- erogeneous data sources on speech-to-text foundation models,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning On the effects of het- erogeneous data sources on speech-to-text foundation models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.611777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.703596Z digest=sha256:52c9aa3479a564088c07291a14df2cd54983fb4b08d004dda0c59131cf237a3d

Observation 77bf07e5-9d02-4ed0-a257-5038fcbd6f38 · outbound

This paper cites SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:19.793022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:19.793022Z digest=sha256:66c9ed5d4b0d8b99c0b67415ccb07074b3c739806214253b53f5de4b8a4e5d76

Observation 4575edf5-26aa-40ee-8c78-c593b0634b56 · outbound

This paper cites MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.463609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.884354Z digest=sha256:5c7f72e644477a04cb94b40d6258400537e3751d6982955653d1776ddcdff808

Observation c8bfbcdb-e401-4f64-8bbb-f01fba88ba5a · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.270870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:19.960379Z digest=sha256:2c963fca268a0c5ff094f30d20048f637479f9bd0df5001ba130180b16b0d537

Observation 405fdaac-7213-4d65-b898-1e5404abd152 · outbound

This paper cites GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.068090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.068090Z digest=sha256:cd519b4d9f750a08fb22f3f2d604587ad8a2d122a0a1c6e3e1cfb679e924362d

Observation b61a362a-029f-4875-a106-8c302d761e8b · outbound

This paper cites MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Founda- tion Model Training on EU Languages,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Founda- tion Model Training on EU Languages,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.100262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:20.156850Z digest=sha256:cdbc6e67bea6f35a3cb89ad0bf482e14d98728bd925efa304b79b550e0a9a38a

Observation 037e1b2f-f68f-4680-9328-3b76ff9c7ab4 · outbound

This paper cites CTC-Segmentation of Large Corpora for German End-to-End Speech Recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning CTC-Segmentation of Large Corpora for German End-to-End Speech Recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.956806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:20.255300Z digest=sha256:e0152c0c041eb1785a1fc8ee13933faf81c6970cb8cf7504731fa38b3f1fe9fa

Observation ff89a244-e405-41d6-a38f-b4a7e9909921 · outbound

This paper cites Bag of Tricks for Efficient Text Classification.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Bag of Tricks for Efficient Text Classification

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.365163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.365163Z digest=sha256:1ccf34d10ec61ff63c396d1c8f55c3248dc041ad34b657f225cebda34eded43a

Observation 2207002f-9773-4956-b514-6e8c6148b10c · outbound

This paper cites FastText.zip: Compressing text classification models.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FastText.zip: Compressing text classification models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.479161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.479161Z digest=sha256:9e7eaadc448e161c027d4c1e2fb00c074dfc969d51eee5dd18860ca182162bc5

Observation a1278efc-5e62-4b40-95e9-1fc712b23bd1 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning SpeechBrain: A General-Purpose Speech Toolkit

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.555853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.555853Z digest=sha256:068dcb3ccac4d23306fabf2e2079c5fe2388603d77ba2bd91ac21952014c35b1

Observation 39e7fa83-9333-474a-b8a6-85f1307ad3fb · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Common Voice: A Massively-Multilingual Speech Corpus

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.661052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.661052Z digest=sha256:0e12d09b3c00289d04d574d63bfe0eabec7bb50c9059d699bca6e82fc8ec7f49

Observation 2b5be670-959e-4d36-82bc-9f0fba6c716b · outbound

This paper cites Pytorch: An imperative style, high- performance deep learning library,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Pytorch: An imperative style, high- performance deep learning library,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.792853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:20.741492Z digest=sha256:aefeb30c72e12d0b8037c17cc0deeade0260ea318eb4c11efdd6bdd9dc643c8e

Observation 8e7d6a8a-75ae-48f8-ab33-a35fe2579cab · outbound

This paper cites FlashAttention-2: Faster Attention with Better Paral- lelism and Work Partitioning,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FlashAttention-2: Faster Attention with Better Paral- lelism and Work Partitioning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.648375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:20.824229Z digest=sha256:1a450eec612e0734f9730ce0cd6cc875b8aaf3ea13dddc9058a7d277ceea48ca

Observation 59c95f76-2645-43e3-8463-5e7bbf4706eb · outbound

This paper cites Decoupled weight decay regular- ization,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Decoupled weight decay regular- ization,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.469508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:20.908318Z digest=sha256:975dc467fa184e1dfecc769aaa5809036263b18aef21d7ac9e250b0f16d9fbeb

Observation 9ab61f24-357c-4dfc-915a-db9a72d7dfc7 · outbound

This paper cites CoV oST 2 and Massively Multilingual Speech Translation,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning CoV oST 2 and Massively Multilingual Speech Translation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.314269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:21.008812Z digest=sha256:e5e8884a2a8d0552f5f4ec0c3481e13b2c2a27334107213789e05ce8a732bf72

Observation 820967f7-1ead-4b55-b5d4-46b2b45fe252 · outbound

This paper cites FLEURS: Few-Shot Learning Evaluation of Universal Representations of Speech,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FLEURS: Few-Shot Learning Evaluation of Universal Representations of Speech,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.131499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:21.049823Z digest=sha256:fac6bad51da0415e9f401aa31c8540284fdff56156b5397a3310c74f49904b39

Observation d196b4a7-095d-443c-b774-7b1f45f97f0c · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:21.094758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:21.094758Z digest=sha256:83251851946bbd38d38e5834ad453efadba9238354e18136bdab6ced8d11d9f2

Observation 48a6f4de-4db2-4861-8697-d222a52175a2 · outbound

This paper cites Srivastav, S.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Srivastav, S

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:21.979286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:21.174309Z digest=sha256:b73b9a380714638faf06d2f89b613d4ff6c0ddea21eafdd7691a74a0c9b598fa

Pith citing papers

Observation da98eed8-8609-4fd1-87f4-9a9d2c0507e2 · inbound

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning cites this paper.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-08-07T12:12:21.828937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:12:17.177212Z digest=sha256:fbfa1fcf9af3df16d3e7d694827df3cf02069e84a200591af958514af1b4365f