Pith. sign in

Paper Citation Record · LEDGER

The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2111.09344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09344 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:13:57.274992Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:10:07.966776Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 424a669b-2016-4d62-a083-84d59e362c72 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.274992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.274992Z digest=sha256:1339febd002da1313d4494fec8c47bca2371cd00e422e8b4ba88119453ca47fd

Observation 52b343b8-f89f-43d4-9dfa-e35a1fa2c00c · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:48.463964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:48.463964Z digest=sha256:1b017bacf77000119f5ca1560c499c5b779354f2aedf7c2f134375319d811642

Observation d94107cd-9278-4f9c-abfd-72870ccb72d7 · inbound

Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use cites this paper.

Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:20.652179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:20.652179Z digest=sha256:1d7e49d6f5f17a0b52269f07422d98ac8b1677b5df5ccdbd39d9b2797bd51fef

Observation 9ec2465d-b686-4696-a457-74c0a22ec0aa · inbound

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning cites this paper.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:19.512688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:19.512688Z digest=sha256:c43ac29c4abf3d6ae51dc24a1840e1ba34d9d0c2a1016280ddf3b45d5ba860b9

Observation 7b443f9f-0199-4440-9394-983dbf2e8ece · inbound

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data cites this paper.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:13.381090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:13.381090Z digest=sha256:63f6c54a3e159fdc681aa38b4d3dcb63101ac4522883cb076485907cf66c7fb6

Observation 007423dc-82e9-40c5-be88-0607471da412 · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.524656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.524656Z digest=sha256:739d8ff500b6adc8a4c739c75463fcb416f95aca1a229a36f41fa0ceaa8a2214

Observation 82dc9500-d820-4672-8a3b-a1b4f3ff12c7 · inbound

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning cites this paper.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.432916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.432916Z digest=sha256:9e2cff696a85378c772004e2535cffd4409c227711e4faa18fe13b302553f886

Observation 1f9a1363-5163-4ac3-aa36-633587a2a7a7 · inbound

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit cites this paper.

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:12.458369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:12.458369Z digest=sha256:9c4d207296aaf8bf43b66f597862aac5384603fa8fde6d5d21c8a9dd338cc735

Observation 1116d9ff-1bb6-45fe-8c3b-2d690506a98d · inbound

TTS-1 Technical Report cites this paper.

TTS-1 Technical Report The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.830714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.830714Z digest=sha256:789aaae286380cf1771b09c2a3a032c40e4416122eb7b9e906aeb7bb66ac33ce

Observation 3401423f-1608-4c4e-bf02-fad5fe8224fa · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.415152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.415152Z digest=sha256:acde19051c850ec41ad4769be356bf6063501af203928364986ddf74c37fc19d

Observation 11f03851-e677-4215-b44e-f17e27cdb100 · inbound

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training cites this paper.

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:53:38.926567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:53:38.926567Z digest=sha256:ae852f018daaf56559212bf24487de157987b2e52451cf1787423c51018b3dac

Observation 6f3f1d5d-1fc1-4078-aac8-50ad37a2817d · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:01:24.397871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:e34156da24a8a59d837e3bd0dfb72f38399fd62a3ea4666cb2100f486426312a

Observation f08ae745-8c0c-45c9-8e5d-5c4ce0d8d8ab · inbound

Swivuriso: The South African Next Voices Multilingual Speech Dataset cites this paper.

Swivuriso: The South African Next Voices Multilingual Speech Dataset The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T19:05:08.543038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:05:08.543038Z digest=sha256:c43812e55e523fb4c80e8740edf0e0a656c9de31949cc680f18508b86bae1c8f

Observation d7b854e1-cb19-4514-8f8b-5dff47f5ff84 · inbound

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs cites this paper.

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:30:57.407031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:06:50.408402Z digest=sha256:a76c0536101d3108f30083cfbfe1c33f53c1f3cefa5877a2a7989331097b2673

Observation bb045c2c-b638-4482-a623-086d0037e76f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.964075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:76572c92a9525860da21419a8342d271852b5885611df8b3ae6f8af1e7d225b5

Observation 068edde5-1384-476f-b6df-1f05db66a483 · inbound

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations cites this paper.

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:31:27.811824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T02:34:01.865434Z digest=sha256:c9d322443d3ba423487ac094fd8ba853f11c505e63e4a7a6b4be8dd9795abab5

Observation be8a2da7-1b27-480e-87ce-879cc3e59f4b · inbound

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations cites this paper.

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:59:11.692624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T22:57:59.897674Z digest=sha256:1a86683bd6b959343cf6f35f99a18db35fde93c5c4df661e8077d16209c8f732

Observation 49cb625b-4304-4618-9094-15bb842ea5f8 · inbound

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations cites this paper.

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:45:45.419554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T22:53:18.321056Z digest=sha256:c52373332949cf1c19a52a7f9f0736c2171d38fa7a0995214560e1e1dea8d2d9

Observation 167ab0d4-6cc9-48fb-9d09-f00e78cea41b · inbound

A Semi-Supervised Framework for Speech Confidence Detection using Whisper cites this paper.

A Semi-Supervised Framework for Speech Confidence Detection using Whisper The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:02:13.237407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T04:00:49.889795Z digest=sha256:855b18a6738c2508f853479c812ae952528af3ce571d5dcb9b3a9aa1167aa091

Observation 0d7c2b33-d648-40b7-8bcc-e3d509e1b74f · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.507165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:38be60d487eaedd0e8b123ba0e5bd730690d0af8a7f77bfd9b56ff8bc93cfe03

Observation a8d9ac97-39d3-446e-bb53-a2f48a6c3582 · inbound

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails cites this paper.

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:47:19.541629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T21:19:56.932689Z digest=sha256:76c97a8b62da9a1f44cfcf73d4ad6557a5e0c46b64e456764fa8cc51219d1fdc

Observation df4de4f9-255a-4315-90ca-f4423eb35d0b · inbound

Interleaved Speech Language Models Latently Work In Text cites this paper.

Interleaved Speech Language Models Latently Work In Text The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:59:42.903183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T10:41:19.777779Z digest=sha256:b73ced44bc1b115f5eeb76077c5002e4a5be0182af3f0f2af8763d00881a228b

Observation 6a8b02e2-f9f1-4c90-8e35-fb7561cdf825 · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.969024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:7a725d03bad70f1ee6567cbbdbc5b8efe02dd83a20f165144129a54db9eff086

Observation b47761c7-f254-4d3a-9afa-934c9feee09d · inbound

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR cites this paper.

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T09:35:44.381743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T09:35:44.381743Z digest=sha256:e217538728e577fed4f00aea5af1821848e110e4c5ecb182e0c16fc661407034