Pith. sign in

Paper Citation Record · LEDGER

Common Voice: A Massively-Multilingual Speech Corpus

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 98 inbound Pith citation observations for arXiv:1912.06670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.06670 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 98 of 98 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:42:39.900727Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T20:34:09.715401Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 87d70c28-3dcf-4566-9713-5d67ff848de3 · inbound

High Fidelity Neural Audio Compression cites this paper.

High Fidelity Neural Audio Compression Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:49:52.212568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T21:49:52.184932Z digest=sha256:f37b601d8d2e1e2ad0033a7efce106fa052412eddb159f27edf7e2b01bbef02e

Observation d7ebdee7-0e41-4c7c-89d6-d2cb89f158a1 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.422301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:23c1c031a806bff26413301b0997451166332307f54cb34683f8eb1086fbbdac

Observation 0193b1ef-cf8c-4438-bfde-1bb999c3fee4 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Common Voice: A Massively-Multilingual Speech Corpus

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.464496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:0545133181710d5ca485fcb07937504ed367e46430771514b10f498891498640

Observation 8484836c-7d02-49a9-b0a0-c48bc5308321 · inbound

Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data cites this paper.

Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data Common Voice: A Massively-Multilingual Speech Corpus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T20:42:39.900727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:42:39.900727Z digest=sha256:bee5143fe53bcee8f633ff1b50d5636293db780097a231c49e6379abf1a2cac5

Observation 34702040-3119-4a93-9df6-9688a1d41e76 · inbound

LIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative Models cites this paper.

LIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:27:34.226624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:27:34.226624Z digest=sha256:52b908c8b221d1ed4888650bd0bc2f4337a96a19be6f4c0b981e30013802c34c

Observation 956ed686-9594-4924-b517-422a317da9f2 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.996509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.996509Z digest=sha256:8342d8fb051f8a82a0cab655b39b784a68a7bfbbf9556af955828c0afaf00341

Observation cc86c316-5d8f-44fe-a60a-5e5e3d4e7047 · inbound

Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models cites this paper.

Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:53:50.193004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:53:50.193004Z digest=sha256:9200401abf715f2f7fa967890e514a27482fd9fb20d61338cf498b71cd6b0a71

Observation 1ba52b68-e7f6-4832-9721-4ba82c72dc29 · inbound

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch cites this paper.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Common Voice: A Massively-Multilingual Speech Corpus

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.362702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.362702Z digest=sha256:3aaa6b130eeff3ed5a8a9ff24e9bb24b83161723d26cfaf62a15f1e08dcbc508

Observation 433c683d-8d92-4f14-85ee-792f70a05c4a · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:05.972313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:05.972313Z digest=sha256:57091163bf6ee94b5695c7707a6d84be757b7d3a5877b9922c66b027f71947aa

Observation 069f4b31-69c7-48b8-9e71-9f8df2d5b3ec · inbound

Open Universal Arabic ASR Leaderboard cites this paper.

Open Universal Arabic ASR Leaderboard Common Voice: A Massively-Multilingual Speech Corpus

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:51:45.584508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:51:45.584508Z digest=sha256:ee2de5fd5cf79cafa8cefcb793295f25f4152e36f6350ec581459ac5c8115d37

Observation c6431d14-40bc-4618-97ec-5dacb1f90ac8 · inbound

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch cites this paper.

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch Common Voice: A Massively-Multilingual Speech Corpus

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:34.696506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:34.696506Z digest=sha256:eb740cbe526e26d86a464050590cf09f9ceb6cf244bd21bb5c74400e23cbc9d0

Observation 9bedda1e-26fe-4ac7-9716-f8d7f61661de · inbound

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training cites this paper.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Common Voice: A Massively-Multilingual Speech Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.639543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.639543Z digest=sha256:cdbde9011048e3a2c6bb0ea476f5b155e549fb4f30c8f742c45502dfd908f042

Observation 91b55cb7-7ec2-419a-9915-154f9b0191c5 · inbound

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues cites this paper.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Common Voice: A Massively-Multilingual Speech Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.080449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.080449Z digest=sha256:2eaa163fc97f1eb03a65392a0aa8fb4950fe3d9b3dde62049691a39c7ffd28f3

Observation 4bad7737-d010-4bdb-8af7-b4f872455d9b · inbound

Bridging the Data Provenance Gap Across Text, Speech and Video cites this paper.

Bridging the Data Provenance Gap Across Text, Speech and Video Common Voice: A Massively-Multilingual Speech Corpus

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:18:37.455322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:18:37.455322Z digest=sha256:99e6d35be8c1f396ce07262269611156f8e2f6d5e844a2ce16cbde73d8abe1ac

Observation eb003fc7-c6f4-4ad0-8706-355b6141c441 · inbound

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis cites this paper.

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Common Voice: A Massively-Multilingual Speech Corpus

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:50:14.391612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:50:14.391612Z digest=sha256:970855cf183a6344c6525591bf63224f3363a42f5d325854415823df920fd216

Observation aa485003-a78d-4b0c-920b-5c604ec2c12c · inbound

Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages cites this paper.

Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages Common Voice: A Massively-Multilingual Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:54:41.272845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:54:41.272845Z digest=sha256:5b388f4919167ba4199284335ac773de643bfc77a4932c4af5a4bdc897292504

Observation 8822f781-c07d-4f46-8fbb-655b88e1ed4a · inbound

Benchmarking Rotary Position Embeddings for Automatic Speech Recognition cites this paper.

Benchmarking Rotary Position Embeddings for Automatic Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:52.761839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:52.761839Z digest=sha256:d56484b7fe19cfc3e04b462af0a4e80166c51025d298e51f7a1c9e0e35f3bec7

Observation c7a839c8-45f7-4aad-af85-79b48f825d85 · inbound

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction cites this paper.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Common Voice: A Massively-Multilingual Speech Corpus

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.614639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.614639Z digest=sha256:cdc90ed09604466dc593c2318f62fb9bfb9beb7868b7fb1102eaf7d404d891bd

Observation e76a2dad-6511-442d-98a5-0d93d246afa1 · inbound

Continual Learning with Embedding Layer Surgery and Task-wise Beam Search using Whisper cites this paper.

Continual Learning with Embedding Layer Surgery and Task-wise Beam Search using Whisper Common Voice: A Massively-Multilingual Speech Corpus

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:20.595842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:20.595842Z digest=sha256:b1948eb63ee6d9330fb5d4889fcddf69982cb9f2f61977b2e89dc2608c927d2f

Observation b2eb1122-4a9a-44b3-9555-1dc409910eeb · inbound

Deep Learning-Based Feature Fusion for Emotion Analysis and Suicide Risk Differentiation in Chinese Psychological Support Hotlines cites this paper.

Deep Learning-Based Feature Fusion for Emotion Analysis and Suicide Risk Differentiation in Chinese Psychological Support Hotlines Common Voice: A Massively-Multilingual Speech Corpus

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:24:34.024941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:24:34.024941Z digest=sha256:e33e10c4cb80501023c573de9ede58c94f6c2f8c7a28ee75c5a05387589ecaf1

Observation beb4ae5f-7384-4bd6-ab7c-16cfd379af47 · inbound

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis cites this paper.

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:03:12.651296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:03:12.651296Z digest=sha256:ab93203d3cbb5a49d7692c3ec7841d5f49482dbdcb2f716d2d3f7ee8c964a49a

Observation fb80b939-254c-44d0-a589-6401fb6b00cf · inbound

PIER: A Novel Metric for Evaluating What Matters in Code-Switching cites this paper.

PIER: A Novel Metric for Evaluating What Matters in Code-Switching Common Voice: A Massively-Multilingual Speech Corpus

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:00:24.493040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:00:24.493040Z digest=sha256:0b4015f2bf5704ec8dbb1daf10af87b3b65207d1ae616fd8593e3b5d422b193e

Observation e45163ec-a30c-4ba3-9c74-163985ba9022 · inbound

GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems cites this paper.

GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems Common Voice: A Massively-Multilingual Speech Corpus

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:00.182562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:06:00.182562Z digest=sha256:9f4086d907dfeb09e564092e7eb5bdf4336ddc00fdea6dc9e2d17b26d8c90897

Observation 9d73cbc9-6125-4053-9972-51c881cd0a5e · inbound

Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets cites this paper.

Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets Common Voice: A Massively-Multilingual Speech Corpus

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:43:47.039372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:43:47.039372Z digest=sha256:fb4ace4ed83015ed0f86756993ee38e9272271d8adae7078932d2787683603a8

Observation 9794c380-5dd9-4569-a695-f8ce81a29f3f · inbound

Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages cites this paper.

Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:35:03.015272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:35:03.015272Z digest=sha256:c2c3450e2dfc5875c3b5325beb7af77040528695ec5a031710bf8fbfdeffebdd

Observation f20de216-c0fa-4ae2-9415-b9ade29e7593 · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.679058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.679058Z digest=sha256:78be0877a744c784ba7d1116c94e2840832511342a1e1c9faae14eaa6202d0f9

Observation a820f333-86f3-46f8-b3e8-c2ec82ea625a · inbound

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators cites this paper.

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators Common Voice: A Massively-Multilingual Speech Corpus

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T12:30:52.006659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:30:52.006659Z digest=sha256:460938ab815d7e33555fc0c96c9739295a5f90f04c9eac6177f98ca021cdfdd3

Observation 02251048-9c56-4018-b48c-6bc4fd013b8d · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.398909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.398909Z digest=sha256:07cb6d970040dd49fbb2e6efec0c3f0a7a6891a4f9d0a5a5e3e728cebb6a1b79

Observation 278b1ffd-223a-4ed5-b5d0-e5c554b3a460 · inbound

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System cites this paper.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Common Voice: A Massively-Multilingual Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.575954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.575954Z digest=sha256:542000ca18faedce6ce72b8099f30912d604d1d95aea71b2b6c87e8b89fbb764

Observation f9fc609c-8aac-4ee2-b428-e7c5142a8e11 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Common Voice: A Massively-Multilingual Speech Corpus

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:21:27.174181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:44a1ff0370f37b96700b22e86e3de6bbb0287857ccfa9cf1575f09bbfaf28135

Observation 2530e5b1-52e4-485b-a0c5-2e93f5dc27e0 · inbound

Word Level Timestamp Generation for Automatic Speech Recognition and Translation cites this paper.

Word Level Timestamp Generation for Automatic Speech Recognition and Translation Common Voice: A Massively-Multilingual Speech Corpus

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:34.180674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:34.180674Z digest=sha256:4eb57a5054b8653f5b14cd8365699ac2cb34a92a712b80951d26dbfed17b2b20

Observation 49016cc0-d922-442c-819a-db771d2bb707 · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Common Voice: A Massively-Multilingual Speech Corpus

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:55.919987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:55.919987Z digest=sha256:728ed43964b0105d180aeb2b9365dc86bba1ef2bc0969689720369c26ba0c3bb

Observation 51065def-f9b0-4641-bc23-c1fe90918aba · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:53.700409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:53.700409Z digest=sha256:e001f148c3cdfa6f0a705047e3c25a256ed74a305e2c4a76ba781fb57c8745ac

Observation ead80c36-0dca-484d-8c5c-2884a276e6e4 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training Common Voice: A Massively-Multilingual Speech Corpus

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:27:25.601064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:ba00656825cd6fcb8aecd7e7aa59af9f5394f08ad5fb4067a72923aae6e84ec1

Observation 822593e5-f9d0-44e6-af48-be096195ce24 · inbound

TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation cites this paper.

TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation Common Voice: A Massively-Multilingual Speech Corpus

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:03.646387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:03.646387Z digest=sha256:ea60c70f254ac16e13c2426ace669236209245e22b92712a88a527fc7e757b60

Observation e30882b3-e03d-401d-80b8-addefa9bd630 · inbound

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving cites this paper.

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving Common Voice: A Massively-Multilingual Speech Corpus

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:33.640309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:33.640309Z digest=sha256:0b7bd8a4e62caa6be6128a4be6aae58369e6611c5bb52058dc0621ae569715c5

Observation 77e558dc-80be-43d2-aa3a-a3b4ef0300ee · inbound

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning cites this paper.

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:01.031221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:01.031221Z digest=sha256:7b17f7b4df906b92a3b1922ea07c424fa7d0828e64ed55656454e77b27415d04

Observation 8907619b-305c-480c-bbdf-d9ac1a0a2de7 · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation Common Voice: A Massively-Multilingual Speech Corpus

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:48.384304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:48.384304Z digest=sha256:0dea2cc8dba7a5eeb7077bb5150bc620fb2d9ea7a84c177a205ac0476672b038

Observation 511f7089-1cab-4aa8-82d5-f460ace0b40d · inbound

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems cites this paper.

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems Common Voice: A Massively-Multilingual Speech Corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:53.941399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:53.941399Z digest=sha256:6a415216ae8e061a7840e0224aa756e8a620899b825a2b6ef79ec64ace1a097a

Observation 9dc565a7-c1be-46a6-9342-50d71a9f83fe · inbound

Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use cites this paper.

Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use Common Voice: A Massively-Multilingual Speech Corpus

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:21.198546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:21.198546Z digest=sha256:d622b38c8184ca8d6771b8a9cfac040f23e01c3b924c6dcba1c22de59d9273a2

Observation 41670f95-1e15-4599-984b-269336e035e4 · inbound

FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing System cites this paper.

FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing System Common Voice: A Massively-Multilingual Speech Corpus

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.436359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.436359Z digest=sha256:e904e9eb68ac57039f76695e41e04fd5a036b57e8fb1f4c1ed8a9eeb91914c2a

Observation 0b96b370-57c2-41bb-860b-08283607a3cf · inbound

SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization cites this paper.

SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization Common Voice: A Massively-Multilingual Speech Corpus

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:12:18.238078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T13:10:14.839742Z digest=sha256:9f5f4ade97d30a98378a10e1d24944925996afdfddb22110656bab258ef746d2

Observation 39e7fa83-9333-474a-b8a6-85f1307ad3fb · inbound

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning cites this paper.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Common Voice: A Massively-Multilingual Speech Corpus

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.661052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.661052Z digest=sha256:85e3b8da43a1119147aa448a70661ff63927454370e97b820015777c8e24be2a

Observation a1636c54-3f43-496b-9d53-5b339f20af2f · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.997075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.997075Z digest=sha256:94dd92da2bc2475ed3d6dfe03d2518a4de6bc95716caeb67db76a4b05b51c3f4

Observation 801486c9-2bc7-4312-845a-bd11e17f897f · inbound

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models cites this paper.

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:50.628216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:50.628216Z digest=sha256:e701959ad1e5dd533508a9e5bb6568f016c54eb47eb75c79a68d7f6649d35f76

Observation b37e89de-19da-4527-81ff-4772d24a45b4 · inbound

WAKE: Watermarking Audio with Key Enrichment cites this paper.

WAKE: Watermarking Audio with Key Enrichment Common Voice: A Massively-Multilingual Speech Corpus

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:28.765296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:28.765296Z digest=sha256:6a10c2af6cfd6560586642efade295644d86d64e76705809ac2b707973ee1792

Observation 1c04a5d4-acfa-4ff3-9d3a-f6b83235485a · inbound

DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction cites this paper.

DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction Common Voice: A Massively-Multilingual Speech Corpus

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:55.440956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:37:55.440956Z digest=sha256:f22a42e43bf710a6bf007f035866035af41c103c774213e927bf26e735f0118f

Observation 68cb72eb-e85f-4ef2-beb9-ec13ed9b7923 · inbound

Unified Semi-Supervised Pipeline for Automatic Speech Recognition cites this paper.

Unified Semi-Supervised Pipeline for Automatic Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:13.879732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:13.879732Z digest=sha256:dbf4cf12db8a692074cbcfa156d89e34794cfecdf383ca3a24cb9b7db10edc99

Observation 106a614b-fb86-4a1a-8522-122a8f3ffcd0 · inbound

Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study cites this paper.

Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study Common Voice: A Massively-Multilingual Speech Corpus

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:32:39.891475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:32:39.891475Z digest=sha256:eeb15b351e12d6fdf499c16d2a31285472d97b0935878c2cc750575a4a92e4ec

Observation 2962894f-7c94-4b4a-a343-4ee1a5ee70f1 · inbound

A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset cites this paper.

A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset Common Voice: A Massively-Multilingual Speech Corpus

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:22:11.038788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T08:18:48.936867Z digest=sha256:e989da5901849f328831d4973363f28c46e3d1f932d33c4fa9c5d171af78a38f

Observation 076ba73f-35e0-40f5-b804-1de34f299adb · inbound

Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages cites this paper.

Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages Common Voice: A Massively-Multilingual Speech Corpus

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:14:37.444832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:14:37.444832Z digest=sha256:b45fce6c99abde568b80647efc11a937fa853958119ff780df0428861a60c6c7

Observation 8ccdeccb-30cc-4b4d-8ca3-85e4de1a9633 · inbound

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit cites this paper.

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:11.785431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:11.785431Z digest=sha256:d6c1019823e2dee2a4929e850d2426b41b3b5e2acc997b6fa41283497f09bbfa

Observation f262486c-a9d8-4205-a021-39119f2fe0fb · inbound

Word stress in self-supervised speech models: A cross-linguistic comparison cites this paper.

Word stress in self-supervised speech models: A cross-linguistic comparison Common Voice: A Massively-Multilingual Speech Corpus

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:26.890986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:26.890986Z digest=sha256:ec85a64e91b707e20dc9682a7dfee2f9fa1de8f917d5cb0ba6bc91e7ebdfe5a3

Observation 6bc46688-aa9c-4cc9-8319-b95a644da6ff · inbound

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models cites this paper.

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:16.667135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:16.667135Z digest=sha256:aee361344594a6abd2f48d8ca8e92d09885283b857192de731c18f9677c90888

Observation 6e892657-92d8-449e-a339-711e1027ca6e · inbound

ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition cites this paper.

ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:32.006992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:23:32.006992Z digest=sha256:3d92bf9a8f80151730c07ec7780e33db70396a3fd9d34a0c498d7affef3de51a

Observation 1c8bd800-123e-4fda-8558-b0dd1ccf24ed · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Common Voice: A Massively-Multilingual Speech Corpus

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.974182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.974182Z digest=sha256:26ee2cfd8b5b3bef5a0d3f5cf7c4be3cc4b8331230a7e1894176c530b38ea605

Observation a56c7731-77f3-4b30-bc4d-519e45c14942 · inbound

PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy cites this paper.

PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy Common Voice: A Massively-Multilingual Speech Corpus

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:30.385565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:30.385565Z digest=sha256:2db517d4520945e3572d39fdb05f711e5833a9723cb215ec6bd088c8637c520d

Observation a231c882-4e11-4c31-b255-050eb910fe15 · inbound

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages cites this paper.

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:29.110045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:14:29.110045Z digest=sha256:88a047ccf1403d5e3f5db5961e0849797414f92853fdfd37653afb951857c965

Observation 8dc625b3-666a-4a08-9606-f8e326765323 · inbound

WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes cites this paper.

WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:35.064625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:35.064625Z digest=sha256:906e41b4d0ba3ff8cab7da400e5698a0b40ef607c3d8c57d0fd36fc3f0fc755b

Observation c9987e9c-dc00-46ec-9860-5cfaef5566e5 · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Common Voice: A Massively-Multilingual Speech Corpus

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:16.109262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:16.109262Z digest=sha256:91f4fa43bc32f1cc002d490a069a95af2198ee8ced5fb64a6bffc34d5e84b757

Observation a6762119-6952-44d1-ab89-8349a4d19939 · inbound

Large Language Model Data Generation for Enhanced Intent Recognition in German Speech cites this paper.

Large Language Model Data Generation for Enhanced Intent Recognition in German Speech Common Voice: A Massively-Multilingual Speech Corpus

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:52:07.674548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:52:07.674548Z digest=sha256:edadd168978fa15949966ceba6123c6b5025a4f2a03475e77414d97a9f45f304

Observation e7584f03-f959-4394-9e4f-b1583112fd27 · inbound

Non-Intrusive Automatic Speech Recognition Refinement: A Survey cites this paper.

Non-Intrusive Automatic Speech Recognition Refinement: A Survey Common Voice: A Massively-Multilingual Speech Corpus

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:35:45.863198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T23:34:52.249143Z digest=sha256:7efea7076a3c6bb70cde6bab71d4f05e55a987b2eaaffa1635a72c538afdf078

Observation fe1ca944-0ce8-466f-ad70-fe641d73ffa2 · inbound

Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge cites this paper.

Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge Common Voice: A Massively-Multilingual Speech Corpus

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:01:03.616409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:01:03.616409Z digest=sha256:044f91fafabfeff983a7b90da1cbfceb47997ea744a7330802d3242b3156661d

Observation 422dde6d-0940-4250-a76b-d943a18ec09d · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model Common Voice: A Massively-Multilingual Speech Corpus

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:49.739464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:49.739464Z digest=sha256:5007bddd59529146373bda6accedf2e46bde161d7ae96c6bdc84395844a8ebbc

Observation a2426604-3b6c-4168-a590-8cb736b79234 · inbound

AHELM: A Holistic Evaluation of Audio-Language Models cites this paper.

AHELM: A Holistic Evaluation of Audio-Language Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:33.923565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:33.923565Z digest=sha256:376a819fae59ee3123cb6966194df8764259d99d25e8160c16502ac8be9a5bc2

Observation ae85fb1a-5d31-4372-9235-8e5873baa5fc · inbound

Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal cites this paper.

Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal Common Voice: A Massively-Multilingual Speech Corpus

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:37:05.779244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:37:05.779244Z digest=sha256:c7ea7507bea0abc8dbfd8c2c2d37ccc1660334c8fe35d10d9785a0ed58e3f398

Observation 1eb22567-2f8d-49a6-ad62-29353a903ebe · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.418220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.418220Z digest=sha256:440ab5caad0b1805f351e0dc9ccda3ae87041b96c2fef76f802ae686b37fd060

Observation ea11c561-6f19-4e29-8379-56498d8eca7d · inbound

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation cites this paper.

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation Common Voice: A Massively-Multilingual Speech Corpus

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:58.811779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:58.811779Z digest=sha256:81c1e50a270bf4e9bce9c6b832c6f6eaf443593909c2af7024cb5ea4899f85a0

Observation dcb45fbf-76b2-4b23-9282-7c5e7b5d1de8 · inbound

DarkStream: real-time speech anonymization with low latency cites this paper.

DarkStream: real-time speech anonymization with low latency Common Voice: A Massively-Multilingual Speech Corpus

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T06:01:28.511141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:01:28.511141Z digest=sha256:4fd5bfa2e558ea9c537b175d0a255788aab509f7069e7b48e0fdde85b47076b4

Observation b032847c-892e-4e51-8698-b22636ca72b7 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs Common Voice: A Massively-Multilingual Speech Corpus

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:01:24.413278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:08fde0d57fe7bc2d4a6e6accf112a6090a50b85783df900b69b7a3975131abe4

Observation 45a4395a-97e2-4646-939d-8f26357d7598 · inbound

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis cites this paper.

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis Common Voice: A Massively-Multilingual Speech Corpus

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:17:43.634745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:17:43.634745Z digest=sha256:9acd87950e408cff504bc5196ab4db2fd7efcf870d8b79dc2bb5932d82802856

Observation d6b5de18-f890-4206-93bc-4bd68033e922 · inbound

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation cites this paper.

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Common Voice: A Massively-Multilingual Speech Corpus

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:07.490691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:07.490691Z digest=sha256:da6f22005eb3ccd6bcf1347bc25146a06d1c940c5e33fc3fdb7c38cf67df833c

Observation 7d369fb6-a3c9-4bf7-8fdd-983a32285256 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding Common Voice: A Massively-Multilingual Speech Corpus

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T18:20:29.243328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:26fbe5ef5caaf9ffbdc2b6b5ddf4c94177665cba6077df16099d7e14397fe2ef

Observation 7f940f57-c615-4b29-971a-b7a253b8a8ee · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Common Voice: A Massively-Multilingual Speech Corpus

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T12:01:59.035213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:01:59.035213Z digest=sha256:af0a4b1f6a093804292b6053c40b73149ac4fd5bd00b604a4e87e26962b627ea

Observation 63171c25-1032-49cb-b3f3-1a3814d33a8f · inbound

Interfaze: The Future of AI is built on Task-Specific Small Models cites this paper.

Interfaze: The Future of AI is built on Task-Specific Small Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:48:05.360003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:48:05.360003Z digest=sha256:97fb459e1b231ab6c26376b0ec1635d56787f101cbe6fe666dcbedb8f87d3813

Observation 5524b788-b888-421a-b675-65cc93c9bf0e · inbound

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition cites this paper.

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T00:03:31.986628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T00:03:31.986628Z digest=sha256:49de6340464c07dda34bcb4dc0d242aedb20e1b3cfa0defe6cf2c8c3b2271e02

Observation 7de6052c-8007-4a3a-ab65-7e072c95854e · inbound

Membership Inference for Contrastive Pre-training Models with Text-only PII Queries cites this paper.

Membership Inference for Contrastive Pre-training Models with Text-only PII Queries Common Voice: A Massively-Multilingual Speech Corpus

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:09:59.551459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T12:09:32.204944Z digest=sha256:7c1911a97960ef06cf39e367beeb3f0942006140727b7ab93c4214c8850ca41b

Observation 9e4644c7-a385-4738-a293-765d5fa09f60 · inbound

IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA) cites this paper.

IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA) Common Voice: A Massively-Multilingual Speech Corpus

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:18:30.051019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T00:15:10.797922Z digest=sha256:706d7231a4b4d72d904c36718653c8142767af2395ad67a013ad2dc4073fe18f

Observation 10bc4ec0-7b9e-4e25-b5cc-045f77a37c64 · inbound

BlasBench: An Open Benchmark for Irish Speech Recognition cites this paper.

BlasBench: An Open Benchmark for Irish Speech Recognition Common Voice: A Massively-Multilingual Speech Corpus

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.711718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T15:53:54.092426Z digest=sha256:1318148e452eb8e0244ccf8d99bfda3df06e99c122fa527126f772051afcd3ad

Observation 1b71ebac-e480-422c-a3f4-c5c1c6a696fc · inbound

HARNESS: Lightweight Distilled Arabic Speech Foundation Models cites this paper.

HARNESS: Lightweight Distilled Arabic Speech Foundation Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:56:14.209808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T02:17:14.045133Z digest=sha256:6932dfb5198dbc5c33519855d240a9386ebd855475164a27c8ac70ab77ac9777

Observation 522f620e-e0e9-4be2-b730-883459543836 · inbound

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions cites this paper.

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions Common Voice: A Massively-Multilingual Speech Corpus

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:35:26.720991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:25:50.524448Z digest=sha256:538fbc192adb1ce5f938aedc19513073cf2b0996e10c22bc2538ca6c688d33b1

Observation dcb5a8e6-1466-41e5-8cd7-eec2813cb4b3 · inbound

Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR cites this paper.

Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR Common Voice: A Massively-Multilingual Speech Corpus

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:27.996529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:59:47.796221Z digest=sha256:10bad7d8d7d9898298e2625803d338053759aed6784e76a8d4af9fea04a2a034

Observation f136c8b9-50e0-49a3-9d51-62b02c8a5089 · inbound

Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs cites this paper.

Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs Common Voice: A Massively-Multilingual Speech Corpus

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:23.375304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-09T15:17:58.020698Z digest=sha256:784f552a80d4480f891b9efe629529ebe450363fc5bd7d799177fdbafbe4529c

Observation 872903ab-40f3-421c-b4df-b32d44ece1d0 · inbound

Keyword spotting using convolutional neural network for speech recognition in Hindi cites this paper.

Keyword spotting using convolutional neural network for speech recognition in Hindi Common Voice: A Massively-Multilingual Speech Corpus

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:05.900745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-09T20:06:30.395199Z digest=sha256:d4b93c4a1c861461a55f3f8abce469a3f88cef362a696d30a440e38158c258e9

Observation fc5c1070-dd68-4004-8cd4-c0fb07e20c0d · inbound

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues cites this paper.

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues Common Voice: A Massively-Multilingual Speech Corpus

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:37:42.805412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-19T18:36:34.351533Z digest=sha256:b4ae73548345c5f17855cbc782a202bdaf5a64dc4ee405efc877af2dbe5dd643

Observation 78f7ef53-1272-47b8-8317-1d3307895569 · inbound

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking cites this paper.

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking Common Voice: A Massively-Multilingual Speech Corpus

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:03:14.633064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T07:55:17.832799Z digest=sha256:15b678c6e92cdb07cf56d6fdefc84a9fe92c5f84c79af70c0fbd51c53d9364b7

Observation 6f04cc74-060f-4da5-ad2e-dc517202f40d · inbound

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi cites this paper.

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi Common Voice: A Massively-Multilingual Speech Corpus

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:39:39.270772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T13:07:37.187581Z digest=sha256:397cb218efd7a4d9822df98758302cf8c75140a6abba3970695f13d65e6c2c46

Observation f6011e6e-ed6b-4391-8251-7d67b61172d2 · inbound

Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis cites this paper.

Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis Common Voice: A Massively-Multilingual Speech Corpus

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:29:41.251618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T11:43:27.916617Z digest=sha256:5a545eba1baccdf552bc6b2b9f456b5e6c299a7b133a67a54189d83e9077ba03

Observation ab8dc48c-42bd-4bf8-b2af-870fa22a983d · inbound

Learning to Evade: Adaptive Attacks on Audio Watermarking cites this paper.

Learning to Evade: Adaptive Attacks on Audio Watermarking Common Voice: A Massively-Multilingual Speech Corpus

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.958897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T10:10:18.351233Z digest=sha256:4f6263e755073498ded52526e21a628eba8bedce57686bd2f03442c7dbc08c8f

Observation 7047bc05-c243-4be1-81a6-650780d4da94 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Common Voice: A Massively-Multilingual Speech Corpus

Reference 144

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.265773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:d037c238599b04655fc4d429ae2550459be0164aca6c0c22d3bfa7c913548a34

Observation 1a6a08bd-188d-432f-b825-39d22a18e583 · inbound

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions cites this paper.

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions Common Voice: A Massively-Multilingual Speech Corpus

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T20:34:09.719817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-07T20:30:53.546267Z digest=sha256:081cd9845d2fc3509bc721672298437ff0336c79828afcc8c9bf0881b354a13b

Observation c790d5c2-fdba-44e9-a9d2-a68a4f1238c0 · inbound

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing cites this paper.

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T14:53:55.894823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-07T14:52:35.127533Z digest=sha256:4f8164a83789c3ea92b5b5c25e90093cdfe2e1b93e2891a6897818ffec24548c

Observation b3cbbfa1-9c6c-4a0a-ac99-8ca4a6c40ae0 · inbound

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing cites this paper.

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing Common Voice: A Massively-Multilingual Speech Corpus

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:34:02.641142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:34:02.641142Z digest=sha256:f4590976b329fd3ee47a776d69b0c388a8cd2b680c8e370e0b58084bd183beac

Observation e4380b62-0936-4e47-82a3-4bb94c3e055c · inbound

Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models cites this paper.

Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:37.320627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:53:37.320627Z digest=sha256:448c1be32810121f7bd4ddc2660ea58abb616372f2b261e01c61d3aecae0028d

Observation 95199efc-9c44-46ac-830e-f2c320f21ca5 · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision Common Voice: A Massively-Multilingual Speech Corpus

Reference 185

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.600447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.600447Z digest=sha256:a98cbbadab42281e1bb74776a43414c291b6475c367e8e5ab06b85559ad8db90

Observation ecd4ba63-4d6f-4dcf-9137-29f26b6533f3 · inbound

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm cites this paper.

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm Common Voice: A Massively-Multilingual Speech Corpus

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:25.843132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:25.843132Z digest=sha256:716547addfd6f447595677fd7064b9eb17b5cccad6d6de44b42d8bc24549e299

Observation 0e9e35a7-7d5f-4727-a126-579c37965c3c · inbound

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection cites this paper.

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection Common Voice: A Massively-Multilingual Speech Corpus

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T15:10:14.178271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:10:14.178271Z digest=sha256:6eee532f68444d32144bb3ded9798f28bb313d95ca3b3c8bbb52acac8955e17a

Observation 3476e971-3e6f-40ba-8731-b3c333b0d9d3 · inbound

Leveraging Beam Search Information for Confidence Estimation in E2E ASR cites this paper.

Leveraging Beam Search Information for Confidence Estimation in E2E ASR Common Voice: A Massively-Multilingual Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T09:46:27.897409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:46:27.897409Z digest=sha256:46127f9bce5e5d89dc9b1971c7cebf060794f30a21e2e2453e2d10e27a8afe48