Pith. sign in

Paper Citation Record · LEDGER

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2101.00390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2101.00390 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:03:32.780424Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:59:42.926035Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b5f067a1-5f74-48e5-bc36-15b8f732aaea · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.473790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:e226ba3f46d0f05c9493a8acabcca83b5c5091123fdf1d9c5720912312d8f21f

Observation df8b7a56-ea67-465b-8f34-2edfe29b78dc · inbound

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention cites this paper.

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:30:31.517132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T08:30:15.011210Z digest=sha256:f93fa7e905d6402fbfe5a7cdc0ce8e6e0e5ef0f6f91fd63990f02c4d52f2e4ee

Observation 444d5cec-fedb-4c44-87c4-6ac9730a923b · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.071481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:6b545da8240334b00f7ebea13ec292f7ff16689f4faa7c93dea56068a54e9afe

Observation 3cfacb16-af02-45e6-9163-d43e839a6056 · inbound

HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification cites this paper.

HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:32.780424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:32.780424Z digest=sha256:f5646a4b9f3b3553e07972f9d26a32379f7c18c76534965b59d6ca71edafe679

Observation ee506292-b81c-45b4-8312-f35493fea5df · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.237399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.237399Z digest=sha256:3792b2bc65f58420d8130907adb618eb9722b2adac721646305cf3485883005d

Observation b12b32fd-de54-4556-817b-b2c42d11dca1 · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.846781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.846781Z digest=sha256:a2a6e740721e8297d19b97620e7866bb93d01f1b2652ec957bb69e246daf271d

Observation fc9a8ab5-343b-4ba1-ae6c-dedcb0b7b62b · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:48.668479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:48.668479Z digest=sha256:d2f63a7be88f3994fed910d8f9d36c245cd1b13d4ebf1805b3818c614978e9b0

Observation 62cf6209-dec3-431b-98bc-12855fdd7f48 · inbound

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition cites this paper.

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:26.212501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:02:26.212501Z digest=sha256:abd2aa896f0836243381b148afc805203983a96d3fcd78242e5f33b92fe09e36

Observation 51c5bdbc-2ebc-4811-8a2f-a6b72e0b28e1 · inbound

Unified Semi-Supervised Pipeline for Automatic Speech Recognition cites this paper.

Unified Semi-Supervised Pipeline for Automatic Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:13.888662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:13.888662Z digest=sha256:81f05b25e8aab22299e249dd9a3e8279635ec755d5b48871412895a773bc4a3b

Observation 9d90560f-7ae2-47ac-a5cf-b99fa37eaee7 · inbound

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning cites this paper.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.173525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.173525Z digest=sha256:8fc5b018940d7a8815831b8a164b2e2923254c0fef272727802dd81796c798d5

Observation 2e9d6ea3-8a1f-4ed8-8865-22e14f8480c2 · inbound

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models cites this paper.

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:18.489814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:18.489814Z digest=sha256:d2edaf4f5a8d07a0636728c0d10a09427b3d67ea99caedab990b13b85f7aa852

Observation 3ece12cd-30a7-43f6-b941-35aafc10cc84 · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:45.248930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:6019e4b0c80c6fefcf6f2bc0a44b2ec0a157a8068e36e74a9a948f08043cfc86

Observation 6d9537f6-57df-417c-a1d3-c98e3510fb15 · inbound

On Barriers to Archival Audio Processing cites this paper.

On Barriers to Archival Audio Processing VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:08.902551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:14:08.902551Z digest=sha256:f667e789767665c7d215192e7509eabd30d1dd39211a7146c7cd00d3530ba2c1

Observation 6404bd70-32fa-466e-b654-406b91370c41 · inbound

An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications cites this paper.

An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:13:09.439905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:13:09.439905Z digest=sha256:2df8698f7083e75fb0ee51c087e5392d81e5fe8974584aee0315d6794227813b

Observation bf6f6e49-c096-41c8-86db-03287d382039 · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.412121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.412121Z digest=sha256:7f98a94d44e7c4d1d79311368062413c05878b7a1ff23d54d850b662edf69060

Observation 1d292078-0891-4d23-92c2-4d347df25a3b · inbound

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings cites this paper.

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:53:25.480535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:53:25.480535Z digest=sha256:e3e6b8e040e43d85b045524267f6e3d2419d91b39a446d24e082c01b0f2021fa

Observation caa0a545-4f09-45b7-a604-0ee6e3a1e41a · inbound

From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model cites this paper.

From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:16.274239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:02:16.274239Z digest=sha256:c4592712aaf702fd2ce0a1ea4fdf6aaf2b27654bb36c5fa5080cc6bd240da990

Observation cda344ad-35fd-40b0-8bde-87d3c6790e7e · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.405536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:4049728fc7372fad90b706d5097a5c1b7f8bf5e65d04a0790f67eac103239e20

Observation 512e650d-05d0-4566-8218-f216c51451e2 · inbound

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis cites this paper.

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:17:43.934752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:17:43.934752Z digest=sha256:f9e38acd12aa5a26796dcfb93119e5b7ee39906d4e760982310674ff65b59f09

Observation 9f150e90-e3a9-48ec-90b8-7a5747de2ada · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:02:02.010823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:02:02.010823Z digest=sha256:3134475223f004abc587b23b293b1bb6cb4df20a3839567a0db7935753d00293

Observation 21093ffe-41f4-445d-8d76-c11f59339e69 · inbound

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition cites this paper.

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-15T00:03:31.986628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T00:03:31.986628Z digest=sha256:fcf3365b1aaba4b2cfe72dcfd5b7b59fd7c728ba1d16a9cb8839537453d054cf

Observation fa3ec273-1c9f-4c28-a4d4-a5f4b4cada9b · inbound

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions cites this paper.

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.707984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:25:50.524448Z digest=sha256:693762957ae2c1005797555b7fb92ec1147a541bed9bf94f30fb52a579cba971

Observation c1ff7800-3715-43f2-8791-515dfcb4e63f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 126

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.995475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:1fcba8ebf80710a2602b758589390af3860dc1badf671c2d8fc65ba0e8b45fef

Observation d8f15238-338e-4b4a-bcd5-70bc8f4a5c2d · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.398729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:e976f5cf6f74657ae56b0bf4e4d69baf30d22045acf8313e579781092a5b11af

Observation afe61cb0-722c-4aac-b777-01fe72d86fb1 · inbound

A Unified and Reproducible Experimentation Framework for Speech Understanding cites this paper.

A Unified and Reproducible Experimentation Framework for Speech Understanding VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:16:12.228406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:20:16.428207Z digest=sha256:c32a63c9ccd13ca7894324f39dc0d5341711b1ce30ca376cc496b4ca16bb689e

Observation 50d2ac01-4d92-4ee5-8c71-936d4a543597 · inbound

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation cites this paper.

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:29.117785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:59:29.020530Z digest=sha256:cd17e8846e5b418803a13e7d4962b63cee7e046359407e522bdb31b1007b79af

Observation 796dd2b9-b815-40d1-9b73-63d8316d7d75 · inbound

Interleaved Speech Language Models Latently Work In Text cites this paper.

Interleaved Speech Language Models Latently Work In Text VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:59:42.927597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T10:41:19.777779Z digest=sha256:efa26ed7266e2db49298b5145e2c46c98bde75d5cbdaf5c625f8fa9771d7535e

Observation 2611f189-4158-47dd-b79a-961d6fd9cc4a · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.517800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.517800Z digest=sha256:070cdbb9f5f222adb3d408f6807707d48c5c15b1b961340c2aeb7dfc647fd47d