Pith. sign in

Paper Citation Record · LEDGER

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2101.00390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2101.00390 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:52:53.410512Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:59:42.926035Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b5f067a1-5f74-48e5-bc36-15b8f732aaea · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.473790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:bc4c8e38c8ecdfacac20b58d6d4d9e372dbad05cbc027611a40cbafab509793f

Observation eee5a22d-e625-4e09-90e3-1b350bbe1090 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:58.097144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:58.097144Z digest=sha256:b17df60672817ae33a0af468b5da9871ae6fa0e4bff4f62ab334df8a8905d828

Observation 0416c9cf-68ee-4316-8720-57bb6deee458 · inbound

Scaling Speech-Text Pre-training with Synthetic Interleaved Data cites this paper.

Scaling Speech-Text Pre-training with Synthetic Interleaved Data VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:02:40.541338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:02:40.541338Z digest=sha256:75c2aca1486944c7c66441e25bab2eeeb619983d95f8284f3cc1601f8a19fe80

Observation 3d464ebc-e6c3-4f03-9102-50a4f111c769 · inbound

CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing cites this paper.

CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:10.982319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:10.982319Z digest=sha256:6d0ffd5c780f19da63c5348486792146583fe1df26dfd0db99783c5259d1b0ab

Observation b3ea748a-b7de-4c56-8c13-5ab7b14e3b4f · inbound

A Survey on Spoken Italian Datasets and Corpora cites this paper.

A Survey on Spoken Italian Datasets and Corpora VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:59:35.364417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:59:35.364417Z digest=sha256:f1dad0bae071302cc1983efd695543bac9d6588a71ae7c891d2860eb3f57f439

Observation 365fdb58-83d0-4756-99da-81ddc14db160 · inbound

When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation cites this paper.

When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T19:16:37.888931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:16:37.888931Z digest=sha256:57b2d860d3f57135fba8e1fb57371f9bdbfe1ec7a66d0a5d9b0b38a87ca583a1

Observation df8b7a56-ea67-465b-8f34-2edfe29b78dc · inbound

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention cites this paper.

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:30:31.517132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-25T08:30:15.011210Z digest=sha256:3b27a3ac6223c58bf06e24e02b7e62a1ae3f6d225c5209e5679313a578b5e020

Observation 8c466509-5f86-4db4-80f3-df8a3dd8b06d · inbound

Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models cites this paper.

Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:49:20.110515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:49:20.110515Z digest=sha256:7ef179bdd49b406bc9f6351e0ab32a8bd639dc2b0bb31895010b20003d13cabd

Observation b4ca6cc9-34e4-439c-b240-26c98d0a3493 · inbound

On the use of Performer and Agent Attention for Spoken Language Identification cites this paper.

On the use of Performer and Agent Attention for Spoken Language Identification VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T17:49:21.497522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:49:21.497522Z digest=sha256:86a18b5aac73ce76cc7278b5328f59a3b02a01b3b523ab02c6b532b4a9da5302

Observation 1c453f19-e2a0-4c06-8853-114f6e5ab342 · inbound

Speech to Speech Translation with Translatotron: A State of the Art Review cites this paper.

Speech to Speech Translation with Translatotron: A State of the Art Review VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T17:12:07.468458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:12:07.468458Z digest=sha256:7c6df4e81ce9440c4e1f3e304ed4d517e053a3c196d622ba86718192f9287a9d

Observation b8dc698b-945b-4e0b-bc4d-6653c2e744d6 · inbound

Evaluation of Deep Audio Representations for Hearables cites this paper.

Evaluation of Deep Audio Representations for Hearables VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:47:08.009413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:47:08.009413Z digest=sha256:23b7a9075bd37ab79cfe5bab7fbf25ebf4dfeb3304528d1c920c0652c00fbbdb

Observation 444d5cec-fedb-4c44-87c4-6ac9730a923b · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.071481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:19fce50a183f0617b032ff06f525b91c9ec711e2c959ec7d47026139fa79313f

Observation 4dbce808-0594-4b3e-8ae8-6d608a832352 · inbound

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities cites this paper.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.410512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.410512Z digest=sha256:5eeb7b059a5fc71640296de63e01429c12e30c40e305cd50a529bc8de70decf1

Observation 339a71f8-d533-4f2b-8879-123ffeacb751 · inbound

Inclusivity of AI Speech in Healthcare: A Decade Look Back cites this paper.

Inclusivity of AI Speech in Healthcare: A Decade Look Back VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:20:25.387876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:20:25.387876Z digest=sha256:829f5b4758a64300c7b1d1b31d22eb10df1fd11ce2ff05023a3d12d94a7d4432

Observation 3cfacb16-af02-45e6-9163-d43e839a6056 · inbound

HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification cites this paper.

HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:32.780424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:32.780424Z digest=sha256:ddeb9c3e10893cb9bb9325aae7b164dec341c3531af42c872e76ae9fb36cbd67

Observation ee506292-b81c-45b4-8312-f35493fea5df · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.237399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.237399Z digest=sha256:4ee358a965d0d44a9e6f0842126630ad9403253b2d55da96632294cc13c815aa

Observation b12b32fd-de54-4556-817b-b2c42d11dca1 · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.846781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.846781Z digest=sha256:e3e0e49da9a0a35c533179dd83947f6d39cfbce5138bd0166b37b47610fb54fa

Observation fc9a8ab5-343b-4ba1-ae6c-dedcb0b7b62b · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:48.668479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:48.668479Z digest=sha256:74132d6dd8d20524771817e9595248f57a44cb071a8b2bf5ed47afc5bbb178f5

Observation 62cf6209-dec3-431b-98bc-12855fdd7f48 · inbound

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition cites this paper.

MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:26.212501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:02:26.212501Z digest=sha256:51d957bd4fb6b0e7278d1f9c57cc19730974619caf14c2e6baff3bbacf0b0be5

Observation 51c5bdbc-2ebc-4811-8a2f-a6b72e0b28e1 · inbound

Unified Semi-Supervised Pipeline for Automatic Speech Recognition cites this paper.

Unified Semi-Supervised Pipeline for Automatic Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:13.888662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:13.888662Z digest=sha256:b8c1eea19d578991854fb068274cc16676b428d7313178cff8eb4348c06213c5

Observation 2ae6d9e7-e12c-4cd5-98cb-8ce49c00d530 · inbound

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching cites this paper.

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:52.578363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:52.578363Z digest=sha256:ccd90a4f79c2f537965b660e26d937b598e0fe5437bddabf247989ef8815b5f3

Observation 9d90560f-7ae2-47ac-a5cf-b99fa37eaee7 · inbound

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning cites this paper.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.173525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.173525Z digest=sha256:dfbe3a38e80bc5eed24c5e283cf073e6a663d3ccb8539ffa5e59f29846081dc5

Observation 2e9d6ea3-8a1f-4ed8-8865-22e14f8480c2 · inbound

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models cites this paper.

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:18.489814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:18.489814Z digest=sha256:6e7bff38f0e5839a7aec91544e3d6a2813fca491883488ec45c48e1cf064686b

Observation 3ece12cd-30a7-43f6-b941-35aafc10cc84 · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:45.248930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:7c483708148a11a2067d57fd18fb4432d788ace70bbe69380d4244bda9aaaa2d

Observation 6d9537f6-57df-417c-a1d3-c98e3510fb15 · inbound

On Barriers to Archival Audio Processing cites this paper.

On Barriers to Archival Audio Processing VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:08.902551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:14:08.902551Z digest=sha256:a6ce4e113f090a246f812b161406fb0245f6914c198f645df22a6309408ace31

Observation 6404bd70-32fa-466e-b654-406b91370c41 · inbound

An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications cites this paper.

An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:13:09.439905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:13:09.439905Z digest=sha256:11816d473d06eb0a479ad5b8eb229451ae78a7efa09b9a89431a863045aa159a

Observation bf6f6e49-c096-41c8-86db-03287d382039 · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.412121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.412121Z digest=sha256:a846ac4ef769f290557bb97ceede38fe5c957559b7a1514f587a236eab553d9c

Observation 1d292078-0891-4d23-92c2-4d347df25a3b · inbound

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings cites this paper.

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:53:25.480535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:53:25.480535Z digest=sha256:7b2dbd91440b56a1b9f60625082c639fea0179e277a952aede3481b7c3fe1f0f

Observation caa0a545-4f09-45b7-a604-0ee6e3a1e41a · inbound

From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model cites this paper.

From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:16.274239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:02:16.274239Z digest=sha256:2aa4b99cf08a559d4b8110e301cf0b17c6577a3b59d599a43c2575c29e5387cf

Observation cda344ad-35fd-40b0-8bde-87d3c6790e7e · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.405536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:1ce06ffb8c86a93e9cc70afb0b8dc6d832341dd93379b77ad678b837c3e4603f

Observation 512e650d-05d0-4566-8218-f216c51451e2 · inbound

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis cites this paper.

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:17:43.934752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:17:43.934752Z digest=sha256:049936e99d10ca12691b83c27c3b5d223149b3c16c15e3fe9ff0904fc1cba95d

Observation 9f150e90-e3a9-48ec-90b8-7a5747de2ada · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:02:02.010823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:02:02.010823Z digest=sha256:3583cf109b6646b6695b12011468ec2404f54aab8bdd2dfe54962103069156dd

Observation 21093ffe-41f4-445d-8d76-c11f59339e69 · inbound

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition cites this paper.

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-15T00:03:31.986628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T00:03:31.986628Z digest=sha256:97968d3dd076bf0299282397c28733df9385591adc985c49c20f694800731ebd

Observation fa3ec273-1c9f-4c28-a4d4-a5f4b4cada9b · inbound

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions cites this paper.

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.707984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:25:50.524448Z digest=sha256:948d692ef937c8a84b7edc833be1a07c23875b30d1263d6eaf6423e00d066d15

Observation c1ff7800-3715-43f2-8791-515dfcb4e63f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 126

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.995475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:8af89d8459ed183ce05de839923c21cd8050b643f96f7a60e703be2b8bf123dc

Observation d8f15238-338e-4b4a-bcd5-70bc8f4a5c2d · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.398729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:fcf2f999717b7ceac8bd9a1d5f49cd107a7dafebd1b11cd63f4a504db1ddba3a

Observation afe61cb0-722c-4aac-b777-01fe72d86fb1 · inbound

A Unified and Reproducible Experimentation Framework for Speech Understanding cites this paper.

A Unified and Reproducible Experimentation Framework for Speech Understanding VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:16:12.228406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T21:20:16.428207Z digest=sha256:52ed634030945450eac7ec68875742ac9e1397030bdd66c12e0a1adb77aec723

Observation 50d2ac01-4d92-4ee5-8c71-936d4a543597 · inbound

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation cites this paper.

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:29.117785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:59:29.020530Z digest=sha256:000d927811c3bffd009e76b4f7697d4f06e9b9ebbb94bb68f8a319c35af29282

Observation 796dd2b9-b815-40d1-9b73-63d8316d7d75 · inbound

Interleaved Speech Language Models Latently Work In Text cites this paper.

Interleaved Speech Language Models Latently Work In Text VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:59:42.927597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T10:41:19.777779Z digest=sha256:339a500a8ad4b571977d4531d48956a56678091c433fdf372ed9026b51dac3c4

Observation 2611f189-4158-47dd-b79a-961d6fd9cc4a · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.517800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.517800Z digest=sha256:bbad1a3d91300f3f590d15a163a60ad6b2af1c8d6fe5fbea7f9ab86dc8e74bf1