Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 72 inbound Pith citation observations for arXiv:2106.06909.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:38:53.974812Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-07T20:34:09.701797Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8a18530a-2ebd-472d-880f-03b06f2bb154 · inbound
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab94963-c9e0-4eb6-93b1-b26583bf40eb · inbound
Whisper Finetuning on Nepali Language GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db6cd1f-8b9a-4688-a0af-bcaaf51a1b75 · inbound
WavChat: A Survey of Spoken Dialogue Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b85e59-2d2d-4431-a9e5-654e150edb9e · inbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c048bb-6d54-4e1b-8abd-1e3a9e20b0c1 · inbound
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1a21545-2c5e-419e-929b-ef44887ff30c · inbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24beb63f-628e-4a50-ad59-36a4a5a335d0 · inbound
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c846a0c2-7942-4035-9215-f616b19fe0a9 · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93415f77-b6eb-4f7d-a13a-7e99a0715fb7 · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f0b1b95-91c9-4d0e-8c6e-06143f7f07cb · inbound
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c68a647c-86b7-4515-9d19-a420d719dd5a · inbound
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf0e655-85f2-4d3f-9b8b-37945160f7b0 · inbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 339dcc89-68b2-4e8b-ab24-4856fa4e70ca · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0f1a0c-da2c-42f6-9c52-6db19d93f9c0 · inbound
Evaluation of Deep Audio Representations for Hearables GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1260c168-220b-46b3-905e-6764d7681cbe · inbound
Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c60449-ed93-428c-8f9b-b216f18f5e12 · inbound
Kimi-Audio Technical Report GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c62cb1a9-98d8-40f0-9cf9-153aa3c703e3 · inbound
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7402db6e-24e3-42a7-a81c-af7fabe7ee52 · inbound
Inclusivity of AI Speech in Healthcare: A Decade Look Back GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1424ff3-c562-4105-8186-aa8cf416a15e · inbound
Granary: Speech Recognition and Translation Dataset in 25 European Languages GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd816b8-bd3f-4403-b5fd-9e9f06e09840 · inbound
HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ba5252f-5ee5-4872-bef6-831bfa81de8b · inbound
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6c69078-b779-436a-8e7e-6dc2e647e0ad · inbound
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f8d9f6-776d-4b89-85bd-820eb4c85620 · inbound
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9409cea-226f-4a71-8934-78f224063019 · inbound
CASPER: A Large Scale Spontaneous Speech Dataset GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4953a05e-06d3-4de5-8514-3180d6571e33 · inbound
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d25116b-720e-4bed-b17b-d3100c804aa7 · inbound
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c4cdf49-992a-4d22-a6ff-09981ad38bb5 · inbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d1b093-dbc2-468e-83ea-b92226d3d12a · inbound
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3470d9a5-a1c5-459d-9c1d-a7bd6143cace · inbound
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb48705-2f65-4847-94ff-a778753b0bd1 · inbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16578a72-185e-49e9-96be-25c9a678c37e · inbound
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b71415f-4c51-4d89-a66a-02a530104bb8 · inbound
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6883eb5-7f96-4603-9ce0-1d88626dbdbe · inbound
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3da493c0-b1b5-47cc-afeb-a8ec8a9fa604 · inbound
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5156ee4-8fc9-44d8-9f7d-ac6a252ac9bb · inbound
The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a05ab93-c5b3-4b3c-9891-049835cbe534 · inbound
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52c13a8d-bb18-4c53-8442-e9b31c5ffe1c · inbound
VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 501501f7-ad34-4fb3-a70f-02e0dd2f7e04 · inbound
Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef66db8-5f36-4677-96d1-51de0a0f3860 · inbound
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d58ead-f7e1-4f16-a889-aa6095d78d70 · inbound
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e896d81-0b83-4218-a502-514d47561471 · inbound
Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ad6b17-a9c7-4b66-a224-3486e2a55fb8 · inbound
Audio Deepfake Verification GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f991931-c0a6-40ee-a174-58eef507336e · inbound
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 67ef9509-bed9-4fba-a2c6-3335ddaeeb53 · inbound
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a29ff68-8741-4f7c-9515-2653e64cea0d · inbound
Sharp spectral estimates for free boundary problems arising in plasma physics GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f5d046-cb19-44bc-ba56-f6dbee7fc1bc · inbound
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a88ddcd8-5c1b-4e75-821f-7f5efe39d2f3 · inbound
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation befac34a-ccf9-4788-82e2-266637a6f6f6 · inbound
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 50615cdb-2663-4062-b444-913c348a35bf · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 19713299-624b-4ae6-98cc-34828054b3a6 · inbound
HARNESS: Lightweight Distilled Arabic Speech Foundation Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e5a2f0f4-37b2-4d07-90fc-62420c64bad7 · inbound
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35d924f2-8145-43ef-9cf0-e82ec750015f · inbound
V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a16c3f7f-e9ca-40e2-bb3e-6f23844d1146 · inbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 03e116d0-4672-4b0f-9492-ee2972551b38 · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cc7acf4a-e9b9-45bc-9fb3-58dcb36f1766 · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cfe9b6dc-fde0-4cc0-82a5-b1a57952a8d8 · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a8cc4277-ef35-4b6e-9809-1634e7d965f4 · inbound
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ef5668f7-499d-475e-bb4d-8c27062d0c79 · inbound
Toward Native Multimodal Modeling: A Roadmap GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7b8c78ff-3fb0-44c9-8187-4d9844e2d9c2 · inbound
Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 02b22fe6-8a57-4ebc-866d-17c377016d92 · inbound
MURMUR: An Efficient Inference System for Long-Form ASR GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation df26543a-e4ae-420f-b718-dd98e6bece8f · inbound
Continuous Audio Thinking for Large Audio Language Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b621edce-4dd5-46e3-9a4f-e21b0914857d · inbound
Learning to Evade: Adaptive Attacks on Audio Watermarking GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a789fa27-f67e-4aab-abe5-b2853c6704ff · inbound
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ea645a66-611d-4f17-8f19-4fe61d3755fd · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec5937e-80d3-493e-a287-688527d91673 · inbound
Teffic-Audio: Tell Fact from Fiction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419dafd1-6ce6-42e6-b268-393854088f08 · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8863186-fa0b-4565-b00b-518392991fb4 · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a755dd70-14c7-4da9-a1b6-0ed318841856 · inbound
Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b16eb8-10f0-4ab0-b52d-b78e3a58f35a · inbound
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad93df12-96db-4d22-8d52-56e995c2a678 · inbound
Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21430d1-f07b-4d58-b5af-fc516e071a1f · inbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef19d6b0-298f-446a-838c-2426a92e094b · inbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.