Pith. sign in

Paper Citation Record · LEDGER

GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 72 inbound Pith citation observations for arXiv:2106.06909.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.06909 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 72 of 72 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:38:53.974812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T20:34:09.701797Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8a18530a-2ebd-472d-880f-03b06f2bb154 · inbound

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models cites this paper.

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:18.836309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:18.836309Z digest=sha256:dc7bb192c93d7e09c17419dcc89e575661940b45bb25f1abad7cd592b99544b8

Observation 7ab94963-c9e0-4eb6-93b1-b26583bf40eb · inbound

Whisper Finetuning on Nepali Language cites this paper.

Whisper Finetuning on Nepali Language GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:27:49.083893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:27:49.083893Z digest=sha256:e14cff6e482d0da83c67038254f5317353c32a7cc286b23b573265ba5163778d

Observation 4db6cd1f-8b9a-4688-a0af-bcaaf51a1b75 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.074793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.074793Z digest=sha256:e36c9dbe81b6c5376de315dde30aea8bef32d404b9eb53aac3cbb4e45c37a745

Observation 30b85e59-2d2d-4431-a9e5-654e150edb9e · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.012180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.012180Z digest=sha256:45cf003812eb12050e9c5a349e5e6ad7f368303fd42ec46a6c9913cfc9497467

Observation b4c048bb-6d54-4e1b-8abd-1e3a9e20b0c1 · inbound

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch cites this paper.

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:33.321155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:33.321155Z digest=sha256:6ce39f3cb45f248c2e14d2739ec8be39e9b9d3972f554de357d7e293f53a75b7

Observation c1a21545-2c5e-419e-929b-ef44887ff30c · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.931762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.931762Z digest=sha256:3b37f21b01ac1fd7f63e1476c51bae92be4a7b2ef9e2c27d7d0a330d9335b374

Observation 24beb63f-628e-4a50-ad59-36a4a5a335d0 · inbound

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues cites this paper.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.127633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.127633Z digest=sha256:f3340d1ccc69023bdfd8ef26138d28f7cffe36540f4ab7557842a95859112769

Observation c846a0c2-7942-4035-9215-f616b19fe0a9 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.322175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.322175Z digest=sha256:d81365f65167652f01e076f5f926377bede727069e7f4fbf32b7efd82e1214c3

Observation 93415f77-b6eb-4f7d-a13a-7e99a0715fb7 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.775669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:337a939cb886597bdcca5f0c927ce6978300f8518f16722e56b5b0d4d9124894

Observation 9f0b1b95-91c9-4d0e-8c6e-06143f7f07cb · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.164421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.164421Z digest=sha256:8190ec899d58f1cc92ff971e2b1d8378a03a37ee330c5c17d0024b133b3cd01f

Observation c68a647c-86b7-4515-9d19-a420d719dd5a · inbound

When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation cites this paper.

When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T19:16:37.885007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:16:37.885007Z digest=sha256:c3ce66a122d48352148e7d2f2c00a7cbf26fa9600e175b6eb0fbd7d359aece1c

Observation 3bf0e655-85f2-4d3f-9b8b-37945160f7b0 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.440103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.440103Z digest=sha256:8ed9718d8491e3ff25a69a12cd5f4bab84875baedd511297d97d99f9aec3dc89

Observation 339dcc89-68b2-4e8b-ab24-4856fa4e70ca · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.053243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.053243Z digest=sha256:69676f95a548ece2affccc0ee24e00f5ac8fb47f8cd23b3d2a42b84fafb8da92

Observation dc0f1a0c-da2c-42f6-9c52-6db19d93f9c0 · inbound

Evaluation of Deep Audio Representations for Hearables cites this paper.

Evaluation of Deep Audio Representations for Hearables GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:47:08.004692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:47:08.004692Z digest=sha256:ae790069b6e195516fb1e1bea90bbbf63f6a2527c2404b6fb2ef88dac9db0c1a

Observation 1260c168-220b-46b3-905e-6764d7681cbe · inbound

Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning cites this paper.

Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:38:53.974812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:38:53.974812Z digest=sha256:05effe06fe3968469c73e42590032c39609764603f03093b4c5e7a53eacaa3a3

Observation 42c60449-ed93-428c-8f9b-b216f18f5e12 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.160148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:e1bb70e19bf7f4ddc9ae384e0db6ff28b56060b796ce81efa79e1c7916f2195f

Observation c62cb1a9-98d8-40f0-9cf9-153aa3c703e3 · inbound

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities cites this paper.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.397615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.397615Z digest=sha256:2dde560032cfca8b2347b50746e74eb0cfdf26861ef9bf9484a9065b5a4908fa

Observation 7402db6e-24e3-42a7-a81c-af7fabe7ee52 · inbound

Inclusivity of AI Speech in Healthcare: A Decade Look Back cites this paper.

Inclusivity of AI Speech in Healthcare: A Decade Look Back GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:20:25.251918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:20:25.251918Z digest=sha256:a5b00bc8e775fd126472106c23e00f25ce0284072e327f97dce532c9c7e26c9f

Observation f1424ff3-c562-4105-8186-aa8cf416a15e · inbound

Granary: Speech Recognition and Translation Dataset in 25 European Languages cites this paper.

Granary: Speech Recognition and Translation Dataset in 25 European Languages GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:19:15.068168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:19:15.068168Z digest=sha256:dad2e66295f62816d6517e155bc7b1da14a7e35e107db8cbd2f39490d35214ab

Observation ddd816b8-bd3f-4403-b5fd-9e9f06e09840 · inbound

HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification cites this paper.

HPP-Voice: A Large-Scale Evaluation of Speech Embeddings for Multi-Phenotypic Classification GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:28.008270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:28.008270Z digest=sha256:00940eef6f8d780ec134fc047f3b9908ba106511d83ecc68a2d128cd9bb8c4ea

Observation 1ba5252f-5ee5-4872-bef6-831bfa81de8b · inbound

TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation cites this paper.

TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:04.923875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:04.923875Z digest=sha256:cbd6b6bc0cfd67123e0a3d7d15a57901042fbfbc034848d785e6963605b0239e

Observation b6c69078-b779-436a-8e7e-6dc2e647e0ad · inbound

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval cites this paper.

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.327895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.327895Z digest=sha256:1f810b9a2708864ad75228f5fbdcc3be83ab7f294d0332586cc745553854161a

Observation 48f8d9f6-776d-4b89-85bd-820eb4c85620 · inbound

Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use cites this paper.

Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:21.093321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:21.093321Z digest=sha256:29b7406d3dd8943a8c52117847fb52f08be400d2d4f55591d42706c8d079bb63

Observation b9409cea-226f-4a71-8934-78f224063019 · inbound

CASPER: A Large Scale Spontaneous Speech Dataset cites this paper.

CASPER: A Large Scale Spontaneous Speech Dataset GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.821940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.821940Z digest=sha256:8231beababc7575b6734d975813f431af4544ce453398813ac82232ad45f771d

Observation 4953a05e-06d3-4de5-8514-3180d6571e33 · inbound

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling cites this paper.

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:23.318170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:23.318170Z digest=sha256:659dfba68748e6c2bd73026cf80681aca447703ff9925630d94071a91b49e023

Observation 3d25116b-720e-4bed-b17b-d3100c804aa7 · inbound

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion cites this paper.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:41.066944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:41.066944Z digest=sha256:6965b2706b6949bac60f408cf27ea001f0867c96802a8f660d0d1fa326b02833

Observation 2c4cdf49-992a-4d22-a6ff-09981ad38bb5 · inbound

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation cites this paper.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.762687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.762687Z digest=sha256:1886c6edae386a72f876cd91c92c3c326149d399423cb863b6cd104b858adce2

Observation b1d1b093-dbc2-468e-83ea-b92226d3d12a · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.342184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.342184Z digest=sha256:54e6e253a26953904ce43ac52d3ee84e2b41805cd6ccb31d0f87d6ff53cdf5aa

Observation 3470d9a5-a1c5-459d-9c1d-a7bd6143cace · inbound

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching cites this paper.

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:52.575179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:52.575179Z digest=sha256:6fbe7c5b6d66036c80509adafa77de2da0d5b173d911bfc52c1589a8a736356b

Observation beb48705-2f65-4847-94ff-a778753b0bd1 · inbound

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs cites this paper.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.185729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.185729Z digest=sha256:f1a24d1fba30a7e738685c3e642b802640d556ce734b9833801d3c9c1f6b53f1

Observation 16578a72-185e-49e9-96be-25c9a678c37e · inbound

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit cites this paper.

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:12.222702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:12.222702Z digest=sha256:92860f1b4f2b1d80f556781e4d1f549667d1df94e223c8f9006d7d85378b8c35

Observation 2b71415f-4c51-4d89-a66a-02a530104bb8 · inbound

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models cites this paper.

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:16.997764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:16.997764Z digest=sha256:e118eb30f200c66fbb691153faa204823f6dda506d260a25d543356f9c17f02a

Observation b6883eb5-7f96-4603-9ce0-1d88626dbdbe · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:44.738231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:edccdc746ca82b9e9eb2064b1291040abf5986eacd582850c8a79f0cde0c2fed

Observation 3da493c0-b1b5-47cc-afeb-a8ec8a9fa604 · inbound

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection cites this paper.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:53.805528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:53.805528Z digest=sha256:27175c72b71a2bf18cff94f58705832d7fc1e324ef0f9d40fe60f020a920d6a4

Observation d5156ee4-8fc9-44d8-9f7d-ac6a252ac9bb · inbound

The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge cites this paper.

The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:26.487908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:26.487908Z digest=sha256:5b88fbcda8f2ebe810f2797a9e82ddcedb5a3cdfe28edb4e6c2e489e6b6786c4

Observation 0a05ab93-c5b3-4b3c-9891-049835cbe534 · inbound

Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems cites this paper.

Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:45.946423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:45.946423Z digest=sha256:5bea20757ea70b8d7602c54658baa6b4d54451a30cea0f8175f8ba27cda392db

Observation 52c13a8d-bb18-4c53-8442-e9b31c5ffe1c · inbound

VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks cites this paper.

VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:28:10.217367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:28:10.217367Z digest=sha256:2acb59595d244983683c4681a37849c88d184eb0070abd6dd73c14f7483d1e75

Observation 501501f7-ad34-4fb3-a70f-02e0dd2f7e04 · inbound

Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge cites this paper.

Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:01:03.619245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:01:03.619245Z digest=sha256:643f3375757f6446cba09b32455221dcf2457557092ec4f085173b0519a3834c

Observation 2ef66db8-5f36-4677-96d1-51de0a0f3860 · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:50.891846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:50.891846Z digest=sha256:90a243927a922e43da6f2f62057cf878e1e927ca95037db58e78637d46d38afa

Observation 59d58ead-f7e1-4f16-a889-aa6095d78d70 · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:47.691984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:47.691984Z digest=sha256:9c1b1b411bfab9d31bec853b944c0f86a4bbb8b7d3e8aa7aacf5b730b6cae0f8

Observation 3e896d81-0b83-4218-a502-514d47561471 · inbound

Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models cites this paper.

Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:39:22.341037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:39:22.341037Z digest=sha256:b6ee6dee6068063c838adc0d0881d1f67d1ec487707847ac32456c644ad61f37

Observation 45ad6b17-a9c7-4b66-a224-3486e2a55fb8 · inbound

Audio Deepfake Verification cites this paper.

Audio Deepfake Verification GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:12:07.877383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:12:07.877383Z digest=sha256:e1d63c219335f7d1385210315cdff1f82d9cb02b90ce91c81e33f58f17785ad4

Observation 3f991931-c0a6-40ee-a174-58eef507336e · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.309499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:07e5a696dc5776fbb3d4686a9e8c64dde6bd85714b11be06d382f6ecf0db6c4c

Observation 67ef9509-bed9-4fba-a2c6-3335ddaeeb53 · inbound

ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition cites this paper.

ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T12:01:01.032776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:01:01.032776Z digest=sha256:252417e603dad71981098e54a25802ec9f0dce694ae08c599be6ae26baff30fd

Observation 1a29ff68-8741-4f7c-9515-2653e64cea0d · inbound

Sharp spectral estimates for free boundary problems arising in plasma physics cites this paper.

Sharp spectral estimates for free boundary problems arising in plasma physics GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T19:56:33.793167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:56:33.793167Z digest=sha256:6ca569da8c9f68d0e979e4d3f44f54772396116056826849c69aa6a5d6827910

Observation 49f5d046-cb19-44bc-ba56-f6dbee7fc1bc · inbound

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection cites this paper.

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.985611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T20:55:09.141906Z digest=sha256:fdff4cd3dfe820437b361aeb18453b3ae3ae2ba10fd1b0d5ba0497acc0c4dd78

Observation a88ddcd8-5c1b-4e75-821f-7f5efe39d2f3 · inbound

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs cites this paper.

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:57.573353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:06:50.408402Z digest=sha256:680303daf47f7c7b0cf2faf05621580c05ae80578ef35f5125ae722ccd52312b

Observation befac34a-ccf9-4788-82e2-266637a6f6f6 · inbound

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition cites this paper.

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:07.146906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:22:08.670559Z digest=sha256:372006e5834a242f6efaa337373e94a9c33008554f93f269f36c8cfcc3547a74

Observation 50615cdb-2663-4062-b444-913c348a35bf · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.956560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:9e2045972360c79342e0910982405c57c70c6d7e9906448faf5bfae7d75150b3

Observation 19713299-624b-4ae6-98cc-34828054b3a6 · inbound

HARNESS: Lightweight Distilled Arabic Speech Foundation Models cites this paper.

HARNESS: Lightweight Distilled Arabic Speech Foundation Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:56:14.218117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T02:17:14.045133Z digest=sha256:e774c96e831b8dbac33733b833dc10f8777e1a2132b30b3b798cbc2e095da819

Observation e5a2f0f4-37b2-4d07-90fc-62420c64bad7 · inbound

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation cites this paper.

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.468559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T05:36:07.090627Z digest=sha256:9c9a39c5dfdf410e65b0c3e9dd9778d8431b54dfdea66806b688da71b2db9364

Observation 35d924f2-8145-43ef-9cf0-e82ec750015f · inbound

V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data cites this paper.

V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:14.576247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T07:46:23.931059Z digest=sha256:bbac071fa578e1b94f753f4d12bf6751e1999e942642c3ef7d0cb4eef08efcd1

Observation a16c3f7f-e9ca-40e2-bb3e-6f23844d1146 · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:46.275282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:239143030c19dd2020cdcbfa3248a28065dc4c4f6df367213c76771a8df529d8

Observation 03e116d0-4672-4b0f-9492-ee2972551b38 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.384531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:7862f15572c25571b6b7171c830345f7d9332055e1ff3cb8d7bae15a9edea560

Observation cc7acf4a-e9b9-45bc-9fb3-58dcb36f1766 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.600593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:cde34242282879563eb44cde7d5a707f3870e3278eb89f76a778766dc90ec46d

Observation cfe9b6dc-fde0-4cc0-82a5-b1a57952a8d8 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.222701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:6d1e797d33859190fa64c476110df538b9cf8559d4c100f11f205ba90436bad0

Observation a8cc4277-ef35-4b6e-9809-1634e7d965f4 · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.423216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:0fec9c08052a2081d52fc898849a7aaae616ea5640a4667531c070311c08e57d

Observation ef5668f7-499d-475e-bb4d-8c27062d0c79 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.787035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:63f604ddc6d0f1049cd85a2f3988beb4803dfcda8b0ab5ecd9851750cc310424

Observation 7b8c78ff-3fb0-44c9-8187-4d9844e2d9c2 · inbound

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation cites this paper.

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.728960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T07:30:11.718647Z digest=sha256:8234760a2df024c70f79ba580c82109fa306a0cb5a3021f408bb7ee966ea59e6

Observation 02b22fe6-8a57-4ebc-866d-17c377016d92 · inbound

MURMUR: An Efficient Inference System for Long-Form ASR cites this paper.

MURMUR: An Efficient Inference System for Long-Form ASR GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:12:25.104029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T17:06:02.978205Z digest=sha256:1f2e0e0b4031228437905cf440fab094c8709fbe898525de8354464c333c578a

Observation df26543a-e4ae-420f-b718-dd98e6bece8f · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:18.017535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:453548a51d331eff7f7e79b1d64c991b56d727c7377a04692c95a352415f1513

Observation b621edce-4dd5-46e3-9a4f-e21b0914857d · inbound

Learning to Evade: Adaptive Attacks on Audio Watermarking cites this paper.

Learning to Evade: Adaptive Attacks on Audio Watermarking GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:19:43.956144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T10:10:18.351233Z digest=sha256:d19b2dabd8528d8d1b6683ee805bbf01aee6ecb547c5cbf9d4cb123271432694

Observation a789fa27-f67e-4aab-abe5-b2853c6704ff · inbound

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions cites this paper.

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-07T20:34:09.705579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-07T20:30:53.546267Z digest=sha256:d93b02cb05824a2af72f52b676b35e1246a7a47e2c0703a3a26dff9ad5677a63

Observation ea645a66-611d-4f17-8f19-4fe61d3755fd · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.158490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.158490Z digest=sha256:63239577c5ded3013017cd847144be4d64c432e5d7d026e7b079b48a1732ffa0

Observation 4ec5937e-80d3-493e-a287-688527d91673 · inbound

Teffic-Audio: Tell Fact from Fiction cites this paper.

Teffic-Audio: Tell Fact from Fiction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T10:28:23.121360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:28:23.121360Z digest=sha256:09715710880a2cf7a13d59d9a29d0e7232ba896e2886ee8944046832f5d2a3a1

Observation 419dafd1-6ce6-42e6-b268-393854088f08 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:18.068131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:18.068131Z digest=sha256:794462d6e5016a5030c5f82d14f09eef4c503bf6bd110b23b2546e8f82b050f7

Observation b8863186-fa0b-4565-b00b-518392991fb4 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.762978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:41.762978Z digest=sha256:9624f429e37cec1d3919ae5d1ef1b3cb0da8267f7ed7cf5548dc5195bea18982

Observation a755dd70-14c7-4da9-a1b6-0ed318841856 · inbound

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing cites this paper.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.734798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.734798Z digest=sha256:779210ad5b3c1943c051841e7c63feaddacc0b35f67a81f67cf9a421f920118d

Observation d8b16eb8-10f0-4ab0-b52d-b78e3a58f35a · inbound

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis cites this paper.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.898571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.898571Z digest=sha256:6156c8a109fc2ceefcf171bb78ee3ca6dc99f3741017c0a2e774d58a2eb9f5e2

Observation ad93df12-96db-4d22-8d52-56e995c2a678 · inbound

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition cites this paper.

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:29:58.817680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:29:58.817680Z digest=sha256:222cb3f8bcb8977cbc03007445775f558aa05241d37cb49ce307572b60829f71

Observation d21430d1-f07b-4d58-b5af-fc516e071a1f · inbound

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction cites this paper.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.708566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.708566Z digest=sha256:71df826949570274fa1f5fa6b43d4ee4721eb9eded7ac3d46afa28046f072786

Observation ef19d6b0-298f-446a-838c-2426a92e094b · inbound

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization cites this paper.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.864550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.864550Z digest=sha256:10aa8bd4de2aa385bab46ccb597ca6c352991fe2545419fb83c04e518b394636