Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2305.11013.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T06:05:11.464027Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T11:27:03.142196Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3ad2844e-c9f5-4714-bff0-1a4b7959d5fd · inbound
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 029a82d4-079b-478f-abb3-40d13ae9a8f8 · inbound
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e4641d32-5c61-4246-ac0a-482b5848c79b · inbound
Qwen2-Audio Technical Report FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f19d96cf-cee9-4b9d-a175-f7cc473c622b · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1abb4ed1-d7bc-4696-a293-e4bab6d300a5 · inbound
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02405732-ba93-47b7-b6f4-96fc2629f5f4 · inbound
Real-Time Textless Dialogue Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b04887-867e-437a-b4fc-f3046be5db61 · inbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fb6bf2-234e-4f52-8519-7e19cf27c6fb · inbound
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629376f1-90e3-4113-994c-f94bc4b53cd7 · inbound
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 884fbf92-8d19-4944-975d-4e40dafb05b1 · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 70a0d441-835b-48cb-9407-933b084066dc · inbound
Ming-Omni: A Unified Multimodal Model for Perception and Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b16968f-f463-4b34-8c15-fb9c2251b1bb · inbound
RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf1f8fd-29cf-44d3-b2d5-a4c8a7b6e62e · inbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38cbefd-6d7e-4e97-b19d-251648fe0990 · inbound
SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1f3415-c111-4879-a6d0-e4066fe59202 · inbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af261e11-cd76-4626-a8df-d031500f5e34 · inbound
Inference-time Scaling for Diffusion-based Audio Super-resolution FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b679bf31-2d50-4682-94bf-78ffa0bfacc2 · inbound
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7dee25-57bb-44b1-8618-6dca5ab4f017 · inbound
OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2303f6b9-62c9-446c-981c-2b260579df84 · inbound
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 76e7137b-e86d-491d-abc7-8aeff764eece · inbound
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233cc8ab-de9b-4889-88be-9f91a35479ce · inbound
AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c9b3087a-ded2-474f-a585-9c328d49f5ae · inbound
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d235f828-8034-4223-a655-4712c715573e · inbound
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6ec8d7f1-8f17-47db-9598-021d3bb1e87a · inbound
Audio Editing in the Era of Foundation Models: A Survey FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 914d7818-2ee7-424b-8676-1cf91c95e82c · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f6548111-d76a-43b9-a2f2-93c09fac96e8 · inbound
Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ce05492e-c9bd-43a6-9817-f50342cdc1cf · inbound
Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ab0643-701f-4c6b-9525-15fb7ab44527 · inbound
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee9a78e7-0c36-4084-b44b-9198b2f3fd84 · inbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.