Pith. sign in

Paper Citation Record · LEDGER

FunASR: A Fundamental End-to-End Speech Recognition Toolkit

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2305.11013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.11013 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:05:11.464027Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T11:27:03.142196Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3ad2844e-c9f5-4714-bff0-1a4b7959d5fd · inbound

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models cites this paper.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.775426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:60703d70341543242c6309ce723b9ee5782d916df4f39373b0c4a88cf22dcda4

Observation 029a82d4-079b-478f-abb3-40d13ae9a8f8 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.427555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:20ef23fd13e5dc7042a99bfba1a2e5f58e6fdc08f411bff0fe27462cd22c3ad8

Observation e4641d32-5c61-4246-ac0a-482b5848c79b · inbound

Qwen2-Audio Technical Report cites this paper.

Qwen2-Audio Technical Report FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:14:45.481814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T02:14:45.371564Z digest=sha256:de3d1e8ba1141708ec26a61c042d770fde6494e9854cb0f3eb0eff06ad5cb695

Observation f19d96cf-cee9-4b9d-a175-f7cc473c622b · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.549207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:0cf8dfe28752cafb324fef2964e7466e0e8d6b57eb9a818a792ffeeb6b4b69e0

Observation 1abb4ed1-d7bc-4696-a293-e4bab6d300a5 · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.354349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.354349Z digest=sha256:6a122b6ee818bdcfad4af2c0e8cde3fbd1876db4dd31b7a7dc05e08b928bb800

Observation 02405732-ba93-47b7-b6f4-96fc2629f5f4 · inbound

Real-Time Textless Dialogue Generation cites this paper.

Real-Time Textless Dialogue Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:27:45.551451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:27:45.551451Z digest=sha256:7df397c34e855e582877721a683dc94f8aecbef4c3455b6fbf9ff4fdca4064fe

Observation e4b04887-867e-437a-b4fc-f3046be5db61 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.511185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.511185Z digest=sha256:54c969802a7d84a70f92d024f16bc0ec16347ddc58aef4d9c0c68d7b2a6b4cb3

Observation 42fb6bf2-234e-4f52-8519-7e19cf27c6fb · inbound

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget cites this paper.

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T06:05:11.464027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:05:11.464027Z digest=sha256:a77f5a034023ceb56d1aa4f272ee6b939a9215f5f916e215a20a33f7a274fbad

Observation 629376f1-90e3-4113-994c-f94bc4b53cd7 · inbound

A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model cites this paper.

A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:08.470461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:08.470461Z digest=sha256:48e511a46744009f5121282ce103d48947e821f9df619315533fe222e759980f

Observation 884fbf92-8d19-4944-975d-4e40dafb05b1 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.458153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:60c9e6e3c931701e759c1ed98e483be27a4816ab6b41e33dad38e1de266ef58d

Observation 70a0d441-835b-48cb-9407-933b084066dc · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.617554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.617554Z digest=sha256:1fd9dae28b31f1e020011c71d142e59d65519b2c6b5ffe2b3ba1cec815a7b354

Observation 5b16968f-f463-4b34-8c15-fb9c2251b1bb · inbound

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture cites this paper.

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.728503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.728503Z digest=sha256:cdff422de952652915f0bbb4c5893802691329a6cb7f5f543bbf1a397259fbc7

Observation 5bf1f8fd-29cf-44d3-b2d5-a4c8a7b6e62e · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:23.415106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:23.415106Z digest=sha256:41a0b581f8446a3542cbdeeee76ef9d4fe13c2910448dd634d54f1d3bc67780a

Observation d38cbefd-6d7e-4e97-b19d-251648fe0990 · inbound

SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling cites this paper.

SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:26.953516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:26.953516Z digest=sha256:c42cc520834e7dbeed327358a3061dc2b1ea519186c634971a30589ba51a34f5

Observation ac1f3415-c111-4879-a6d0-e4066fe59202 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.876089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.876089Z digest=sha256:0f0a81cb34b6cca6a6b6b754dd078f74a554358684051bcc318646fa14ef65e4

Observation af261e11-cd76-4626-a8df-d031500f5e34 · inbound

Inference-time Scaling for Diffusion-based Audio Super-resolution cites this paper.

Inference-time Scaling for Diffusion-based Audio Super-resolution FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:26.342534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:26.342534Z digest=sha256:b210dcb96cbe9419612f95fd4cd0a8e2925850433585f9bc2c021f0d5e0ad868

Observation b679bf31-2d50-4682-94bf-78ffa0bfacc2 · inbound

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations cites this paper.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.009070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.009070Z digest=sha256:12cc85754ac49812c5b001104baa73dbde91fa1dc17167869b0318446732bcc8

Observation 3e7dee25-57bb-44b1-8618-6dca5ab4f017 · inbound

OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection cites this paper.

OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:21.154166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:21.154166Z digest=sha256:7390868ffc3769316aa50abfafdafffbeb6c400f5cc239dfac4b4e7f2ee9b4c7

Observation 2303f6b9-62c9-446c-981c-2b260579df84 · inbound

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering cites this paper.

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:45:24.300261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T22:44:49.949759Z digest=sha256:321d4147ef592698c5464fa188b21ff8505e3c6d9b35770bdb6ca20f5d6e9625

Observation 76e7137b-e86d-491d-abc7-8aeff764eece · inbound

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models cites this paper.

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:58.971096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:58.971096Z digest=sha256:da7ebb6a0f0d281aad80fc02bb78ca51477e6b092f658210af4ab8c31d5921a6

Observation 233cc8ab-de9b-4889-88be-9f91a35479ce · inbound

AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models cites this paper.

AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.273904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:51:29.385520Z digest=sha256:aecaa5ce688dfe70ab2ae747a7e31cdc6e4d6e0c3a9c29389c1de16e753dafa2

Observation c9b3087a-ded2-474f-a585-9c328d49f5ae · inbound

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models cites this paper.

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:56:05.281227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T04:54:35.815868Z digest=sha256:5d211b5d167d34c90d1eceb60d059827da56199b352d89742def9229919fcae4

Observation d235f828-8034-4223-a655-4712c715573e · inbound

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing cites this paper.

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:06:23.860879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T12:54:20.815371Z digest=sha256:a7ff5f4b6f77c0d9a3fdb70204254907f1d45f55312fb911e650a8beffefbda9

Observation 6ec8d7f1-8f17-47db-9598-021d3bb1e87a · inbound

Audio Editing in the Era of Foundation Models: A Survey cites this paper.

Audio Editing in the Era of Foundation Models: A Survey FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:49.451429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:8ef765bb7e1399c0ac58675b0c51609e84f9b48f2d7cdadc902612f4f1a41623

Observation 914d7818-2ee7-424b-8676-1cf91c95e82c · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.188685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:839aace230d9095f8ccd90785e32e89ca168636e25766b0470e999f0526864f0

Observation f6548111-d76a-43b9-a2f2-93c09fac96e8 · inbound

Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech cites this paper.

Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:27:03.143835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T11:18:43.887987Z digest=sha256:eb33f28e037ddda2bad7287e805a2ce59d840ba56e8c7b7c9224a4fc8a878d1d

Observation ce05492e-c9bd-43a6-9817-f50342cdc1cf · inbound

Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech cites this paper.

Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T07:58:45.439616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:58:45.439616Z digest=sha256:29e81fb15be69f73c900be1e83c7d3974231f90a942d734fb2e75768f5a378b8

Observation a1ab0643-701f-4c6b-9525-15fb7ab44527 · inbound

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm cites this paper.

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:23.937463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:23.937463Z digest=sha256:669820466260d1f9f2a17a0d7d813bd150c7bdbd57fd919558013ed2b34ebaca

Observation ee9a78e7-0c36-4084-b44b-9198b2f3fd84 · inbound

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization cites this paper.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.911143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.911143Z digest=sha256:f6ddb2c96f89061117e4e76daa9929e12f496b4f918acd66c70a29d0f472b0ca