Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:30:28.634356Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2608.11804.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:30:28.634356Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 49513efb-d097-47fb-b871-d9a39cea64a6 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Qwen3-TTS Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f71a3a-1b2a-483b-a1d8-229a4b02c2ab · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd007df7-276f-474c-83e2-e61e28ac6a4c · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Simple and controllable music generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a10e29e7-2f4c-46bf-95c2-09b9deee3180 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ebcbf914-76d8-4740-8bf8-243e6e424550 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Plumbley
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 57b22c66-ff37-4d62-80bf-81d474c6a82c · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60677b62-7417-4406-aba3-835b57ab905a · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Uniaudio: An audio foundation model toward universal audio generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 77868152-0428-4b36-afb3-d429fd9569fa · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Uniflow-audio: Unified flow matching for audio generation from omni-modalities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cae037ee-52d5-4000-84d6-e6b0176c55a7 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 796550f7-6ab1-49e1-928b-9aa1b2377ad4 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Denoising diffusion probabilistic models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d7d8ba55-c3b2-46a0-8519-4afd5d5ea09d · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Flow Matching for Generative Modeling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36ef1c8c-d0e2-4081-bc38-5d86f77e6cbe · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Make-an-audio: Text-to-audio generation with prompt- enhanced diffusion models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1bad4996-86e6-441c-92aa-77005dadf0e2 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Plumbley
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 631d52ea-1269-42cb-9413-5192c6206da9 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaecdde2-311a-49e3-8d02-fc7b577efd91 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, Rif A
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d73faa25-f3c3-4d75-bc6d-3707a6aa2f05 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Fastspeech: Fast, robust and controllable text to speech
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b019842-1e32-4dfa-9ee5-318bd915d314 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b0c8f966-b356-460a-a49a-5603eb964d90 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8321aa47-6ff7-445e-8ea1-1c43ba56aee3 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Soundstorm: Efficient parallel audio generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b8ed3be-f1d0-4fc9-9555-fdace64326e2 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e123b37-f8cc-4fa7-836b-5a7229c2ad20 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ec0dec-25ef-4c2b-b013-10c15af19ffc · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c0e266-9abb-4532-b032-9591ba55ea7b · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching AudioX: A Unified Framework for Anything-to-Audio Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b29ad2ff-f47b-42ed-b075-6cf5e0977a17 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48aa3da9-8602-4426-b95d-9ed4dd885a1e · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Dashengtokenizer: One layer is enough for unified audio understanding and generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2f8247-8e77-4c18-abd8-2d0c12de5c9a · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Qwen3 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37be04c0-5fb6-46e9-ae4a-785c43c5bdcc · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Midashenglm: Efficient audio understanding with general audio captions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 469c3f21-0d6c-4c8f-87ba-d32002ac8415 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c811b625-a07e-4eba-a17f-496f7967e13f · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Llm.int8(): 8-bit matrix multiplication for transformers at scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 61d182ff-7b49-421d-a8f7-f79e7b32f9e0 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Acavcaps: Enabling large-scale training for fine-grained and diverse audio understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fe1dbb88-ccba-4ffc-9a7b-b43e914c5418 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Acav100m: Automatic curation of large-scale datasets for audio-visual video represen- tation learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c113222-091f-4011-b810-1a24b3d474cd · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 17f418fc-35eb-4d64-bd47-9bbe2cb43f24 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Weiss, Viet Dang, Ye Jia, Yonghui Wu, Yu Zhang, and Zhifeng Chen
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bffdba0b-d3e2-4685-960e-41d456e1b5a4 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching The LJ speech dataset
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 47a7f023-72fc-43ca-a8b0-275e440e8d0d · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching AISHELL-3: A multi-speaker mandarin TTS corpus and the baselines
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1b034a7c-808d-4d31-9264-1b997031e92e · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching WenetSpeech4TTS: A 12,800-hour mandarin TTS corpus for large-scale speech generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 61559e0b-027d-483f-b059-51f8342b1cf2 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Audiocaps: Gener- ating captions for audios in the wild
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ecd7bb0f-f42f-4a99-880a-25ba3510719f · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Funasr: A fundamental end-to-end speech recognition toolkit
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 66393fa3-c722-4d97-8490-3fe132374c71 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Robust speech recognition via large-scale weak supervision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 73cdd2f1-b259-455a-af2d-98b2b7c70361 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5052416e-0a7f-4ad4-9d34-374d2109c4d6 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed95e110-236d-41b6-8119-0b99312a354c · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation be771b3b-814f-45dd-bc59-50befd136b65 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching CLAP: Learn- ing audio concepts from natural language supervision
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec1babc5-bebc-41e9-b24d-f2cce33d0a61 · outbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching Diffusion Transformers with Representation Autoencoders
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.