Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:58.967310Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2505.22053.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:58.967310Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e59564b0-faa7-4fa2-b063-1cfdf8a78654 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64ecccab-f79b-48ec-9a8c-1a3a7fc068e1 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8bda85d4-81a5-46aa-8b0f-46e61aeb39e2 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation df47d99a-1d52-4dd4-9dad-0b517ca5f2f0 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ae9c5823-7e1b-47ea-96ea-dbe1f43f8c2f · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation be4f4d1f-8775-4a22-91e8-4e49ee6ee2dc · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f33d15-5961-462a-a3c1-2596da605e46 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 928f9db6-b39d-41bc-8a1f-97fbde115821 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0685359b-7a42-4457-841b-712cee014186 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a09098c8-da34-4399-b0d1-2434d0bf4d40 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c2beeb1f-e6d0-4a78-b4b4-fed8e7dcf683 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5b2fb02c-38de-4b22-823a-b4850735da3c · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 964a4fce-8618-41b9-907e-7504d1e34885 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 78cc656d-022f-48c5-a905-22c95be6291c · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6997051a-2caa-4e04-b942-94e18220a095 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a208771-1651-4bbf-a981-84ba2a535a8c · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 193f254f-3b9a-4d84-9195-336cfade55d6 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7425d4dc-2ea1-4ce5-9d9b-c5eae321ef72 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d24af8ac-82f1-484e-9ffd-09890d7b1ed3 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Gotta Hear Them All: Towards Sound Source Aware Audio Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 921043d0-f78f-4408-9f73-4bbae1471b78 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 29723086-15d4-47f6-a5e0-d48190012ec4 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1140f2-2b3c-4537-9ac9-4a4f46771a08 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b8748545-5261-4b8b-8699-f94a5310b7bb · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360dcbc0-7289-4af2-b9c0-c7166dd44bbf · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0bbc7f0e-0522-47ed-815a-f4588693dde8 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1d212153-f057-4e3a-85de-b412b5b71553 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2ee6235-ff7f-4d09-b72c-2083eda3acd3 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d692022e-2746-40d5-9d11-31017794ea97 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bd40c5e1-a6ce-43c3-b3ea-431b489a486e · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 06fc2d77-1508-4223-baf9-b59640f87fa7 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9d153604-3a49-4782-899a-23fbe1ff70e2 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8cc796-369b-47e0-a23d-e545379c2e1d · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3a31b0b9-a32c-4fb7-ad95-26b0b4e209d1 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 96db99ca-bf26-4cb7-af47-25339983acc0 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f33974-41b7-43ab-8725-76e5cc46a5cc · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation abcdffa6-0467-4f04-94e8-a8f31291579b · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Mustango: Toward Controllable Text-to-Music Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12aecfa3-8505-4034-927e-37d77f62dbac · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b6fffd-e5a6-45f6-9d95-7f4a8f3da74d · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3656fd69-9d59-415e-ab82-c11e341767cf · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation ChatDev: Communicative Agents for Software Development
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc161f4-586e-45ca-9711-d9f7328f422c · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1d9045c3-606d-4115-bc9d-ce538e1310e4 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a6ac12b9-1b3b-4cf7-b8ca-409ada56db0f · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f1387cbc-73fd-4394-bae2-6c5afeb6b1ed · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation In IEEE International Conference on Acoustics, Speech and Signal Processing
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 326d0a29-9d5f-440f-bc76-9f74d33031b3 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation AudioX: A Unified Framework for Anything-to-Audio Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e247ced-3b1e-4915-8d1c-826888a83921 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa97fce-f1a4-4996-a476-ee774eb43c53 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4776d5b9-a466-4f0b-a1e8-cac4f9bcf3ce · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf3f454-60bd-4915-b3f2-42f6460e0291 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ee1da94b-53f1-4c30-9612-d9d262826f7e · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f5116c3-0ee9-4230-a820-f812c357912b · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e230f18e-81fa-4d4d-8444-6c259f26b889 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a373711-a316-4cf0-803a-08a4eba23553 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1de571c7-cfc9-4e41-97cf-c729a7956329 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0917101-9ef7-4b39-a02f-bba861ffe3f6 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 139aa635-a8ea-41e0-9d05-f85eccea2155 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation FilmComposer: LLM-Driven Music Production for Silent Film Clips
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 69032484-0ce4-4eda-94c1-ce8dd8ca1928 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba027057-682c-4a1c-ab32-830c7d5f8b01 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 25181938-4807-45da-87c9-0fc54f7ef63a · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b264ce73-22ed-4027-9115-481e8ed68cd8 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d65f4609-06db-4fb8-aea3-36be82c8ab74 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Long-Video Audio Synthesis with Multi-Agent Collaboration
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e5da6fe-d576-431e-9d93-14d25401335d · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 507b188f-e57e-4594-b01a-ecc653d8148a · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8dfead22-de67-4f01-b641-00b7cddd2890 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66786ed3-3f7d-473f-99a3-9763ff2a6414 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f64a50f9-8283-4b54-a678-867818026ec2 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe3256f-d477-46a4-b171-6e4a6dfa6bf7 · outbound
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.