Pith. sign in

Paper Citation Record · LEDGER

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation

As of 20 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2505.22053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22053 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:58.967310Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e59564b0-faa7-4fa2-b063-1cfdf8a78654 · outbound

This paper cites Qwen2.5-VL Technical Report.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:52.288197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:52.288197Z digest=sha256:09142769ff07f5ef4184918f110710f1785cdba7bf9cb3ff4d650765e39f8f5c

Observation 64ecccab-f79b-48ec-9a8c-1a3a7fc068e1 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:05.496317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:52.354675Z digest=sha256:d40286b8498e19432bba5853ac488f460cfbf38ab81ca31d5936c2f635d6666c

Observation 8bda85d4-81a5-46aa-8b0f-46e61aeb39e2 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:05.393667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:52.455828Z digest=sha256:6f6a247c813cd136e033f99c460468836f4e1f9843ec80fdea63ab94ed2f3ed6

Observation df47d99a-1d52-4dd4-9dad-0b517ca5f2f0 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:05.212173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:52.587724Z digest=sha256:2b2b71a8b39ccb7755728765e265a646e8c24ee4b4757ea88cc2e578dd528bf5

Observation ae9c5823-7e1b-47ea-96ea-dbe1f43f8c2f · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:05.079914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:52.725127Z digest=sha256:5a95f4aba60323b1d3e06f4d85a53bc0a007a74e3c8281608cd2d92737ecc056

Observation be4f4d1f-8775-4a22-91e8-4e49ee6ee2dc · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:52.835786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:52.835786Z digest=sha256:0b3c389c3096debf7d5afdfdfe91cc6f40a79d2892516b455df41e6e7158bd0d

Observation 30f33d15-5961-462a-a3c1-2596da605e46 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:53.035618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:53.035618Z digest=sha256:704d17581cebe7d54bee4f6e8b3791b189252a537409182716f21a8eea59ca29

Observation 928f9db6-b39d-41bc-8a1f-97fbde115821 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:04.933514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:53.217158Z digest=sha256:1ed45de231716868e5d77dfa521f4b99a8e11ffa1208b6f90756aed8079aef59

Observation 0685359b-7a42-4457-841b-712cee014186 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:04.775608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:53.359002Z digest=sha256:aef1cbb0f54e719b784eb23a9f7bb9ef5e128d318f25ca018cd1f7e48a601e39

Observation a09098c8-da34-4399-b0d1-2434d0bf4d40 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:04.610186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:53.464057Z digest=sha256:aa217db2f9cb26aa1dc494e65069b350855e09de8dc31bb0c268328a2ffab4b6

Observation c2beeb1f-e6d0-4a78-b4b4-fed8e7dcf683 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:04.479355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:53.581105Z digest=sha256:5430733126b2930b42590d111727c614e088fa159f9815ce95e87d09ee123dcc

Observation 5b2fb02c-38de-4b22-823a-b4850735da3c · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:04.259543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:53.685587Z digest=sha256:73a2d65f251fa6399cc59e2881ed1f42a911c4e99132425432e1bc2ad85815e9

Observation 964a4fce-8618-41b9-907e-7504d1e34885 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:04.129945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:53.840769Z digest=sha256:01289ec0ca2cbcc6b2943528b5ea874e0ab609a5556297fa5d685c6c306d286e

Observation 78cc656d-022f-48c5-a905-22c95be6291c · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:03.965988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:54.033089Z digest=sha256:0267787df445de72b8b4a37baa2c46fdb75b845f9a4adcb919186e633f65ec91

Observation 6997051a-2caa-4e04-b942-94e18220a095 · outbound

This paper cites SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:54.170893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:54.170893Z digest=sha256:6f931cc338d1b4df74f3e2583baa0100cade1b183b6c6dc6ae8654031eebbd9b

Observation 6a208771-1651-4bbf-a981-84ba2a535a8c · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:54.288844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:54.288844Z digest=sha256:588f626cd68ceddddde804e753c80456358346a06358f66a49fe613a1b04feca

Observation 193f254f-3b9a-4d84-9195-336cfade55d6 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:03.816769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:54.478056Z digest=sha256:412acf4bbd3dccc00bdb26a32d1b8c4f93aef7c23b050e19933020bc951a9962

Observation 7425d4dc-2ea1-4ce5-9d9b-c5eae321ef72 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:54.561667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:54.561667Z digest=sha256:91d58a496e0bd560811e4b9b7765abccb38f43b547c9ee4f966923d402733311

Observation d24af8ac-82f1-484e-9ffd-09890d7b1ed3 · outbound

This paper cites Gotta Hear Them All: Towards Sound Source Aware Audio Generation.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Gotta Hear Them All: Towards Sound Source Aware Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:54.732077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:54.732077Z digest=sha256:ae320ffb618f87375085ce5927b0c9b7d616a36c4c4e7f0f38d698fd80fa8fa1

Observation 921043d0-f78f-4408-9f73-4bbae1471b78 · outbound

This paper cites Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:21:59.717273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:54.818470Z digest=sha256:b0a8279c0c8a3bdbdec70601f44b4df9968b47248a4bb2d871ad13f7b019cd40

Observation 29723086-15d4-47f6-a5e0-d48190012ec4 · outbound

This paper cites Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:54.970330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:54.970330Z digest=sha256:ee8f45def2794055d19bbd16534beaecc4c9355c539b6e44f528084e0a80530e

Observation 7d1140f2-2b3c-4537-9ac9-4a4f46771a08 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:03.692016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:55.087295Z digest=sha256:436751aa9714f8640ef2c4bc52b3d2de39a293d23724806b14849c6150545c25

Observation b8748545-5261-4b8b-8699-f94a5310b7bb · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.133752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.133752Z digest=sha256:ebf90933abe632e5a57dc9fa760c7070fcf6a55ceaec32cd38604c8be492aa75

Observation 360dcbc0-7289-4af2-b9c0-c7166dd44bbf · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:03.541309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:55.167463Z digest=sha256:511e16cd9d2ea16c06eca8ff09200ed70086b0a9209601db9cda62490161fa87

Observation 0bbc7f0e-0522-47ed-815a-f4588693dde8 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:03.375323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:55.358455Z digest=sha256:94da6b42a6e1319df5139c695875588ac582be8fe6d0aeb917bbb53f04a46bd2

Observation 1d212153-f057-4e3a-85de-b412b5b71553 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.448953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.448953Z digest=sha256:cdf90800bf45b4d78cb8a277738bfb23cbdc0092e3fdf71b58acc1b20c604a4c

Observation e2ee6235-ff7f-4d09-b72c-2083eda3acd3 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.563873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.563873Z digest=sha256:bc7e0ed25cbf46946fa55da7913b694ac30854ceeb553627e35aede7edd3794c

Observation d692022e-2746-40d5-9d11-31017794ea97 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:03.214456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:55.718021Z digest=sha256:24d925ed5097bcadeb6076e392630b06a22393bb4edd6d3d2520e953fa9090d9

Observation bd40c5e1-a6ce-43c3-b3ea-431b489a486e · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:03.043249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:55.777053Z digest=sha256:58f18fdfe8a1649834b11e2e7558eb116bfcb581c028caf36c3db378178b45e5

Observation 06fc2d77-1508-4223-baf9-b59640f87fa7 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:02.926757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:55.849731Z digest=sha256:7576b0a4f79db89fa53f706ff2eab255747475a56d1e81244d5938efc42a71c3

Observation 9d153604-3a49-4782-899a-23fbe1ff70e2 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.918085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.918085Z digest=sha256:9c2f6faaba5bdcf52449a7ce2bc1922d84c5d93046ba8baf1f0fc18bec41f2a2

Observation 9f8cc796-369b-47e0-a23d-e545379c2e1d · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:02.759310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:55.992831Z digest=sha256:946c8e887ab44c9b08645914e31c6791f27b02caec46159bc23aff45d3ac255c

Observation 3a31b0b9-a32c-4fb7-ad95-26b0b4e209d1 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:02.562454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:56.053796Z digest=sha256:cfdc772941a9f839e83c73ef45f677b735a1184cbc3f7eeb78fb3b929e6fc0ce

Observation 96db99ca-bf26-4cb7-af47-25339983acc0 · outbound

This paper cites SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:56.134180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:56.134180Z digest=sha256:db5831eb126c9194ff461a42b73890b4fb895a9a1a8f0551a4ab774438d468b9

Observation 97f33974-41b7-43ab-8725-76e5cc46a5cc · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:02.358953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:56.257027Z digest=sha256:efc54c44e2d021c3335ac1c10a3a8ef71b1bf103196960f35b493db519ff3a8b

Observation abcdffa6-0467-4f04-94e8-a8f31291579b · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:56.333453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:56.333453Z digest=sha256:3d76fdc823905cfdd558e15a38857c52f7e97eb79c5410e1ce87a5d48fbd7866

Observation 12aecfa3-8505-4034-927e-37d77f62dbac · outbound

This paper cites DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:56.399053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:56.399053Z digest=sha256:e4fbd55e3e42d4c5a1bf2aed11a13e9572c2bc71161e3f027f7deb13c608e248

Observation f1b6fffd-e5a6-45f6-9d95-7f4a8f3da74d · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:56.474901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:56.474901Z digest=sha256:f849082996a5e2ed9eec12c83495ae58d403888fc18d8d5ded013391d7d3838a

Observation 3656fd69-9d59-415e-ab82-c11e341767cf · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation ChatDev: Communicative Agents for Software Development

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:56.590613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:56.590613Z digest=sha256:8c3be8478e3890aa03ba51b046fea743ff21a6dcc3673679dafb5f6fe5b764ee

Observation 4cc161f4-586e-45ca-9711-d9f7328f422c · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:02.204467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:56.660286Z digest=sha256:20474fb2f356d2c3501ba499fec437c35d1d13592bea2ec1922d09f12492df6f

Observation 1d9045c3-606d-4115-bc9d-ce538e1310e4 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:01.795737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:56.812707Z digest=sha256:cd237e88cb576670f617c6e3b0041a4f4ce3f31e296d338fe6a2963e818adc14

Observation a6ac12b9-1b3b-4cf7-b8ca-409ada56db0f · outbound

This paper cites Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:21:59.519543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:56.900297Z digest=sha256:161760f85e29a39f7e1106141fc0f7e6f3a996d08409f434ea72f50e8f17942f

Observation f1387cbc-73fd-4394-bae2-6c5afeb6b1ed · outbound

This paper cites In IEEE International Conference on Acoustics, Speech and Signal Processing.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation In IEEE International Conference on Acoustics, Speech and Signal Processing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:01.978267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:56.747070Z digest=sha256:6ed344d3920fde18806581f025ed0b1ac3b0739f15885a5e23ef629736cdff17

Observation 326d0a29-9d5f-440f-bc76-9f74d33031b3 · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.043843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.043843Z digest=sha256:3522f47e6305ec5b6c2b24eed27c2519cb2d28814e1f35171f55b92efb3027c6

Observation 3e247ced-3b1e-4915-8d1c-826888a83921 · outbound

This paper cites VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.116700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.116700Z digest=sha256:c459af1b2ed7ebaa914155637cc64da4a982c9aaefe1f3292129a6c0232a9f87

Observation afa97fce-f1a4-4996-a476-ee774eb43c53 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:01.624561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:56.983667Z digest=sha256:33ea41050757e80008d5abb796cd8cee5c06085bdc428787cd26ae031d644221

Observation 4776d5b9-a466-4f0b-a1e8-cac4f9bcf3ce · outbound

This paper cites SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.279182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.279182Z digest=sha256:94e887dec414cc1d4dfc4782923347bd9225ad0bf7bf874d5b9aea30558eda79

Observation 4cf3f454-60bd-4915-b3f2-42f6460e0291 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:01.407498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:57.362014Z digest=sha256:1fba084ed23a37f1d909d49fd49c7b6069bc3782b6b1270f70e2d69a8f12fb3a

Observation ee1da94b-53f1-4c30-9612-d9d262826f7e · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.205800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.205800Z digest=sha256:cdcd3d9a788014db71fcb6c711ee326296b78817d03740dac188c5f5007340f4

Observation 6f5116c3-0ee9-4230-a820-f812c357912b · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.516006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.516006Z digest=sha256:d44804093ab4cf0b5c5635274bf27bb108673cbffd007fd781f6210bd9556113

Observation e230f18e-81fa-4d4d-8444-6c259f26b889 · outbound

This paper cites Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.611094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.611094Z digest=sha256:896f5d468b3f034035689cacbcddc4a3dcfda509c90751020647d08f7a3911b3

Observation 8a373711-a316-4cf0-803a-08a4eba23553 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:01.207554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:57.433467Z digest=sha256:50636d1afd50dc9a39cfb03b73a782b0ebd529ede9763cfe7d59543c787a5b1a

Observation 1de571c7-cfc9-4e41-97cf-c729a7956329 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:57.885000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:57.885000Z digest=sha256:82b1ab2f4d67599f855dbd7371afa945dd819aa49726b5eb94251a1357acff23

Observation b0917101-9ef7-4b39-a02f-bba861ffe3f6 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:00.936519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:57.958512Z digest=sha256:615f0fddbdac74129affb4075dcf9d0f7f1e86ec11bf36c82182112080eee98e

Observation 139aa635-a8ea-41e0-9d05-f85eccea2155 · outbound

This paper cites FilmComposer: LLM-Driven Music Production for Silent Film Clips.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation FilmComposer: LLM-Driven Music Production for Silent Film Clips

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:21:59.315620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:57.720909Z digest=sha256:82fa204d52fdff9afe8415e70946ceb1feb63ba43d02257757ef3a012a9893b4

Observation 69032484-0ce4-4eda-94c1-ce8dd8ca1928 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:00.537761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:58.249634Z digest=sha256:c9981a3c1842805852612c503a94727646364a9160a512e577250234f55ed875

Observation ba027057-682c-4a1c-ab32-830c7d5f8b01 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:00.340024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:58.368130Z digest=sha256:941c5cc2654d4fde9544a503b68e75ccd3904cae3fa2384a2a5457da68a57dc7

Observation 25181938-4807-45da-87c9-0fc54f7ef63a · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:00.671072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:58.120464Z digest=sha256:88cce2634da50ec291ebf5e428dd4b55687fe15a8ff486cf6b5a9c4fab6d8054

Observation b264ce73-22ed-4027-9115-481e8ed68cd8 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:58.600648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:58.600648Z digest=sha256:129476f1dd24d57d69cf15dbbb1bcd84a49fc4f7256bb87177f100dde9742e45

Observation d65f4609-06db-4fb8-aea3-36be82c8ab74 · outbound

This paper cites Long-Video Audio Synthesis with Multi-Agent Collaboration.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Long-Video Audio Synthesis with Multi-Agent Collaboration

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:58.660971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:58.660971Z digest=sha256:54ee679c59d375456469246c1d296ffba27f097b2c6b38521b1775fd20dc5c60

Observation 4e5da6fe-d576-431e-9d93-14d25401335d · outbound

This paper cites InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:58.501923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:58.501923Z digest=sha256:adfe653a22560c8d75be6e2ab25ca854a8c9079a45387381309bf95d8ca4781a

Observation 507b188f-e57e-4594-b01a-ecc653d8148a · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:21:59.932941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:58.843271Z digest=sha256:2eba81098c13402e9727bf8177b6ade98f7ab8f826a01888d418b82a69b66a54

Observation 8dfead22-de67-4f01-b641-00b7cddd2890 · outbound

This paper cites GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:58.967310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:58.967310Z digest=sha256:00bf4f73c8ee26f987c87c78d59e4939df9bcd21c966165b5e6e09102c7c788c

Observation 66786ed3-3f7d-473f-99a3-9763ff2a6414 · outbound

This paper cites an unresolved cited work.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:22:00.143804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:21:58.735968Z digest=sha256:b3e040436f66f21164f702978ff44da4c54c313975dbad8a3bd2a87d82753405

Observation f64a50f9-8283-4b54-a678-867818026ec2 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.658153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.658153Z digest=sha256:10c744a708685f82148486cb4fe3ceefa426f603360e754796093118f70a0f68

Observation bfe3256f-d477-46a4-b171-6e4a6dfa6bf7 · outbound

This paper cites From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.261294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.261294Z digest=sha256:f343b16ff26ca64266e25776c27d2d441d5a10313f58a22ea3c17f27b4a1e4e4

Pith citing papers

No inbound Pith citation observations are available.