Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:32.585912Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2509.06389.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:32.585912Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:32.435054Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-04T23:46:32.951229Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3600fb9b-cce2-4bb8-a054-c5e367bbf70a · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 03f5461a-ad39-404e-a58c-f6d106d9c3ad · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation On top of this backbone, MeanFlow formulation is introduced which directly models the average velocity to enable native one-step generation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3fd227ec-169f-479e-8daf-41354811e4b9 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Multimodal Dataset The proposed MF-MJT is trained on multimodal datasets comprising both audio-video-text and audio-text pairs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5f345c19-dbae-41e8-8956-f0fc65fd41f9 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Comparison with Baselines Table 1 summarizes the performance of the proposed MF-MJT against representative VTA synthesis baselines on the VGGSound test set
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5f5c05e2-f170-4adf-a37d-ab118d5c469f · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a67df04d-e494-46b0-a727-aed4387b61e6 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Frieren: Efficient video- to-audio generation network with rectified flow matching,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b2172030-e743-4194-8873-e560323e78b9 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation LoV A: Long-form video- to-audio generation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 042929a9-f966-4f2d-a360-4236041ad7a7 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioLDM: Text-to-audio generation with latent diffusion models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 47b97ae1-b526-4cfb-a4a5-55c4653ba99c · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation ImageBind: One embedding space to bind them all,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84baceec-4072-4255-a76d-e439ee6b142b · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Seeing and Hearing: Open- domain visual-audio generation with diffusion latent aligners,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0dcbd26a-08fd-4884-adeb-b9529976d844 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b767ada-39d8-4d89-adaf-751159b082f7 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation TA-V2A: Textually assisted video- to-audio generation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 41ae4ce0-e219-4ab1-9cbe-e6fbf2058de3 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MMAudio: Tam- ing multimodal joint training for high-quality video-to-audio synthesis,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 30f5277f-ba65-4eef-bdb6-289bf375842d · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Kling-Foley: Multimodal diffusion transformer for high-quality video-to-audio genera- tion,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e0faab3-cb83-47e5-9dbc-ff112b26a7bd · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Denoising diffusion probabilis- tic models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2bb2793f-c683-49cd-bbe6-6032eff702ec · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Flow straight and fast: Learn- ing to generate and transfer data with rectified flow,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a9688b03-26c1-46ab-b37e-091350255898 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Instaflow: One step is enough for high-quality diffusion-based text-to-image generation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 896d1435-b761-40ad-a5be-728234741047 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Mean Flows for One-step Generative Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1debeb-1c45-4715-bfe5-e3a9505e361b · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Classifier-Free Diffusion Guidance
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 992acc49-5a1e-412d-86bd-554aa35a69ab · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d78d1673-8707-46fe-967c-4e491df61e79 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Scaling rectified flow transformers for high-resolution image synthesis,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 56f499bf-339a-48f6-a196-eb479603b0a3 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Scalable diffusion models with trans- formers,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a1303cec-94e9-40a7-8d4f-5d6781368649 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Learning transfer- able visual models from natural language supervision,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5079f05a-0c9a-4aeb-b6fc-e0f069be909e · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation A versatile diffusion transformer with mixture of noise levels for audiovisual gen- eration,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 358179d5-ffae-4870-b6d1-f346144c9806 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Synchformer: Efficient syn- chronization from sparse cues,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 72ba48a8-6683-4bfb-93da-3286d88a10e2 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Score-based generative modeling through stochastic differential equations,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f9c9b6c0-aecb-4e90-bd51-744f0d1fa999 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation VGGSound: A large-scale audio-visual dataset,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c7ca1149-03fd-48b4-b82f-a105d0d9e8b4 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioCaps: Generating cap- tions for audios in the wild,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba099949-7db9-43f2-9403-2a65f35fabf5 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation WavCaps: A ChatGPT- assisted weakly-labelled audio captioning dataset for audio- language multimodal research,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 11b40170-8da9-4314-948a-d1aadc9f118e · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Decoupled Weight Decay Regularization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42521d5-2ffb-4f22-a5c0-9e33a9c06816 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Audio Set: An ontology and human-labeled dataset for audio events,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b4f63100-b62e-409a-9b3b-bb3b38fa46af · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b1ebeab0-251d-4093-bf90-e8ce74829df5 · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Efficient train- ing of audio transformers with patchout,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d9a819d7-f750-4095-bbfc-7dc47566a11d · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation CLAP: Learn- ing audio concepts from natural language supervision,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a7287d51-c426-474d-a050-35f35fa93fdf · outbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioLCM: Efficient and high-quality text-to-audio generation with minimal inference steps,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3600fb9b-cce2-4bb8-a054-c5e367bbf70a · inbound
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.