A MeanFlow-based one-step generator with scaled classifier-free guidance synthesizes video-to-audio and text-to-audio about 2x-500x faster than prior iterative methods with comparable automatic-metric quality.
Comparison with Baselines Table 1 summarizes the performance of the proposed MF-MJT against representative VTA synthesis baselines on the VGGSound test set
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation
A MeanFlow-based one-step generator with scaled classifier-free guidance synthesizes video-to-audio and text-to-audio about 2x-500x faster than prior iterative methods with comparable automatic-metric quality.