REVIEW 3 cited by
Video-Guided Foley Sound Generation with Multimodal Controls
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFoley, a model designed for video-guided sound generation that supports multimodal conditioning through text, audio, and video. Given a silent video and a text prompt, MultiFoley allows users to create clean sounds (e.g., skateboard wheels spinning without wind noise) or more whimsical sounds (e.g., making a lion's roar sound like a cat's meow). MultiFoley also allows users to choose reference audio from sound effects (SFX) libraries or partial videos for conditioning. A key novelty of our model lies in its joint training on both internet video datasets with low-quality audio and professional SFX recordings, enabling high-quality, full-bandwidth (48kHz) audio generation. Through automated evaluations and human studies, we demonstrate that MultiFoley successfully generates synchronized high-quality sounds across varied conditional inputs and outperforms existing methods. Please see our project page for video results: https://ificl.github.io/MultiFoley/
Forward citations
Cited by 3 Pith papers
-
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
Video-to-audio synthesis can be made incremental by subtracting an audio-conditioned prediction from the text-guided prediction, producing complementary layers without extra multi-reference training data.
-
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
A shared-backbone multi-scale diffusion transformer with motion-score conditioning generates competitive videos at reduced compute and with adjustable dynamics.
-
Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
The paper proposes an open-source 10-hour video-to-piano benchmark with four-level Chain-of-Perform annotations, but supplies only preliminary, incomplete baseline results.
Discussion (0). Continue with ORCID to comment.