Pith. sign in

REVIEW 3 cited by

Video-Guided Foley Sound Generation with Multimodal Controls

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.17698 v4 pith:OTKAL72K submitted 2024-11-26 cs.CV cs.MMcs.SDeess.AS

classification cs.CVcs.MMcs.SDeess.AS
keywords soundmultifoleyaudiovideoeffectsgenerationsoundsallows
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFoley, a model designed for video-guided sound generation that supports multimodal conditioning through text, audio, and video. Given a silent video and a text prompt, MultiFoley allows users to create clean sounds (e.g., skateboard wheels spinning without wind noise) or more whimsical sounds (e.g., making a lion's roar sound like a cat's meow). MultiFoley also allows users to choose reference audio from sound effects (SFX) libraries or partial videos for conditioning. A key novelty of our model lies in its joint training on both internet video datasets with low-quality audio and professional SFX recordings, enabling high-quality, full-bandwidth (48kHz) audio generation. Through automated evaluations and human studies, we demonstrate that MultiFoley successfully generates synchronized high-quality sounds across varied conditional inputs and outperforms existing methods. Please see our project page for video results: https://ificl.github.io/MultiFoley/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-to-audio synthesis can be made incremental by subtracting an audio-conditioned prediction from the text-guided prediction, producing complementary layers without extra multi-reference training data.

  2. Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A shared-backbone multi-scale diffusion transformer with motion-score conditioning generates competitive videos at reduced compute and with adjustable dynamics.

  3. Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks

    cs.SD 2025-05 reject novelty 4.0 of 10

    The paper proposes an open-source 10-hour video-to-piano benchmark with four-level Chain-of-Perform annotations, but supplies only preliminary, incomplete baseline results.

Pith tools