Pith. sign in

REVIEW 4 cited by

Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.06135 v1 pith:WABEMXTT submitted 2024-09-10 cs.SD cs.CVcs.MMeess.AS

classification cs.SDcs.CVcs.MMeess.AS
keywords audiovideoloudnessdrawsynthesisvideo-to-audiochallengesconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley task, presents inherent challenges related to audio-visual synchronization. These challenges encompass maintaining the content consistency between the input video and the generated audio, as well as the alignment of temporal and loudness properties within the video. To address these issues, we construct a controllable video-to-audio synthesis model, termed Draw an Audio, which supports multiple input instructions through drawn masks and loudness signals. To ensure content consistency between the synthesized audio and target video, we introduce the Mask-Attention Module (MAM), which employs masked video instruction to enable the model to focus on regions of interest. Additionally, we implement the Time-Loudness Module (TLM), which uses an auxiliary loudness signal to ensure the synthesis of sound that aligns with the video in both loudness and temporal dimensions. Furthermore, we have extended a large-scale V2A dataset, named VGGSound-Caption, by annotating caption prompts. Extensive experiments on challenging benchmarks across two large-scale V2A datasets verify Draw an Audio achieves the state-of-the-art. Project page: https://yannqi.github.io/Draw-an-Audio/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    AV-Link unifies video-to-audio and audio-to-video generation by aligning frozen diffusion-model activations with temporally matched rotary position embeddings in a shared Fusion Block.

  2. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single flow-matching transformer trained jointly on audio-video and audio-text data produces state-of-the-art public video-to-audio synthesis with a frame-level synchronization module.

  3. Gotta Hear Them All: Towards Sound Source Aware Audio Generation

    cs.MM 2024-11 conditional novelty 6.0 of 10

    A sound-source-aware image-to-audio generator that detects objects, disambiguates their audio semantics in a learned cross-modal manifold, and mixes them to synthesize audio.

  4. YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

    cs.SD 2024-12 reject novelty 4.0 of 10

    A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.

Pith tools