Pith. sign in

REVIEW 2 cited by

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06594 v2 pith:7BPN54QB submitted 2023-05-11 cs.SD cs.CVcs.LGcs.MMeess.AS

classification cs.SDcs.CVcs.LGcs.MMeess.AS
keywords musicaudiogenerationv2meowvideosignaturesvideo-to-musicvisual
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced audio codecs, the exploration of video-acoustic signatures has been confined to specific visual scenarios. In contrast, our research confronts the challenge of learning globally aligned signatures between video and music directly from paired music and videos, without explicitly modeling domain-specific rhythmic or semantic relationships. We propose V2Meow, a video-to-music generation system capable of producing high-quality music audio for a diverse range of video input types using a multi-stage autoregressive model. Trained on 5k hours of music audio clips paired with video frames mined from in-the-wild music videos, V2Meow is competitive with previous domain-specific models when evaluated in a zero-shot manner. It synthesizes high-fidelity music audio waveforms solely by conditioning on pre-trained general-purpose visual features extracted from video frames, with optional style control via text prompts. Through both qualitative and quantitative evaluations, we demonstrate that our model outperforms various existing music generation systems in terms of visual-audio correspondence and audio quality. Music samples are available at tinyurl.com/v2meow.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Let Your Video Listen to Your Music!

    cs.CV 2025-06 reject novelty 6.0 of 10

    MVAA aligns a video's motion peaks to music beats via keyframe re-timing and diffusion-based inpainting, aiming to preserve the original content while improving rhythmic synchronization.

  2. GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

    cs.SD 2025-01 conditional novelty 6.0 of 10

    GVMGen generates background music from video using spatial and temporal cross-attention to condition a MusicGen decoder, reporting state-of-the-art correspondence and diversity.

Pith tools