REVIEW 10 cited by
Temporally Aligned Audio for Video with Autoregression
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feature fusion strategy to capture fine-grained visual motion events and ensure precise temporal alignment. Additionally, we propose VisualSound, a benchmark dataset with high audio-visual relevance. VisualSound is based on VGGSound, a video dataset consisting of in-the-wild samples extracted from YouTube. During the curation, we remove samples where auditory events are not aligned with the visual ones. V-AURA outperforms current state-of-the-art models in temporal alignment and semantic relevance while maintaining comparable audio quality. Code, samples, VisualSound and models are available at https://v-aura.notion.site
Forward citations
Cited by 10 Pith papers
-
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
AV-Link unifies video-to-audio and audio-to-video generation by aligning frozen diffusion-model activations with temporally matched rotary position embeddings in a shared Fusion Block.
-
OmniAudio: Generating Spatial Audio from 360-Degree Video
OmniAudio generates First-order Ambisonics audio directly from 360-degree video using dual-branch video encoding and flow-matching pre-training, and it introduces the Sphere360 dataset and benchmark.
-
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
A single flow-matching transformer trained jointly on audio-video and audio-text data produces state-of-the-art public video-to-audio synthesis with a frame-level synchronization module.
-
Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
A GPT-2 mapper over dual visual encoders claims 16% training cost and better alignment, but test-time use of true class labels makes the comparison invalid for V2A.
-
LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
Dual lightweight adapters on a video-to-audio backbone improve long-form audio generation quality and reduce splicing artifacts, backed by a newly released clean sound-effect dataset.
-
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
Fine-tuning a video encoder with self-distillation on cropped and shifted clips makes video-to-audio generation robust to partially visible Foley targets.
-
Video-Guided Foley Sound Generation with Multimodal Controls
A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.
-
Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
The paper proposes an open-source 10-hour video-to-piano benchmark with four-level Chain-of-Perform annotations, but supplies only preliminary, incomplete baseline results.
-
Sound Scene Synthesis at the DCASE 2024 Challenge
Four text-to-audio systems were evaluated against a human reference in the DCASE 2024 Task 7 challenge, with a 36% quality gap and strong but small-sample FAD-to-human correlation.
-
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.
Discussion (0). Continue with ORCID to comment.