Pith. sign in

REVIEW 10 cited by

Temporally Aligned Audio for Video with Autoregression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.13689 v1 pith:2HQFOKEX submitted 2024-09-20 cs.CV cs.MMcs.SDeess.AS

classification cs.CVcs.MMcs.SDeess.AS
keywords v-auraalignmentrelevancesamplestemporalvisualvisualsoundaligned
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feature fusion strategy to capture fine-grained visual motion events and ensure precise temporal alignment. Additionally, we propose VisualSound, a benchmark dataset with high audio-visual relevance. VisualSound is based on VGGSound, a video dataset consisting of in-the-wild samples extracted from YouTube. During the curation, we remove samples where auditory events are not aligned with the visual ones. V-AURA outperforms current state-of-the-art models in temporal alignment and semantic relevance while maintaining comparable audio quality. Code, samples, VisualSound and models are available at https://v-aura.notion.site

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    AV-Link unifies video-to-audio and audio-to-video generation by aligning frozen diffusion-model activations with temporally matched rotary position embeddings in a shared Fusion Block.

  2. OmniAudio: Generating Spatial Audio from 360-Degree Video

    eess.AS 2025-04 conditional novelty 6.0 of 10

    OmniAudio generates First-order Ambisonics audio directly from 360-degree video using dual-branch video encoding and flow-matching pre-training, and it introduces the Sphere360 dataset and benchmark.

  3. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single flow-matching transformer trained jointly on audio-video and audio-text data produces state-of-the-art public video-to-audio synthesis with a frame-level synchronization module.

  4. Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

    cs.CV 2025-09 reject novelty 5.0 of 10

    A GPT-2 mapper over dual visual encoders claims 16% training cost and better alignment, but test-time use of true class labels makes the comparison invalid for V2A.

  5. LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    Dual lightweight adapters on a video-to-audio backbone improve long-form audio generation quality and reduce splicing artifacts, backed by a newly released clean sound-effect dataset.

  6. Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fine-tuning a video encoder with self-distillation on cropped and shifted clips makes video-to-audio generation robust to partially visible Foley targets.

  7. Video-Guided Foley Sound Generation with Multimodal Controls

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.

  8. Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks

    cs.SD 2025-05 reject novelty 4.0 of 10

    The paper proposes an open-source 10-hour video-to-piano benchmark with four-level Chain-of-Perform annotations, but supplies only preliminary, incomplete baseline results.

  9. Sound Scene Synthesis at the DCASE 2024 Challenge

    cs.AI 2025-01 conditional novelty 4.0 of 10

    Four text-to-audio systems were evaluated against a human reference in the DCASE 2024 Task 7 challenge, with a 36% quality gap and strong but small-sample FAD-to-human correlation.

  10. YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

    cs.SD 2024-12 reject novelty 4.0 of 10

    A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.

Pith tools