Pith. sign in

REVIEW 2 cited by

STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.08601 v2 pith:CUGEH4ME submitted 2024-09-13 cs.SD cs.MMeess.AS

classification cs.SDcs.MMeess.AS
keywords semantictemporalaudioalignmentgenerationproposevideovideo-to-audio
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned Video-to-Audio (STA-V2A), an approach that enhances audio generation from videos by extracting both local temporal and global semantic video features and combining these refined video features with text as cross-modal guidance. To address the issue of information redundancy in videos, we propose an onset prediction pretext task for local temporal feature extraction and an attentive pooling module for global semantic feature extraction. To supplement the insufficient semantic information in videos, we propose a Latent Diffusion Model with Text-to-Audio priors initialization and cross-modal guidance. We also introduce Audio-Audio Align, a new metric to assess audio-temporal alignment. Subjective and objective metrics demonstrate that our method surpasses existing Video-to-Audio models in generating audio with better quality, semantic consistency, and temporal alignment. The ablation experiment validated the effectiveness of each module. Audio samples are available at https://y-ren16.github.io/STAV2A.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

    cs.CV 2025-09 reject novelty 5.0 of 10

    A GPT-2 mapper over dual visual encoders claims 16% training cost and better alignment, but test-time use of true class labels makes the comparison invalid for V2A.

  2. Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fine-tuning a video encoder with self-distillation on cropped and shifted clips makes video-to-audio generation robust to partially visible Foley targets.

Pith tools