Pith. sign in

REVIEW 1 cited by

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09168 v1 pith:BS33F62Z submitted 2024-12-12 cs.SD cs.CVcs.MMeess.AS

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

classification cs.SD cs.CVcs.MMeess.AS
keywords soundyingsoundgenerationaudioeffectsfew-shothigh-qualitymodule
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge of limited labeled data in real-world scenes, we introduce YingSound, a foundation model designed for video-guided sound generation that supports high-quality audio generation in few-shot settings. Specifically, YingSound consists of two major modules. The first module uses a conditional flow matching transformer to achieve effective semantic alignment in sound generation across audio and visual modalities. This module aims to build a learnable audio-visual aggregator (AVA) that integrates high-resolution visual features with corresponding audio features at multiple stages. The second module is developed with a proposed multi-modal visual-audio chain-of-thought (CoT) approach to generate finer sound effects in few-shot settings. Finally, an industry-standard video-to-audio (V2A) dataset that encompasses various real-world scenarios is presented. We show that YingSound effectively generates high-quality synchronized sounds across diverse conditional inputs through automated evaluations and human studies. Project Page: \url{https://giantailab.github.io/yingsound/}

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

    cs.CV 2025-12 conditional novelty 6.0

    The paper introduces an event-level hierarchical control benchmark and a training-free agentic pipeline for video-to-audio generation, claiming 40.7% better controllability and 12.5% better perceptual quality than pri...