REVIEW 2 cited by
Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interaction are required. The previous method applies multi-frame cross-modal attention to conduct pixel-level interactions between audio features and visual features of multiple frames simultaneously, which is both redundant and implicit. In this paper, we propose an Audio-Queried Transformer architecture, AQFormer, where we define a set of object queries conditioned on audio information and associate each of them to particular sounding objects. Explicit object-level semantic correspondence between audio and visual modalities is established by gathering object information from visual features with predefined audio queries. Besides, an Audio-Bridged Temporal Interaction module is proposed to exchange sounding object-relevant information among multiple frames with the bridge of audio features. Extensive experiments are conducted on two AVS benchmarks to show that our method achieves state-of-the-art performances, especially 7.1% M_J and 7.6% M_F gains on the MS3 setting.
Forward citations
Cited by 2 Pith papers
-
Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation
A new audio-visual instance segmentation architecture, using audio separation and an audio-modulated Mamba, reaches 48.54 mAP on AVISeg with a COCO-pretrained ResNet50.
-
Implicit Counterfactual Learning for Audio-Visual Segmentation
Implicit text features and diffusion-based counterfactual samples improve audio-visual segmentation, achieving state-of-the-art results on AVS-Object and AVS-Semantic.
Discussion (0). Sign in to comment.