Pith. sign in

REVIEW 2 cited by

Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09501 v1 pith:Q3KHO5XO submitted 2023-09-18 cs.CV

classification cs.CV
keywords audiosoundingvisualfeaturesobjectsinformationqueriescorrespondence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interaction are required. The previous method applies multi-frame cross-modal attention to conduct pixel-level interactions between audio features and visual features of multiple frames simultaneously, which is both redundant and implicit. In this paper, we propose an Audio-Queried Transformer architecture, AQFormer, where we define a set of object queries conditioned on audio information and associate each of them to particular sounding objects. Explicit object-level semantic correspondence between audio and visual modalities is established by gathering object information from visual features with predefined audio queries. Besides, an Audio-Bridged Temporal Interaction module is proposed to exchange sounding object-relevant information among multiple frames with the bridge of audio features. Extensive experiments are conducted on two AVS benchmarks to show that our method achieves state-of-the-art performances, especially 7.1% M_J and 7.6% M_F gains on the MS3 setting.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation

    cs.MM 2026-08 conditional novelty 6.0 of 10

    A new audio-visual instance segmentation architecture, using audio separation and an audio-modulated Mamba, reaches 48.54 mAP on AVISeg with a COCO-pretrained ResNet50.

  2. Implicit Counterfactual Learning for Audio-Visual Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Implicit text features and diffusion-based counterfactual samples improve audio-visual segmentation, achieving state-of-the-art results on AVS-Object and AVS-Semantic.

Pith tools