Pith. sign in

REVIEW 1 cited by

Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.01595 v1 pith:3OALKTCZ submitted 2020-05-29 eess.AS cs.LGcs.SDeess.IVstat.ML

classification eess.AScs.LGcs.SDeess.IVstat.ML
keywords soundsaudioaudiovisualfusionmodelmodelsaudiosetlabeled
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds using both the audio and visual modalities as opposed to one or the other. We present an audiovisual fusion model that learns to recognize sounds from weakly labeled video recordings. The proposed fusion model utilizes an attention mechanism to dynamically combine the outputs of the individual audio and visual models. Experiments on the large scale sound events dataset, AudioSet, demonstrate the efficacy of the proposed model, which outperforms the single-modal models, and state-of-the-art fusion and multi-modal models. We achieve a mean Average Precision (mAP) of 46.16 on Audioset, outperforming prior state of the art by approximately +4.35 mAP (relative: 10.4%).

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition

    cs.CV 2025-07 reject novelty 5.0 of 10

    DICCAE dynamically weights a confusion loss using measured inter-class overlap and reports 65.5% audio-visual top-1 on VGGSound, but the evaluation protocol uses test data during training.

Pith tools