Pith. sign in

REVIEW 1 cited by

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11729 v1 pith:GEW2PGJT submitted 2024-09-18 cs.MM cs.CVcs.SDeess.AS

classification cs.MMcs.CVcs.SDeess.AS
keywords audio-visualobjectlearningmethodrepresentationanimalscategoriesclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current audio-visual representation learning can capture rough object categories (e.g., ``animals'' and ``instruments''), but it lacks the ability to recognize fine-grained details, such as specific categories like ``dogs'' and ``flutes'' within animals and instruments. To address this issue, we introduce DETECLAP, a method to enhance audio-visual representation learning with object information. Our key idea is to introduce an audio-visual label prediction loss to the existing Contrastive Audio-Visual Masked AutoEncoder to enhance its object awareness. To avoid costly manual annotations, we prepare object labels from both audio and visual inputs using state-of-the-art language-audio models and object detectors. We evaluate the method of audio-visual retrieval and classification using the VGGSound and AudioSet20K datasets. Our method achieves improvements in recall@10 of +1.5% and +1.2% for audio-to-visual and visual-to-audio retrieval, respectively, and an improvement in accuracy of +0.6% for audio-visual classification.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LG-CAV-MAE trains a tri-modal audio-visual-text model on CLAP-filtered auto-generated captions and beats existing audio-visual masked autoencoders on retrieval and classification.

Pith tools