Pith. sign in

REVIEW 6 cited by

Contrastive Audio-Visual Masked Autoencoder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07839 v4 pith:FTPAGX2H submitted 2022-10-02 cs.MM cs.CVcs.SDeess.AS

classification cs.MMcs.CVcs.SDeess.AS
keywords audio-visualcontrastivemaskedmodelcav-maelearningpretrainedauto-encoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE) by combining contrastive learning and masked data modeling, two major self-supervised learning frameworks, to learn a joint and coordinated audio-visual representation. Our experiments show that the contrastive audio-visual correspondence learning objective not only enables the model to perform audio-visual retrieval tasks, but also helps the model learn a better joint representation. As a result, our fully self-supervised pretrained CAV-MAE achieves a new SOTA accuracy of 65.9% on VGGSound, and is comparable with the previous best supervised pretrained model on AudioSet in the audio-visual event classification task. Code and pretrained models are at https://github.com/yuangongnd/cav-mae.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio-Visual Continual Test-Time Adaptation without Forgetting

    cs.LG 2026-02 conditional novelty 6.0 of 10

    By adapting only the fusion layer and retrieving past good parameter states via raw input statistics, AV-CTTA outperforms existing audio-visual continual test-time adaptation methods and forgets far less.

  2. Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

    eess.AS 2025-11 conditional novelty 6.0 of 10

    A 10.7M-pair audio-caption corpus and systematic comparison show contrastive pretraining is more data-efficient while captioning scales better, and supervised initialization yields diminishing returns.

  3. Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LG-CAV-MAE trains a tri-modal audio-visual-text model on CLAP-filtered auto-generated captions and beats existing audio-visual masked autoencoders on retrieval and classification.

  4. DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

    cs.MM 2025-07 conditional novelty 6.0 of 10

    A single multimodal language model can generate intelligible speech and synchronized background audio jointly from a silent video, transcript, and reference voice, outperforming separately concatenated speech and audi...

  5. Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new interactive method lets users click on an object in a video and generates audio for just that object, using mask-conditioned contrastive fine-tuning and latent diffusion.

  6. Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition

    cs.CV 2025-07 reject novelty 5.0 of 10

    DICCAE dynamically weights a confusion loss using measured inter-class overlap and reports 65.5% audio-visual top-1 on VGGSound, but the evaluation protocol uses test data during training.

Pith tools