Pith. sign in

REVIEW 2 cited by

Audiovisual Masked Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.05922 v3 pith:YNQB3XH5 submitted 2022-12-09 cs.CV cs.SD

classification cs.CVcs.SD
keywords audiovisualpretrainingdownstreamleveragemaskedstate-of-the-arttasksachieve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Can we leverage the audiovisual information already present in video to improve self-supervised representation learning? To answer this question, we study various pretraining architectures and objectives within the masked autoencoding framework, motivated by the success of similar methods in natural language and image understanding. We show that we can achieve significant improvements on audiovisual downstream classification tasks, surpassing the state-of-the-art on VGGSound and AudioSet. Furthermore, we can leverage our audiovisual pretraining scheme for multiple unimodal downstream tasks using a single audiovisual pretrained model. We additionally demonstrate the transferability of our representations, achieving state-of-the-art audiovisual results on Epic Kitchens without pretraining specifically for this dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video

    cs.CV 2025-12 unverdicted novelty 8.0 of 10

    Passive scene audio, especially direction-of-arrival cues, improves relative camera pose estimation when combined with vision in real-world videos.

  2. A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

    cs.MM 2024-11 conditional novelty 3.0 of 10

    A review that categorizes deep audio-visual correlation learning methods by architectures, objective functions, datasets, and evaluation metrics, and points to missing standardized benchmarks.

Pith tools