Pith. sign in

REVIEW 3 cited by

MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.16691 v1 pith:6JRHYYSI submitted 2022-03-30 eess.AS cs.AIcs.CLcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.LGcs.SD
keywords audiopretrainingssastencoderdownstreamfindinputmask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose a simple yet powerful improvement over the recent Self-Supervised Audio Spectrogram Transformer (SSAST) model for speech and audio classification. Specifically, we leverage the insight that the SSAST uses a very high masking ratio (75%) during pretraining, meaning that the vast majority of self-attention compute is performed on mask tokens. We address this by integrating the encoder-decoder architecture from Masked Autoencoders are Scalable Vision Learners (MAE) into the SSAST, where a deep encoder operates on only unmasked input, and a shallow decoder operates on encoder outputs and mask tokens. We find that MAE-like pretraining can provide a 3x speedup and 2x memory usage reduction over the vanilla SSAST using current audio pretraining strategies with ordinary model and input sizes. When fine-tuning on downstream tasks, which only uses the encoder, we find that our approach outperforms the SSAST on a variety of downstream tasks. We further conduct comprehensive evaluations into different strategies of pretraining and explore differences in MAE-style pretraining between the visual and audio domains.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SSLAM pre-trains audio transformers on partially mixed audio clips with a source retention loss, improving polyphonic sound tagging while keeping monophonic benchmark scores.

  2. Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A multi-view SSL framework with combined reconstruction and separation-based contrastive losses obtains disentangled pitch and instrument subspaces without the accuracy loss seen with contrastive-only training on NSynth.

  3. OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A shared-backbone transformer with pairwise modality training reports top results across 25 datasets spanning 12 modalities.

Pith tools