Pith. sign in

REVIEW 4 cited by

EAT: Self-Supervised Pre-Training with Efficient Audio Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.03497 v1 pith:SHJZZE5Q submitted 2024-01-07 eess.AS cs.AIcs.CLcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.LGcs.SD
keywords audiopre-trainingself-supervisedefficientmodalitymodelsrepresentationssignificant
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress. However, the extensive computational demands during pre-training pose a significant barrier to the potential application and optimization of audio SSL models. In this paper, inspired by the success of data2vec 2.0 in image modality and Audio-MAE in audio modality, we introduce Efficient Audio Transformer (EAT) to further improve the effectiveness and efficiency in audio SSL. The proposed EAT adopts the bootstrap self-supervised training paradigm to the audio domain. A novel Utterance-Frame Objective (UFO) is designed to enhance the modeling capability of acoustic events. Furthermore, we reveal that the masking strategy is critical in audio SSL pre-training, and superior audio representations can be obtained with large inverse block masks. Experiment results demonstrate that EAT achieves state-of-the-art (SOTA) performance on a range of audio-related tasks, including AudioSet (AS-2M, AS-20K), ESC-50, and SPC-2, along with a significant pre-training speedup up to ~15x compared to existing audio SSL models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection

    eess.AS 2026-01 conditional novelty 6.0 of 10

    Better proxy-task performance does not generally improve anomalous sound detection; only source separation showed a strong, consistent positive correlation.

  2. Hidden-Domain Routing for All-Type Audio Deepfake Detection

    cs.SD 2026-08 accept novelty 5.0 of 10

    A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.

  3. EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge

    cs.SD 2026-05 unverdicted novelty 4.0 of 10

    EnvTriCascade is a tri-stage cascaded framework using mix-consistency detection followed by dual SSL-based five-class classifiers with cross-branch attention and RawBoost augmentation, achieving 0.8266 Macro-F1 on the...

  4. From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning

    eess.AS 2026-07 unverdicted novelty 3.0 of 10

    A survey that organizes audio SSL into five objective paradigms, relates their demands to architectural biases, and interprets downstream applications as tests of generalization.

Pith tools