Pith. sign in

REVIEW 6 cited by

Can Masked Autoencoders Also Listen to Birds?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.12880 v4 pith:P23XUSCG submitted 2025-04-17 cs.LG cs.SDeess.AS

Can Masked Autoencoders Also Listen to Birds?

classification cs.LG cs.SDeess.AS
keywords birdsetaudiobird-maeclassificationfine-tuningprobingprototypicalautoencoders
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Masked Autoencoders (MAEs) learn rich semantic representations in audio classification through an efficient self-supervised reconstruction task. However, general-purpose models fail to generalize well when applied directly to fine-grained audio domains. Specifically, bird-sound classification requires distinguishing subtle inter-species differences and managing high intra-species acoustic variability, revealing the performance limitations of general-domain Audio-MAEs. This work demonstrates that bridging this domain gap domain gap requires full-pipeline adaptation, not just domain-specific pretraining data. We systematically revisit and adapt the pretraining recipe, fine-tuning methods, and frozen feature utilization to bird sounds using BirdSet, a large-scale bioacoustic dataset comparable to AudioSet. Our resulting Bird-MAE achieves new state-of-the-art results in BirdSet's multi-label classification benchmark. Additionally, we introduce the parameter-efficient prototypical probing, enhancing the utility of frozen MAE representations and closely approaching fine-tuning performance in low-resource settings. Bird-MAE's prototypical probes outperform linear probing by up to 37 percentage points in mean average precision and narrow the gap to fine-tuning across BirdSet downstream tasks. Bird-MAE also demonstrates robust few-shot capabilities with prototypical probing in our newly established few-shot benchmark on BirdSet, highlighting the potential of tailored self-supervised learning pipelines for fine-grained audio domains.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MetaPerch: Learning from metadata for bioacoustics foundation models

    cs.LG 2026-07 conditional novelty 6.0

    Adding location, season, and background-species prediction as auxiliary training tasks improves bioacoustic species identification transfer across acoustic, species, and geographic domain shifts, with modest average g...

  2. A Self-Supervised Approach for Minimal-Annotation Hydroacoustic Data Exploration

    cs.SD 2026-07 conditional novelty 6.0

    Event-level MAE embeddings plus UMAP/HDBSCAN or K-Means clustering recover 15 hydroacoustic classes from multi-year Mayotte data with ~1 hour of annotation and detector-comparable F1.

  3. Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

    cs.CL 2026-05 unverdicted novelty 6.0

    Meow-Omni 1 is a quad-modal MLLM that fuses video, audio, physiological time-series, and text to achieve 71.16% accuracy on feline intent recognition in the new MeowBench benchmark.

  4. Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

    cs.SD 2026-07 conditional novelty 5.0

    For BirdCLEF+ 2026, a frozen Perch-v2 probe plus a trained HGNetV2-B0 SED net and non-bird prototype heads reach private LB 0.936, while WavTokenizer codec tokens collapse and four general audio transformers lag under...

  5. AVEX: What Matters for Animal Vocalization Encoding

    cs.SD 2025-08 unverdicted novelty 5.0

    Large empirical study finds self-supervised pre-training then supervised post-training on mixed bioacoustics and general audio data produces the strongest encoders across 26 datasets for species classification, detect...

  6. Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

    cs.SD 2026-07 conditional novelty 4.5

    On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...