Pith. sign in

REVIEW 4 cited by

BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.10380 v6 pith:MBI7MWXX submitted 2024-03-15 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords classificationaudiodatasetbenchmarkbirdsetcasesevaluationlarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Deep learning (DL) has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While AudioSet is a pivotal step to bridge this gap as a universal-domain dataset, its restricted accessibility and limited range of evaluation use cases challenge its role as the sole resource. Therefore, we introduce BirdSet, a large-scale benchmark dataset for audio classification focusing on avian bioacoustics. BirdSet surpasses AudioSet with over 6,800 recording hours ($\uparrow\!17\%$) from nearly 10,000 classes ($\uparrow\!18\times$) for training and more than 400 hours ($\uparrow\!7\times$) across eight strongly labeled evaluation datasets. It serves as a versatile resource for use cases such as multi-label classification, covariate shift or self-supervised learning. We benchmark six well-known DL models in multi-label classification across three distinct training scenarios and outline further evaluation use cases in audio classification. We host our dataset on Hugging Face for easy accessibility and offer an extensive codebase to reproduce our results.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Output-space adversarial training improved clean-data performance and adversarial robustness of two bird sound classifiers across seven soundscape test sets, and stabilized prototype-based explanations.

  2. The iNaturalist Sounds Dataset

    cs.SD 2025-05 accept novelty 6.0 of 10

    A new large-scale, weakly labeled audio dataset of 230K recordings across 5,569 species, with benchmarks showing that models trained on it transfer to downstream bioacoustic classification.

  3. Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

    cs.SD 2026-07 conditional novelty 5.0 of 10

    For BirdCLEF+ 2026, a frozen Perch-v2 probe plus a trained HGNetV2-B0 SED net and non-bird prototype heads reach private LB 0.936, while WavTokenizer codec tokens collapse and four general audio transformers lag under...

  4. Distilling Spectrograms into Tokens: Fast and Lightweight Bioacoustic Classification for BirdCLEF+ 2025

    cs.SD 2025-07 conditional novelty 4.0 of 10

    The authors report that a Word2Vec-style model on K-means spectrogram tokens classifies BirdCLEF+ 2025 soundscapes in about 6 minutes, reaching a public ROC-AUC of 0.559, far below transfer-learning baselines.

Pith tools