Pith. sign in

REVIEW 2 cited by

AVES: Animal Vocalization Encoder based on Self-Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.14493 v1 pith:OTFMEZCU submitted 2022-10-26 cs.SD eess.AS

classification cs.SDeess.AS
keywords audioavesmodelsanimaltasksannotatedbioacousticsclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The lack of annotated training data in bioacoustics hinders the use of large-scale neural network models trained in a supervised way. In order to leverage a large amount of unannotated audio data, we propose AVES (Animal Vocalization Encoder based on Self-Supervision), a self-supervised, transformer-based audio representation model for encoding animal vocalizations. We pretrain AVES on a diverse set of unannotated audio datasets and fine-tune them for downstream bioacoustics tasks. Comprehensive experiments with a suite of classification and detection tasks have shown that AVES outperforms all the strong baselines and even the supervised "topline" models trained on annotated audio classification datasets. The results also suggest that curating a small training subset related to downstream tasks is an efficient way to train high-quality audio representation models. We open-source our models at \url{https://github.com/earthspecies/aves}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study

    eess.AS 2025-02 conditional novelty 6.0 of 10

    SSL convolutional audio models pre-trained on speech, non-speech, or both perform nearly equally well across speech and non-speech downstream tasks, while domain-specific baselines struggle outside their domains.

  2. FinchGPT: a Transformer based language model for birdsong analysis

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A transformer language model trained on tokenized Bengalese finch songs outperforms Markov, RNN, and LSTM baselines and provides evidence for long-range syllable dependencies.

Pith tools