Pith. sign in

REVIEW 1 cited by

Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.00791 v1 pith:7FFZEWYX submitted 2024-07-17 eess.AS cs.SD

Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training

classification eess.AS cs.SD
keywords trainingaudiodatasetssecondstagetransformersatstcrnn
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This technical report describes the CP-JKU team's submission for Task 4 Sound Event Detection with Heterogeneous Training Datasets and Potentially Missing Labels of the DCASE 24 Challenge. We fine-tune three large Audio Spectrogram Transformers, PaSST, BEATs, and ATST, on the joint DESED and MAESTRO datasets in a two-stage training procedure. The first stage closely matches the baseline system setup and trains a CRNN model while keeping the large pre-trained transformer model frozen. In the second stage, both CRNN and transformer are fine-tuned using heavily weighted self-supervised losses. After the second stage, we compute strong pseudo-labels for all audio clips in the training set using an ensemble of all three fine-tuned transformers. Then, in a second iteration, we repeat the two-stage training process and include a distillation loss based on the pseudo-labels, boosting single-model performance substantially. Additionally, we pre-train PaSST and ATST on the subset of AudioSet that comes with strong temporal labels, before fine-tuning them on the Task 4 datasets.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis

    eess.AS 2025-09 unverdicted novelty 7.0

    DeepASA unifies source separation, dereverberation, SED, classification, and DoAE via object-oriented processing, chain-of-inference, and temporal coherence matching, reporting SOTA on ASA2, MC-FUSS, and STARSS23.