Pith. sign in

REVIEW 2 cited by

Rethinking CNN Models for Audio Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.11154 v2 pith:XJONYQJT submitted 2020-07-22 cs.CV cs.SDeess.AS

classification cs.CVcs.SDeess.AS
keywords pretrainedaudiomodelsweightsaccuracyclassificationimagenetmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we show that ImageNet-Pretrained standard deep CNN models can be used as strong baseline networks for audio classification. Even though there is a significant difference between audio Spectrogram and standard ImageNet image samples, transfer learning assumptions still hold firmly. To understand what enables the ImageNet pretrained models to learn useful audio representations, we systematically study how much of pretrained weights is useful for learning spectrograms. We show (1) that for a given standard model using pretrained weights is better than using randomly initialized weights (2) qualitative results of what the CNNs learn from the spectrograms by visualizing the gradients. Besides, we show that even though we use the pretrained model weights for initialization, there is variance in performance in various output runs of the same model. This variance in performance is due to the random initialization of linear classification layer and random mini-batch orderings in multiple runs. This brings significant diversity to build stronger ensemble models with an overall improvement in accuracy. An ensemble of ImageNet pretrained DenseNet achieves 92.89% validation accuracy on the ESC-50 dataset and 87.42% validation accuracy on the UrbanSound8K dataset which is the current state-of-the-art on both of these datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Label Smoothing++: Enhanced Label Regularization for Training Neural Networks

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Label Smoothing++ learns a class-wise non-target probability distribution to replace the uniform distribution in label smoothing, yielding accuracy gains across many benchmarks.

  2. Charting 15 years of progress in deep learning for speech emotion recognition: A replication study

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Newer, larger deep learning models show no consistent gains over older architectures for speech emotion recognition across two naturalistic benchmarks, with results sensitive to model selection and hyperparameters.

Pith tools