Pith. sign in

REVIEW 1 cited by

Receptive-field-regularized CNN variants for acoustic scene classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.02859 v1 pith:4NAAKS2F submitted 2019-09-05 eess.AS cs.LGcs.SDstat.ML

Receptive-field-regularized CNN variants for acoustic scene classification

classification eess.AS cs.LGcs.SDstat.ML
keywords cnnsfrequencytasksacousticaudioclassificationdcasedifferent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Acoustic scene classification and related tasks have been dominated by Convolutional Neural Networks (CNNs). Top-performing CNNs use mainly audio spectograms as input and borrow their architectural design primarily from computer vision. A recent study has shown that restricting the receptive field (RF) of CNNs in appropriate ways is crucial for their performance, robustness and generalization in audio tasks. One side effect of restricting the RF of CNNs is that more frequency information is lost. In this paper, we perform a systematic investigation of different RF configuration for various CNN architectures on the DCASE 2019 Task 1.A dataset. Second, we introduce Frequency Aware CNNs to compensate for the lack of frequency information caused by the restricted RF, and experimentally determine if and in what RF ranges they yield additional improvement. The result of these investigations are several well-performing submissions to different tasks in the DCASE 2019 Challenge.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. C2GA: A Class-Controllable Generative Augmentation Framework for Respiratory Sound Classification

    cs.SD 2026-06 unverdicted novelty 4.0

    C2GA uses conditional VQ-VAE with decoupled local tokens and global class prototypes plus a Transformer prior to generate high-fidelity label-consistent Mel-spectrograms for respiratory sound data augmentation.