Pith. sign in

REVIEW 2 cited by

Device-Robust Acoustic Scene Classification Based on Two-Stage Categorization and Data Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.08389 v2 pith:H26RUG7S submitted 2020-07-16 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords taskdataclassestwo-stageaccuracyacousticaugmentationclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this technical report, we present a joint effort of four groups, namely GT, USTC, Tencent, and UKE, to tackle Task 1 - Acoustic Scene Classification (ASC) in the DCASE 2020 Challenge. Task 1 comprises two different sub-tasks: (i) Task 1a focuses on ASC of audio signals recorded with multiple (real and simulated) devices into ten different fine-grained classes, and (ii) Task 1b concerns with classification of data into three higher-level classes using low-complexity solutions. For Task 1a, we propose a novel two-stage ASC system leveraging upon ad-hoc score combination of two convolutional neural networks (CNNs), classifying the acoustic input according to three classes, and then ten classes, respectively. Four different CNN-based architectures are explored to implement the two-stage classifiers, and several data augmentation techniques are also investigated. For Task 1b, we leverage upon a quantization method to reduce the complexity of two of our top-accuracy three-classes CNN-based architectures. On Task 1a development data set, an ASC accuracy of 76.9\% is attained using our best single classifier and data augmentation. An accuracy of 81.9\% is then attained by a final model fusion of our two-stage ASC classifiers. On Task 1b development data set, we achieve an accuracy of 96.7\% with a model size smaller than 500KB. Code is available: https://github.com/MihawkHu/DCASE2020_task1.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer

    eess.AS 2025-01 conditional novelty 5.0 of 10

    A variational Bayesian method adapts deep acoustic models by estimating distributions over hidden features, with a Gaussian mean-field variant for parallel data and an empirical Bayes variant for non-parallel data, an...

  2. Improving Acoustic Scene Classification in Low-Resource Conditions

    eess.AS 2024-12 conditional novelty 4.0 of 10

    DS-FlexiNet achieves 58.25% accuracy after int8 quantization on TAU22 Task 1A with 30.69K parameters and 8.27M MACs, using residual normalization, ADIR augmentation, and 12-teacher knowledge distillation.

Pith tools