Multi-stream Network With Temporal Attention For Environmental Sound Classification

Katrin Kirchhoff; Venkata Chebiyyam; Xinyu Li

arxiv: 1901.08608 · v1 · pith:OPLHYZXEnew · submitted 2019-01-24 · 💻 cs.SD · cs.MM· eess.AS

Multi-stream Network With Temporal Attention For Environmental Sound Classification

Xinyu Li , Venkata Chebiyyam , Katrin Kirchhoff This is my paper

classification 💻 cs.SD cs.MMeess.AS

keywords classificationnetworksoundtemporalattentionaudioenvironmentalchanges

0 comments

read the original abstract

Environmental sound classification systems often do not perform robustly across different sound classification tasks and audio signals of varying temporal structures. We introduce a multi-stream convolutional neural network with temporal attention that addresses these problems. The network relies on three input streams consisting of raw audio and spectral features and utilizes a temporal attention function computed from energy changes over time. Training and classification utilizes decision fusion and data augmentation techniques that incorporate uncertainty. We evaluate this network on three commonly used data sets for environmental sound and audio scene classification and achieve new state-of-the-art performance without any changes in network architecture or front-end preprocessing, thus demonstrating better generalizability.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Attention based Convolutional Recurrent Neural Network for Environmental Sound Classification
cs.SD 2019-07 unverdicted novelty 4.0

A CRNN model with frame-level attention achieves state-of-the-art accuracy on ESC-10 and ESC-50 environmental sound classification datasets.