Pith. sign in

REVIEW 2 cited by

R-CRNN: Region-based Convolutional Recurrent Neural Network for Audio Event Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.06627 v1 pith:BIGCK5HC submitted 2018-08-20 cs.SD eess.AS

classification cs.SDeess.AS
keywords networkconvolutionalmethodregion-basedaudiodetectioneventlevel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper proposes a Region-based Convolutional Recurrent Neural Network (R-CRNN) for audio event detection (AED). The proposed network is inspired by Faster-RCNN, a well known region-based convolutional network framework for visual object detection. Different from the original Faster-RCNN, a recurrent layer is added on top of the convolutional network to capture the long-term temporal context from the extracted high level features. While most of the previous works on AED generate predictions at frame level first, and then use post-processing to predict the onset/offset timestamps of events from a probability sequence; the proposed method generates predictions at event level directly and can be trained end-to-end with a multitask loss, which optimizes the classification and localization of audio events simultaneously. The proposed method is tested on DCASE 2017 Challenge dataset. To the best of our knowledge, R-CRNN is the best performing single-model method among all methods without using ensembles both on development and evaluation sets. Compared to the other region-based network for AED (R-FCN) with an event-based error rate (ER) of 0.18 on the development set, our method reduced the ER to half.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Recognizing Ornaments in Vocal Indian Art Music with Active Annotation

    eess.AS 2025-05 reject novelty 5.0 of 10

    A new ROD dataset and an ED-TCN model with don't-care chunking reportedly detect six Hindustani vocal ornaments with F1 near 90 on the in-domain split.

  2. Sound Event Detection in Multichannel Audio using Convolutional Time-Frequency-Channel Squeeze and Excitation

    eess.AS 2019-08 conditional novelty 4.0 of 10

    A tfc-SE attention block inserted after CNN layers of a CRNN reduces sound event detection error rate from 0.2538 to 0.2026 on the synthetic CRESIM overlap-3 dataset.

Pith tools