Pith. sign in

REVIEW

Mask Detection and Breath Monitoring from Speech: on Data Augmentation, Feature Representation and Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.05175 v2 pith:3BO4SR4T submitted 2020-08-12 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords maskdatafeaturesaugmentationbreathbreathingdeepdetection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces our approaches for the Mask and Breathing Sub-Challenge in the Interspeech COMPARE Challenge 2020. For the mask detection task, we train deep convolutional neural networks with filter-bank energies, gender-aware features, and speaker-aware features. Support Vector Machines follows as the back-end classifiers for binary prediction on the extracted deep embeddings. Several data augmentation schemes are used to increase the quantity of training data and improve our models' robustness, including speed perturbation, SpecAugment, and random erasing. For the speech breath monitoring task, we investigate different bottleneck features based on the Bi-LSTM structure. Experimental results show that our proposed methods outperform the baselines and achieve 0.746 PCC and 78.8% UAR on the Breathing and Mask evaluation set, respectively.

Discussion (0). Continue with ORCID to comment.

Pith tools