Pith. sign in

REVIEW 1 cited by

Mask scalar prediction for improving robust automatic speech recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.12092 v1 pith:23N3TMVD submitted 2022-04-26 eess.AS cs.SD

classification eess.AScs.SD
keywords speechmaskacousticadditionalautomaticdistortionenhancementhand-tuned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Using neural network based acoustic frontends for improving robustness of streaming automatic speech recognition (ASR) systems is challenging because of the causality constraints and the resulting distortion that the frontend processing introduces in speech. Time-frequency masking based approaches have been shown to work well, but they need additional hyper-parameters to scale the mask to limit speech distortion. Such mask scalars are typically hand-tuned and chosen conservatively. In this work, we present a technique to predict mask scalars using an ASR-based loss in an end-to-end fashion, with minimal increase in the overall model size and complexity. We evaluate the approach on two robust ASR tasks: multichannel enhancement in the presence of speech and non-speech noise, and acoustic echo cancellation (AEC). Results show that the presented algorithm consistently improves word error rate (WER) without the need for any additional tuning over strong baselines that use hand-tuned hyper-parameters: up to 16% for multichannel enhancement in noisy conditions, and up to 7% for AEC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Magnitude strength, not estimated phase, drives SE-induced ASR degradation, and the optimal strength is recognizer-dependent (strong for wav2vec 2.0, mild for Whisper).

Pith tools