Pith. sign in

REVIEW

Towards Personalization of CTC Speech Recognition Models with Contextual Adapters and Adaptive Boosting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.09510 v3 pith:YFOUISKI submitted 2022-10-18 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modelsapproachattentionboostingpredictionsrarerecognitionspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregressive CTC decoder is often used at inference time due to its speed and simplicity. However, such models are hard to personalize because of their conditional independence assumption that prevents output tokens from previous time steps to influence future predictions. To tackle this, we propose a novel two-way approach that first biases the encoder with attention over a predefined list of rare long-tail and out-of-vocabulary (OOV) words and then uses dynamic boosting and phone alignment network during decoding to further bias the subword predictions. We evaluate our approach on open-source VoxPopuli and in-house medical datasets to showcase a 60% improvement in F1 score on domain-specific rare words over a strong CTC baseline.

Discussion (0). Sign in to comment.

Pith tools