REVIEW 2 cited by
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability to dynamically adjust the computational load of neural models during inference is crucial for on-device processing scenarios characterised by limited and time-varying computational resources. A promising solution is presented by early-exit architectures, in which additional exit branches are appended to intermediate layers of the encoder. In self-attention models for automatic speech recognition (ASR), early-exit architectures enable the development of dynamic models capable of adapting their size and architecture to varying levels of computational resources and ASR performance demands. Previous research on early-exiting ASR models has relied on pre-trained self-supervised models, fine-tuned with an early-exit loss. In this paper, we undertake an experimental comparison between fine-tuning pre-trained backbones and training models from scratch with the early-exiting objective. Experiments conducted on public datasets reveal that early-exit models trained from scratch not only preserve performance when using fewer encoder layers but also exhibit enhanced task accuracy compared to single-exit or pre-trained models. Furthermore, we explore an exit selection strategy grounded in posterior probabilities as an alternative to the conventional frame-based entropy approach. Results provide insights into the training dynamics of early-exit architectures for ASR models, particularly the efficacy of training strategies and exit selection methods.
Forward citations
Cited by 2 Pith papers
-
Input Conditioned Layer Dropping in Speech Foundation Models
An input-conditioned layer selector, trained with top-k gating, lets speech foundation models drop encoder layers per sample while outperforming random dropping and matching early exit on four audio benchmarks.
-
An Effective Training Framework for Light-Weight Automatic Speech Recognition Models
A representation-learning pretraining step, followed by brief CTC fine-tuning, yields lightweight Conformer ASR models with lower WER than from-scratch training in the paper's reported setup.
Discussion (0). Continue with ORCID to comment.