Pith. sign in

Training dynamic models using early exits for automatic speech recognition on resource-constrained devices

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The ability to dynamically adjust the computational load of neural models during inference is crucial for on-device processing scenarios characterised by limited and time-varying computational resources. A promising solution is presented by early-exit architectures, in which additional exit branches are appended to intermediate layers of the encoder. In self-attention models for automatic speech recognition (ASR), early-exit architectures enable the development of dynamic models capable of adapting their size and architecture to varying levels of computational resources and ASR performance demands. Previous research on early-exiting ASR models has relied on pre-trained self-supervised models, fine-tuned with an early-exit loss. In this paper, we undertake an experimental comparison between fine-tuning pre-trained backbones and training models from scratch with the early-exiting objective. Experiments conducted on public datasets reveal that early-exit models trained from scratch not only preserve performance when using fewer encoder layers but also exhibit enhanced task accuracy compared to single-exit or pre-trained models. Furthermore, we explore an exit selection strategy grounded in posterior probabilities as an alternative to the conventional frame-based entropy approach. Results provide insights into the training dynamics of early-exit architectures for ASR models, particularly the efficacy of training strategies and exit selection methods.

citation-role summary

background 1

citation-polarity summary

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Input Conditioned Layer Dropping in Speech Foundation Models

cs.SD · 2025-07-10 · conditional · novelty 4.0

An input-conditioned layer selector, trained with top-k gating, lets speech foundation models drop encoder layers per sample while outperforming random dropping and matching early exit on four audio benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Input Conditioned Layer Dropping in Speech Foundation Models cs.SD · 2025-07-10 · conditional · none · ref 18 · internal anchor

    An input-conditioned layer selector, trained with top-k gating, lets speech foundation models drop encoder layers per sample while outperforming random dropping and matching early exit on four audio benchmarks.