A fine-tuned Whisper tiny.en model reaches 15.9% WER on children's speech and runs in real time on a Raspberry Pi, with low-rank compression trading an 11% relative WER rise for lighter computation.
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Modern automatic speech recognition (ASR) models, such as OpenAI's Whisper, rely on deep encoder-decoder architectures, and their encoders are a critical bottleneck for efficient deployment due to high computational intensity. We introduce LiteASR, a low-rank compression scheme for ASR encoders that significantly reduces inference costs while maintaining transcription accuracy. Our approach leverages the strong low-rank properties observed in intermediate activations: by applying principal component analysis (PCA) with a small calibration dataset, we approximate linear transformations with a chain of low-rank matrix multiplications, and further optimize self-attention to work in reduced dimensionality. Evaluation results show that our method can compress Whisper large-v3's encoder size by over 50%, matching Whisper medium's size with better transcription accuracy, thereby establishing a new Pareto frontier of accuracy and efficiency. The code of LiteASR is available at https://github.com/efeslab/LiteASR.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
A fine-tuned Whisper tiny.en model reaches 15.9% WER on children's speech and runs in real time on a Raspberry Pi, with low-rank compression trading an 11% relative WER rise for lighter computation.