By turning each layer's width into a continuous, noise-smoothed parameter, the authors train speech models whose sizes shrink during training, reducing FLOPs and size by roughly 80–90% in their case studies.
Nas-scae: Searching compact attention-based encoders for end-to-end automatic speech recognition,
1 Pith paper cite this work, alongside 1 external citations. Polarity classification is still indexing.
1
Pith paper citing it
1
external citations · OpenAlex
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Performance and Complexity Trade-off Optimization of Speech Models During Training
By turning each layer's width into a continuous, noise-smoothed parameter, the authors train speech models whose sizes shrink during training, reducing FLOPs and size by roughly 80–90% in their case studies.