An input-conditioned layer selector, trained with top-k gating, lets speech foundation models drop encoder layers per sample while outperforming random dropping and matching early exit on four audio benchmarks.
For each input sample, the LS block selects the finest combination of encoder layers achieving optimal performance for various resource settings
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Input Conditioned Layer Dropping in Speech Foundation Models
An input-conditioned layer selector, trained with top-k gating, lets speech foundation models drop encoder layers per sample while outperforming random dropping and matching early exit on four audio benchmarks.