Fine-tuning an LLM on text instructions converted to Whisper semantic tokens enables it to understand spoken instructions at inference, bypassing TTS and speech instruction data.
This is also clear when see that with added noise (VBD noisy), the Whisper encoder starts to generate tokens that show poorer WER in comparison
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
Fine-tuning an LLM on text instructions converted to Whisper semantic tokens enables it to understand spoken instructions at inference, bypassing TTS and speech instruction data.