A separately-trained speech projector and text LoRA can be merged with only 1K steps of multimodal fine-tuning to yield a competitive ASR, ST, and spoken QA system.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
NAVER LABS Europe Submission to the Instruction-following Track
A separately-trained speech projector and text LoRA can be merged with only 1K steps of multimodal fine-tuning to yield a competitive ASR, ST, and spoken QA system.