Adding language-specific prompts and bi-directional conversational context to a speech LLM cuts validation error by 18% relative and edges out a model trained on four times more data.
Conversational speech recognition by learning audio-textual cross-modal contex- tual representation,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
baseline 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
Adding language-specific prompts and bi-directional conversational context to a speech LLM cuts validation error by 18% relative and edges out a model trained on four times more data.