Locate-and-Focus localizes the audio span of a terminology in an utterance and uses the located clip, a matched audio replacement, and a special <Term> cue to make speech LLMs translate the terminology correctly.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models
Locate-and-Focus localizes the audio span of a terminology in an utterance and uses the located clip, a matched audio replacement, and a special <Term> cue to make speech LLMs translate the terminology correctly.