A single audio-plus-text LLM jointly performs voice trigger detection, device-directed speech detection, dialog act classification, and ASR, with reported EER reductions of 64% and 22% over dedicated baselines.
A multimodal approach to device-directed speech detection with large language models,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions
A single audio-plus-text LLM jointly performs voice trigger detection, device-directed speech detection, dialog act classification, and ASR, with reported EER reductions of 64% and 22% over dedicated baselines.