arXiv preprint arXiv:2306.06819 , year=

Multimodal audio-textual architecture for robust spoken language understanding , author= · 2023 · arXiv 2306.06819

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

representative citing papers

CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents

cs.HC · 2026-06-19 · unverdicted · novelty 6.0

CORTIS is a text-only adaptation method for spoken language models that enables direct speech-to-structured-output generation for task-oriented agents and matches or exceeds ASR-LLM cascades under acoustic degradation.

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

eess.AS · 2026-06-19 · unverdicted · novelty 6.0

An end-to-end SLU architecture with frozen SSL acoustic encoder, LSTM classification head, and cross-modal distillation achieves 93% accuracy on simple commands and 82% on spontaneous speech at 7 ms latency on the new VoiceStick corpus, outperforming cascade baselines.

citing papers explorer

Showing 1 of 1 citing paper after filters.

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users eess.AS · 2026-06-19 · unverdicted · none · ref 25
An end-to-end SLU architecture with frozen SSL acoustic encoder, LSTM classification head, and cross-modal distillation achieves 93% accuracy on simple commands and 82% on spontaneous speech at 7 ms latency on the new VoiceStick corpus, outperforming cascade baselines.

arXiv preprint arXiv:2306.06819 , year=

fields

years

verdicts

representative citing papers

citing papers explorer