The paper presents an LLM-based system architecture that fuses speech transcription, gesture recognition, and beat detection to generate and execute robotic action sequences.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music
The paper presents an LLM-based system architecture that fuses speech transcription, gesture recognition, and beat detection to generate and execute robotic action sequences.