For a simple industrial pick-and-place task, a scripted robot interaction matches LLM-enhanced interaction on objective efficiency and focus, while subjective ratings only marginally favor the LLM.
Bidirectional Intent Communication: A Role for Large Foundation Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Integrating multimodal foundation models has significantly enhanced autonomous agents' language comprehension, perception, and planning capabilities. However, while existing works adopt a \emph{task-centric} approach with minimal human interaction, applying these models to developing assistive \emph{user-centric} robots that can interact and cooperate with humans remains underexplored. This paper introduces ``Bident'', a framework designed to integrate robots seamlessly into shared spaces with humans. Bident enhances the interactive experience by incorporating multimodal inputs like speech and user gaze dynamics. Furthermore, Bident supports verbal utterances and physical actions like gestures, making it versatile for bidirectional human-robot interactions. Potential applications include personalized education, where robots can adapt to individual learning styles and paces, and healthcare, where robots can offer personalized support, companionship, and everyday assistance in the home and workplace environments.
citation-role summary
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Evaluating Efficiency and Engagement in Scripted and LLM-Enhanced Human-Robot Interactions
For a simple industrial pick-and-place task, a scripted robot interaction matches LLM-enhanced interaction on objective efficiency and focus, while subjective ratings only marginally favor the LLM.