Shake-VLA integrates YOLOv8, EasyOCR, Whisper, RAG, and GPT-4o on bimanual robots to prepare cocktails from voice commands, reporting 91-100% component and overall success rates.
Get smart: Collaborative goal setting with cognitively assistive robots,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing
Shake-VLA integrates YOLOv8, EasyOCR, Whisper, RAG, and GPT-4o on bimanual robots to prepare cocktails from voice commands, reporting 91-100% component and overall success rates.