A gaze-plus-language pipeline enables a robot to select tabletop objects from a remote user's monitor and generate human-aware grasps for handover, with real-world tests on YCB objects.
CAGE: Context-Aware Grasping Engine
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Semantic grasping is the problem of selecting stable grasps that are functionally suitable for specific object manipulation tasks. In order for robots to effectively perform object manipulation, a broad sense of contexts, including object and task constraints, needs to be accounted for. We introduce the Context-Aware Grasping Engine, which combines a novel semantic representation of grasp contexts with a neural network structure based on the Wide & Deep model, capable of capturing complex reasoning patterns. We quantitatively validate our approach against three prior methods on a novel dataset consisting of 14,000 semantic grasps for 44 objects, 7 tasks, and 6 different object states. Our approach outperformed all baselines by statistically significant margins, producing new insights into the importance of balancing memorization and generalization of contexts for semantic grasping. We further demonstrate the effectiveness of our approach on robot experiments in which the presented model successfully achieved 31 of 32 suitable grasps. The code and data are available at: https://github.com/wliu88/rail_semantic_grasping
citation-role summary
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Multimodal Human-Intent Modeling for Contextual Robot-to-Human Handovers of Arbitrary Objects
A gaze-plus-language pipeline enables a robot to select tabletop objects from a remote user's monitor and generate human-aware grasps for handover, with real-world tests on YCB objects.