Pith. sign in

REVIEW 5 cited by

GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.12286 v2 pith:QDYGBLZF submitted 2024-11-19 cs.RO cs.CV

classification cs.ROcs.CV
keywords affordancereasoningglovergraspingopen-vocabularyacrossenableestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Inferring affordable (i.e., graspable) parts of arbitrary objects based on human specifications is essential for robots advancing toward open-vocabulary manipulation. Current grasp planners, however, are hindered by limited vision-language comprehension and time-consuming 3D radiance modeling, restricting real-time, open-vocabulary interactions with objects. To address these limitations, we propose GLOVER, a unified Generalizable Open-Vocabulary Affordance Reasoning framework, which fine-tunes the Large Language Models (LLMs) to predict the visual affordance of graspable object parts within RGB feature space. We compile a dataset of over 10,000 images from human-object interactions, annotated with unified visual and linguistic affordance labels, to enable multi-modal fine-tuning. GLOVER inherits world knowledge and common-sense reasoning from LLMs, facilitating more fine-grained object understanding and sophisticated tool-use reasoning. To enable effective real-world deployment, we present Affordance-Aware Grasping Estimation (AGE), a non-parametric grasp planner that aligns the gripper pose with a superquadric surface derived from affordance data. In evaluations across 30 table-top real-world scenes, GLOVER achieves success rates of 86.0% in part identification and 76.3% in grasping, with speeds approximately 29 times faster in affordance reasoning and 40 times faster in grasping pose estimation than the previous state-of-the-art. We also validate the generalization across embodiments, showing effectiveness in humanoid robots with dexterous hands.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Advancing Creative Physical Intelligence in Large Multimodal Models

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    Introduces MM-CreativityBench for affordance-grounded creative tool use and shows that DPO-based alignment with an affordance knowledge base improves entity and part selection while cutting hallucination errors in LMMs.

  2. FrameVGGT: Coherence-Preserving Memory for Bounded Streaming Geometry

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    FrameVGGT replaces token-level KV retention with frame-level segments and prototypes to bound memory while preserving geometric coherence in streaming VGGT.

  3. PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

    cs.RO 2026-01 unverdicted novelty 6.0 of 10

    PALM improves long-horizon robotic manipulation success by distilling affordance representations for object interaction and predicting within-subtask progress in a VLA model.

  4. Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A bidirectional vision-language-depth fusion model, OGRG, outperforms prior baselines in grounding and grasping objects described by spatial language, including with duplicate objects and weak grasp labels.

  5. FrameVGGT: Coherence-Preserving Memory for Bounded Streaming Geometry

    cs.CV 2026-03 unverdicted novelty 5.0 of 10

    FrameVGGT maintains stable long-horizon 3D reconstruction, depth, and pose under fixed memory by organizing history as complementary frame-wise KV prototypes plus sparse anchors.

Pith tools