Pith. sign in

REVIEW 3 cited by

Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07487 v1 pith:GILM5GLK submitted 2024-01-15 cs.RO cs.CV

classification cs.ROcs.CV
keywords affordanceobjectsrobo-abccorrespondencesemanticcategoriescontactenabling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Enabling robotic manipulation that generalizes to out-of-distribution scenes is a crucial step toward open-world embodied intelligence. For human beings, this ability is rooted in the understanding of semantic correspondence among objects, which naturally transfers the interaction experience of familiar objects to novel ones. Although robots lack such a reservoir of interaction experience, the vast availability of human videos on the Internet may serve as a valuable resource, from which we extract an affordance memory including the contact points. Inspired by the natural way humans think, we propose Robo-ABC: when confronted with unfamiliar objects that require generalization, the robot can acquire affordance by retrieving objects that share visual or semantic similarities from the affordance memory. The next step is to map the contact points of the retrieved objects to the new object. While establishing this correspondence may present formidable challenges at first glance, recent research finds it naturally arises from pre-trained diffusion models, enabling affordance mapping even across disparate object categories. Through the Robo-ABC framework, robots may generalize to manipulate out-of-category objects in a zero-shot manner without any manual annotation, additional training, part segmentation, pre-coded knowledge, or viewpoint restrictions. Quantitatively, Robo-ABC significantly enhances the accuracy of visual affordance retrieval by a large margin of 31.6% compared to state-of-the-art (SOTA) end-to-end affordance models. We also conduct real-world experiments of cross-category object-grasping tasks. Robo-ABC achieved a success rate of 85.7%, proving its capacity for real-world tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weakly-Supervised Learning of Dense Functional Correspondences

    cs.CV 2025-09 conditional novelty 7.0 of 10

    A weakly-supervised pipeline that distills VLM functional part knowledge and multi-view spatial structure into a model for dense cross-category functional correspondence, outperforming baselines on new synthetic and r...

  2. BiAssemble: Learning Collaborative Affordance for Bimanual Geometric Assembly

    cs.RO 2025-06 conditional novelty 6.0 of 10

    BiAssemble predicts bimanual grasp and assembly actions for geometric reassembly of fractured objects via point-level collaborative affordance, and reports simulation gains over baselines plus a real-world benchmark.

  3. Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A pseudo-supervised pipeline with an affordance-to-part mapping, label refinement, cross-view alignment, and a reasoning module achieves state-of-the-art weakly supervised affordance grounding on AGD20K.

Pith tools