Pith. sign in

REVIEW 3 cited by

UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20551 v2 pith:R2MNR7C5 submitted 2024-09-30 cs.RO

classification cs.RO
keywords manipulationuniaffroboticunifiedaffordancesarticulatedcategoriescomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified formulation. Specifically, we constructed a dataset labeled with manipulation-related key attributes, comprising 900 articulated objects from 19 categories and 600 tools from 12 categories. Furthermore, we leverage MLLMs to infer object-centric representations for manipulation tasks, including affordance recognition and reasoning about 3D motion constraints. Comprehensive experiments in both simulation and real-world settings indicate that UniAff significantly improves the generalization of robotic manipulation for tools and articulated objects. We hope that UniAff will serve as a general baseline for unified robotic manipulation tasks in the future. Images, videos, dataset, and code are published on the project website at:https://sites.google.com/view/uni-aff/home

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Creative Robot Tool Use by Counterfactual Reasoning

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    Robots discover causal tool features through VLM suggestions and physics-based counterfactual perturbations in simulation, then transfer manipulation skills via conditioned keypoint matching.

  2. RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A multi-domain affordance benchmark with 273k images and 26k reasoning instructions is introduced, together with a VLM-based grasping pipeline that shows strong zero-shot affordance segmentation and real-robot performance.

  3. ArtGS:3D Gaussian Splatting for Interactive Visual-Physical Modeling and Manipulation of Articulated Objects

    cs.RO 2025-07 conditional novelty 6.0 of 10

    ArtGS combines multi-view 3D reconstruction, language-model joint initialization, and closed-loop optimization to improve articulated object manipulation.

Pith tools