Pith. sign in

REVIEW 12 cited by

A3VLM: Actionable Articulation-Aware Vision Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07549 v2 pith:FP7IYUKK submitted 2024-06-11 cs.RO

classification cs.RO
keywords a3vlmlanguageroboticsvisionvlmsactionactionableactions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision Language Models (VLMs) have received significant attention in recent years in the robotics community. VLMs are shown to be able to perform complex visual reasoning and scene understanding tasks, which makes them regarded as a potential universal solution for general robotics problems such as manipulation and navigation. However, previous VLMs for robotics such as RT-1, RT-2, and ManipLLM have focused on directly learning robot-centric actions. Such approaches require collecting a significant amount of robot interaction data, which is extremely costly in the real world. Thus, we propose A3VLM, an object-centric, actionable, articulation-aware vision language model. A3VLM focuses on the articulation structure and action affordances of objects. Its representation is robot-agnostic and can be translated into robot actions using simple action primitives. Extensive experiments in both simulation benchmarks and real-world settings demonstrate the effectiveness and stability of A3VLM. We release our code and other materials at https://github.com/changhaonan/A3VLM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    UnfoldArt uses multi-agent debate grounded in vision-language and video models to infer articulation parameters and reconstruct full 3D objects including occluded parts from text or image inputs.

  2. UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    UnfoldArt uses a two-round structured debate between high-level semantic agents and low-level parameter agents, grounded in generated video, to infer articulation and reconstruct full articulated 3D objects including ...

  3. Sketch2Arti: Sketch-based Articulation Modeling of CAD Objects

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Sketch2Arti is the first sketch-based system that automatically finds movable parts in CAD objects and predicts their motion parameters from 2D user sketches, trained without object category labels and supporting inte...

  4. Sketch2Arti: Sketch-based Articulation Modeling of CAD Objects

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Sketch2Arti is the first category-agnostic sketch-based system that infers movable parts and motion parameters from 2D user sketches on CAD objects and supports sketch-guided internal structure completion for shell models.

  5. ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

    cs.RO 2024-09 conditional novelty 7.0 of 10

    ReKep encodes robotic tasks as optimizable Python functions over 3D keypoints that are generated automatically from language and RGB-D input, enabling real-time hierarchical planning on single- and dual-arm platforms ...

  6. Analytic Concept-Centric Memory for Agentic Embodied Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Proposes a structured concept-centric memory system for embodied agents that connects object, scene, transition, and skill memories to support coarse-to-fine retrieval and improve task performance over baselines.

  7. Action with Visual Primitives

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    AVP architecture has VLM emit visual-primitive tokens to condition flow-matching action expert, yielding 27.61% higher success rate than pi_0.5 on real-robot pick-and-place tasks.

  8. ArtMesh: Part-Aware Articulated Mesh Fields with Motion-Consistent Dynamics

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    ArtMesh presents a mesh-native pipeline for articulated reconstruction that uses restricted Delaunay remeshing and bidirectional motion consistency to outperform 3D Gaussian Splatting methods on joint estimation and p...

  9. Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    Vision-language models generate executable Behavior Tree policies for robots from synthetic vision-language data, with successful transfer demonstrated on two real manipulators.

  10. Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    A 12B-parameter VLM learns to synthesize executable Behavior Tree policies from multimodal inputs via synthetic neuro-symbolic supervision, achieving zero-shot real-world transfer on robotic manipulators.

  11. PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

    cs.RO 2026-01 unverdicted novelty 6.0 of 10

    PALM improves long-horizon robotic manipulation success by distilling affordance representations for object interaction and predicting within-subtask progress in a VLA model.

  12. Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Embodied-R1 uses a pointing-centric representation and reinforced fine-tuning on a 200K dataset to achieve state-of-the-art results on embodied benchmarks plus 56.2% success in SIMPLEREnv and 87.5% on real XArm tasks ...

Pith tools