Pith. sign in

REVIEW 26 cited by

Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17842 v2 pith:HMBANQW4 submitted 2023-11-29 cs.RO cs.AIcs.CLcs.CVcs.LG

Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

classification cs.RO cs.AIcs.CLcs.CVcs.LG
keywords planningroboticllmsmodelsvilavision-languageextensiveknowledge
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this study, we are interested in imbuing robots with the capability of physically-grounded task planning. Recent advancements have shown that large language models (LLMs) possess extensive knowledge useful in robotic tasks, especially in reasoning and planning. However, LLMs are constrained by their lack of world grounding and dependence on external affordance models to perceive environmental information, which cannot jointly reason with LLMs. We argue that a task planner should be an inherently grounded, unified multimodal system. To this end, we introduce Robotic Vision-Language Planning (ViLa), a novel approach for long-horizon robotic planning that leverages vision-language models (VLMs) to generate a sequence of actionable steps. ViLa directly integrates perceptual data into its reasoning and planning process, enabling a profound understanding of commonsense knowledge in the visual world, including spatial layouts and object attributes. It also supports flexible multimodal goal specification and naturally incorporates visual feedback. Our extensive evaluation, conducted in both real-robot and simulated environments, demonstrates ViLa's superiority over existing LLM-based planners, highlighting its effectiveness in a wide array of open-world manipulation tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning

    cs.RO 2026-04 unverdicted novelty 7.0

    KinDER is a new open-source benchmark that demonstrates substantial gaps in current robot learning and planning methods for handling physical constraints.

  2. ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

    cs.RO 2024-09 conditional novelty 7.0

    ReKep encodes robotic tasks as optimizable Python functions over 3D keypoints that are generated automatically from language and RGB-D input, enabling real-time hierarchical planning on single- and dual-arm platforms ...

  3. World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models

    cs.AI 2026-07 conditional novelty 6.0

    Pairing VLM-generated action proposals with rollouts from a pose-image-conditioned video world model yields high success rates in novel simulated manipulation tasks without end-to-end policy retraining.

  4. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  5. CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    CoStream composes semantic, predictive, and reactive behaviors on an SE(3) interface to enable precise, generalizable performance on eight real-world contact-rich manipulation tasks.

  6. SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments

    cs.AI 2026-06 unverdicted novelty 6.0

    SCOPE is a self-adaptive symbolic planning framework that refines plans and evolves symbolic world models via simulator feedback and distilled knowledge to improve long-horizon planning in open-ended embodied environments.

  7. SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning

    cs.AI 2026-06 unverdicted novelty 6.0

    SVoT uses RL with GRPO to train MLLMs on interleaved textual and visual reasoning chains for multi-hop spatial tasks, achieving up to 65% accuracy gains on new domains with quantitative state verification.

  8. ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop

    cs.AI 2026-06 conditional novelty 6.0

    ProcAgent demonstrates a complete step-by-step assembly assistant that runs on a single edge device by pairing cheap continuous perception with on-demand vision-language verification, a task graph, and human confirmation.

  9. Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning

    cs.AI 2026-05 unverdicted novelty 6.0

    BISON learns bilevel policies over symbolic world models to generalize long-horizon robotic planning beyond VLA and end-to-end baselines while remaining efficient even at 10,000-object scale.

  10. dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model

    cs.RO 2026-04 unverdicted novelty 6.0

    A discrete diffusion model tokenizes multimodal robotic data and uses a progress token to predict future states and task completion for scalable policy evaluation.

  11. ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making

    cs.RO 2026-03 unverdicted novelty 6.0

    ThermoAct integrates thermal imaging into VLA models via a VLM planner to enable robots to perceive physical properties like heat and improve safety over vision-only systems.

  12. Visual-Language-Guided Task Planning for Horticultural Robots

    cs.RO 2026-01 conditional novelty 6.0

    A vision-language model drives a simulated greenhouse robot through simple crop-inspection tasks with ~87% success, but long multi-target tasks collapse to under 10% success.

  13. SkillWrapper: Generative Predicate Invention for Task-level Robot Planning

    cs.RO 2025-11 reject novelty 6.0

    SkillWrapper learns human-readable skill models from images via VLM predicate invention, but its provable soundness/completeness claim is not established.

  14. SkillWrapper: Generative Predicate Invention for Task-level Robot Planning

    cs.RO 2025-11 unverdicted novelty 6.0

    SkillWrapper learns human-interpretable symbolic representations of robot skills from images via foundation models, yielding operators for provably sound and complete planning on long-horizon tasks.

  15. $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

    cs.LG 2025-04 unverdicted novelty 6.0

    π_{0.5} is a VLA model that achieves long-horizon dexterous manipulation in entirely new homes through co-training on heterogeneous tasks and multi-source data including web and semantic predictions.

  16. Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

    cs.RO 2025-02 unverdicted novelty 6.0

    A hierarchical VLA architecture lets robots follow complex instructions and situated feedback by separating high-level reasoning from low-level control.

  17. DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

    cs.RO 2025-02 unverdicted novelty 6.0

    DexVLA combines a scaled diffusion action expert with embodiment curriculum learning to achieve better generalization and performance than prior VLA models on diverse robot hardware and long-horizon tasks.

  18. NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

    cs.CV 2024-02 unverdicted novelty 6.0

    NaVid, a video-based VLM trained on 510k navigation and 763k web samples, achieves SOTA VLN performance using only monocular RGB video for next-step action planning in sim and real environments.

  19. Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

    cs.RO 2026-07 accept novelty 5.5

    Bimanual VLA coordination strategies, training recipes, and continuous action chunking transfer to unmanned aerial systems; the survey maps 183 works and lists fourteen shared research directions.

  20. CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

    cs.RO 2026-06 unverdicted novelty 5.0

    Complex manipulation emerges from composing semantic, predictive, and reactive behaviors on a shared SE(3) interface executed by a compliant controller.

  21. Make Your VLA More Robust Without More Data By Interleaving Motion Planning

    cs.RO 2026-05 unverdicted novelty 5.0

    MPVI interleaves model-based motion planning with VLAs via VLM completion checking to achieve 113% higher task progress on BEHAVIOR-1K without extra data.

  22. RoboAgent: Chaining Basic Capabilities for Embodied Task Planning

    cs.RO 2026-04 unverdicted novelty 5.0

    RoboAgent chains basic vision-language capabilities inside a single VLM via a scheduler and trains it in three stages (behavior cloning, DAgger, RL) to improve embodied task planning.

  23. Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

    cs.RO 2025-08 unverdicted novelty 5.0

    This survey organizes large VLM-based VLA models for robotic manipulation into monolithic and hierarchical paradigms, reviews their integrations and datasets, and outlines future directions.

  24. VeriGraph: Scene Graphs for Execution Verifiable Robot Planning

    cs.RO 2024-11 unverdicted novelty 5.0

    VeriGraph integrates VLMs with scene-graph verification to raise robot task success rates by 30-58% over baselines in manipulation scenarios.

  25. Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own

    cs.RO 2023-10 unverdicted novelty 5.0

    RLFP and the FAC algorithm combine foundation-model priors for policy, value, and rewards to produce sample-efficient robotic RL that reaches 86% real-robot success after one hour and 100% success on 7/8 Meta-world ta...

  26. Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning

    cs.CV 2025-06 unverdicted novelty 4.0

    Vision-EKIPL injects high-quality actions from external models into RL training to expand exploration and raise the reasoning ceiling of MLLMs, reporting up to 5% gains on the Reason-RFT-CoT benchmark.