REVIEW 2 cited by
GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce a pipeline that enhances a general-purpose Vision Language Model, GPT-4V(ision), to facilitate one-shot visual teaching for robotic manipulation. This system analyzes videos of humans performing tasks and outputs executable robot programs that incorporate insights into affordances. The process begins with GPT-4V analyzing the videos to obtain textual explanations of environmental and action details. A GPT-4-based task planner then encodes these details into a symbolic task plan. Subsequently, vision systems spatially and temporally ground the task plan in the videos. Objects are identified using an open-vocabulary object detector, and hand-object interactions are analyzed to pinpoint moments of grasping and releasing. This spatiotemporal grounding allows for the gathering of affordance information (e.g., grasp types, waypoints, and body postures) critical for robot execution. Experiments across various scenarios demonstrate the method's efficacy in enabling real robots to operate from one-shot human demonstrations. Meanwhile, quantitative tests have revealed instances of hallucination in GPT-4V, highlighting the importance of incorporating human supervision within the pipeline. The prompts of GPT-4V/GPT-4 are available at this project page: https://microsoft.github.io/GPT4Vision-Robot-Manipulation-Prompts/
Forward citations
Cited by 2 Pith papers
-
SAGP: Semantic Affordance-Guided Grasp Planning via Coarse-Zone VLM Reasoning
Coarse-zone VLM ratings re-rank geometric grasps to prefer functionally appropriate regions, preserving physical success while improving semantic appropriateness on handle-bearing objects.
-
LA-RCS: LLM-Agent-Based Robot Control System
LA-RCS reports that a dual-agent LLM system controls a small car robot to complete 18 of 20 self-designed commands with the GPT-4o variant, but the supporting evaluation is inconsistent and not reproducible.
Discussion (0). Continue with ORCID to comment.