Pith. sign in

REVIEW 6 cited by

Reasoning Grasping via Multimodal Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06798 v3 pith:JRRYBXGS submitted 2024-02-09 cs.RO

classification cs.RO
keywords graspingreasoningmodelbenchmarkenvironmentsgrasphumanimplicit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite significant progress in robotic systems for operation within human-centric environments, existing models still heavily rely on explicit human commands to identify and manipulate specific objects. This limits their effectiveness in environments where understanding and acting on implicit human intentions are crucial. In this study, we introduce a novel task: reasoning grasping, where robots need to generate grasp poses based on indirect verbal instructions or intentions. To accomplish this, we propose an end-to-end reasoning grasping model that integrates a multimodal Large Language Model (LLM) with a vision-based robotic grasping framework. In addition, we present the first reasoning grasping benchmark dataset generated from the GraspNet-1 billion, incorporating implicit instructions for object-level and part-level grasping. Our results show that directly integrating CLIP or LLaVA with the grasp detection model performs poorly on the challenging reasoning grasping tasks, while our proposed model demonstrates significantly enhanced performance both in the reasoning grasping benchmark and real-world experiments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

    cs.CV 2025-11 conditional novelty 7.0 of 10

    CFG-Bench uses 19,562 questions across four cognitive tiers to show that vision-language models are weak at fine-grained physical action understanding, and that SFT on its data improves scores on external embodied benchmarks.

  2. Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Auras, a perception-generation disaggregation framework with a public context buffer and asynchronous pipeline executor, raises embodied-agent throughput by 2.54x on average without losing accuracy (102.7%).

  3. Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A bidirectional vision-language-depth fusion model, OGRG, outperforms prior baselines in grounding and grasping objects described by spatial language, including with duplicate objects and weak grasp labels.

  4. GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A multi-agent system with planner, coder, and observer agents achieves zero-shot language-driven grasp detection that outperforms existing baselines on benchmarks and robots.

  5. Transforming Remanufacturing Automation with Large Language Models: A Forward-Looking Analysis with Case Studies

    eess.SY 2026-08 conditional novelty 5.0 of 10

    The authors propose ReManGPT, a conceptual orchestration framework for applying LLMs to remanufacturing, and illustrate it with case studies in disassembly planning, repair guidance, and robotic execution.

  6. Foundation Model Driven Robotics: A Comprehensive Review

    cs.RO 2025-07 conditional novelty 2.0 of 10

    A review of foundation-model-driven robotics that synthesizes recent work across perception, planning, control, HRI, simulation, and sim-to-real transfer, and highlights open challenges.

Pith tools