Pith. sign in

REVIEW 2 cited by

GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.14403 v1 pith:IQ6EF7BQ submitted 2024-09-22 cs.RO cs.CV

classification cs.ROcs.CV
keywords detectiongraspgraspmambafeaturesinferencelanguage-drivenmamba-basedapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Grasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with cluttered images, lengthy textual descriptions, or slow inference speed. We introduce GraspMamba, a new language-driven grasp detection method that employs hierarchical feature fusion with Mamba vision to tackle these challenges. By leveraging rich visual features of the Mamba-based backbone alongside textual information, our approach effectively enhances the fusion of multimodal features. GraspMamba represents the first Mamba-based grasp detection model to extract vision and language features at multiple scales, delivering robust performance and rapid inference time. Intensive experiments show that GraspMamba outperforms recent methods by a clear margin. We validate our approach through real-world robotic experiments, highlighting its fast inference speed.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A multi-agent system with planner, coder, and observer agents achieves zero-shot language-driven grasp detection that outperforms existing baselines on benchmarks and robots.

  2. MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A two-stage language-driven grasping system that pools visual features inside a predicted object mask improves grasp accuracy and training efficiency versus CLIP baselines, supported by a new 219M-grasp dataset.

Pith tools