Pith. sign in

REVIEW 3 cited by

GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.13441 v3 pith:OHCWK6DI submitted 2025-05-19 cs.RO

classification cs.RO
keywords graspmolmomodelsyntheticdatadatasetgeneralizablegraspinggrasps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present GrasMolmo, a generalizable open-vocabulary task-oriented grasping (TOG) model. GraspMolmo predicts semantically appropriate, stable grasps conditioned on a natural language instruction and a single RGB-D frame. For instance, given "pour me some tea", GraspMolmo selects a grasp on a teapot handle rather than its body. Unlike prior TOG methods, which are limited by small datasets, simplistic language, and uncluttered scenes, GraspMolmo learns from PRISM, a novel large-scale synthetic dataset of 379k samples featuring cluttered environments and diverse, realistic task descriptions. We fine-tune the Molmo visual-language model on this data, enabling GraspMolmo to generalize to novel open-vocabulary instructions and objects. In challenging real-world evaluations, GraspMolmo achieves state-of-the-art results, with a 70% prediction success on complex tasks, compared to the 35% achieved by the next best alternative. GraspMolmo also successfully demonstrates the ability to predict semantically correct bimanual grasps zero-shot. We release our synthetic dataset, code, model, and benchmarks to accelerate research in task-semantic robotic manipulation, which, along with videos, are available at https://abhaybd.github.io/GraspMolmo/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training

    cs.RO 2025-07 conditional novelty 6.0 of 10

    GraspGen shows that training a grasp-scoring discriminator on the generator's own simulated outputs, plus a large new multi-gripper dataset, improves 6-DOF grasping across simulation and a real robot.

  2. Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    Embodied-R1.5 is an 8B EFM achieving SOTA on 16 of 24 embodied VLM benchmarks, fine-tunable to outperform leading VLAs, with claimed zero-shot real-robot generalization.

  3. MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A two-stage language-driven grasping system that pools visual features inside a predicted object mask improves grasp accuracy and training efficiency versus CLIP baselines, supported by a new 219M-grasp dataset.

Pith tools