Pith. sign in

REVIEW 2 cited by

Learning Efficient and Robust Language-conditioned Manipulation using Textual-Visual Relevancy and Equivariant Language Mapping

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15677 v2 pith:D6U4KJ5E submitted 2024-06-21 cs.RO

classification cs.RO
keywords languageequivariantmanipulationrobottasksapproachcompareddata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Controlling robots through natural language is pivotal for enhancing human-robot collaboration and synthesizing complex robot behaviors. Recent works that are trained on large robot datasets show impressive generalization abilities. However, such pretrained methods are (1) often fragile to unseen scenarios, and (2) expensive to adapt to new tasks. This paper introduces Grounded Equivariant Manipulation (GEM), a robust yet efficient approach that leverages pretrained vision-language models with equivariant language mapping for language-conditioned manipulation tasks. Our experiments demonstrate GEM's high sample efficiency and generalization ability across diverse tasks in both simulation and the real world. GEM achieves similar or higher performance with orders of magnitude fewer robot data compared with major data-efficient baselines such as CLIPort and VIMA. Finally, our approach demonstrates greater robustness compared to large VLA model, e.g, OpenVLA, at correctly interpreting natural language commands on unseen objects and poses. Code, data, and training details are available https://saulbatman.github.io/gem_page/

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop Robotic Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    EquAct embeds SE(3) equivariance into a multi-task keyframe manipulation transformer with language conditioning, improving spatial generalization over non-equivariant baselines.

  2. Improving Generalization of Language-Conditioned Robot Manipulation

    cs.RO 2025-08 unverdicted novelty 5.0 of 10

    A two-stage fine-tuning framework with instance-level semantic fusion lets language-conditioned robots learn object-arrangement tasks from a few demonstrations and generalize to unseen environments.

Pith tools