Pith. sign in

REVIEW 4 cited by

CAGE: Causal Attention Enables Data-Efficient Generalizable Robotic Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14974 v2 pith:XN6OBJCY submitted 2024-10-19 cs.RO

classification cs.RO
keywords cagemanipulationattentioncausalenvironmentsgeneralizationraterobotic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generalization in robotic manipulation remains a critical challenge, particularly when scaling to new environments with limited demonstrations. This paper introduces CAGE, a novel robotic manipulation policy designed to overcome these generalization barriers by integrating a causal attention mechanism. CAGE utilizes the powerful feature extraction capabilities of the vision foundation model DINOv2, combined with LoRA fine-tuning for robust environment understanding. The policy further employs a causal Perceiver for effective token compression and a diffusion-based action prediction head with attention mechanisms to enhance task-specific fine-grained conditioning. With as few as 50 demonstrations from a single training environment, CAGE achieves robust generalization across diverse visual changes in objects, backgrounds, and viewpoints. Extensive experiments validate that CAGE significantly outperforms existing state-of-the-art RGB/RGB-D approaches in various manipulation tasks, especially under large distribution shifts. In similar environments, CAGE offers an average of 42% increase in task completion rate. While all baselines fail to execute the task in unseen environments, CAGE manages to obtain a 43% completion rate and a 51% success rate in average, making a huge step towards practical deployment of robots in real-world settings. Project website: cage-policy.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Cross-Embodiment World Models for Dexterous Manipulation

    cs.RO 2025-11 conditional novelty 6.0 of 10

    A single particle-based world model trained on many simulated robot hands and real human hands can plan dexterous manipulation on robot hands it never trained on.

  2. Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Gaze-guided foveated patch tokenization reduces ViT tokens by 94%, accelerates training 7x and inference 3x, and improves robustness to distractors in bimanual manipulation policies.

  3. Knowledge-Driven Imitation Learning: Enabling Generalization Across Diverse Conditions

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A semantic keypoint graph matched to novel objects lets imitation-learned manipulation policies generalize with a quarter of the demonstrations.

  4. CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World

    cs.RO 2025-02 conditional novelty 6.0 of 10

    CordViP achieves strong real-world dexterous manipulation by feeding a diffusion policy with pose-tracked 3D object models and hand point clouds, pretrained on contact maps and arm-hand coordination.

Pith tools