Pith. sign in

REVIEW 3 cited by

Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15815 v2 pith:4YL7SYKG submitted 2024-07-22 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords maniwherevisuallearninggeneralizationacrossapproachframeworkgeneralizable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Can we endow visuomotor robots with generalization capabilities to operate in diverse open-world scenarios? In this paper, we propose \textbf{Maniwhere}, a generalizable framework tailored for visual reinforcement learning, enabling the trained robot policies to generalize across a combination of multiple visual disturbance types. Specifically, we introduce a multi-view representation learning approach fused with Spatial Transformer Network (STN) module to capture shared semantic information and correspondences among different viewpoints. In addition, we employ a curriculum-based randomization and augmentation approach to stabilize the RL training process and strengthen the visual generalization ability. To exhibit the effectiveness of Maniwhere, we meticulously design 8 tasks encompassing articulate objects, bi-manual, and dexterous hand manipulation tasks, demonstrating Maniwhere's strong visual generalization and sim2real transfer abilities across 3 hardware platforms. Our experiments show that Maniwhere significantly outperforms existing state-of-the-art methods. Videos are provided at https://gemcollector.github.io/maniwhere/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A hybrid-attention latent-guided online RL adapter lifts frozen world-action models from 26.4% to 87.1% average success on four real precision manipulation tasks in 45–75 minutes each.

  2. DeGuV: Depth-Guided Visual Reinforcement Learning for Generalization and Interpretability in Manipulation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Depth-guided masking improves visual RL generalization, sample efficiency, and interpretability on manipulation tasks.

  3. RoboPearls: Editable Video Simulation for Robot Manipulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RoboPearls is a 3D Gaussian Splatting based framework that edits demonstration videos into varied photorealistic simulations, and training on them improves robot manipulation success rates on RLBench and COLOSSEUM.

Pith tools