Pith. sign in

REVIEW 8 cited by

PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00278 v2 pith:TPQYNOPT submitted 2024-06-29 cs.RO cs.AIcs.CVcs.LG

PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks

classification cs.RO cs.AIcs.CVcs.LG
keywords bimanualmanipulationtasksbenchmarkcodecoordinationlearningperact2
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Bimanual manipulation is challenging due to precise spatial and temporal coordination required between two arms. While there exist several real-world bimanual systems, there is a lack of simulated benchmarks with a large task diversity for systematically studying bimanual capabilities across a wide range of tabletop tasks. This paper addresses the gap by extending RLBench to bimanual manipulation. We open-source our code and benchmark comprising 13 new tasks with 23 unique task variations, each requiring a high degree of coordination and adaptability. To kickstart the benchmark, we extended several state-of-the art methods to bimanual manipulation and also present a language-conditioned behavioral cloning agent -- PerAct2, which enables the learning and execution of bimanual 6-DoF manipulation tasks. Our novel network architecture efficiently integrates language processing with action prediction, allowing robots to understand and perform complex bimanual tasks in response to user-specified goals. Project website with code is available at: http://bimanual.github.io

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RIO: Flexible Real-Time Robot I/O for Cross-Embodiment Robot Learning

    cs.RO 2026-05 unverdicted novelty 7.0

    RIO introduces a lightweight open-source framework that abstracts real-time robot I/O to support easy switching between embodiments and platforms for collecting data and deploying VLAs.

  2. DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

    cs.CV 2025-07 unverdicted novelty 6.0

    DreamVLA uses dynamic-region-guided world knowledge prediction, block-wise attention to disentangle information types, and a diffusion transformer for actions, reaching 76.7% success on real robot tasks and 4.44 avera...

  3. RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

    cs.RO 2025-07 unverdicted novelty 6.0

    RoboEval is a new benchmark providing eight bimanual tasks, thousands of expert demonstrations, and standardized metrics for efficiency, coordination, safety, and failure localization in robotic manipulation.

  4. 3D-DLP: Self-Supervised 3D Object-Centric Scene Representation Learning

    cs.LG 2026-06 unverdicted novelty 5.0

    3D-DLP decomposes 3D scenes into controllable latent particles via self-supervised reconstruction for improved robotic tasks.

  5. VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation

    cs.RO 2025-09 unverdicted novelty 5.0

    VLBiMan framework enables generalizable bimanual manipulation from single human demonstrations via vision-language anchored task decomposition and adaptation without retraining.

  6. ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation

    cs.RO 2025-09 unverdicted novelty 5.0

    ROPA augments bimanual imitation learning datasets by generating synthetic RGB-D observations and actions via fine-tuned diffusion models with physical consistency constraints.

  7. Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

    cs.RO 2025-08 unverdicted novelty 5.0

    This survey organizes large VLM-based VLA models for robotic manipulation into monolithic and hierarchical paradigms, reviews their integrations and datasets, and outlines future directions.

  8. JOIN: Anchor-Grasp-Conditioned Joining via Opposition, Inference, and Navigation for Bimanual Assistive Manipulation

    cs.RO 2026-06 unverdicted novelty 4.0

    JOIN decomposes bimanual joining into plan-drive-grasp phases and uses a VLM to let a mobile manipulator complete tasks with a pre-grasped anchor arm, achieving 19/20 success versus 14/20 for baselines on representative ADLs.