Pith. sign in

REVIEW 19 cited by

ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02292 v1 pith:YD4CXOW3 submitted 2024-02-07 cs.RO cs.LG

ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation

classification cs.RO cs.LG
keywords alohahardwarebimanualenhancedrobustnessteleoperationaccelerateadvances
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Diverse demonstration datasets have powered significant advances in robot learning, but the dexterity and scale of such data can be limited by the hardware cost, the hardware robustness, and the ease of teleoperation. We introduce ALOHA 2, an enhanced version of ALOHA that has greater performance, ergonomics, and robustness compared to the original design. To accelerate research in large-scale bimanual manipulation, we open source all hardware designs of ALOHA 2 with a detailed tutorial, together with a MuJoCo model of ALOHA 2 with system identification. See the project website at aloha-2.github.io.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Dexora: Open-source VLA for High-DoF Bimanual Dexterity

    cs.RO 2026-05 unverdicted novelty 7.0

    Dexora is the first open-source VLA system for dual-arm dual-hand high-DoF manipulation, trained on 100K simulated and 10K real teleoperated trajectories with a discriminator-weighted diffusion policy, achieving 66.7%...

  2. BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination

    cs.RO 2026-04 conditional novelty 7.0

    BiCoord is a new benchmark for long-horizon tightly coordinated bimanual manipulation that includes quantitative metrics and shows existing policies like DP, RDT, Pi0 and OpenVLA-OFT struggle on such tasks.

  3. TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance

    cs.RO 2026-01 unverdicted novelty 7.0

    TouchGuide improves contact-rich robot manipulation by steering diffusion or flow-matching visuomotor policies with tactile feasibility scores from a contrastively trained Contact Physical Model.

  4. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  5. IOI: Decoupling Kinematics and Physics for Interactive World Models

    cs.RO 2026-06 unverdicted novelty 6.0

    IOI decouples deterministic kinematics from stochastic physics in interactive world models by rendering forward-kinematics trajectories into multi-view projections that guide a video generator, achieving SOTA fidelity...

  6. MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    MotionWAM conditions a policy on intermediate features from a video world model to predict unified whole-body motion tokens, enabling real-time humanoid loco-manipulation that outperforms VLA baselines by over 30% on ...

  7. Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

    cs.RO 2026-05 unverdicted novelty 6.0

    Retrieve-then-steer stores successful observation-action segments in memory, retrieves relevant chunks, filters them, and uses an elite prior with confidence-adaptive guidance to steer a flow-matching action sampler f...

  8. Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

    cs.RO 2026-05 unverdicted novelty 6.0

    A retrieve-then-steer method stores successful robot actions in memory and uses them to steer a frozen VLA's flow-matching sampler for better test-time reliability without parameter updates.

  9. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

    cs.RO 2025-03 unverdicted novelty 6.0

    GR00T N1 is a new open VLA foundation model for humanoid robots that outperforms imitation learning baselines in simulation and shows strong performance on real-world bimanual manipulation tasks.

  10. $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

    cs.LG 2024-10 unverdicted novelty 6.0

    π₀ is a vision-language-action flow model trained on diverse multi-platform robot data that supports zero-shot task performance, language instruction following, and efficient fine-tuning for dexterous tasks.

  11. MEVION: Low-Cost Open-Source Data Collection System for Powerful and High-Speed Dual-Arm Manipulation

    cs.RO 2026-07 conditional novelty 5.0

    MEVION is an open-source dual-arm teleoperation platform with 60 Nm joint torque, built from e-commerce parts for about $14,000 per four-arm system, enabling heavier, faster manipulation data collection.

  12. FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

    cs.CV 2026-06 unverdicted novelty 5.0

    FOCA improves few-shot VLA adaptation by explicitly predicting future interaction embeddings and implicitly aligning to goal observations, yielding up to 26% gains on real robots with only 20 demonstrations.

  13. GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments

    cs.RO 2026-06 unverdicted novelty 5.0

    GASE automates high-fidelity simulation scene reconstruction from multi-view panoramic videos via Gaussian splatting, object extraction, and inpainting, yielding robot policies with under 10% performance gap versus re...

  14. HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

    cs.RO 2026-05 unverdicted novelty 5.0

    HumanEgo reports 92.5% average success on four real robot tasks using only 15-30 minutes of human video per task and zero robot data, with zero-shot transfer to new robots and cameras.

  15. VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation

    cs.RO 2025-09 unverdicted novelty 5.0

    VLBiMan framework enables generalizable bimanual manipulation from single human demonstrations via vision-language anchored task decomposition and adaptation without retraining.

  16. DexTeleop-0: Force-Aware Bimanual Dexterous Teleoperation with Ego-Centric Perception towards Shared Autonomy

    cs.RO 2026-06 unverdicted novelty 4.0

    DexTeleop-0 adds a tactile-driven adaptation loop to bimanual dexterous teleoperation that estimates contact points and applies localized force-compliant corrections via operational-space Jacobian updates.

  17. A Reproducible and Physically Feasible Dynamic Parameter Identification Framework for a Low-Cost Robot Arm

    cs.RO 2026-05 unverdicted novelty 4.0

    A pipeline reduces a robot arm's rigid-body parameters from 65 to 39 via symmetry, fits them with OLS+SDP+CLIE on hand-designed trajectories, selects a central model via PCA, and audits inertia positive-definiteness t...

  18. A Reproducible and Physically Feasible Dynamic Parameter Identification Framework for a Low-Cost Robot Arm

    cs.RO 2026-05 conditional novelty 4.0

    A pipeline that reduces a robot arm's dynamic model to 39 parameters, fits them via OLS on hand-designed motions, projects to feasible values with SDP, refines with CLIE, and selects a central physically valid model f...

  19. TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

    cs.RO 2024-09 unverdicted novelty 4.0

    TinyVLA achieves faster inference and higher data efficiency than OpenVLA on robotic manipulation tasks by initializing from high-speed multimodal models and adding a diffusion policy decoder, without any pre-training phase.