Pith. sign in

REVIEW 19 cited by

Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08001 v3 pith:32JR5VRR submitted 2024-10-10 cs.RO cs.AI

classification cs.ROcs.AI
keywords generalistpolicyspecialistdatarobodualactiondual-systemhigh-level
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The increasing demand for versatile robotic systems to operate in diverse and dynamic environments has emphasized the importance of a generalist policy, which leverages a large cross-embodiment data corpus to facilitate broad adaptability and high-level reasoning. However, the generalist would struggle with inefficient inference and cost-expensive training. The specialist policy, instead, is curated for specific domain data and excels at task-level precision with efficiency. Yet, it lacks the generalization capacity for a wide range of applications. Inspired by these observations, we introduce RoboDual, a synergistic dual-system that supplements the merits of both generalist and specialist policy. A diffusion transformer-based specialist is devised for multi-step action rollouts, exquisitely conditioned on the high-level task understanding and discretized action output of a vision-language-action (VLA) based generalist. Compared to OpenVLA, RoboDual achieves 26.7% improvement in real-world setting and 12% gain on CALVIN by introducing a specialist policy with merely 20M trainable parameters. It maintains strong performance with 5% of demonstration data only, and enables a 3.8 times higher control frequency in real-world deployment. Code would be made publicly available. Our project page is hosted at: https://opendrivelab.com/RoboDual/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A compact latent interaction policy trained with bidirectional masked trajectory prediction matches or exceeds large VLA and world-action-model baselines on robot manipulation benchmarks while using far less compute.

  2. TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A two-timescale RL post-training method that updates the semantic projection layer rarely and the action expert often improves VLA policy success on long-horizon manipulation tasks.

  3. Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection

    cs.RO 2026-08 conditional novelty 6.0 of 10

    EMS, a dual-system VLA framework with RL-trained switching, achieves near-large-model success on LIBERO at high effective command rate.

  4. Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Aligning a VLA's latent features with instruction-selected target-object tri-views (VAE and VGGT) improves manipulation success, especially under target occlusion, with a compact 345M backbone.

  5. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  6. Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Preserving pretrained VLM features with layer-wise distillation plus supervising the language head on discretized action directions improves OOD generalization of VLA policies on LIBERO, CALVIN, and a real xArm7.

  7. UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Uncertainty-aware observation reinjection into FFN layers improves VLA manipulation success rates across LIBERO, SIMPLER, CALVIN, and real-robot tasks with no training.

  8. FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A compact 950-million-parameter robot policy trained in about 200 GPU-hours matches or beats multi-billion-parameter baselines on most manipulation benchmarks, including a new best score on CALVIN ABC.

  9. ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An asynchronous dual-system architecture forecasts large-model features into the current frame and adds a small-model update to drive in near real time, scoring 69.53 on Bench2Drive at 50 ms.

  10. ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

    cs.RO 2025-05 conditional novelty 6.0 of 10

    ChatVLA-2 uses dynamic mixture-of-experts and a two-stage training recipe to let a vision-language-action model retain pretrained reasoning while following robot instructions.

  11. Hume: Introducing System-2 Thinking in Visual-Language-Action Model

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A dual-system vision-language-action model that improves robot control by ranking multiple sampled action chunks with a learned value function before fast execution.

  12. WorldEval: World Model as Real-World Robot Policies Evaluator

    cs.RO 2025-05 conditional novelty 6.0 of 10

    WorldEval conditions a video generation model on a policy's internal action embeddings (Policy2Vec) and shows generated-video success rates correlate with real-world robot success rates.

  13. Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A robot policy generates its own language reasoning before acting and injects it into a diffusion action decoder, outperforming several VLA baselines on real-robot manipulation.

  14. TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A 0.5B VLA with bridge-conditioned discrete diffusion and 2D temporal–spatial action masking reaches 95.7% LIBERO success and 4.19 CALVIN average length.

  15. Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A dual-system framework with a structured subtask interface, event-balanced training, and inference harness enables VLM-guided long-horizon robotic manipulation, achieving 95.5% on LIBERO-Long and 65% on real-world ch...

  16. EnerVerse-AC: Envisioning Embodied Environments with Action Condition

    cs.RO 2025-05 conditional novelty 5.0 of 10

    EnerVerse-AC generates realistic multi-view robot videos conditioned on action sequences and shows early evidence it can augment training data and rank policy performance like a real robot.

  17. OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

    cs.RO 2025-05 conditional novelty 5.0 of 10

    OpenHelix shows that a frozen vision-language model with a prompt-tuned token and an auxiliary action-prediction head beats full fine-tuning on CALVIN language generalization while training far fewer parameters.

  18. StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation

    cs.RO 2026-02 reject novelty 4.0 of 10

    StemVLA supervises a GPT-2-based VLA with predicted future 3D-geometry features (VGGT) and temporally aggregated history, reporting 86.0% on LIBERO-Long - but its CALVIN results and equations are placeholders.

  19. Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey

    cs.RO 2025-10 conditional novelty 4.0 of 10

    A survey that groups VLA efficiency techniques into four categories: model architecture, perception features, action generation, and training/inference strategies.

Pith tools