Pith. sign in

REVIEW 25 cited by

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14420 v2 pith:W4XUSKKN submitted 2025-02-20 cs.RO cs.CVcs.LG

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

classification cs.RO cs.CVcs.LG
keywords understandingchatvlacontrolmultimodalperformancerobottrainingunified
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in vision-language-action models (VLA), we identify two key challenges: spurious forgetting, where robot training overwrites crucial visual-text alignments, and task interference, where competing control and understanding tasks degrade performance when trained jointly. To overcome these limitations, we propose ChatVLA, a novel framework featuring Phased Alignment Training, which incrementally integrates multimodal data after initial control mastery, and a Mixture-of-Experts architecture to minimize task interference. ChatVLA demonstrates competitive performance on visual question-answering datasets and significantly surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks. Notably, it achieves a six times higher performance on MMMU and scores 47.2% on MMStar with a more parameter-efficient design than ECoT. Furthermore, ChatVLA demonstrates superior performance on 25 real-world robot manipulation tasks compared to existing VLA methods like OpenVLA. Our findings highlight the potential of our unified framework for achieving both robust multimodal understanding and effective robot control.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

    cs.LG 2026-04 unverdicted novelty 7.0

    π₀.₇ is a steerable generalist robotic model that uses rich multimodal prompts including language, subgoal images, and performance metadata to achieve out-of-the-box generalization across tasks and robot bodies.

  2. UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

    cs.RO 2026-02 unverdicted novelty 7.0

    UniLACT improves VLA models by adding depth-aware unified latent action pretraining that outperforms RGB-only baselines on seen and unseen manipulation tasks.

  3. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  4. Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

    cs.RO 2026-07 conditional novelty 6.0

    Preserving pretrained VLM features with layer-wise distillation plus supervising the language head on discretized action directions improves OOD generalization of VLA policies on LIBERO, CALVIN, and a real xArm7.

  5. OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform

    cs.RO 2026-06 conditional novelty 6.0

    OpenEAI-Platform delivers an open-source low-cost robotic arm and VLA model that outperforms commercial arms and matches large pretrained baselines on four real-world manipulation tasks using limited open data.

  6. VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving

    cs.CV 2026-05 unverdicted novelty 6.0

    VL-DPO uses a VLM as a zero-shot reasoner to generate preference pairs from pretrained model rollouts, then finetunes via DPO on the Waymo Open End-to-End Driving Dataset, yielding 11.94% higher rater feedback score a...

  7. Nautilus: From One Prompt to Plug-and-Play Robot Learning

    cs.RO 2026-05 conditional novelty 6.0

    A typed-contract harness with containerized 'chambers' and robotics-specific agent skills lets a coding LLM turn a single natural-language prompt into working reproduction, evaluation, and deployment workflows for rob...

  8. UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation

    cs.RO 2026-05 unverdicted novelty 6.0

    UniSteer unifies human corrective actions and noise-space RL for VLA adaptation by inverting actions to noise targets, raising success rates from 20% to 90% in 66 minutes across four real-world manipulation tasks.

  9. UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation

    cs.RO 2026-05 conditional novelty 6.0

    By inverting a frozen flow-matching robot policy with fixed-point iteration, UniSteer converts human corrections into noise-space supervision and lifts real-world VLA adaptation success from 20% to 90% in about 66 min...

  10. Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

    cs.RO 2025-12 conditional novelty 6.0

    A 6K-question embodied-reasoning benchmark plus a flow-matching action tokenizer let one 3B vision-language model reason and manipulate better than continuous- or discrete-action VLA baselines.

  11. VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

    cs.RO 2025-12 conditional novelty 6.0

    An open benchmark with 170 graded manipulation tasks shows current VLA robot policies memorize their training settings, degrade sharply under visual shifts, ignore safety constraints, and fail to compose long-horizon skills.

  12. Contrastive Representation Regularization for Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 6.0

    Adding a robot-state-aware contrastive loss to VLA training improves manipulation success on RoboCasa-Kitchen and real-robot tasks.

  13. FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 6.0

    A VLM trained on auto-generated failure trajectories with executable correction actions helps VLA models detect and fix manipulation errors, lifting success rates by up to 22.6 percentage points.

  14. LLaDA-VLA: Vision Language Diffusion Action Models

    cs.RO 2025-09 conditional novelty 6.0

    LLaDA-VLA applies a masked diffusion vision-language model to robot control with localized action-token classification and hierarchical decoding, achieving SOTA success rates on SimplerEnv, CALVIN, and real-robot tasks.

  15. DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

    cs.CV 2025-07 unverdicted novelty 6.0

    DreamVLA uses dynamic-region-guided world knowledge prediction, block-wise attention to disentangle information types, and a diffusion transformer for actions, reaching 76.7% success on real robot tasks and 4.44 avera...

  16. Nautilus: From One Prompt to Plug-and-Play Robot Learning

    cs.RO 2026-05 unverdicted novelty 5.0

    NAUTILUS is a prompt-driven harness that automates plug-and-play adapters, typed contracts, and validation for policies, benchmarks, and robots in learning research.

  17. ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning

    cs.RO 2026-04 unverdicted novelty 5.0

    ReFineVLA adds teacher-generated reasoning steps to VLA training and reports state-of-the-art success rates on SimplerEnv WidowX and Google Robot benchmarks.

  18. Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation

    cs.RO 2026-02 conditional novelty 5.0

    A bounding-box-conditioned diffusion policy shows a power-law improvement with the number of object classes in training data, reaching about 85% success on four semantic manipulation tasks.

  19. Causal World Modeling for Robot Control

    cs.CV 2026-01 unverdicted novelty 5.0

    LingBot-VA combines video world modeling with policy learning via Mixture-of-Transformers, closed-loop rollouts, and asynchronous inference to improve robot manipulation in simulation and real settings.

  20. RynnVLA-002: A Unified Vision-Language-Action and World Model

    cs.RO 2025-11 conditional novelty 5.0

    A single model that jointly predicts robot actions and future images outperforms separate action-only and video-only models on LIBERO and real SO100 manipulation tasks.

  21. Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

    cs.RO 2025-09 conditional novelty 5.0

    A plug-and-play fine-tuning method using two VAEs and a latent-distance guidance loss improves cross-embodiment and cross-task success rates of diffusion- and flow-based VLA policies.

  22. Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

    cs.RO 2025-08 unverdicted novelty 5.0

    This survey organizes large VLM-based VLA models for robotic manipulation into monolithic and hierarchical paradigms, reviews their integrations and datasets, and outlines future directions.

  23. SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning

    cs.RO 2025-03 unverdicted novelty 5.0

    SafeVLA applies constrained reinforcement learning via CMDP min-max optimization to VLAs, cutting safety violation costs by 83.58% while preserving task success on long-horizon mobile manipulation tasks.

  24. General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

    cs.CV 2026-05 unverdicted novelty 4.0

    GAM framework uses arc-length parameterization for temporal invariance and schema-affine factorization for geometric invariance to build a covariant action manifold integrated into VLA models for improved generalizati...

  25. Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges

    cs.RO 2025-08 conditional novelty 4.0

    A survey that taxonomizes robotic manipulation policies trained by imitation learning, traces their evolution, and compiles benchmark comparisons.