Pith. sign in

REVIEW 17 cited by

Unified Vision-Language-Action Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.19850 v1 pith:MAJZ2LYI submitted 2025-06-24 cs.CV cs.RO

classification cs.CVcs.RO
keywords modelsunivlaactioncausalliberomanipulationmodelmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on the general comprehension capabilities of vision-language models (VLMs) to generate action signals, often overlooking the rich temporal and causal structure embedded in visual observations. In this paper, we present UniVLA, a unified and native multimodal VLA model that autoregressively models vision, language, and action signals as discrete token sequences. This formulation enables flexible multimodal tasks learning, particularly from large-scale video data. By incorporating world modeling during post-training, UniVLA captures causal dynamics from videos, facilitating effective transfer to downstream policy learning--especially for long-horizon tasks. Our approach sets new state-of-the-art results across several widely used simulation benchmarks, including CALVIN, LIBERO, and Simplenv-Bridge, significantly surpassing previous methods. For example, UniVLA achieves 95.5% average success rate on LIBERO benchmark, surpassing pi0-FAST's 85.5%. We further demonstrate its broad applicability on real-world ALOHA manipulation and autonomous driving.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MobileManiBench: Simplifying Model Verification for Mobile Manipulation

    cs.RO 2026-02 conditional novelty 7.0 of 10

    MobileManiBench is a 300K-trajectory, multi-robot, multi-camera simulated benchmark for VLA model training and evaluation in mobile manipulation.

  2. TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A two-timescale RL post-training method that updates the semantic projection layer rarely and the action expert often improves VLA policy success on long-horizon manipulation tasks.

  3. CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

    cs.AI 2026-07 conditional novelty 6.0 of 10

    CoTinyVLA, a 0.9B vision-language-action model, outperforms 3-7B baselines on all four LIBERO-Plus robustness suites.

  4. Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    LaMem-VLA reconstructs robotic history into dual short-term and long-term latent memory tokens that are woven directly into a VLA model's reasoning sequence to improve long-horizon manipulation.

  5. Learning 4D Geometric Priors for Inference-Efficient World Action Models

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Training-only multi-expert co-training with decayed 4D read-mask attention and action-aware temporal geometric distillation improves WAM manipulation success while keeping the original lightweight inference graph.

  6. ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A single 8B backbone unifies spatial perception, decision making, navigation/manipulation, and progress estimation with SSR+ merging, reporting gains on most spatial benchmarks and competitive action/progress results.

  7. Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments

    cs.RO 2026-06 conditional novelty 6.0 of 10

    Fluent expert demonstrations under-supervise the short alignment phase that decides success, and a compact spatio-temporal dynamic feature (STAIR) recovers most of the deliberate-demonstration gain from fluent data alone.

  8. Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

    cs.RO 2026-05 conditional novelty 6.0 of 10

    GTA-VLA lets humans steer a robot policy with spatial cues (points, boxes, traces) that condition the model's visual chain-of-thought, improving OOD robustness and recovering about 20% of failed episodes.

  9. History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation

    cs.RO 2026-03 conditional novelty 6.0 of 10

    History-conditioned A-MMR pruning of current and past visual tokens beats prior training-free pruners on R2R/RxR at 70–90% drop rates and runs onboard a Unitree Go2.

  10. Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

    cs.RO 2026-03 conditional novelty 6.0 of 10

    NIAF turns robot action chunks into a continuous SIREN function modulated by a VLM, enabling analytic velocity/jerk supervision and state-of-the-art CALVIN/LIBERO results.

  11. Mixture of Horizons in Action Chunking

    cs.RO 2025-11 conditional novelty 6.0 of 10

    A gated mixture of multiple action-chunk horizons in a shared full-attention transformer improves VLA manipulation success and enables consensus-based early stopping.

  12. LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes

    cs.RO 2025-08 reject novelty 6.0 of 10

    A staged teacher-student pipeline lets a simulated humanoid relocate two objects in sequence without resets, from egocentric RGB and language, beating the single-task baseline on 350 training and 66 unseen layouts.

  13. LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding

    cs.RO 2026-08 conditional novelty 5.0 of 10

    A local cross-layer routing mechanism for VLA models, feeding each action-decoder block features from a small window of adjacent VLM layers, improves success rates on LIBERO, CALVIN, and LIBERO-Plus.

  14. Native Video-Action Pretraining for Generalizable Robot Control

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A video-action foundation model pretrained natively with a causal diffusion transformer and semantic visual-action tokenizer reports improved few-shot robot manipulation and 225 Hz asynchronous closed-loop control.

  15. When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

    cs.CV 2026-02 conditional novelty 5.0 of 10

    VLAs fail most counterfactual instructions because vision shortcuts dominate language; the new LIBERO-CF benchmark quantifies this, and CAG, an inference-time action mixer, improves grounding and success.

  16. Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey

    cs.RO 2025-10 conditional novelty 4.0 of 10

    A survey that groups VLA efficiency techniques into four categories: model architecture, perception features, action generation, and training/inference strategies.

  17. ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

    cs.RO 2025-08 unverdicted novelty 4.0 of 10

    Adding a reconstruction target that redraws the object region makes a vision-language-action model focus its attention on the right object and manipulate more precisely.

Pith tools