REVIEW 17 cited by
Unified Vision-Language-Action Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on the general comprehension capabilities of vision-language models (VLMs) to generate action signals, often overlooking the rich temporal and causal structure embedded in visual observations. In this paper, we present UniVLA, a unified and native multimodal VLA model that autoregressively models vision, language, and action signals as discrete token sequences. This formulation enables flexible multimodal tasks learning, particularly from large-scale video data. By incorporating world modeling during post-training, UniVLA captures causal dynamics from videos, facilitating effective transfer to downstream policy learning--especially for long-horizon tasks. Our approach sets new state-of-the-art results across several widely used simulation benchmarks, including CALVIN, LIBERO, and Simplenv-Bridge, significantly surpassing previous methods. For example, UniVLA achieves 95.5% average success rate on LIBERO benchmark, surpassing pi0-FAST's 85.5%. We further demonstrate its broad applicability on real-world ALOHA manipulation and autonomous driving.
Forward citations
Cited by 17 Pith papers
-
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
MobileManiBench is a 300K-trajectory, multi-robot, multi-camera simulated benchmark for VLA model training and evaluation in mobile manipulation.
-
TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models
A two-timescale RL post-training method that updates the semantic projection layer rarely and the action expert often improves VLA policy success on long-horizon manipulation tasks.
-
CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
CoTinyVLA, a 0.9B vision-language-action model, outperforms 3-7B baselines on all four LIBERO-Plus robustness suites.
-
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
LaMem-VLA reconstructs robotic history into dual short-term and long-term latent memory tokens that are woven directly into a VLA model's reasoning sequence to improve long-horizon manipulation.
-
Learning 4D Geometric Priors for Inference-Efficient World Action Models
Training-only multi-expert co-training with decayed 4D read-mask attention and action-aware temporal geometric distillation improves WAM manipulation success while keeping the original lightweight inference graph.
-
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI
A single 8B backbone unifies spatial perception, decision making, navigation/manipulation, and progress estimation with SSR+ merging, reporting gains on most spatial benchmarks and competitive action/progress results.
-
Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments
Fluent expert demonstrations under-supervise the short alignment phase that decides success, and a compact spatio-temporal dynamic feature (STAIR) recovers most of the deliberate-demonstration gain from fluent data alone.
-
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
GTA-VLA lets humans steer a robot policy with spatial cues (points, boxes, traces) that condition the model's visual chain-of-thought, improving OOD robustness and recovering about 20% of failed episodes.
-
History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation
History-conditioned A-MMR pruning of current and past visual tokens beats prior training-free pruners on R2R/RxR at 70–90% drop rates and runs onboard a Unitree Go2.
-
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models
NIAF turns robot action chunks into a continuous SIREN function modulated by a VLM, enabling analytic velocity/jerk supervision and state-of-the-art CALVIN/LIBERO results.
-
Mixture of Horizons in Action Chunking
A gated mixture of multiple action-chunk horizons in a shared full-attention transformer improves VLA manipulation success and enables consensus-based early stopping.
-
LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes
A staged teacher-student pipeline lets a simulated humanoid relocate two objects in sequence without resets, from egocentric RGB and language, beating the single-task baseline on 350 training and 66 unseen layouts.
-
LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding
A local cross-layer routing mechanism for VLA models, feeding each action-decoder block features from a small window of adjacent VLM layers, improves success rates on LIBERO, CALVIN, and LIBERO-Plus.
-
Native Video-Action Pretraining for Generalizable Robot Control
A video-action foundation model pretrained natively with a causal diffusion transformer and semantic visual-action tokenizer reports improved few-shot robot manipulation and 225 Hz asynchronous closed-loop control.
-
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
VLAs fail most counterfactual instructions because vision shortcuts dominate language; the new LIBERO-CF benchmark quantifies this, and CAG, an inference-time action mixer, improves grounding and success.
-
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
A survey that groups VLA efficiency techniques into four categories: model architecture, perception features, action generation, and training/inference strategies.
-
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
Adding a reconstruction target that redraws the object region makes a vision-language-action model focus its attention on the right object and manipulate more precisely.
Discussion (0). Continue with ORCID to comment.