Long-VLA uses phase-aware camera masking and a phase token to let a single end-to-end diffusion VLA handle 10-step manipulation, reporting large gains on its own L-CALVIN benchmark and on real-world tasks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
Long-VLA uses phase-aware camera masking and a phase token to let a single end-to-end diffusion VLA handle 10-step manipulation, reporting large gains on its own L-CALVIN benchmark and on real-world tasks.