Pith. sign in

ROSA: Harnessing Robot States for Vision-Language and Action Alignment

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Vision-Language-Action (VLA) models have recently made significant advance in multi-task, end-to-end robotic control, due to the strong generalization capabilities of Vision-Language Models (VLMs). A fundamental challenge in developing such models is effectively aligning the vision-language space with the robotic action space. Existing approaches typically rely on directly fine-tuning VLMs using expert demonstrations. However, this strategy suffers from a spatio-temporal gap, resulting in considerable data inefficiency and heavy reliance on human labor. Spatially, VLMs operate within a high-level semantic space, whereas robotic actions are grounded in low-level 3D physical space; temporally, VLMs primarily interpret the present, while VLA models anticipate future actions. To overcome these challenges, we propose a novel training paradigm, ROSA, which leverages robot state estimation to improve alignment between vision-language and action spaces. By integrating robot state estimation data obtained via an automated process, ROSA enables the VLA model to gain enhanced spatial understanding and self-awareness, thereby boosting performance and generalization. Extensive experiments in both simulated and real-world environments demonstrate the effectiveness of ROSA, particularly in low-data regimes.

fields

cs.RO 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

How Should Vision-Language-Action Models Use Proprioceptive State?

cs.RO · 2026-08-04 · conditional · novelty 6.0

In a flow-matching VLA, single-frame proprioceptive state works best in the VLM prefix, while 8-frame state history works best in the action prefix, with gains from genuine temporal content rather than extra tokens.

citing papers explorer

Showing 1 of 1 citing paper.

  • How Should Vision-Language-Action Models Use Proprioceptive State? cs.RO · 2026-08-04 · conditional · none · ref 2020 · internal anchor

    In a flow-matching VLA, single-frame proprioceptive state works best in the VLM prefix, while 8-frame state history works best in the action prefix, with gains from genuine temporal content rather than extra tokens.