ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training

Chiming Liu; Chuheng Zhang; Hecheng Wang; Lizhe Qi; Maoqing Yao; Rushuai Yang; Shuoyu Yue; Wei Shan; Xiaohan Yan; Xuan Du

arxiv: 2602.12691 · v3 · pith:AZMN26M3new · submitted 2026-02-13 · 💻 cs.RO · cs.AI

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training

Rushuai Yang , Hecheng Wang , Zhichao Wu , Chiming Liu , Xiaohan Yan , Xuan Du , Shuoyu Yue , Chuheng Zhang

show 6 more authors

Yunlong Wang Yongcheng Liu Lizhe Qi Yi Chen Wei Shan Maoqing Yao

This is my paper

classification 💻 cs.RO cs.AI

keywords aloevaluelearningpolicypost-trainingreal-worldevaluationoff-policy

0 comments

read the original abstract

We study how to improve large foundation vision-language-action (VLA) systems through human-in-the-loop reinforcement learning (RL) in real-world environments. A key challenge is learning reliable value functions from heterogeneous real-world experience, as value estimation provides the primary learning signal for VLA training. In practice, replay buffers contain trajectories collected from historical policies, online rollouts, demonstrations, and intermittent human interventions. Because replay buffers mix trajectories generated by different behaviors, the observed returns can be mismatched with the quality of the current policy. Prior VLA post-training methods often rely on progress-style value signals, which reflect the average quality of historical behaviors, leading to mismatched learning signals for the current policy. In this paper, we propose ALOE, an off-policy evaluation framework whose value function directly evaluates current-policy behavior for each iteration. Specifically, ALOE combines chunked temporal-difference bootstrapping and conservative value aggregation to perform stable current-policy evaluation, then uses these estimates for advantage-weighted policy improvement. This design improves credit assignment to critical action chunks under sparse rewards and supports stable policy improvement. We evaluate ALOE on four real-world manipulation tasks encompassing long-horizon and high-precision scenarios: smartphone packing, laundry folding, multi-object sorting, and phone assembly. Across all tasks, ALOE outperforms other VLA post-training methods, highlighting the benefit of off-policy value estimates for real-world VLA post-training. Videos are available at our project website https://rooshy-yang.github.io/aloe.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Unified Noise Steering for Efficient Human-Guided VLA Adaptation
cs.RO 2026-05 unverdicted novelty 6.0

UniSteer unifies human corrective actions and noise-space RL for VLA adaptation by inverting actions to noise targets, raising success rates from 20% to 90% in 66 minutes across four real-world manipulation tasks.