Pith. sign in

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The ability to efficiently and reliably learn new tasks has been a foundational challenge in robotics. Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse manipulation tasks, yet pretrained policies consistently fall short of the reliability required for real-world deployment. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches either train from scratch without fully leveraging pretrained priors, or fine-tune VLAs without achieving the sample efficiency and success rates that practical deployment demands. We present EXPO-FT, a system for stable, sample-efficient RL finetuning of pretrained VLA policies that closes this gap. Our system solves a suite of challenging manipulation tasks, including routing string lights and inserting the plug to light it up, striking a pool ball into a pocket, and inserting a flower into a wine bottle, each requiring combinations of high precision, dynamic actions, and robustness to varied initial states. Our system achieves perfect task performance (30/30 successes) across all evaluated tasks within an average of 19.1 minutes of online robot data, outperforming both prior RL-from-scratch and VLA finetuning approaches. We release an open-source codebase with the aim of facilitating broader adoption of RL finetuning of VLA models in robotics.

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

cs.LG · 2026-07-29 · conditional · novelty 6.0

On six robot-manipulation tasks, offline Q-pretraining does not accelerate online RL fine-tuning from a pretrained policy, while seeding the replay buffer with rollouts from an ensemble of policies (IPE) improves final performance by about 26%.

citing papers explorer

Showing 1 of 1 citing paper.

  • Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cs.LG · 2026-07-29 · conditional · none · ref 5 · internal anchor

    On six robot-manipulation tasks, offline Q-pretraining does not accelerate online RL fine-tuning from a pretrained policy, while seeding the replay buffer with rollouts from an ensemble of policies (IPE) improves final performance by about 26%.