A reinforcement learning framework that shares a decomposed critic between arm and gripper policies improves real-world pick-and-place success rates by 20 to 70 percentage points over a strong baseline.
Improving vision-language-action model with online reinforcement learning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition
A reinforcement learning framework that shares a decomposed critic between arm and gripper policies improves real-world pick-and-place success rates by 20 to 70 percentage points over a strong baseline.