RoboMonkey shows that test-time sampling with Gaussian perturbation and a VLM-based action verifier improves the success rate of vision-language-action models on manipulation tasks.
Neural Scaling Laws in Robotics
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Neural scaling laws have driven significant advancements in machine learning, particularly in domains like language modeling and computer vision. However, the exploration of neural scaling laws within robotics has remained relatively underexplored, despite the growing adoption of foundation models in this field. This paper represents the first comprehensive study to quantify neural scaling laws for Robot Foundation Models (RFMs) and Large Language Models (LLMs) in robotics tasks. Through a meta-analysis of 327 research papers, we investigate how data size, model size, and compute resources influence downstream performance across a diverse set of robotic tasks. Consistent with previous scaling law research, our results reveal that the performance of robotic models improves with increased resources, following a power-law relationship. Promisingly, the improvement in robotic task performance scales notably faster than language tasks. This suggests that, while performance on downstream robotic tasks today is often moderate-to-poor, increased data and compute are likely to signficantly improve performance in the future. Also consistent with previous scaling law research, we also observe the emergence of new robot capabilities as models scale.
citation-role summary
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
RoboMonkey shows that test-time sampling with Gaussian perturbation and a VLM-based action verifier improves the success rate of vision-language-action models on manipulation tasks.