OpenHelix shows that a frozen vision-language model with a prompt-tuned token and an auxiliary action-prediction head beats full fine-tuning on CALVIN language generalization while training far fewer parameters.
Minivla: A better vla with a smaller footprint
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
OpenHelix shows that a frozen vision-language model with a prompt-tuned token and an auxiliary action-prediction head beats full fine-tuning on CALVIN language generalization while training far fewer parameters.