OE-VLA extends vision-language-action models to follow open-ended instructions embedded in images, videos, and goal snapshots, matching text-only performance on the CALVIN benchmark.
Visual Instruction Tuning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
OE-VLA extends vision-language-action models to follow open-ended instructions embedded in images, videos, and goal snapshots, matching text-only performance on the CALVIN benchmark.