An MLLM trained with auxiliary goal-prediction tasks and multi-token prediction achieves SOTA on COIN and CrossTask visual planning and matches SOTA on Ego4D LTA.
Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction
An MLLM trained with auxiliary goal-prediction tasks and multi-token prediction achieves SOTA on COIN and CrossTask visual planning and matches SOTA on Ego4D LTA.