Space Tokens distill 3D reconstruction and object bounding box knowledge into continuous latent tokens that a VLM can reason over, improving VSI-Bench scores, though the RL stage's evaluation independence is not established.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models
Space Tokens distill 3D reconstruction and object bounding box knowledge into continuous latent tokens that a VLM can reason over, improving VSI-Bench scores, though the RL stage's evaluation independence is not established.