← back to paper
arxiv: 2608.01899 · 2 revisions
SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models