{"total":16,"items":[{"citing_arxiv_id":"2607.08317","ref_index":72,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models","primary_cat":"cs.AI","submitted_at":"2026-07-09T09:56:50+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A 235-item multimodal stress-test shows frontier closed models outpace open-weight peers by ~10% and leaves shared failures on counting, spatial, and character-level tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07251","ref_index":37,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models","primary_cat":"cs.CL","submitted_at":"2026-07-08T10:34:54+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Open vision-language models fail to select spatial demonstratives based on object distance in a human-like manner across four languages.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06165","ref_index":25,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"EAGOR: Embodied Reasoning in Omni-direction","primary_cat":"cs.RO","submitted_at":"2026-07-07T11:39:57+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":7.0,"formal_verification":"none","one_line_summary":"EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.31285","ref_index":67,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Spatial Reasoning via Modality Switching Between Language and Symbolic Representation","primary_cat":"cs.AI","submitted_at":"2026-06-30T08:02:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Introduces a modality-switching mechanism for LLMs on spatial reasoning tasks using a trustworthiness and complexity based metric, showing up to 42% performance improvement.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.11719","ref_index":34,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning","primary_cat":"cs.CV","submitted_at":"2026-06-10T06:49:21+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A self-evolving training loop that generates its own spatial QA data with executable code and difficulty feedback lifts Qwen3-VL-4B/8B to 62.7/63.3 on VSI-Bench using an order of magnitude less data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.09669","ref_index":45,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks","primary_cat":"cs.AI","submitted_at":"2026-06-08T15:51:51+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SpatialWorld is a new multi-simulator benchmark showing top multimodal agents achieve under 18% success on interactive spatial tasks requiring active exploration and long-horizon planning.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.31148","ref_index":13,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes","primary_cat":"cs.CV","submitted_at":"2026-05-29T10:59:26+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SpatialAct benchmark shows VLMs handle isolated spatial reasoning but fail to maintain coherent spatial beliefs and produce reliable actions in multi-turn 3D interactions, underperforming humans.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.28490","ref_index":25,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs","primary_cat":"cs.CV","submitted_at":"2026-05-27T13:45:34+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SSR3D-LLM improves fine-grained 3D grounding in unified 3D-LLMs by generating and scoring sequences of latent spatial reasoning steps from the query using fixed Mask3D proposals.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22570","ref_index":52,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis","primary_cat":"cs.CV","submitted_at":"2026-05-21T14:48:35+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"VGenST-Bench is a new video benchmark for MLLM spatio-temporal reasoning built via generative synthesis, a multi-agent pipeline with human oversight, a 3x2x2 taxonomy, and hierarchical tasks separating perception from reasoning.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"[50] Xiongkun Linghu, Jiangyong Huang, Xuesong Niu, Xiaojian Ma, Baoxiong Jia, and Siyuan Huang. Multi-modal situated reasoning in 3d scenes.Advances in Neural Information Process- ing Systems, 37:140903-140936, 2024. [51] Fangyu Liu, Guy Emerson, and Nigel Collier. Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635-651, 2023. [52] Weichen Liu, Qiyao Xue, Haoming Wang, Xiangyu Yin, Boyuan Yang, and Wei Gao. Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods. arXiv preprint arXiv:2511.15722, 2025. [53] Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around"},{"citing_arxiv_id":"2605.22100","ref_index":65,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing","primary_cat":"cs.AI","submitted_at":"2026-05-21T07:36:41+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MPDocBench-Parse provides 433 annotated multi-page documents and an evaluation protocol covering text/table/formula extraction, merging, figure extraction, reading order, and heading hierarchy for realistic document parsing.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.20837","ref_index":22,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models","primary_cat":"cs.CV","submitted_at":"2026-05-20T07:27:08+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"ArchSIBench is a new benchmark dataset and evaluation suite that measures vision-language models on architectural spatial intelligence across 17 subtasks, showing most models lag human baselines especially in transformation and configuration.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.13169","ref_index":29,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PanoWorld: Towards Spatial Supersensing in 360$^\\circ$ Panorama World","primary_cat":"cs.CV","submitted_at":"2026-05-13T08:31:22+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"PanoWorld adds spherical spatial cross-attention and pano-native training data to MLLMs for improved spatial reasoning on ERP panoramas, outperforming baselines on new and existing benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.10106","ref_index":30,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models","primary_cat":"cs.CV","submitted_at":"2026-05-11T07:20:09+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"ViSRA boosts MLLM 3D spatial reasoning performance by up to 28.9% on unseen tasks via a plug-and-play video-based agent that extracts explicit spatial cues from expert models without any post-training.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding.arXiv preprint arXiv:2507.07984, 2025. [29] Benlin Liu, Yuhao Dong, Yiqin Wang, Zixian Ma, Yansong Tang, Luming Tang, Yongming Rao, Wei-Chiu Ma, and Ranjay Krishna. Coarse correspondences boost spatial-temporal reasoning in multimodal language model. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 3783-3792, 2025. [30] Weichen Liu, Qiyao Xue, Haoming Wang, Xiangyu Yin, Boyuan Yang, and Wei Gao. Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods.arXiv preprint arXiv:2511.15722, 2025. [31] Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models."},{"citing_arxiv_id":"2605.07148","ref_index":10,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models","primary_cat":"cs.CV","submitted_at":"2026-05-08T02:32:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"VLMs possess a latent 3D scene topology subspace corresponding to Laplacian eigenmaps that can be causally shaped via Dirichlet energy regularization to improve spatial task performance by up to 12.1%.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"but precise spatial understanding and reasoning. Nowadays, Vision-Language Models (VLMs) [5, 6] can succeed on simple spatial reasoning tasks given only egocentric visual inputs instead of dedicated 3D geometric information (e.g., depth estimation) [ 7, 8, 9], but they fail on more difficult tasks in complex scenes with many objects, partial views and cluttered layouts [10, 11]. The natural question, then, is whether VLMs have topology- Preprint. arXiv:2605.07148v1 [cs.CV] 8 May 2026 Egocentric visual input VLM Activation of tokens from different objects Probing of latent representation Recovered scene topology Ground truth topology Diagnostic signal Improving VLM's ability in spatial perception and reasoning"},{"citing_arxiv_id":"2605.01333","ref_index":2,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice","primary_cat":"cs.CL","submitted_at":"2026-05-02T09:08:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"OralMLLM-Bench reveals performance gaps between multimodal large language models and clinicians on cognitive tasks for dental radiographic analysis across periapical, panoramic, and cephalometric images.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.22409","ref_index":45,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments","primary_cat":"cs.CV","submitted_at":"2026-04-24T10:06:41+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SpaMEM is a diagnostic benchmark showing that current vision-language models exhibit a sharp collapse in spatial reasoning when transitioning from text-aided state tracking to purely visual memory in dynamic environments.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"existing simulators, while supporting active movement, operate instatic envi- ronmentswhere the spatial layout remains immutable throughout an episode [9, 10, 13, 16, 27, 28]. This allows models to heavily rely on \"statistical co- occurrence biases\"-predicting object locations based on 2D semantic priors rather than genuine 3D geometric understanding [45]. SpaMEM diverges from these conventions by introducingdynamic scene evolution. By performing in- tentional object manipulations-such asspawn,place, andremove-we create a fluid environment that forces models to constantly revise their internal spatial beliefs. This setting is closely related to rearrangement-style embodied tasks that explicitly change object poses/states within an episode [29, 33, 37], and com-"}],"limit":50,"offset":0}