Pith. sign in

Paper Citation Record · LEDGER

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

As of 12 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 8 inbound Pith citation observations for arXiv:2509.02359.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02359 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:40:25.147462Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:33:30.837244Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:10.271272Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d43cbc69-ffeb-4b7c-a279-03935eadac1f · outbound

This paper cites GPT-4 Technical Report.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:20.917876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:20.917876Z digest=sha256:90ae9fd660743fca9d64809f2d1d0454835c2687b399e04406a0908d96f57591

Observation 1fadd94f-a23d-4006-a4ca-231ea67b3297 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.184222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.002919Z digest=sha256:90c707704026ca6a6d4ba45e6177d7ce2d125aba2d0aef6e46f1a5a2a36a92bf

Observation c0d61a00-b585-478c-aabb-3d8897018d75 · outbound

This paper cites Qwen2.5-VL Technical Report.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.062992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.062992Z digest=sha256:381fa445e5e9197f0b8ffc37be86c87f85ddafc1028ffabcbe444f7eebfa3947

Observation 9fdec59d-0450-4b24-9941-57f9a9bdc608 · outbound

This paper cites C.; Geva, M.; He, J.; Wu, J.; and Li, M.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture C.; Geva, M.; He, J.; Wu, J.; and Li, M

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:40:26.169659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.124329Z digest=sha256:1e1efa5e46a8bd67c0887fb48623eb8009948b8d250a55236ac45667a957e77f

Observation 850362c7-dec3-4fde-8ee5-ebee38b93b0a · outbound

This paper cites Multi-Object Hallucination in Vision-Language Models.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Multi-Object Hallucination in Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.185699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.185699Z digest=sha256:490e46593862b46f50f972db785a28b3def58b4cc596d1d287462a76ba1ca6fe

Observation 9f49fa7c-3e12-4290-aef7-767d57209960 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.155733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.262546Z digest=sha256:704e163b498fbeb5a04f7b5fc95f9c7d235236f71e84bd141eff24e1036b1fdd

Observation ff00f254-9fce-4018-87cd-dad413f06d7e · outbound

This paper cites Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:40:25.837858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.327031Z digest=sha256:774df3c9b1530ee0ea376bb103566bf408f45bd49efb662ce6b15b40f01765ac

Observation 08d70087-a7c7-41a2-a144-805825c4e51f · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.408465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.408465Z digest=sha256:c278d013fda972a6199b11d73f9d1a160aa150999910d7a2eac3e76c0cb1d526

Observation 3c7b937c-c260-4d00-b0f2-9190e4cf32e9 · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.474439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.474439Z digest=sha256:abb4b84111b7ab94bbd826f673964bd09bdfffa2cac0db2d626e4ca0d9ae1800

Observation a7546be4-08d4-4789-aea9-58029f5ab5c9 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.539233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.539233Z digest=sha256:b9bc8573e9a50292c4a5aa36aaac73a03e55ee3d5d358a1abb30d0c444b4e3e4

Observation 163291e9-bbce-4c0d-81d4-877fd88e0107 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.141801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.609690Z digest=sha256:4aa21e0245282791095cd94a04fb019c813a6672d0d1ccd30415970ea0afae21

Observation c7ecdfc5-0bc4-4898-84be-61894d17628a · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.124659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.703116Z digest=sha256:456ead247ad31725d52cd36cdaf1556780c472ab5b5b2fe6550173f8cda00a63

Observation 568b43a4-4766-4cfd-8a54-cf81b7105324 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.757867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.757867Z digest=sha256:71d26302905aacfea3b6f2e8d114b6eb7a357b492679cdb673215fb6d007cb42

Observation c3293966-faaa-4a51-aec1-94670e677bca · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.108894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.847352Z digest=sha256:e0f874dd65db7ed5b8359e9bc601e75b916d03904621d248be85589ac33533db

Observation 1db48a6c-7491-4470-a3f8-6bfa46112aa5 · outbound

This paper cites J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:40:26.093096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:21.902977Z digest=sha256:58a8e92d86a0b42e256ecca67b6f6f92d6259462137017dee593850b94b68739

Observation f8d1a677-5f44-486a-b010-114fcca076dc · outbound

This paper cites GPT-4o System Card.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:21.985438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:21.985438Z digest=sha256:4971bb896c3441194b9778c143baf2be6ffb1b24bf969f9656b8e2f370762453

Observation fcc756ef-b175-4e56-92b9-42691d5bb884 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.040115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.040115Z digest=sha256:ca3cd1e2ef2b3a9a94b3919405a298e9f4f8c8d18a88f98199f687d0a88f149f

Observation 0b96c9d4-86fa-4a0a-a84d-f15b19b42214 · outbound

This paper cites "Well, Keep Thinking": Enhancing LLM Reasoning with Adaptive Injection Decoding.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture "Well, Keep Thinking": Enhancing LLM Reasoning with Adaptive Injection Decoding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.153841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.153841Z digest=sha256:27d7fc584dc0fc82f3e66f9b1de1dd3706a6a90ca00d089d0becf6f743e88f3b

Observation 9b2f27ef-2772-4c7b-9afd-8c80d6d9b29f · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.078596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:22.201294Z digest=sha256:31ea4bf8b1955a793bb1e2f8d3ea9b3a0bcd7a823a3e458b719c54bc1792abe7

Observation 321e73f9-a5cb-4c18-9f98-8102aeccbeeb · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.265805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.265805Z digest=sha256:f3cc0c6b7a50cea82eac852834cce1929594c3860855bc6b64d6ec0153dec516

Observation 437e3c59-89c7-4fee-8faa-de9207f17962 · outbound

This paper cites A.; et al.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture A.; et al

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:40:26.064579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:22.336424Z digest=sha256:e28f291d2df71283970be799c0aeb99a7b051f46b57b40a94c667e8dd1e46cf5

Observation c70964dc-271f-4a34-9fc1-3a19a1fa9d34 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.438908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.438908Z digest=sha256:f4bb1bf9f2dc20a63d722af31fbe810756a3c10bc72fdb540073a41edb8b065a

Observation e2e0e38c-2833-4be1-aa0f-1cecfe54f569 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.049526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:22.541628Z digest=sha256:d56a5810716c72cf83e854e04b049ea61a4fdd9fdcd0d7979aef866e1631d3d6

Observation 6023ceb1-fe25-4dd7-9265-2df50acdd989 · outbound

This paper cites ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.609304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.609304Z digest=sha256:bb00656b2b0b85387f2fd233c57780216f78427af6eda9c8e3cb80eefa7237f1

Observation c3a40141-58cc-4308-b02d-db22c41b6f97 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.714354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.714354Z digest=sha256:5bfdf7880146ac15bb13cb6ddc9a774491bfa8bc6983b31897c5324cf3d7a531

Observation 1e0205e3-8c82-4a78-993b-45b3ead780f1 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.773906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.773906Z digest=sha256:919490cededcc3782f91ae3c17a0d58aed96cf23a670fcdbe322d14f1538d6cb

Observation ede1b8b4-47f2-4291-b58b-c48bb33fa726 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.835023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.835023Z digest=sha256:2e165c91bef525534309a949f2f3c63ab33999bf4bbd530c11c5220f4aec8ade

Observation 19b493c7-82ac-4842-9a67-268051338242 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.893886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.893886Z digest=sha256:49e5b235c01c21dbe57a34600f65f91ce4b07079eb27b766a43e36646ed5a1aa

Observation 6075b8c2-d58a-4fb9-975d-1df9b2df2341 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.027800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.027800Z digest=sha256:ab198eb728cafc4d9f083f4996c451dd82ceba3f09aa4acf5c3b7c28b36328b1

Observation 7ee2fe4f-6421-460a-b40b-9b26b89109b7 · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.101934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.101934Z digest=sha256:ccdbae132e9ef07508d69520bbf173059b113ed030453591ccb72e1f20038789

Observation 3f266c45-5168-4cfd-8d3d-37bd68848f6c · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.014987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:23.186680Z digest=sha256:ebba30e7c832b70b095c111250713ad6972edbe8684b403a328b1649a93f6e33

Observation cbcd96db-b8ce-418d-b8bd-df236d371313 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.287038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.287038Z digest=sha256:4decd318963dd95794faaceb9dc7214c76fd6c67456a497643c4d82792b1a915

Observation de4b4843-70f8-4554-b263-123776f5cb38 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.409237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.409237Z digest=sha256:b0dd6a27035db41ff831ea397d5da5652c0986e45a68c77819eb3ae75fc9c989

Observation 97553819-fd5a-44e4-9949-f61828487942 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:26.000797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:23.532237Z digest=sha256:fe8c1b9973a8a4a8c781d9637f657eb0faf38e21c86e981c8425125c23bc50b8

Observation bfdee345-9c44-4db4-8531-9b1bcd881ded · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:25.987127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:23.645591Z digest=sha256:616065003c8b1161dccdf2513e592374862686d20fd41113052706fc31448bf9

Observation f8855729-854d-4693-86b5-e93008d60952 · outbound

This paper cites V.; Zhou, D.; et al.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture V.; Zhou, D.; et al

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.706928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.706928Z digest=sha256:420dd2216a40784797ddfefb60ace7aab0eb11324220097a2c6adb772654f65c

Observation 403e88a9-0bf6-44aa-8100-3224b1e886ed · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.835237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.835237Z digest=sha256:518b434b0b0a48b88a5014bbb19d63bb5aa9915a21c925575a318cee7956cdb2

Observation ef897e87-fbc2-49b5-94b3-bc2454b30aaa · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:23.949207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:23.949207Z digest=sha256:a95707f57e0411821b51c734bf73a3635983b424fb88ec333f4837071720db2c

Observation f2774bf3-af99-407d-af29-21f65c5c471d · outbound

This paper cites Grounded Chain-of-Thought for Multimodal Large Language Models.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.078659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.078659Z digest=sha256:d2c135cd537dc975b4c2dfd61e9fe3678be5cac4533c415c82a0af1f4c71fa1b

Observation d704a65f-af9c-4d3c-afde-1b14391faacc · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:25.963261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:24.164678Z digest=sha256:43bda6b582d7e4bed17468b534eb0090fd24ff330fa4f168d421b58121e0d6e4

Observation b9820517-4e69-4ed2-8b76-c3d73ac9ccbe · outbound

This paper cites Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.298029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.298029Z digest=sha256:82ae296787a6932ef23957d9b47b4b388311576bc6bd6ac1b0035325677b087c

Observation af90493a-917d-49dd-990a-54b38a3e4d24 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.437770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.437770Z digest=sha256:5182208bec24eca550865114f4c299a25fa38cf369d5e9ee649d91f1b50d9fff

Observation 5b34fe1e-369b-4e4f-8a31-1f4ff23b8b7d · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:25.949071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:24.532215Z digest=sha256:6fdf966188fbc7f9c91fb669ccb175e92c091f14468cfea3bec06dd0db8f743e

Observation 8956414d-be47-464a-8362-24872587b1fc · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:25.934472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:24.634793Z digest=sha256:51ddbea1714a29dd181e7573afba2f862b842b2208a081156e858714d494510f

Observation 2d0cd3b5-f8a4-4fb0-95d8-c5b13fa3842e · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.704510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.704510Z digest=sha256:29b6cef559f6948d9a8fbdb39963460d630d68a5f2c559d743dc78e5d446ca41

Observation c9dd781a-e214-460d-9585-0f1e0eddc8b5 · outbound

This paper cites an unresolved cited work.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-05T11:40:25.919917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T11:40:24.841237Z digest=sha256:a6dc5e844fba3e9a3a6896e4ba365c3e7fdd7b698675caca621fd2c1ef582bcb

Observation 45e98c5a-18a5-468e-8f65-9ca4a5819e59 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.932938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.932938Z digest=sha256:2e47099cf8d0a1b02d6c4db19582603df57cd48c670aa284af584574d3ec1a82

Observation beec9b9e-010f-42fa-a22b-4ea66791e88f · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture , " * write output.state after.block = add.period write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:25.030350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:25.030350Z digest=sha256:ae8c2d96d96100fb0137b10caefcc44c6e336a2a1d4c7d7a007bd81b261f43ab

Observation f1de502a-8bf3-41f3-8753-cc32b6943997 · outbound

This paper cites write newline.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture write newline

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:25.147462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:25.147462Z digest=sha256:90573f7f7adc5595d9b4dcb49f4e430e418797862ecc75db0cbfafc1a4985d40

Pith citing papers

Observation ca2531c2-ce95-4d7d-83cd-3b5390d70603 · inbound

SCP: Spatial Causal Prediction in Video cites this paper.

SCP: Spatial Causal Prediction in Video Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:50:11.305915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T16:47:44.523606Z digest=sha256:97efb25bb2767b63bd4882c39a19514f9bceb384c63b2016ef1480db0952db72

Observation f8784c33-9dd8-429b-956d-07c8ea5c2d13 · inbound

Why MLLMs Struggle to Determine Object Orientations cites this paper.

Why MLLMs Struggle to Determine Object Orientations Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:04.597092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:19:44.979076Z digest=sha256:de91f9af3ac2b5d6e6bf058ab41e07232fb40487f8787c4f891cffc45f2ab0fd

Observation 6be88569-d713-4f35-b426-8593bf50f9a3 · inbound

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing cites this paper.

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:01:08.940684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T05:58:04.055855Z digest=sha256:ac243f92297aa04497231fb8db2bdbe3f440c1bbedeb1b7c5ccf5c4aa651d717

Observation 9ee60d47-bac9-4064-951c-b7ec93060950 · inbound

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing cites this paper.

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.241818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T17:37:33.750306Z digest=sha256:2ef1d27c20f258f53bae6d8d22976e9b79d7b1bfc115724d763da087718e34a5

Observation aa78aa4f-a42b-476b-b5e6-c66dfc775739 · inbound

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models cites this paper.

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.654177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T08:06:32.403727Z digest=sha256:627793cebcd0927c3d25fbd55850ca177d5f2ef56289ce91e846ea2ee573c71d

Observation ed6f41d5-a745-45da-95ca-f5c128022165 · inbound

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning cites this paper.

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.932385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T09:59:02.899488Z digest=sha256:cbf7bc9105c76d59461931e570f1379113b23994ee9b0751df22c855a79f3b67

Observation c0e47212-e3d1-4316-8fe1-d51a03f0261d · inbound

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity cites this paper.

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:10.273001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T21:02:44.441202Z digest=sha256:df56694921e498abc17caa1a77c82a6604f4f0645df24f65ec157cbb6e60b55f

Observation 800777b3-396f-4f56-848a-6d90e020a154 · inbound

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence cites this paper.

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:30.837244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:30.837244Z digest=sha256:4827541d95a317b1074ab2f4798f3fb34fc0c510ac42fad7f0ae22502d78157b