Pith. sign in

Paper Citation Record · LEDGER

Enhancing Spatial Reasoning through Visual and Textual Thinking

As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2507.20529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20529 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:46:28.348279Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T12:53:40.783281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:03:26.885433Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e8b5468-f395-419d-9244-d1ac0239f2ad · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

Enhancing Spatial Reasoning through Visual and Textual Thinking SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.171941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.171941Z digest=sha256:a736e82e863d1bdcd68081cdfd3ea8a72ea18e580c4ef6feae1a91db275a03a1

Observation da9656cb-7018-45af-8910-01e3f889940e · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Enhancing Spatial Reasoning through Visual and Textual Thinking PaLM-E: An Embodied Multimodal Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.177268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.177268Z digest=sha256:e281b89aa3b373dfa7b7e28afd5f521172232fcd2a3c9a0d320e7b06f6c1176d

Observation 099994c9-42f5-4528-910a-d66c3848e413 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

Enhancing Spatial Reasoning through Visual and Textual Thinking Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.181982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.181982Z digest=sha256:925a02389f917fcd94ae4adb4007cd827fb7aa9e809d33398ea08d914c9ad306

Observation f95802d2-3734-4585-b2b8-ceb59cbbe0d0 · outbound

This paper cites Visual instruction tuning, 2023.

Enhancing Spatial Reasoning through Visual and Textual Thinking Visual instruction tuning, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.186135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.186135Z digest=sha256:f52cf86192f9c8a963a70188b313847e0c9a0d82eccacc705c6d3670309693dc

Observation 56676944-487c-4b55-a134-bf3c2fb30629 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Enhancing Spatial Reasoning through Visual and Textual Thinking Improved baselines with visual instruction tuning, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.190265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.190265Z digest=sha256:68bc52ef11618565a52a121577f1338a5317b509b3ea3fa1fb1b0f6a97ed62b3

Observation bee5b45e-1f2f-4785-a9b7-29c546f32e90 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Enhancing Spatial Reasoning through Visual and Textual Thinking LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.194309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.194309Z digest=sha256:e748899fd98a0d534777ff0f62fd5c94578ac16e79c3504a476fcd7e1b8dd117

Observation 3d7ac006-2f0c-4d91-8c1e-68376eef744d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Enhancing Spatial Reasoning through Visual and Textual Thinking LLaVA-OneVision: Easy Visual Task Transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.199206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.199206Z digest=sha256:7f7bcdd953026682fd724d1dc23e33bb031f77207963d8b08e3b11a4e3d91552

Observation e59a3a24-09fe-458f-b91c-64ec11564fba · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Enhancing Spatial Reasoning through Visual and Textual Thinking Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.203594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.203594Z digest=sha256:bdeefa78e25f1e7cb948a3d754269ec756e2636bfe80234fa085f953f6d31927

Observation b5437c73-d56e-4b0d-b296-1c66f7b1bebf · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Enhancing Spatial Reasoning through Visual and Textual Thinking Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.207857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.207857Z digest=sha256:8590041093d2a6b21ee63f273ec93bfdcc12e686d7c1bf49b5ed74d829d58eb1

Observation b5c453ef-d1d7-470a-a9c1-cacebbafaabd · outbound

This paper cites Qwen2.5-VL Technical Report.

Enhancing Spatial Reasoning through Visual and Textual Thinking Qwen2.5-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.212022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.212022Z digest=sha256:dad12754e9a958f943e61eb24840e9b07a53afe8f836f419d6c553cd2d1a7b8f

Observation 0cd1231b-ee16-445b-be4a-9b70a2e6fcf5 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Enhancing Spatial Reasoning through Visual and Textual Thinking Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.216425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.216425Z digest=sha256:c0c4a9068c4cf8b478debaee2a70ea2379bb806f6a2631b4df2ae1a9f4707c66

Observation 5dcba726-8366-4b2e-8504-a8cfdc74383d · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Enhancing Spatial Reasoning through Visual and Textual Thinking How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.220483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.220483Z digest=sha256:2c4fa690b001566b2ac41f488fc462a7bd5d9fe2d47e54eb614b0dcc5ca4df86

Observation 4b70842f-48f6-4f53-bfda-9a1d33186f6a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Enhancing Spatial Reasoning through Visual and Textual Thinking Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.224690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.224690Z digest=sha256:f6f148e1b7dd34270cd6b810755221f360c8008ff0d235067cf6675802071d1a

Observation a3e73067-eae9-4a4e-b07d-89de312aa068 · outbound

This paper cites GPT-4 Technical Report.

Enhancing Spatial Reasoning through Visual and Textual Thinking GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.228774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.228774Z digest=sha256:fa5ac3e832750b6b03c7d4bf8ae1681d0c9cd982651950d9590ce82c1246b34d

Observation fc50dd51-038b-4e7f-9e5f-3f577e593fa6 · outbound

This paper cites GPT-4o System Card.

Enhancing Spatial Reasoning through Visual and Textual Thinking GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.232915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.232915Z digest=sha256:d25f387615dda868d9a37a68b433dae0cf299d30d0fa52becaab07358ae6b08a

Observation 094c6690-75ac-47f4-a864-07e5aa36828e · outbound

This paper cites The Llama 3 Herd of Models.

Enhancing Spatial Reasoning through Visual and Textual Thinking The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.237063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.237063Z digest=sha256:011a11c5b1b20b98cccc1db393b8483f8dc63410cdb5177976e005695909d956

Observation d4ba45ae-580e-4051-8e4f-ddf3b43c7064 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Enhancing Spatial Reasoning through Visual and Textual Thinking Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.241257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.241257Z digest=sha256:d46b2433f3f9b2a30a009901ab38ca6233a936124de476b0cde8fa71d31c3e2e

Observation 25302f15-b159-4451-b2d7-9aff49fd01c2 · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023.

Enhancing Spatial Reasoning through Visual and Textual Thinking Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.245396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.245396Z digest=sha256:a87df7a00bfbc8d21f3427542bff69ee75e2e7e4523154606e58df5bc43bfd96

Observation 56c4b939-34b3-4254-a79d-651cf201efa1 · outbound

This paper cites Navigating to objects in the real world.Science Robotics, 8(79):eadf6991, 2023.

Enhancing Spatial Reasoning through Visual and Textual Thinking Navigating to objects in the real world.Science Robotics, 8(79):eadf6991, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:46:28.824661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:46:28.249399Z digest=sha256:971fa336bf8addd8a3caaef372886c083d42f39e57929f7fb673b1e71b8092cb

Observation ca75790d-085c-48a3-b343-fd257bef3a9d · outbound

This paper cites Iqa: Visual question answering in interactive environments.

Enhancing Spatial Reasoning through Visual and Textual Thinking Iqa: Visual question answering in interactive environments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:46:28.809394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:46:28.253504Z digest=sha256:aadab5f57f048536d44690486d23673dae0a5dabf2b3467f3f9b8610ba7a9136

Observation 05a2a6cc-8ed2-4ed9-8117-4199f6d06b65 · outbound

This paper cites Learning spatial- semantic representations from natural language descriptions and scene classifications.

Enhancing Spatial Reasoning through Visual and Textual Thinking Learning spatial- semantic representations from natural language descriptions and scene classifications

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:46:28.795505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:46:28.257562Z digest=sha256:e680abfb9974d113d1b0c3a01107b1a34a7777e86e2798ccb90ee730245501a7

Observation dcc62d49-e406-4fbd-a5a5-86eb271e3aee · outbound

This paper cites Scene Graph Reasoning for Visual Question Answering.

Enhancing Spatial Reasoning through Visual and Textual Thinking Scene Graph Reasoning for Visual Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.261466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.261466Z digest=sha256:e78c0f89ebc2d010fb8187fcdfd186d9115d05b8d86b84058846b7093c40d8c0

Observation 321cced8-79c5-46c8-af7d-778c96ece3cc · outbound

This paper cites Learning 3d semantic scene graphs from 3d indoor reconstructions.

Enhancing Spatial Reasoning through Visual and Textual Thinking Learning 3d semantic scene graphs from 3d indoor reconstructions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.265752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.265752Z digest=sha256:f0887e0194f530c35daf8b5e0f7b365f7f5ca4ac699111ecee23f582e1f05bc2

Observation c2c0b0c4-eb03-4e61-91c1-12c6f0a5970c · outbound

This paper cites Learning semantic maps from natural language descriptions.

Enhancing Spatial Reasoning through Visual and Textual Thinking Learning semantic maps from natural language descriptions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:46:28.773978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:46:28.269692Z digest=sha256:01b8fa18d5a2ffbf51ff338b36209f83419a357cd2b24a65b94b4bd9ddd26779

Observation 7a560d01-9f50-4fd0-9e98-a3e2778cf7e3 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Enhancing Spatial Reasoning through Visual and Textual Thinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.273459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.273459Z digest=sha256:47562ec86f37f62c5c105c92c889273e4cad37df35bc9512229927117e298a4c

Observation 189351b5-4cc7-498e-b7e1-e4079555bf4f · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Enhancing Spatial Reasoning through Visual and Textual Thinking Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.277717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.277717Z digest=sha256:9bd7f2042274283da73d68960763ce17674074cd52bdc94ee0da201ffcf8234e

Observation 7863d53f-35bb-4280-817e-5b2c08add317 · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Enhancing Spatial Reasoning through Visual and Textual Thinking Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.282024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.282024Z digest=sha256:275ba9b6a3f097f6f1f25b1cbd550944bce2ef973de982958e298de8cdd638b0

Observation a062d889-abe5-422c-94e7-72ab3453cb2f · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Enhancing Spatial Reasoning through Visual and Textual Thinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.286350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.286350Z digest=sha256:1e78e6547d0cbf398e6444faf49a5530dd00019de572d8accff3ced4f0899d74

Observation d81247c2-0357-465e-9de4-3b592c3c7ec4 · outbound

This paper cites Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models.arXiv e-prints, pages arXiv–2403, 2024.

Enhancing Spatial Reasoning through Visual and Textual Thinking Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models.arXiv e-prints, pages arXiv–2403, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:46:28.760632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:46:28.290597Z digest=sha256:643c3e0727291bf147ce4d3acb9edff9c63f799b51dc267e5246a0a50150ca16

Observation 84e91172-2b8c-49e1-999f-6a8e8e1064df · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, 2025.

Enhancing Spatial Reasoning through Visual and Textual Thinking Llava-cot: Let vision language models reason step-by-step, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.294545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.294545Z digest=sha256:9ecac26c23c0dd350f75ead69981fbb9a7e51a6baace33883a6179918e73fda0

Observation 4c9b82e6-839c-49f1-9544-ad7b2b8ada81 · outbound

This paper cites Introducing Visual Perception Token into Multimodal Large Language Model.

Enhancing Spatial Reasoning through Visual and Textual Thinking Introducing Visual Perception Token into Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.298491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.298491Z digest=sha256:7c44bf3875420e3a617a61a559d3497a2c88909fa2bc7ee59a9edd88e2360f8f

Observation 4eddae57-3463-4415-93d3-c60e004f968c · outbound

This paper cites Regiongpt: Towards region understanding vision language model.

Enhancing Spatial Reasoning through Visual and Textual Thinking Regiongpt: Towards region understanding vision language model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:46:28.739890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:46:28.302881Z digest=sha256:9c390246539abbe1494a3eaa5d3a513db82d90afdee8f7fcbf3270a0be078bb2

Observation 1466eacb-af93-4762-b77b-9c1a39e367c7 · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

Enhancing Spatial Reasoning through Visual and Textual Thinking Osprey: Pixel understanding with visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.306997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.306997Z digest=sha256:cae0101977bf69360e30c1fef45779d1c3fb25e2708d948726b5f2045db13260

Observation 6882f827-8468-4b93-b3d7-3d0f1219eade · outbound

This paper cites OpenAI o1 System Card.

Enhancing Spatial Reasoning through Visual and Textual Thinking OpenAI o1 System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.311068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.311068Z digest=sha256:fff4e6899fd3e58a29fd70acd419af9fd862d769ad8c83c1c165913591f606ee

Observation 79d2ab21-4c76-47e1-81b8-161f199a2d23 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Enhancing Spatial Reasoning through Visual and Textual Thinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.315715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.315715Z digest=sha256:d35c776e2ba12dbcf2a4a58edbca8a4a727d4f6908ca43c6b9402a23b0631cbe

Observation 45ffc4a0-b53b-430f-a3ae-8c4f201f4493 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

Enhancing Spatial Reasoning through Visual and Textual Thinking Qvq: To see the world with wisdom, December 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.319857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.319857Z digest=sha256:1100d4fce303d82613eb8bc68f2f457fb22c1ba2783785a1c4d999f83a3bcebd

Observation 6cf28b03-1ef5-4798-9aff-99a59808d1f2 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Enhancing Spatial Reasoning through Visual and Textual Thinking DINOv2: Learning Robust Visual Features without Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.323821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.323821Z digest=sha256:3dfa200ed062c4caa7fae3170a53e9bde31e2ced104486c34f9472ef19aa970d

Observation c4708994-b8f1-43bd-8fb8-62b31a7d5867 · outbound

This paper cites Relational inductive biases, deep learning, and graph networks.

Enhancing Spatial Reasoning through Visual and Textual Thinking Relational inductive biases, deep learning, and graph networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.327569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.327569Z digest=sha256:e77f6b9b70658ad289fa0aa98998ee798b81bc5ec64b4eee0cf0f550874eb996

Observation 79123037-c741-4491-9870-694884410744 · outbound

This paper cites What’s “up” with vision-language models? investigating their struggle with spatial reasoning.

Enhancing Spatial Reasoning through Visual and Textual Thinking What’s “up” with vision-language models? investigating their struggle with spatial reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.331770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.331770Z digest=sha256:d0511cea345baeb990c258f831c8614a25bffa660a7f158c8f9361a63f8189e1

Observation ded39013-a1bd-406f-b551-ef0bb18caf89 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Enhancing Spatial Reasoning through Visual and Textual Thinking BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.335901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.335901Z digest=sha256:daf7c8b877068b45a2c124f7d3c04501c0d41eb36a767279722f0284e9c367a6

Observation 7d8fdc5b-c957-45c2-8d45-d219b6d34b5c · outbound

This paper cites Reasoning paths with reference objects elicit quantitative spatial reasoning in large vision-language models, 2024.

Enhancing Spatial Reasoning through Visual and Textual Thinking Reasoning paths with reference objects elicit quantitative spatial reasoning in large vision-language models, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:46:28.703158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:46:28.340169Z digest=sha256:c971d464dec5d669882677711d69a2dd5723c9f20cc43478a4b393da7914f788

Observation e97793f1-65ec-43b2-b073-5ae4f9502403 · outbound

This paper cites Deepseek-v3 technical report, 2024.

Enhancing Spatial Reasoning through Visual and Textual Thinking Deepseek-v3 technical report, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.344025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.344025Z digest=sha256:ff65d7bb83f2aef8f126f8c172fd0e2eb6c7a4b82beaef401a24c78fc67df3c0

Observation baa91c5d-606d-464f-893a-1a16ca412b6d · outbound

This paper cites Vqasynth, 2023.

Enhancing Spatial Reasoning through Visual and Textual Thinking Vqasynth, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:46:28.679723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:46:28.348279Z digest=sha256:0d12def0637eed20e54158fb291a004cb5beec875e68755d1c7be3f96715b2d7

Pith citing papers

Observation 000d1a38-48e5-4aba-b1d8-9b8defed9e68 · inbound

Self-Prophetic Decoding to Unlock Visual Search in LVLMs cites this paper.

Self-Prophetic Decoding to Unlock Visual Search in LVLMs Enhancing Spatial Reasoning through Visual and Textual Thinking

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.890094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T12:53:40.783281Z digest=sha256:de87a8de35b109086d76ad12c61493401b5731cc3048f92ed6b390ee7b987fe9