Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:46:28.348279Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2507.20529.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:46:28.348279Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T12:53:40.783281Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T13:03:26.885433Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6e8b5468-f395-419d-9244-d1ac0239f2ad · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da9656cb-7018-45af-8910-01e3f889940e · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking PaLM-E: An Embodied Multimodal Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 099994c9-42f5-4528-910a-d66c3848e413 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f95802d2-3734-4585-b2b8-ceb59cbbe0d0 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Visual instruction tuning, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56676944-487c-4b55-a134-bf3c2fb30629 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Improved baselines with visual instruction tuning, 2023
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee5b45e-1f2f-4785-a9b7-29c546f32e90 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d7ac006-2f0c-4d91-8c1e-68376eef744d · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking LLaVA-OneVision: Easy Visual Task Transfer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e59a3a24-09fe-458f-b91c-64ec11564fba · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5437c73-d56e-4b0d-b296-1c66f7b1bebf · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c453ef-d1d7-470a-a9c1-cacebbafaabd · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Qwen2.5-VL Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cd1231b-ee16-445b-be4a-9b70a2e6fcf5 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dcba726-8366-4b2e-8504-a8cfdc74383d · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b70842f-48f6-4f53-bfda-9a1d33186f6a · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e73067-eae9-4a4e-b07d-89de312aa068 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking GPT-4 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc50dd51-038b-4e7f-9e5f-3f577e593fa6 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking GPT-4o System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094c6690-75ac-47f4-a864-07e5aa36828e · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4ba45ae-580e-4051-8e4f-ddf3b43c7064 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25302f15-b159-4451-b2d7-9aff49fd01c2 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56c4b939-34b3-4254-a79d-651cf201efa1 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Navigating to objects in the real world.Science Robotics, 8(79):eadf6991, 2023
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ca75790d-085c-48a3-b343-fd257bef3a9d · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Iqa: Visual question answering in interactive environments
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 05a2a6cc-8ed2-4ed9-8117-4199f6d06b65 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Learning spatial- semantic representations from natural language descriptions and scene classifications
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dcc62d49-e406-4fbd-a5a5-86eb271e3aee · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Scene Graph Reasoning for Visual Question Answering
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321cced8-79c5-46c8-af7d-778c96ece3cc · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Learning 3d semantic scene graphs from 3d indoor reconstructions
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c0b0c4-eb03-4e61-91c1-12c6f0a5970c · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Learning semantic maps from natural language descriptions
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7a560d01-9f50-4fd0-9e98-a3e2778cf7e3 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 189351b5-4cc7-498e-b7e1-e4079555bf4f · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7863d53f-35bb-4280-817e-5b2c08add317 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a062d889-abe5-422c-94e7-72ab3453cb2f · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81247c2-0357-465e-9de4-3b592c3c7ec4 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models.arXiv e-prints, pages arXiv–2403, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84e91172-2b8c-49e1-999f-6a8e8e1064df · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Llava-cot: Let vision language models reason step-by-step, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9b82e6-839c-49f1-9544-ad7b2b8ada81 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Introducing Visual Perception Token into Multimodal Large Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eddae57-3463-4415-93d3-c60e004f968c · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Regiongpt: Towards region understanding vision language model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1466eacb-af93-4762-b77b-9c1a39e367c7 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Osprey: Pixel understanding with visual instruction tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6882f827-8468-4b93-b3d7-3d0f1219eade · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking OpenAI o1 System Card
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d2ab21-4c76-47e1-81b8-161f199a2d23 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ffc4a0-b53b-430f-a3ae-8c4f201f4493 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Qvq: To see the world with wisdom, December 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf28b03-1ef5-4798-9aff-99a59808d1f2 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking DINOv2: Learning Robust Visual Features without Supervision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4708994-b8f1-43bd-8fb8-62b31a7d5867 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Relational inductive biases, deep learning, and graph networks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79123037-c741-4491-9870-694884410744 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking What’s “up” with vision-language models? investigating their struggle with spatial reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ded39013-a1bd-406f-b551-ef0bb18caf89 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d8fdc5b-c957-45c2-8d45-d219b6d34b5c · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Reasoning paths with reference objects elicit quantitative spatial reasoning in large vision-language models, 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e97793f1-65ec-43b2-b073-5ae4f9502403 · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Deepseek-v3 technical report, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa91c5d-606d-464f-893a-1a16ca412b6d · outbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Vqasynth, 2023
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 000d1a38-48e5-4aba-b1d8-9b8defed9e68 · inbound
Self-Prophetic Decoding to Unlock Visual Search in LVLMs Enhancing Spatial Reasoning through Visual and Textual Thinking
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.