Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2406.01584.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:37:03.235740Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 068ff728-464b-464d-a9d8-881d138a6835 · inbound
WildLMa: Long Horizon Loco-Manipulation in the Wild SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f808913b-6ba7-437c-b9a2-5e19cafa662d · inbound
FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ada32df-b4a2-4556-8cc0-ffa8e2d7a687 · inbound
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8110f618-5fc2-478b-9e0a-cbfeec625a61 · inbound
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 621224e9-6890-40f9-94c7-cf596e8ae6f8 · inbound
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d464e9-0e00-4eb9-902b-c872c8c75760 · inbound
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130f8898-d5e3-45b6-ab5c-43f72c5e6c5d · inbound
3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db059a5a-634b-4ba9-aff2-7ec41000d777 · inbound
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 517a0395-b94f-455c-ac51-15f0b0de6eec · inbound
Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f02065-ef9a-4bc1-b2ec-9e5d976aa0b8 · inbound
SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a1efa2e-8da3-4cf0-b416-121c54c617d4 · inbound
Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 279db72f-026a-4e35-a91b-dd96ba8541ec · inbound
Vision language models are unreliable at trivial spatial cognition SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e64fe9d-a63e-4cbd-9703-ca5374b9bfe3 · inbound
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee237961-8d51-49cb-87d3-00b205a249fb · inbound
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 771d326f-3b91-491c-8196-7d2d99db9355 · inbound
Preliminary Explorations with GPT-4o(mni) Native Image Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90357e8e-50c1-4175-bbaa-18a80186731e · inbound
Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb3764b-8f06-4347-b9f1-de1549c1d379 · inbound
Vision language models have difficulty recognizing virtual objects SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee4ee5c-9a39-4345-90d7-00518300fbf6 · inbound
Can Multimodal Large Language Models Understand Spatial Relations? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · inbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf3a9d2-5463-4376-a6bc-5326c43abcfe · inbound
GenSpace: Benchmarking Spatially-Aware Image Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61d2f55d-3e29-4091-92b3-383a5c5938ae · inbound
BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e85307f9-bf8d-42ae-aaf8-e44fdc65d69a · inbound
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f89895c-f4bf-432a-8c13-6e42b8eabd64 · inbound
Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e0ac93-8634-4490-81b6-3162bc71aa63 · inbound
RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb6f7bd-1dfe-4715-881b-fbe525c78b9a · inbound
Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca1bd479-40fe-484f-a892-1beea35e5999 · inbound
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b71a8db-4c94-48c2-b200-c1b04c17a942 · inbound
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff6343f-b388-4618-810d-b99b1296daa9 · inbound
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4cd893-691c-4236-96ff-78c5eb4d968e · inbound
3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a214032c-b48a-448e-b668-586e250128f4 · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cabf21b-16b5-4f30-849a-44cecbdfbe34 · inbound
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8b5468-f395-419d-9244-d1ac0239f2ad · inbound
Enhancing Spatial Reasoning through Visual and Textual Thinking SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e5a4d9-4f94-4b70-b6f5-839031e9bc6e · inbound
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fbbdb503-c116-4f08-9fb4-c62944dd3a0b · inbound
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 701aa8ba-57e0-404b-aacb-18f69344c1be · inbound
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9cb7737-9013-4111-b31c-9d0ec45dee76 · inbound
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 185a0f71-71f6-47b5-b1c0-81a63a0152c0 · inbound
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c11b4357-5a09-4755-be09-f9b025ac2240 · inbound
Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ef9cc4-cba5-4ace-bfdc-b6888ef921e0 · inbound
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a763a5cf-2318-4778-b72d-9a76454f1ad9 · inbound
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac82ce6-25c0-4e8d-b671-7897e406c611 · inbound
TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9a45f726-1e72-4ffd-bf4c-fc25f0c75fab · inbound
TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f57c5c60-8de1-43ac-a63d-0d3b49227210 · inbound
Spatio-Temporal Grounding of Large Language Models from Perception Streams SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 886c959b-7470-4a40-af9b-df541b56054b · inbound
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ead79631-e448-4237-9af1-29b412302b17 · inbound
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6f538434-ae76-48b7-8472-071c30261078 · inbound
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a3a9c864-aa3e-4fbc-9b2e-78d242c02496 · inbound
Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7e319e28-7db0-4b32-a131-154b894762ed · inbound
Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 72c8c96f-3697-4353-bab3-fce7fbdce1e8 · inbound
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c6e2dfe-20c2-49cd-a532-764928834525 · inbound
Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 80dead48-e9dc-418f-9f8d-49be12347821 · inbound
LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d2bd64cf-6304-4103-8a82-22c262e82591 · inbound
The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b9b16f0b-1593-4c89-b43d-c9450caed3c8 · inbound
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b8566166-c3ac-4e6a-b48d-b1de1d6526c3 · inbound
Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 93afacff-302a-46a6-a0f7-3e0e5f9b8d17 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d5f799bb-513d-43f3-aa90-2c83bb44d230 · inbound
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.