Pith. sign in

Paper Citation Record · LEDGER

SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2406.01584.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01584 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:37:03.235740Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 068ff728-464b-464d-a9d8-881d138a6835 · inbound

WildLMa: Long Horizon Loco-Manipulation in the Wild cites this paper.

WildLMa: Long Horizon Loco-Manipulation in the Wild SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:32:58.152242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:32:58.152242Z digest=sha256:6c658b67e40764e79d41507f5150713d86eee25485475f26cdffcc5165c82bee

Observation f808913b-6ba7-437c-b9a2-5e19cafa662d · inbound

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity cites this paper.

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:23:44.093828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:23:44.093828Z digest=sha256:cfc6ee0aff2ce36452c23f3d8f0b35c985d08c078694bd4a816bd83b698e7650

Observation 9ada32df-b4a2-4556-8cc0-ffa8e2d7a687 · inbound

LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models cites this paper.

LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:48:15.262522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:48:15.262522Z digest=sha256:d4efb056cd3fb6939246bd10b97bc7cdb3c91ce59cca4ac94272871980c7f302

Observation 8110f618-5fc2-478b-9e0a-cbfeec625a61 · inbound

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations cites this paper.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.417197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.417197Z digest=sha256:4f00b9dfee75097e82524462cdc97951a64ea9e8e57553930ac076518829a19b

Observation 621224e9-6890-40f9-94c7-cf596e8ae6f8 · inbound

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction cites this paper.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T13:25:27.377004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:25:27.377004Z digest=sha256:f6ae9f01aa77439e3050b6ad53ab162d3698c016b5a3da0e108472b960fa6918

Observation f2d464e9-0e00-4eb9-902b-c872c8c75760 · inbound

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks cites this paper.

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:00:46.632220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:00:46.632220Z digest=sha256:34e19ab0c9dbdfff4f8fe5d1cec788dfbb80be040af195c09f4dbd20504bbbbe

Observation 130f8898-d5e3-45b6-ab5c-43f72c5e6c5d · inbound

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding cites this paper.

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:36.620135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:36.620135Z digest=sha256:070436bc717708228cc52984b56d15f259c1e544f4b4594dbd1e189883d30982

Observation db059a5a-634b-4ba9-aff2-7ec41000d777 · inbound

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought cites this paper.

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:09:34.859572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T23:09:34.805552Z digest=sha256:a3aa3212888593b31b35a84e26ef3e709ebd0db7f767c3ff4e4139a4cef2679a

Observation 517a0395-b94f-455c-ac51-15f0b0de6eec · inbound

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces cites this paper.

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:08.524496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:08.524496Z digest=sha256:2c017c790bc364f429be442c3c2c04a3d1e1ed0b46cc38e7c93f20bd0eb526db

Observation 89f02065-ef9a-4bc1-b2ec-9e5d976aa0b8 · inbound

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning cites this paper.

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:41.167762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:27:41.167762Z digest=sha256:eea9fb30d4b2065c955061093770f18c99d26edbd99bcbfa427fecf84e5753f2

Observation 4a1efa2e-8da3-4cf0-b416-121c54c617d4 · inbound

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes cites this paper.

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:37:03.235740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:37:03.235740Z digest=sha256:49ccfbe4ec4d621591b8fe8fa70b69b8609b252c7ef66223965407be3598f040

Observation 279db72f-026a-4e35-a91b-dd96ba8541ec · inbound

Vision language models are unreliable at trivial spatial cognition cites this paper.

Vision language models are unreliable at trivial spatial cognition SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.429836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.429836Z digest=sha256:10edc3e1671ee4e813ce312cb5a3cbec0c2f92b85dba3ced669067d81ee4f207

Observation 7e64fe9d-a63e-4cbd-9703-ca5374b9bfe3 · inbound

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning cites this paper.

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:41:30.854068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:41:30.854068Z digest=sha256:34f8571987603a52ebcde5a97caf8f8b495b86e34f739d1d5accc3a09e221d81

Observation ee237961-8d51-49cb-87d3-00b205a249fb · inbound

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models cites this paper.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.594324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.594324Z digest=sha256:4afbac1ab5625a877803a9c371bdb03b3d274ddeae8efaf1334cffb71b28bee1

Observation 771d326f-3b91-491c-8196-7d2d99db9355 · inbound

Preliminary Explorations with GPT-4o(mni) Native Image Generation cites this paper.

Preliminary Explorations with GPT-4o(mni) Native Image Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:46:02.764185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:46:02.764185Z digest=sha256:48eb9b8a2f68310052ab9ac8becd683f8e10986115898237e40abfb7b9562f36

Observation 90357e8e-50c1-4175-bbaa-18a80186731e · inbound

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience cites this paper.

Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-15T23:31:12.518325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:31:12.518325Z digest=sha256:4d31e60f62502a44157dab0a8bec22e2b65078a94b23fbf4f5f836330197a2eb

Observation 2fb3764b-8f06-4347-b9f1-de1549c1d379 · inbound

Vision language models have difficulty recognizing virtual objects cites this paper.

Vision language models have difficulty recognizing virtual objects SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.428655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.428655Z digest=sha256:16258172c6c6e8e53dc32952593bedcbae82f1827d60c64debce094b3e659ff9

Observation 6ee4ee5c-9a39-4345-90d7-00518300fbf6 · inbound

Can Multimodal Large Language Models Understand Spatial Relations? cites this paper.

Can Multimodal Large Language Models Understand Spatial Relations? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:04.777438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:04.777438Z digest=sha256:fc32645f0716182c43b605397688a0624fe2cdd9e2152d510284ff6b0d43db84

Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · inbound

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation cites this paper.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.933165Z digest=sha256:43320a60247fb1bcb843382f51284d16f7dc180acc9f32675d6a05519cad6e13

Observation eaf3a9d2-5463-4376-a6bc-5326c43abcfe · inbound

GenSpace: Benchmarking Spatially-Aware Image Generation cites this paper.

GenSpace: Benchmarking Spatially-Aware Image Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:21:20.914495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:21:20.914495Z digest=sha256:d95d79f450054d978811e0b0ba4f877a15a2694ee01cfcc41a9d6bf3ff10c4a6

Observation 61d2f55d-3e29-4091-92b3-383a5c5938ae · inbound

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models cites this paper.

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:44.249542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:44.249542Z digest=sha256:13e855d616c34abf3d76832cad70670e5b645aa43ebd966b5aaeb76a3bacace2

Observation e85307f9-bf8d-42ae-aaf8-e44fdc65d69a · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.127004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.127004Z digest=sha256:fff4cce7af05adbedad08b78cdb43aaa5c54864816f1b1b32803bfd969bfaafa

Observation 9f89895c-f4bf-432a-8c13-6e42b8eabd64 · inbound

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments cites this paper.

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:05.011646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:05.011646Z digest=sha256:44f2d3053885427fc7e7b7b2a5cc905ceb2bdbdc1b707f47a32d725a5ffb47db

Observation 91e0ac93-8634-4490-81b6-3162bc71aa63 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:29.384793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:29.384793Z digest=sha256:9ea900685bf58a653efed69210dde9995d9f8a2ccba0a55cb438d555dffcd1cd

Observation 1cb6f7bd-1dfe-4715-881b-fbe525c78b9a · inbound

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning cites this paper.

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:12.750753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:12.750753Z digest=sha256:42f00318498d3057014dc9642131f8457a1961f56faccf0a804ae8dc1fb01d25

Observation ca1bd479-40fe-484f-a892-1beea35e5999 · inbound

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding cites this paper.

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:20.579394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:20.579394Z digest=sha256:b111569e28bc4f880cc2f40c76f35480dc9752b6f09cb7d82d93949e823ee50f

Observation 9b71a8db-4c94-48c2-b200-c1b04c17a942 · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.097127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.097127Z digest=sha256:0f50ac1660f2b61f47f8620f7bcd15943806bd34d3847cf70fdeb64316a1a1b8

Observation eff6343f-b388-4618-810d-b99b1296daa9 · inbound

AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning cites this paper.

AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:56:14.629466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:56:14.629466Z digest=sha256:9a46f562861f54ed26303c81845bf8a6299787995cee9c21a5891feb1b0f3ac4

Observation aa4cd893-691c-4236-96ff-78c5eb4d968e · inbound

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds cites this paper.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.638480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.638480Z digest=sha256:457cef577bca74788de72502a993272e651d29fb9482d66547f70a0bb6354827

Observation a214032c-b48a-448e-b668-586e250128f4 · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.450996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.450996Z digest=sha256:533bf1f0aa811c94fe266a570937992b98e4017a3a9706257a7d06ea1d62a9fb

Observation 6cabf21b-16b5-4f30-849a-44cecbdfbe34 · inbound

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models cites this paper.

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:16.913232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:16.913232Z digest=sha256:f86c8765c0dd0edf236004993cceae5cbc7ad160d862b45a27cbada851f6aeeb

Observation 6e8b5468-f395-419d-9244-d1ac0239f2ad · inbound

Enhancing Spatial Reasoning through Visual and Textual Thinking cites this paper.

Enhancing Spatial Reasoning through Visual and Textual Thinking SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.171941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.171941Z digest=sha256:a736e82e863d1bdcd68081cdfd3ea8a72ea18e580c4ef6feae1a91db275a03a1

Observation 49e5a4d9-4f94-4b70-b6f5-839031e9bc6e · inbound

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation cites this paper.

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:51.724391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T22:04:34.235731Z digest=sha256:12e97d7b51c5e21a3d71a9d665ba3c64abb20f4642dbcfd698cc11536409efdf

Observation fbbdb503-c116-4f08-9fb4-c62944dd3a0b · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:08.100667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:08.100667Z digest=sha256:236c322915e118432d9497b10f178a11dc46a95211172adaf7633f0ffa15615f

Observation 701aa8ba-57e0-404b-aacb-18f69344c1be · inbound

SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation cites this paper.

SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-15T16:09:09.577056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:09:09.577056Z digest=sha256:ecebb96a108e8abf715bdb62997c503a4d0dcf806a4ecd348355c78a44d45419

Observation d9cb7737-9013-4111-b31c-9d0ec45dee76 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:20:34.446572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:9c43fe7c4d3df48a25e6415703b6f838bf874dcd603be34d3994f96f139ea79a

Observation 185a0f71-71f6-47b5-b1c0-81a63a0152c0 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:50.465239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:50.465239Z digest=sha256:cd80acd2a65a8ff7b87e17db5e0fd4ae25d145a4f20549877005306fd95215b1

Observation c11b4357-5a09-4755-be09-f9b025ac2240 · inbound

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation cites this paper.

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T21:41:28.734178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:41:28.734178Z digest=sha256:83943fa6e5ac718df09af76f1b4e3569f3f02cea753a97e7fb73a82ecf3ea441

Observation 58ef9cc4-cba5-4ace-bfdc-b6888ef921e0 · inbound

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation cites this paper.

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:32:41.261611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T09:31:39.548562Z digest=sha256:0f10459c236b50020230ab4ec7eb26de91a03877eadca7e008e964f81481aee9

Observation a763a5cf-2318-4778-b72d-9a76454f1ad9 · inbound

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning cites this paper.

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:25:12.110307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:25:12.110307Z digest=sha256:fe9a53786a3636be80505c21af95cf0952aa8e9bb50055e1fe768e8966e092d7

Observation aac82ce6-25c0-4e8d-b671-7897e406c611 · inbound

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization cites this paper.

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:16:09.802485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T15:12:23.459579Z digest=sha256:d6689339af5f00efa410a095ec9b5711b8b5ca036c3dae6cf164a59ca68d1f8a

Observation 9a45f726-1e72-4ffd-bf4c-fc25f0c75fab · inbound

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables cites this paper.

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:23:02.419126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T17:20:28.036531Z digest=sha256:9304c27bde66591a45b7024a5d5871d69e50b45dce94b779fc9a9463e8706589

Observation f57c5c60-8de1-43ac-a63d-0d3b49227210 · inbound

Spatio-Temporal Grounding of Large Language Models from Perception Streams cites this paper.

Spatio-Temporal Grounding of Large Language Models from Perception Streams SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:26:02.689307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:10:45.837684Z digest=sha256:246e2f48186da189f183593793d32c1cbe57e395ead51a1e40f0d26c6b3e9cba

Observation 886c959b-7470-4a40-af9b-df541b56054b · inbound

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark cites this paper.

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.293094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:37:27.926364Z digest=sha256:6a32b50d8f2829c690eb96b739a79c63cdf2d670f0c7b03b714bce1c1e55679a

Observation ead79631-e448-4237-9af1-29b412302b17 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.215752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:52:22.778489Z digest=sha256:03e521813855732311e2533067383c5f6d90cda5ea206a43a7ef688c83b4053c

Observation 6f538434-ae76-48b7-8472-071c30261078 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:47.188163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T18:25:17.831116Z digest=sha256:5848cc0c3f3b4947cced69f85d5c098bb2f90dc287e2b33d852e503acdfe64cb

Observation a3a9c864-aa3e-4fbc-9b2e-78d242c02496 · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.735753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T06:55:04.657347Z digest=sha256:b7f1bed3b21f8fcb8ecb1d8559f4724dfb27b33b8b48cfad5c9ab489dfa86bf3

Observation 7e319e28-7db0-4b32-a131-154b894762ed · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.801081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T17:52:26.785086Z digest=sha256:334fc0ca454884d9b2824e4743c8af9a18e78b7080d56124aa009e0246671032

Observation 72c8c96f-3697-4353-bab3-fce7fbdce1e8 · inbound

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? cites this paper.

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.548789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T07:48:19.295578Z digest=sha256:63544e88322f388661491a96325c0d11397617129e158976e9fe9fba6a15abe2

Observation 6c6e2dfe-20c2-49cd-a532-764928834525 · inbound

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning cites this paper.

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.673197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T15:40:05.730181Z digest=sha256:d4fc5f79e17bfc25e7981577924eb3aa268990ccf700f79a4eecb48dbc8d35a7

Observation 80dead48-e9dc-418f-9f8d-49be12347821 · inbound

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video cites this paper.

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.256251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T02:16:25.730555Z digest=sha256:06972fce18d142eb254ab6b7fbd011775221c23e3b3e630ce77e0f9b5be78dd8

Observation d2bd64cf-6304-4103-8a82-22c262e82591 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.468135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:ad56041654ddb240aec283b10ab2842affd3565f570a5b569aabd8af3abaaf0c

Observation b9b16f0b-1593-4c89-b43d-c9450caed3c8 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.590121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:e54d1d1afdc0dbd55a25a3a931ea6819f55897256b4010c9178d9e523d38a54d

Observation b8566166-c3ac-4e6a-b48d-b1de1d6526c3 · inbound

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation cites this paper.

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-26T10:39:18.313055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T10:37:24.946718Z digest=sha256:56ba311bf86566a08042eb8efcc04009d4fe0364896705ea363287dd8ac622f5

Observation 93afacff-302a-46a6-a0f7-3e0e5f9b8d17 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.087107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:d6f8617d2dd2d59068afe48e369543c185cd98cf6c5e108f8540bacf30633b4b

Observation d5f799bb-513d-43f3-aa90-2c83bb44d230 · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:27.565496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:27.565496Z digest=sha256:481012ce175c67cc3afebf2c364ac49d92fbfc327aa349089a9a0dbb5028b74b