Pith. sign in

Paper Citation Record · LEDGER

SpatialBot: Precise Spatial Understanding with Vision Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2406.13642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.13642 v7

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:51:48.968184Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.306944Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e251e3d7-f2cf-4e7b-a04c-63d90c83f091 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:27:44.172737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:ccf3a248535f9a3181773463aebb5632b6a694f811de727bacbaa19979099ea4

Observation e726471f-2f61-4b87-ae35-606d52eb03d2 · inbound

A Spatial Relationship Aware Dataset for Robotics cites this paper.

A Spatial Relationship Aware Dataset for Robotics SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:24.318444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:24.318444Z digest=sha256:12fb59bfd016495742946da1d8ad4802cad8022a549b12fd65bfc55435858a60

Observation 64ec50b4-7244-4a6f-b574-b2061b2a6d40 · inbound

OscNet v1.5: Energy Efficient Hopfield Network on CMOS Oscillators for Image Classification cites this paper.

OscNet v1.5: Energy Efficient Hopfield Network on CMOS Oscillators for Image Classification SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:48.968184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:48.968184Z digest=sha256:c3e1f753b2041e155377ec9e541e72d377b87dabff624a32324291494b8f49b5

Observation ae73add2-f905-443d-b7aa-abbab6e06216 · inbound

ToSA: Token Merging with Spatial Awareness cites this paper.

ToSA: Token Merging with Spatial Awareness SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.887022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.887022Z digest=sha256:b9c61b7f9ef04df9cd58adcbf14ede82863c21aa8c2ed94cdb294306c9b47778

Observation 27c51cb1-594a-4ea0-9033-28489f4484f4 · inbound

Depth Anything at Any Condition cites this paper.

Depth Anything at Any Condition SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:52:02.423508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:52:02.423508Z digest=sha256:267df768b942b6347e16ec629464c88a3836f948c2f460d11a9f3c32cf61e320

Observation 5404da64-bf32-4d31-9561-f7a31b58e557 · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.079986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.079986Z digest=sha256:cc73ee50b676d95e8e594853038802d15f1cac79cfdcf43f1f33b57191828725

Observation c569df7c-8152-4629-8373-70c74568eca3 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.733114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.733114Z digest=sha256:af9ea8e27793dbdab323b60903ae4c8888405a1489f3fb7bf0171dc0e226d23b

Observation dbb6d299-169c-4016-a284-3201659f70f7 · inbound

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot cites this paper.

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:38.185326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:38.185326Z digest=sha256:b360291f0c54613b27d017d1a89aa571579744c16a2f75f976d28e643cb7363a

Observation 79ba07b9-c5b4-4f00-afe7-84d27c3f0500 · inbound

PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation cites this paper.

PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:48:07.793262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:48:07.793262Z digest=sha256:1525096de6217758d21a2678d5a5de6a39cf4d46704e5e7a4aa65bba262db995

Observation e2ed1f8c-2a8d-4720-84e0-5316589660b5 · inbound

Warehouse Spatial Question Answering with LLM Agent cites this paper.

Warehouse Spatial Question Answering with LLM Agent SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.536084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.536084Z digest=sha256:7166a598742aa42b832ebafd5f1534c8d45a0a6322d74168faa05b104200aaa6

Observation 46a2075a-2a04-4750-af4a-544bd439c12c · inbound

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning cites this paper.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.156627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.156627Z digest=sha256:f1cd6f8bb3ef006750b4ee3fc82baccbc73605e993e4397f90e1691517443b1b

Observation 5957f5e9-d878-47cd-9323-840f7158ce10 · inbound

BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models? cites this paper.

BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models? SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:59.749275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:38:59.749275Z digest=sha256:5b746721e70643b97a81ef37483f280fe5c1fad31625739b00a91ef56e359fe5

Observation 1cfeaa5f-2de8-42bf-b014-002bbc930982 · inbound

Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas cites this paper.

Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:23:09.481651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:23:09.481651Z digest=sha256:4d58d7a31aecb9791b6cab2eb7cd9d403762f1c4c46dca9a9ebfd4e293187260

Observation 452ba9fc-3031-4cb0-aa31-3c394b634645 · inbound

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation cites this paper.

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:51.720215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T22:04:34.235731Z digest=sha256:d53123db8f398aae3e50be9a42b15c8dc194a4da56f14891c4f6010da50b7a4c

Observation 69047047-b5e5-45cd-8174-14a6659374bd · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:07.891723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:07.891723Z digest=sha256:a3f042dd764dc88c15e4d5528c93c879186d0cf7b6f934e99bf30e8930a6071e

Observation 82cd167f-a1ad-452d-a6d8-9646c13e431c · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:11.546062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:11.546062Z digest=sha256:21ecd627df251325d7945ef1bbc06276f606d79a37b6b3a504212a418914fc3d

Observation a73e99a7-7552-4f1a-9e21-59d69faf6d4b · inbound

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation cites this paper.

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:25:32.994831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T00:22:41.611893Z digest=sha256:c87d0bacb9f458e75e52a614bc16d21fa98a2728d9e83941abe45a3e06a00a10

Observation c89204c0-a66f-433b-9146-90382ced24ee · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:49.001243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:49.001243Z digest=sha256:858ba6bcef4f6987be9608cf3a0a07599095290e2c970e4ec3b167f8866e9a5c

Observation f4991661-264f-47e3-b6ef-96f837d1bf93 · inbound

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding cites this paper.

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:04.620503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:21:12.375936Z digest=sha256:5fcae59663c55528e43087bb8c5edd35e13159f74a2ce46d63af19793adfa566

Observation 89fc6ba7-bae3-419e-9de8-09a4738a8d4a · inbound

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images cites this paper.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.925300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.925300Z digest=sha256:bd2e8c5c5b9d7c9d871a5da09b3add7a94c8a7523ec9966310790eef782ee099

Observation 0a578d09-6ef1-47c0-acc8-fb7471255bc4 · inbound

Spatio-Temporal Grounding of Large Language Models from Perception Streams cites this paper.

Spatio-Temporal Grounding of Large Language Models from Perception Streams SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:26:02.678079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:10:45.837684Z digest=sha256:e52945ad5a6b7debe04399e52812cc3caebd8619af474e4691c9ac89418518ae

Observation 64bbe8e8-686b-4ab3-8c2a-8a9ecec42f57 · inbound

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training cites this paper.

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:11:03.815073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:16:08.687340Z digest=sha256:10c61857fedc4410d5afdd580d9fbd972e05eb61c472218bae789754440ce728

Observation 54209770-f05e-4e37-ba5e-3fc297cfd849 · inbound

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding cites this paper.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.692955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:20:08.404090Z digest=sha256:82cc6ade39f33400ca173407247ddec3a82fb7ce8207e64f3d01c53a392effde

Observation fa29e92c-af28-48a0-a4ed-1caaaf27aa2f · inbound

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding cites this paper.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.992535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:4839ddb92362473668641af1886fd368c09ae736caf6d9c62a52deabbaf0a7fc

Observation 9aa7b8c5-3f88-49ab-bdd7-ffe4c9f2a117 · inbound

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence cites this paper.

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.382386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T03:24:41.877312Z digest=sha256:27c80a251c1d98556891d063d1e73734dfea770357fd4dd9710ce1c6bfd59d00

Observation 88753f51-bdb3-4f4a-892d-3971c0915305 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.222326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T10:52:22.778489Z digest=sha256:418e5ec483d95c30afb683f08f7a25bd67df9cf107a3c960bdfdc90dd97918c7

Observation b0018adf-4f46-4670-9535-fda813036c2d · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:47.182481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:25:17.831116Z digest=sha256:b463ebcac8a8a95733ff3f084fa56a364f8584061050980c2a10cdc695d4e396

Observation 016b2ce0-c946-4469-add7-ec1b9b10e22c · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.560592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:02bd7827d88b4e04f79d90e7b8e64c2fbf62408173215427e2d055af824f56ce

Observation effb4d79-bfe3-4ed5-b356-8d1378a6ecba · inbound

VLM3: Vision Language Models Are Native 3D Learners cites this paper.

VLM3: Vision Language Models Are Native 3D Learners SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:14.019914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T07:45:31.978215Z digest=sha256:24f942bd1336c326ec264015fea6d66dd60b477a1cc5eeb7bc17401df749bbe6

Observation 6f97f28b-55ff-41db-b7ee-b719312b5659 · inbound

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks cites this paper.

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:48.424339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T05:59:29.302038Z digest=sha256:50b546ec116647d20ec28924e43d896dc7686a60010dbcf6efbbd4674e4741f3

Observation 443ec8bd-ee16-464b-b4fb-1deb9e64acf5 · inbound

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning cites this paper.

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:41.110274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:23:00.372251Z digest=sha256:8a28bbb51df662d659238b74c7d89a5f4a84309d8e139accab62f0188c280591

Observation bb376c85-f6f5-470b-b01f-51d3b0834949 · inbound

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video cites this paper.

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.308596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T16:16:41.412451Z digest=sha256:540b8f7fcea296fd308ddebbb917a263b19a6c6d22f04d06537fe85153afd42a

Observation 83309848-96fa-4915-a231-5899935daab2 · inbound

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents cites this paper.

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:20.587662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:49:20.587662Z digest=sha256:9538067c67335c5798f96781b3beda7b01b6114be886ea1e0bb3d6e18dc2739f