Pith. sign in

Paper Citation Record · LEDGER

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2401.12168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.12168 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:56:07.947697Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T03:06:43.700595Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 033b7dc6-1001-45df-ae01-844f0722ad04 · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:07.947697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:07.947697Z digest=sha256:2c95e4cda02ce832b2e8c4a6a745c0e78e2a0c9b143233e1c5b0392ac2744b4e

Observation 6e0d4a50-0c88-4a17-9bce-c82ec548426a · inbound

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization cites this paper.

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:16:09.680303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T15:12:23.459579Z digest=sha256:40b2c1649e852f4cc1d292d0a6bf932d7b3866d2bc9871f9223aee69ac41ee37

Observation dc63545e-e5f0-43f7-a984-90ee3e098ed3 · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:a0217a72ce5ad9882030b09125e4f94cb8117165f26c1fde995c678e3e9cdaa2

Observation 1f806e86-4c61-4cc9-b2f2-229ee754e391 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:13:24.784848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:cbae72aace2b2eae260cd7d35aca2011a6292dd6395bbf1929dcad1de4b186ad

Observation 4042bd62-411d-4827-a6a5-d300abfe0de5 · inbound

Exploring Spatial Intelligence from a Generative Perspective cites this paper.

Exploring Spatial Intelligence from a Generative Perspective SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.849225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:45:46.261005Z digest=sha256:251f5c5a37edd306f03f2302a17584a661c588fb1f411af305904bf20f598508

Observation 8115d2dd-7059-40e7-ae26-f6c89343a181 · inbound

Fast Core Identification cites this paper.

Fast Core Identification SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:01:14.738363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T07:11:30.553696Z digest=sha256:b3170913241d683ab63aca268240ddde105f4e3a25d4ff5236966b14b9d76055

Observation 33c744aa-d26f-4a17-a6a6-16932b24b563 · inbound

Latent State Design for World Models under Sufficiency Constraints cites this paper.

Latent State Design for World Models under Sufficiency Constraints SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:04.461629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:55:31.825583Z digest=sha256:62c5dbcf18e2991f91db5fdceb87f8b92c39878ce44bf01649e9dccd5b3681ac

Observation 69da90aa-aaa9-48f4-b247-5247b855e41d · inbound

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cites this paper.

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:57.776205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T20:40:59.877854Z digest=sha256:809a24a38f5997fd15ecb3ca9565c6222de32d0024afd676a629ddb2198dbba7

Observation 23a5b44b-64c5-49ad-be51-9ce0c3695052 · inbound

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cites this paper.

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.149820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T16:57:03.172340Z digest=sha256:55252997814940fc6c86b80b2eed00cd8d199d32d9a1b42312dfed36dc3eb514

Observation b696d70e-7b11-4944-861a-72ea37bed505 · inbound

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning cites this paper.

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.715412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:46:21.821764Z digest=sha256:182bad106d1b49a272e64a4af806d9b7855f737e9f391c1f98261bf71ccebac8

Observation 8b054447-c513-4c7b-96e8-94c20819320c · inbound

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models cites this paper.

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.202939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:04:10.627420Z digest=sha256:9ebd62bea958ae44a2accd19772ea712cf0ea9fb1fae746f9c7f8fc260bd9479

Observation c1ccb55f-fade-44c5-a941-4f5298ab4061 · inbound

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching cites this paper.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.141917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:7367b182ad6fbfb255735adb8654c8e562a1a2c9f5f81c4858b56fa36850a849

Observation 19c5ae1c-b4f1-42ca-830b-a59f7df338e2 · inbound

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video cites this paper.

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.214345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T02:16:25.730555Z digest=sha256:dfc0a368f92ff46edf5cc28898a7928ba2e210b05cb10e7ecc4b4c316b15ba43

Observation 63d4cc74-765a-4972-8a29-1c166fdb8eb9 · inbound

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models cites this paper.

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:57.509150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T02:13:53.219095Z digest=sha256:e68aed705359a3e1f4932cdc49bf0b97ac6a5cbd8aacff5c6d53d737a3691898

Observation e2c309cb-eb6c-44c3-bca1-bf0e98e3f9da · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.661477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:d41a3eb4617d7aee2688041fb741951eeed294ba88e714311b248fa05aa6d755

Observation 71a61b51-b882-40bd-952c-b3fcc22ebe5f · inbound

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory cites this paper.

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.492839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:27:15.889885Z digest=sha256:ab38aee37e3283ae4545d06e3ee5ad6899c1351989303c8552055629222efa74

Observation abacc9ee-aef9-4d82-8fd3-78487bc291c6 · inbound

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping cites this paper.

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:41.219848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:01:37.880943Z digest=sha256:17e7a0e98c97b9b2ba4bc5a41fa1917519a303b74291b88e41b4be3d70dde2c0

Observation 6861a3c7-0b02-429d-b5b7-34abebf4b381 · inbound

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference cites this paper.

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.701881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T02:59:21.353359Z digest=sha256:4371fe09f90c67fdf9e65d5e55e733cf62a9e1021b2f23c9e52a53f4c4b31725

Observation e5a2891b-07d2-407c-b3a2-24b431e394d9 · inbound

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams cites this paper.

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T07:19:14.164642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:19:14.164642Z digest=sha256:ce2c969ad4015d6c9ff43387fb8753328460a9633e17edf73086edb32472817b