Pith. sign in

Paper Citation Record · LEDGER

What's "up" with vision-language models? Investigating their struggle with spatial reasoning

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2310.19785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19785 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:29:05.587498Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:31.500278Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 557c7b6f-1b2d-44cc-86b2-e53dbf62dc5f · inbound

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning cites this paper.

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:05.587498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:05.587498Z digest=sha256:9f6f6d9b8011a2686481f217f171f8e9cfc243437f9d5c38747e52432ce2c107

Observation 18fd7e79-80ef-4760-a495-0665976837fb · inbound

Visual Agentic AI for Spatial Reasoning with a Dynamic API cites this paper.

Visual Agentic AI for Spatial Reasoning with a Dynamic API What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:23:01.255895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:23:01.255895Z digest=sha256:576a712e0e2c90bcfce144804b1f8650e39ae3bd6605a5cdac38ace627842473

Observation cdc2aa98-44b8-4a7a-84fd-17e49a9bb7ad · inbound

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning cites this paper.

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:27:17.819795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T00:26:58.273861Z digest=sha256:4b20f2901814718c870286b985a6d3bad8008eb9eac245cdeb3f98a9a9566427

Observation e32d6721-ca23-4bb3-b7a1-fb4433f8dd39 · inbound

IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth cites this paper.

IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:40.940387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:40.940387Z digest=sha256:25d9757a0987f0cb99359978c721564f1191a4128b0caee5e91af64126315e8d

Observation 95fe4f56-daef-4ca4-a7de-5b3962301881 · inbound

Synthetic Visual Genome cites this paper.

Synthetic Visual Genome What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.253961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.253961Z digest=sha256:4f1920aeebc019968d99eb4c3ce82568b8cb4823b28a123880ba08e9bca59717

Observation 14cc2461-dbfe-4320-a637-c8ec647f4f39 · inbound

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks cites this paper.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.021243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.021243Z digest=sha256:312f483afac85478e79e2d4b3032f059725f56d4b403b9b3e643a9b843da4236

Observation f48f0d58-acf1-479d-8c51-4fe3d9fbbf72 · inbound

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing cites this paper.

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:58:10.274256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T04:58:10.202784Z digest=sha256:b51cdd1de7c00e71521e25a8b66bb0624b5be99385b2e0b0b05fb6583c700084

Observation d5ab4d6f-3172-4960-8696-ad1e559c4e15 · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:59.031874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:59.031874Z digest=sha256:263c0aa155763cd92bd2fedb43fddfe117d53711332c32a50aed1d9f5be6b05e

Observation d82a0fa9-6153-4417-9d18-8c851062505e · inbound

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor cites this paper.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.458422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.458422Z digest=sha256:0fce0489d5b882cc4a5fbb9d9f827038236821a9061147f0d0b4cb825d3a7151

Observation e445760e-2d7b-40a4-99b8-3f4d08f63cc4 · inbound

Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent cites this paper.

Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:05:10.615746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:05:10.615746Z digest=sha256:493cc0b684156573a928d671ec9837e5523bd8337b7b40e5432720b85af01400

Observation bde968c4-e17f-4545-828d-34d1439b738c · inbound

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation cites this paper.

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:06:51.620061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T22:04:34.235731Z digest=sha256:8e2661d6e61297ec436c688281f8c3153309a01800f0ea14e9f35072401e0298

Observation 513f15bc-ecfb-471f-8d8c-1350fc8bcb93 · inbound

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs cites this paper.

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:41:30.121277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T14:41:20.403259Z digest=sha256:f1d498b6fc219b21338cbc5304a494ba17cb47178febc8d067a5d93b249cf0f2

Observation 0a67e95a-d0af-41c6-a213-7ea932f9501a · inbound

Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models cites this paper.

Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:51:13.615332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T09:48:39.943352Z digest=sha256:ed3155fc243797ea703b7c73b3aa47510fb70c88ba141a2835b27ac54f5778a2

Observation a74d0c71-6a7e-4dac-b709-2950e1d0698a · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:51.252391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:51.252391Z digest=sha256:f5f7880d6ecbc3a279dbeed9beaf050766053e03a5de6704ce0ba19c4c084c90

Observation f189b524-7a18-4198-b363-cc1395db5a7b · inbound

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning cites this paper.

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.645102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T22:04:18.591594Z digest=sha256:e8003e11c3c829bf12491be2cf2e4cc19ecfc4a419bb982adab3ea2651f12f8a

Observation 97ef63b7-0aa3-4bc8-b432-d0f4a972e6c8 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:58:49.613832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T17:57:00.909897Z digest=sha256:3fd6401de67a763e27e5deb60cfcefa962eeba0fc84825ff84df37ea9fa87dd0

Observation 1ba9306a-6720-45c0-a11f-1bfb7f5151b3 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.601449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T19:02:05.125937Z digest=sha256:152085f32c88fb02b3085f587dfdbf30e1c111719ae919ee5d049525eebd6353

Observation 43b3f88c-b657-4af1-baa6-fd6a107c9f6f · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:04.389858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:d16886b583cf761a7e107f186ca53e690b400e680bdb5788f37552c8ec0415c5

Observation 826c224a-d034-4193-96e7-0ed9924d3630 · inbound

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images cites this paper.

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:20:25.013156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T05:17:45.034348Z digest=sha256:3fe3c38a35c93585bbdda426efa3e369904ed772b2edbb62481c15fe1657f9c5

Observation be9305c5-d739-4711-8cd3-db43cf1edadf · inbound

PhotoFlow: Agentic 3D Virtual Photography Missions cites this paper.

PhotoFlow: Agentic 3D Virtual Photography Missions What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:35:21.037827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T04:33:24.355622Z digest=sha256:3ae45650b6e4c729373deb33216faed16f01acbf873bf00f05fe91ee5d7856e8

Observation 746f3939-9725-45d9-9779-085f47d937c9 · inbound

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning cites this paper.

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:47:48.101396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T10:41:41.697216Z digest=sha256:e659dbf47e9de0b6519d46606e448ea6f4105d355c12903e26370e55b5a83ce1

Observation c77cc03e-5404-452b-bcb8-d46ae99592b1 · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:31.501685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:7d82a058598ccb4854cb337975226d0c7d4eddf1f0b79bf22e5ed90345feb7e6

Observation c6c3fe19-3c6a-4e47-ac0c-3de9fa4e25c9 · inbound

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning cites this paper.

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:41.057127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T06:23:00.372251Z digest=sha256:f71458cf5247c02a9918e54e0eef706708fe822cb2897a045313590bfb1e9a83

Observation 90abfc1b-b3ca-4c13-82c4-3fb339e1d8ea · inbound

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles cites this paper.

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T03:39:41.724196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:39:41.724196Z digest=sha256:d484a83431f288736a19219f703a56afd23b8b48fdac134567253d3e0b821fcf

Observation a9714ecd-46c7-4709-8dee-e8044ab7beef · inbound

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles cites this paper.

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T04:25:21.491757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:25:21.491757Z digest=sha256:608ef6efe4fc4fe3b499c7e786f3a6e9556fdf49c7d50b0e1acd0d9cd5f6a659