Pith. sign in

Paper Citation Record · LEDGER

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.01709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01709 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:30:23.447065Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87dc45e4-5c3a-4b44-b641-e9305154b74e · outbound

This paper cites MineDojo: Building open- ended embodied agents with internet-scale knowledge,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MineDojo: Building open- ended embodied agents with internet-scale knowledge,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.062789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.283714Z digest=sha256:d16adca510193ff968e742138c35bf7cdf7ccd35b6a3fcc074aea16e21d1aa08

Observation 0298819c-8cbf-4328-a7a1-ce84e788a605 · outbound

This paper cites Guiding long-horizon task and motion planning with vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Guiding long-horizon task and motion planning with vision language models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.048813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.290888Z digest=sha256:386277c433c02b295d5d7e04889c88c7f33298ad2aaf50d7720038c15b3b193a

Observation 0e2cf3ad-b986-4786-b06d-431500771a15 · outbound

This paper cites RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.033839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.296266Z digest=sha256:6d0502a4584041e5afe4b7ebcedb3a939af8dfd6b9fbf6be7213d1205c839d2a

Observation e059d53f-7add-4df2-8310-8f46e39ad889 · outbound

This paper cites GPT-4o System Card.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o System Card

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.301651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.301651Z digest=sha256:3f6eddd7e7cf3e0a7afa414df3670b5352c9d8bcfc250b19e88357c2ac255014

Observation 88cd2bc2-5f58-48c9-b623-af5589a23aea · outbound

This paper cites Qwen2.5-VL Technical Report.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.306951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.306951Z digest=sha256:a360b13448bdff84663ab93f064941bb8adc81ed65d7df5b78bf4c3fd2b2452b

Observation 2c2728c8-37d0-4036-a589-887823782f92 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.311855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.311855Z digest=sha256:42209a97b243f996dfea0b80088c516ad9c76fe9e90e759885efc116d3c599f4

Observation c7ba22f9-c844-40a4-a82c-974a8572ea26 · outbound

This paper cites SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.019844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.317415Z digest=sha256:6b5513cc8ff85fa85cb85feb4214d38fd7b6d34f5798e3c1a2ccd1617098c82a

Observation 7780fd84-6fd8-4081-a9f1-6dc0f4b9312d · outbound

This paper cites SpatialRGPT: Grounded spatial reasoning in vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialRGPT: Grounded spatial reasoning in vision language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.981742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.321859Z digest=sha256:93cc912c46ebe0b396327da4179e7f09fd33884cc9a8df4e9b2d37d676f0aa37

Observation 0af0b425-a73e-4e18-bdec-9a14837c73c1 · outbound

This paper cites Depth Pro: Sharp monocular metric depth in less than a second,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Depth Pro: Sharp monocular metric depth in less than a second,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.966492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.326320Z digest=sha256:567eba402fed171574b3adfd1bdc47c7008733aad3eb44f0e3ed8369a7252de8

Observation c095fd52-af47-432e-bba2-c8b4e30fe240 · outbound

This paper cites SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.951835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.330689Z digest=sha256:9b834ebbb1476b9ecd3d23ba644444fdaad8bfda2fc6485d231e0989e1b2f0c5

Observation d0a5471f-d938-4158-b714-03cfe772a3aa · outbound

This paper cites Spatial reasoning with vision-language models in ego-centric multi-view scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial reasoning with vision-language models in ego-centric multi-view scenes,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.334788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.334788Z digest=sha256:6fbd5e3f14c8efff307c29c4964024373ddb9965a9d4993e62293a52829085dd

Observation 61b5f3fd-bcbb-40fb-8227-3e6125283936 · outbound

This paper cites Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.935532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.339916Z digest=sha256:d0cef39b21f635655cfd5ffa6835ccf2de897545b14be684a82e9c1905dd63b0

Observation 379014f2-0657-4b13-bf6c-f11a96ad612e · outbound

This paper cites BLINK: Multimodal large language models can see but not perceive,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models BLINK: Multimodal large language models can see but not perceive,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.920062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.344222Z digest=sha256:a4a2c0b2735979cf175b26e48677ed0b4777045008618efe5112b5029f05db75

Observation 92576b2f-e40d-428d-b5b6-4d2fece80fc6 · outbound

This paper cites Does spatial cognition emerge in frontier models?.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Does spatial cognition emerge in frontier models?

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.904450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.348928Z digest=sha256:3cff5c9969a7fc380e598b4dc4cf1f77137b919cdc80204cb92ba7d58b1997b9

Observation 10da9520-ed85-4023-899a-c9abdafac454 · outbound

This paper cites 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.886760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.353214Z digest=sha256:cb60f9c823106152d89c72a601736ad670fb73ce16609c7d219bbc35f0096c40

Observation 52425481-1a17-43f5-a128-b77b882d059c · outbound

This paper cites Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.870793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.357336Z digest=sha256:ab95f167456c261da98cf8d193042ca2309ef861c91f571bdd6cc35167f76a09

Observation a188f309-9bbc-4fe2-99cd-3c1766bfa726 · outbound

This paper cites Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.362399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.362399Z digest=sha256:18e163e07883000d3f4ed00cf9cc600bba5cc72c3cd1b8e6e59c5d09f329aa8b

Observation f656a0d9-57bc-4b85-b574-913e22555982 · outbound

This paper cites Perspective- aware reasoning in vision-language models via mental imagery simula- tion,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Perspective- aware reasoning in vision-language models via mental imagery simula- tion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.854558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.367294Z digest=sha256:983cca2af2212a8c9889977da882a2c03651a7b06afe763a89e8887e76e80fb7

Observation 68fe7054-dc93-4cc0-a620-c70ac1b07b1d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.838503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.371626Z digest=sha256:a3ed37f1ed26e8b76f86d1dff767d794d9daf04884500c7abbe6d4d4fe196299

Observation 476a93d6-80f9-4514-bf07-ed793b2fd837 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.376141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.376141Z digest=sha256:c04bb83d41f7bd3f809585453a18e7eb5815d3880685f02e82c43296d8699f20

Observation 7f30990f-4b6e-424d-8950-e472965b64d5 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.822322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.381375Z digest=sha256:f70fc046f3bd6d59e3f5aae2a8421f8fd0b29d86fde12a69a591c0d125b0b005

Observation 519062f1-9c41-454c-9c55-856aa0b2c012 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.385878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.385878Z digest=sha256:696af16220236c0b1b54dc9b4d5e546e455bca9f8d6916c2f06c62eeffc65dc7

Observation ca127aa8-a717-4cf9-9b82-6f0af81548ff · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.804991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.390647Z digest=sha256:392ac8ec81e996ddcc7c99fdfc9f9ed8c5238e2a845666cecda784683321bb21

Observation c9ef6415-a750-4044-bbf2-9d2b8d5436b6 · outbound

This paper cites SAM 3: Segment anything with concepts,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment anything with concepts,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.785581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.395556Z digest=sha256:2ef8958493a387a760c748d184f0f49a0f164bb303174a25471eddbacd16d8c9

Observation 5d7dcb07-d277-4052-902e-303b3de99b80 · outbound

This paper cites MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.768789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.404699Z digest=sha256:886b91e09173db5e45f2ddb9bd5fc1a15285950680f6df2d4d1161fac52275a6

Observation 5ca7e6ad-3183-4a31-91e3-627813761065 · outbound

This paper cites Cubify anything: Scaling indoor 3D object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Cubify anything: Scaling indoor 3D object detection,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.752144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.409810Z digest=sha256:464dfd2d526d6f28abf9c4a6257db9dcc77f21b9987f30b450e78e480e71363f

Observation 50d942a0-c2a3-4ed8-81c3-544d2814d9e5 · outbound

This paper cites Visual spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual spatial reasoning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.733925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.414585Z digest=sha256:3a57a55eed51a8f6b4c8b75ad97d31a351ae3e5f8e1fb74b959d98f6fde3da67

Observation 464a342f-43ad-40ed-87d0-2c1a875fdc33 · outbound

This paper cites SQA3D: Situated question answering in 3D scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SQA3D: Situated question answering in 3D scenes,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.715872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.419014Z digest=sha256:71a3596bd1546e956bff62932999d19d23a7539edf4d36611ff139f1e6190330

Observation 687b4756-2895-497f-a76b-87570acbe809 · outbound

This paper cites EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.423033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.423033Z digest=sha256:0a803cee77ce85378959996b9a9f0235bdff32e124f65741781bc284087dc47b

Observation c9df14e3-ee49-4a38-a992-99480078df7e · outbound

This paper cites GPT-4o mini: Advancing cost-efficient intelligence,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o mini: Advancing cost-efficient intelligence,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.699368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.428269Z digest=sha256:f49906a43270d910378778cde23c1b46e700f2855d545cc5e867a7bbc6fa71e2

Observation cb2cf448-7aa0-448b-b327-272763c575e6 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.433519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.433519Z digest=sha256:3f189dc2e065e3627dd7cf6ed39d506f66cb1eefe1c7017eb7833a58213a2fc1

Observation 3a236af2-8546-427e-8fd4-cb0e601b11cf · outbound

This paper cites SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.683097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.438252Z digest=sha256:af102b777ae7d923d7a1aa45d25c01ed7e3df4caf49ec49372da92796c05e1ac

Observation 6ea41c6d-42eb-4ab2-9905-b6372b374b1b · outbound

This paper cites SpaceOm: Spatial reasoning with extended thinking traces,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceOm: Spatial reasoning with extended thinking traces,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.667452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.442696Z digest=sha256:571dce4eccb92ce634e304e8c6ead382dcb32662ac766a67edcd88172b59e9ff

Observation 4a941714-7708-4e66-a5ee-6ccbe85f1a95 · outbound

This paper cites Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.647404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T22:30:23.447065Z digest=sha256:7a86ddc154ef9a5e039ba50fa4c3d9adc3b2d38d3a2ca0b4adcc9da46fd8c200

Observation d991c080-7989-47ed-92cf-c07e2497874b · outbound

This paper cites SAM 3: Segment Anything with Concepts.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment Anything with Concepts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.400209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.400209Z digest=sha256:360294e2f9ed3ca867099be1165b09176ad40e379f3f997dce2c310840e98b7e

Pith citing papers

No inbound Pith citation observations are available.