Pith. sign in

Paper Citation Record · LEDGER

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.01709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01709 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:30:23.447065Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87dc45e4-5c3a-4b44-b641-e9305154b74e · outbound

This paper cites MineDojo: Building open- ended embodied agents with internet-scale knowledge,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MineDojo: Building open- ended embodied agents with internet-scale knowledge,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.062789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.283714Z digest=sha256:9bc8f46b98a0b552c20162aca688c7989b65ab6041e540a9508c65298bacc847

Observation 0298819c-8cbf-4328-a7a1-ce84e788a605 · outbound

This paper cites Guiding long-horizon task and motion planning with vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Guiding long-horizon task and motion planning with vision language models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.048813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.290888Z digest=sha256:5361e4a2f18e2dd2f772e4a61a8605d02338c54d3db8d17a3b38e5ed27b91fd7

Observation 0e2cf3ad-b986-4786-b06d-431500771a15 · outbound

This paper cites RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.033839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.296266Z digest=sha256:d08ac18db6e4506bc0b99c883472b187cbd143ddf3ff5bcdca0dae9852103d49

Observation e059d53f-7add-4df2-8310-8f46e39ad889 · outbound

This paper cites GPT-4o System Card.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o System Card

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.301651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.301651Z digest=sha256:3f6eddd7e7cf3e0a7afa414df3670b5352c9d8bcfc250b19e88357c2ac255014

Observation 88cd2bc2-5f58-48c9-b623-af5589a23aea · outbound

This paper cites Qwen2.5-VL Technical Report.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.306951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.306951Z digest=sha256:a360b13448bdff84663ab93f064941bb8adc81ed65d7df5b78bf4c3fd2b2452b

Observation 2c2728c8-37d0-4036-a589-887823782f92 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.311855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.311855Z digest=sha256:42209a97b243f996dfea0b80088c516ad9c76fe9e90e759885efc116d3c599f4

Observation c7ba22f9-c844-40a4-a82c-974a8572ea26 · outbound

This paper cites SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.019844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.317415Z digest=sha256:a8639a064f27fb6c3ad6821682c4236b1ee46c5d1c33947116cf42b478ae5c41

Observation 7780fd84-6fd8-4081-a9f1-6dc0f4b9312d · outbound

This paper cites SpatialRGPT: Grounded spatial reasoning in vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialRGPT: Grounded spatial reasoning in vision language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.981742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.321859Z digest=sha256:f56c587cd5100bfba5d88f49d98af4cae1a15aa7bbcfe46d6205c0b410fda150

Observation 0af0b425-a73e-4e18-bdec-9a14837c73c1 · outbound

This paper cites Depth Pro: Sharp monocular metric depth in less than a second,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Depth Pro: Sharp monocular metric depth in less than a second,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.966492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.326320Z digest=sha256:a77d88a09be22ff5927e24fd802ef8ebcbde6213efff6e1213f973a54697316b

Observation c095fd52-af47-432e-bba2-c8b4e30fe240 · outbound

This paper cites SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.951835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.330689Z digest=sha256:fee2c8f172ae139cac7cce4e5b987fa0674e48b20d51f7bd2acbe32873ad5584

Observation d0a5471f-d938-4158-b714-03cfe772a3aa · outbound

This paper cites Spatial reasoning with vision-language models in ego-centric multi-view scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial reasoning with vision-language models in ego-centric multi-view scenes,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.334788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.334788Z digest=sha256:6fbd5e3f14c8efff307c29c4964024373ddb9965a9d4993e62293a52829085dd

Observation 61b5f3fd-bcbb-40fb-8227-3e6125283936 · outbound

This paper cites Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.935532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.339916Z digest=sha256:b8ee054f2640e3b6adf2084d6b72fda56d5a474532c2e08825d042c4e68405fe

Observation 379014f2-0657-4b13-bf6c-f11a96ad612e · outbound

This paper cites BLINK: Multimodal large language models can see but not perceive,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models BLINK: Multimodal large language models can see but not perceive,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.920062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.344222Z digest=sha256:1f287d1edd777997435cf18110c691dcb4890c21fea2a636dd1a09eb73e71cbd

Observation 92576b2f-e40d-428d-b5b6-4d2fece80fc6 · outbound

This paper cites Does spatial cognition emerge in frontier models?.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Does spatial cognition emerge in frontier models?

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.904450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.348928Z digest=sha256:dc6da92a72b2ddf66dc29533abc187236a719e8a7cfde41fb241f108dc7b6460

Observation 10da9520-ed85-4023-899a-c9abdafac454 · outbound

This paper cites 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.886760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.353214Z digest=sha256:cf55493596ad48c68a48f2829fd0814d9370040139a0bcd9bc82e4bcd7582ba5

Observation 52425481-1a17-43f5-a128-b77b882d059c · outbound

This paper cites Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.870793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.357336Z digest=sha256:2a5661f1a241542378bb65297d11e2d267e654d5b86d3f7baf7f14674f37c5a2

Observation a188f309-9bbc-4fe2-99cd-3c1766bfa726 · outbound

This paper cites Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.362399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.362399Z digest=sha256:18e163e07883000d3f4ed00cf9cc600bba5cc72c3cd1b8e6e59c5d09f329aa8b

Observation f656a0d9-57bc-4b85-b574-913e22555982 · outbound

This paper cites Perspective- aware reasoning in vision-language models via mental imagery simula- tion,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Perspective- aware reasoning in vision-language models via mental imagery simula- tion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.854558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.367294Z digest=sha256:a09f6b3209dde1b72af125e747b32e5b3bb64f474a59b8f41ac05455e4e7f281

Observation 68fe7054-dc93-4cc0-a620-c70ac1b07b1d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.838503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.371626Z digest=sha256:e1b063f619a9b788b02042aab056665cc53e379655f24c8376336430c6253a40

Observation 476a93d6-80f9-4514-bf07-ed793b2fd837 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.376141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.376141Z digest=sha256:c04bb83d41f7bd3f809585453a18e7eb5815d3880685f02e82c43296d8699f20

Observation 7f30990f-4b6e-424d-8950-e472965b64d5 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.822322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.381375Z digest=sha256:4c29e6a7511e6e863bf91225628b9a6850b4d0728af37ca8cfd3f7e01968a834

Observation 519062f1-9c41-454c-9c55-856aa0b2c012 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.385878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.385878Z digest=sha256:696af16220236c0b1b54dc9b4d5e546e455bca9f8d6916c2f06c62eeffc65dc7

Observation ca127aa8-a717-4cf9-9b82-6f0af81548ff · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.804991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.390647Z digest=sha256:ba1efc3b7bc0c99704be382fdabea1ed818feed9f9a0b59342c1a1ccbbf952dd

Observation c9ef6415-a750-4044-bbf2-9d2b8d5436b6 · outbound

This paper cites SAM 3: Segment anything with concepts,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment anything with concepts,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.785581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.395556Z digest=sha256:96c35558d0840ac35e35100d138ce46d06f240b33727e9faa312f5010009a1d5

Observation 5d7dcb07-d277-4052-902e-303b3de99b80 · outbound

This paper cites MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.768789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.404699Z digest=sha256:2c08c8d6288edf11bf99731eae3b08d4b075c2806daebb5abf39571a04fe34d0

Observation 5ca7e6ad-3183-4a31-91e3-627813761065 · outbound

This paper cites Cubify anything: Scaling indoor 3D object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Cubify anything: Scaling indoor 3D object detection,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.752144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.409810Z digest=sha256:362732673cf5194f5d3dc64c3be00479b5cee3aa1e996833eeaab1b708af01d0

Observation 50d942a0-c2a3-4ed8-81c3-544d2814d9e5 · outbound

This paper cites Visual spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual spatial reasoning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.733925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.414585Z digest=sha256:ef07c9b58181d36d665a1217ccad18c53279ff42fda88f089a45e936abaade4f

Observation 464a342f-43ad-40ed-87d0-2c1a875fdc33 · outbound

This paper cites SQA3D: Situated question answering in 3D scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SQA3D: Situated question answering in 3D scenes,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.715872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.419014Z digest=sha256:e3102b97b0eb8e21b0662e6b86828d363d530fa9a03a4efc14862ac1e9258ebb

Observation 687b4756-2895-497f-a76b-87570acbe809 · outbound

This paper cites EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.423033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.423033Z digest=sha256:0a803cee77ce85378959996b9a9f0235bdff32e124f65741781bc284087dc47b

Observation c9df14e3-ee49-4a38-a992-99480078df7e · outbound

This paper cites GPT-4o mini: Advancing cost-efficient intelligence,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o mini: Advancing cost-efficient intelligence,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.699368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.428269Z digest=sha256:a746990a434b8d7b078cc1f93dca21577069cde4a870a85b28f923a55257ee26

Observation cb2cf448-7aa0-448b-b327-272763c575e6 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.433519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.433519Z digest=sha256:3f189dc2e065e3627dd7cf6ed39d506f66cb1eefe1c7017eb7833a58213a2fc1

Observation 3a236af2-8546-427e-8fd4-cb0e601b11cf · outbound

This paper cites SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.683097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.438252Z digest=sha256:1e52bd8a07d6374fc62716986c4d7d3a94d97645d9ea28e152b88fe058619c29

Observation 6ea41c6d-42eb-4ab2-9905-b6372b374b1b · outbound

This paper cites SpaceOm: Spatial reasoning with extended thinking traces,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceOm: Spatial reasoning with extended thinking traces,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.667452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.442696Z digest=sha256:e2aabfb7396074e1395b51d0232962c792a58e9b29c633f13dd1d701376005b5

Observation 4a941714-7708-4e66-a5ee-6ccbe85f1a95 · outbound

This paper cites Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.647404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-04T22:30:23.447065Z digest=sha256:af13b5aaec11fd84e17c19a62ec0b8f6fcda2862da0ac8b9ead086e24eb3583b

Observation d991c080-7989-47ed-92cf-c07e2497874b · outbound

This paper cites SAM 3: Segment Anything with Concepts.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment Anything with Concepts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.400209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.400209Z digest=sha256:31468a9ba28e6445dc74dabf20d33962f63c8b3c4aabe2fc360b96f4569e3f73

Pith citing papers

No inbound Pith citation observations are available.