Pith. sign in

Paper Citation Record · LEDGER

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

As of 21 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.01709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01709 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:30:23.447065Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87dc45e4-5c3a-4b44-b641-e9305154b74e · outbound

This paper cites MineDojo: Building open- ended embodied agents with internet-scale knowledge,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MineDojo: Building open- ended embodied agents with internet-scale knowledge,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.062789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.283714Z digest=sha256:e7c29b3209c97742cf88892948338c01678f3cad15d361a69ae4cce2e6c25046

Observation 0298819c-8cbf-4328-a7a1-ce84e788a605 · outbound

This paper cites Guiding long-horizon task and motion planning with vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Guiding long-horizon task and motion planning with vision language models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.048813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.290888Z digest=sha256:7c66d179385a2573dc7bde2de912eaca231637cbe58402069589e51faa41af34

Observation 0e2cf3ad-b986-4786-b06d-431500771a15 · outbound

This paper cites RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.033839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.296266Z digest=sha256:231aac4574bcff539a3689b623cf90ad8b0243afecfbbebdf82ba27cea4cb435

Observation e059d53f-7add-4df2-8310-8f46e39ad889 · outbound

This paper cites GPT-4o System Card.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o System Card

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.301651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.301651Z digest=sha256:85feeb6989b9b57d04ed0869a72f64f7ffa51827b0bca8a4d719340f4742da14

Observation 88cd2bc2-5f58-48c9-b623-af5589a23aea · outbound

This paper cites Qwen2.5-VL Technical Report.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.306951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.306951Z digest=sha256:7fca5c5646cd4acd45935c93a5764111a57b0cbe5222e3542f22d863fa060403

Observation 2c2728c8-37d0-4036-a589-887823782f92 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.311855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.311855Z digest=sha256:57198bbecbe73a8c59b61f2e35146c0036ffdf2768c204a8b82c8a50d1ba5ba2

Observation c7ba22f9-c844-40a4-a82c-974a8572ea26 · outbound

This paper cites SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:25.019844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.317415Z digest=sha256:04c4e3f1f931a132ad14faae5d1ee508ae2a5469aa20d89a8154af5510fa93bd

Observation 7780fd84-6fd8-4081-a9f1-6dc0f4b9312d · outbound

This paper cites SpatialRGPT: Grounded spatial reasoning in vision language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialRGPT: Grounded spatial reasoning in vision language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.981742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.321859Z digest=sha256:3fb5bb8c06d0726d37964e0996b6f57b43848dc355ab9e190174ae8330ff6949

Observation 0af0b425-a73e-4e18-bdec-9a14837c73c1 · outbound

This paper cites Depth Pro: Sharp monocular metric depth in less than a second,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Depth Pro: Sharp monocular metric depth in less than a second,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.966492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.326320Z digest=sha256:748881681cb193f7b9079d9b3eb0cbebba568cf4e369fa4a63a8c5330a3a72b3

Observation c095fd52-af47-432e-bba2-c8b4e30fe240 · outbound

This paper cites SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.951835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.330689Z digest=sha256:46582155d2d016f712322a1d85ae086f3a3926ccc881b301390f44dd860f4914

Observation d0a5471f-d938-4158-b714-03cfe772a3aa · outbound

This paper cites Spatial reasoning with vision-language models in ego-centric multi-view scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial reasoning with vision-language models in ego-centric multi-view scenes,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.334788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.334788Z digest=sha256:24f8df2374336695046ff18b6000174c7bd07caa6a39b1741ae7e4222b5cea05

Observation 61b5f3fd-bcbb-40fb-8227-3e6125283936 · outbound

This paper cites Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.935532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.339916Z digest=sha256:9553f6b38480a0779ddb925ae3a1972a92e0ef6fef6745db655f1131f4cb4ac5

Observation 379014f2-0657-4b13-bf6c-f11a96ad612e · outbound

This paper cites BLINK: Multimodal large language models can see but not perceive,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models BLINK: Multimodal large language models can see but not perceive,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.920062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.344222Z digest=sha256:b2f45a7660a8ceedfaa9bce5d5fb1ef6f8f4ecd2c621b65d13d95f1a1fe1283c

Observation 92576b2f-e40d-428d-b5b6-4d2fece80fc6 · outbound

This paper cites Does spatial cognition emerge in frontier models?.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Does spatial cognition emerge in frontier models?

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.904450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.348928Z digest=sha256:2928c7530fe1678df80c2318214ea799925140da2138a5371d9b32ef7fb51919

Observation 10da9520-ed85-4023-899a-c9abdafac454 · outbound

This paper cites 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.886760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.353214Z digest=sha256:d2a15c244f8ea16cf36a21e04e5e5a682603fe7dc0ee8ba5b6ce8d81d2341424

Observation 52425481-1a17-43f5-a128-b77b882d059c · outbound

This paper cites Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.870793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.357336Z digest=sha256:1383b28e14407689b6f9b61607c973f4d775c7277dd52670ea4d2a941f7a7023

Observation a188f309-9bbc-4fe2-99cd-3c1766bfa726 · outbound

This paper cites Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.362399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.362399Z digest=sha256:c7de45be533eefb617e3635c0d158b15ba1cea37d0bacb72df5a41468570f7c2

Observation f656a0d9-57bc-4b85-b574-913e22555982 · outbound

This paper cites Perspective- aware reasoning in vision-language models via mental imagery simula- tion,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Perspective- aware reasoning in vision-language models via mental imagery simula- tion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.854558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.367294Z digest=sha256:549361015438d9efe937b0e43b1134ec3c4354f6ae01bdadf7e27a490c64f7ef

Observation 68fe7054-dc93-4cc0-a620-c70ac1b07b1d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.838503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.371626Z digest=sha256:d1933dbce2dc4bf50a6ce10a3dc311fba3275096cb4f9b761ca1901957a17374

Observation 476a93d6-80f9-4514-bf07-ed793b2fd837 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.376141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.376141Z digest=sha256:80e7471302adbaf8089eec2ebacc6e061e7e21af8185aab4eac5969025be6f60

Observation 7f30990f-4b6e-424d-8950-e472965b64d5 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.822322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.381375Z digest=sha256:e1ec220efa44fd4f97596f67e385e953e97c29a1f896b709633f5402766a76f1

Observation 519062f1-9c41-454c-9c55-856aa0b2c012 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.385878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.385878Z digest=sha256:ad7d3dbd6b46b2aa8e711cff23a065363243e4ed179f60d3d6e939d8d7728abb

Observation ca127aa8-a717-4cf9-9b82-6f0af81548ff · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.804991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.390647Z digest=sha256:65a2541dcb698e539810a373e1bb9a16b369d8c6b06f6e43fa7b7fcad85a9399

Observation c9ef6415-a750-4044-bbf2-9d2b8d5436b6 · outbound

This paper cites SAM 3: Segment anything with concepts,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment anything with concepts,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.785581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.395556Z digest=sha256:55689cae9b64eea12a602910fb80ec054f0f23371437b11b08a83b58b3162d29

Observation 5d7dcb07-d277-4052-902e-303b3de99b80 · outbound

This paper cites MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.768789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.404699Z digest=sha256:39e74a550f9a709616936ca08dcb2bde467a98191966f239f1d226d268149367

Observation 5ca7e6ad-3183-4a31-91e3-627813761065 · outbound

This paper cites Cubify anything: Scaling indoor 3D object detection,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Cubify anything: Scaling indoor 3D object detection,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.752144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.409810Z digest=sha256:b3aad0c14077c51609f491984c9e415bca7e053cd2f74b522d8c205623b01bd0

Observation 50d942a0-c2a3-4ed8-81c3-544d2814d9e5 · outbound

This paper cites Visual spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual spatial reasoning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.733925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.414585Z digest=sha256:db121c5ca790786bb3156a1bd47ec065d6569ee3b52aa70c03b1407cd7ee2d30

Observation 464a342f-43ad-40ed-87d0-2c1a875fdc33 · outbound

This paper cites SQA3D: Situated question answering in 3D scenes,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SQA3D: Situated question answering in 3D scenes,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.715872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.419014Z digest=sha256:4173cc043a06669e1f00697fc23c05c707432052db8c97f880131ee8e68ad90c

Observation 687b4756-2895-497f-a76b-87570acbe809 · outbound

This paper cites EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.423033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.423033Z digest=sha256:4bccc659e7c0b36d1a037df33e3ada80ddec8f9f207d5f2378fdf9ab91f17c0a

Observation c9df14e3-ee49-4a38-a992-99480078df7e · outbound

This paper cites GPT-4o mini: Advancing cost-efficient intelligence,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o mini: Advancing cost-efficient intelligence,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.699368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.428269Z digest=sha256:95f9ed168d1002a29751871fae5a20a2cbcb9c9da0798dc851104108f028b614

Observation cb2cf448-7aa0-448b-b327-272763c575e6 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.433519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.433519Z digest=sha256:93de4b1954a8867273ba4f427d7e17a775dd2d9c2251ed96fc359712d50f6d02

Observation 3a236af2-8546-427e-8fd4-cb0e601b11cf · outbound

This paper cites SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.683097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.438252Z digest=sha256:d081ebf1ac0dcc2bb2686e1e894d1c1fbb5b183b539d2402f18438df87728cb1

Observation 6ea41c6d-42eb-4ab2-9905-b6372b374b1b · outbound

This paper cites SpaceOm: Spatial reasoning with extended thinking traces,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceOm: Spatial reasoning with extended thinking traces,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.667452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.442696Z digest=sha256:795ce3dffff952673fa70ad4e117a061da9f7e88d1ce00729b0ee2e2fb530c0b

Observation 4a941714-7708-4e66-a5ee-6ccbe85f1a95 · outbound

This paper cites Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:30:23.647404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T22:30:23.447065Z digest=sha256:77b88da0078a18c2f523e156430b450e0dd827b1a8c917b4eaa789fe8b873f4b

Observation d991c080-7989-47ed-92cf-c07e2497874b · outbound

This paper cites SAM 3: Segment Anything with Concepts.

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment Anything with Concepts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:23.400209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:30:23.400209Z digest=sha256:fab0a49acf8eb804a24f945d89ed92bfb8731abd7a31147dd31d57108f745fc7

Pith citing papers

No inbound Pith citation observations are available.