Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:30:23.447065Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.01709.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:30:23.447065Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 87dc45e4-5c3a-4b44-b641-e9305154b74e · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MineDojo: Building open- ended embodied agents with internet-scale knowledge,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0298819c-8cbf-4328-a7a1-ce84e788a605 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Guiding long-horizon task and motion planning with vision language models,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e2cf3ad-b986-4786-b06d-431500771a15 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models RoboSpatial: Teaching spatial understanding to 2D and 3D vision- language models for robotics,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e059d53f-7add-4df2-8310-8f46e39ad889 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o System Card
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88cd2bc2-5f58-48c9-b623-af5589a23aea · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2728c8-37d0-4036-a589-887823782f92 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ba22f9-c844-40a4-a82c-974a8572ea26 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialVLM: Endowing vision-language models with spatial reasoning capabilities,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7780fd84-6fd8-4081-a9f1-6dc0f4b9312d · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialRGPT: Grounded spatial reasoning in vision language models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0af0b425-a73e-4e18-bdec-9a14837c73c1 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Depth Pro: Sharp monocular metric depth in less than a second,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c095fd52-af47-432e-bba2-c8b4e30fe240 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpatialPIN: Enhancing spatial reasoning capabilities of vision-language models through prompting and interacting 3D priors,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0a5471f-d938-4158-b714-03cfe772a3aa · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial reasoning with vision-language models in ego-centric multi-view scenes,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61b5f3fd-bcbb-40fb-8227-3e6125283936 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 379014f2-0657-4b13-bf6c-f11a96ad612e · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models BLINK: Multimodal large language models can see but not perceive,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92576b2f-e40d-428d-b5b6-4d2fece80fc6 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Does spatial cognition emerge in frontier models?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10da9520-ed85-4023-899a-c9abdafac454 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models 3DSRBench: A comprehensive 3D spatial reasoning bench- mark,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52425481-1a17-43f5-a128-b77b882d059c · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a188f309-9bbc-4fe2-99cd-3c1766bfa726 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f656a0d9-57bc-4b85-b574-913e22555982 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Perspective- aware reasoning in vision-language models via mental imagery simula- tion,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68fe7054-dc93-4cc0-a620-c70ac1b07b1d · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 476a93d6-80f9-4514-bf07-ed793b2fd837 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f30990f-4b6e-424d-8950-e472965b64d5 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual sketchpad: Sketching as a visual chain of thought for multimodal language models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 519062f1-9c41-454c-9c55-856aa0b2c012 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca127aa8-a717-4cf9-9b82-6f0af81548ff · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Grounding DINO: Marrying DINO with grounded pre- training for open-set object detection,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9ef6415-a750-4044-bbf2-9d2b8d5436b6 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment anything with concepts,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d7dcb07-d277-4052-902e-303b3de99b80 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models MM-Spatial: Exploring 3D spatial understanding in multimodal LLMs,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ca7e6ad-3183-4a31-91e3-627813761065 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Cubify anything: Scaling indoor 3D object detection,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50d942a0-c2a3-4ed8-81c3-544d2814d9e5 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Visual spatial reasoning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 464a342f-43ad-40ed-87d0-2c1a875fdc33 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SQA3D: Situated question answering in 3D scenes,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 687b4756-2895-497f-a76b-87570acbe809 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9df14e3-ee49-4a38-a992-99480078df7e · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models GPT-4o mini: Advancing cost-efficient intelligence,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb2cf448-7aa0-448b-b327-272763c575e6 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a236af2-8546-427e-8fd4-cb0e601b11cf · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceThinker-Qwen2.5VL-3B: A thinking/reasoning VLM for quantitative spatial reasoning,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ea41c6d-42eb-4ab2-9905-b6372b374b1b · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SpaceOm: Spatial reasoning with extended thinking traces,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a941714-7708-4e66-a5ee-6ccbe85f1a95 · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models Spatial-SSRL: Enhancing spatial understanding via self- supervised reinforcement learning,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d991c080-7989-47ed-92cf-c07e2497874b · outbound
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models SAM 3: Segment Anything with Concepts
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.