Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:15:18.724350Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.05747.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:15:18.724350Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 42253e57-622b-40b3-afd3-45ffb4fe9c95 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6bfcd7c-3d19-448f-ab10-1de98cd9705b · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 522bc332-6aa4-43f1-82c2-8e649d86a955 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48afbab3-510e-4247-8b5a-8ff59c37fd2b · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Mm-spatial: Exploring 3d spatial understanding in multimodal llms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0c2694aa-3f0f-4a19-8dfa-12c24caad030 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Robix: A Unified Model for Robot Interaction, Reasoning and Planning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536c9ea0-0a17-4a01-aa26-dab0ba904fee · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Blink: Multimodal large language models can see but not perceive
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2229948-db60-42f8-8924-5f2e21d61fa4 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Gemini 3: Introducing the latest gemini ai model from google.https://blog.google/products/gemini/ gemini-3/, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e5cd3282-a386-4ca1-aef1-8d3c71cad95f · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? GPT-4o System Card
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f25101-e1f7-444f-95cb-9e77f38f45b0 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Artvip: Articulated digital assets of visual realism, modular interaction, and physical fidelity for robot learning.arXiv preprint arXiv:2506.04941, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 504a08fd-00e3-45b5-a17d-736f68c4a734 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? What’s “up” with vision-language models? investigating their struggle with spatial reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f31ade-1f30-4e7e-97ff-e04324714436 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b64c74b1-9ef2-4f84-835a-8aea39b4162d · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6dab8f6-c1fb-4272-8c69-d8548ed13058 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Sti-bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f644afb7-65c6-4da1-9893-c1236dcb0131 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence.arXiv preprint arXiv:2512.10863, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73350104-8c20-4faa-98d6-3f227d77e263 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Ost-bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 31c8c0ee-1531-4362-a436-6028f7be6b1f · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a554ac96-8048-45ec-8bfe-1dad6dd8746b · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Nvila: Efficient frontier visual language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1be615e0-80f8-437c-9bdc-83ed3070ae71 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? 3dsrbench: A comprehensive 3d spatial reasoning benchmark
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2f3515-ed64-4a4b-afcb-da22c99c892c · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Openeqa: Embodied question answering in the era of foundation models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4455015-2bf4-497e-b61a-b2d99d260117 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Cosmos-reason2
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ab821289-cf90-444a-a8b5-b00d22bd9ba8 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e0904a3a-a40d-4f8b-b990-0cac299f58c8 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Seed1.8 Model Card: Towards Generalized Real-World Agency
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29f45c7e-65c3-45db-8d64-7bef050f313b · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? OpenAI GPT-5 System Card
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f689b408-4562-4dee-b6ea-a9cf268ce6e9 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9576de4-a15a-465d-85f3-fd9f156fc3e9 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d628961d-d841-4acf-b113-9d52bf31cbbe · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a13155f-73e4-4fd4-b883-9ab534fe02ae · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Spatial457: A diagnostic benchmark for 6d spatial reasoning of large mutimodal models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aec5670f-4081-4fa7-a379-d4d33b1979ff · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Thinking in space: How multimodal large language models see, remember, and recall spaces
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54fc181a-dde9-4407-9ec2-c63dd0542529 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Cambrian-s: Towards spatial supersensing in video
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 65ce0823-fb8c-4075-af19-503139ff9bc6 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e5482ab-af6b-4d27-be5f-a376c33dfec6 · outbound
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? From flatland to space: Teach- ing vision-language models to perceive and reason in 3d
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.