Pith. sign in

Paper Citation Record · LEDGER

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

As of 19 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.05747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05747 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:15:18.724350Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42253e57-622b-40b3-afd3-45ffb4fe9c95 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.613690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.613690Z digest=sha256:ea865c6e51259b1330d9d7c6618124e7cd2fa3921c7084f4f70c701f5a7147a6

Observation e6bfcd7c-3d19-448f-ab10-1de98cd9705b · outbound

This paper cites Qwen3-VL Technical Report.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.618117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.618117Z digest=sha256:54d04b18c756ea90478cffe18a507915b2e3f3988a9941dc704290ec0ea64d6a

Observation 522bc332-6aa4-43f1-82c2-8e649d86a955 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.622060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.622060Z digest=sha256:1162718e490abd921cf6b59da428e01846837613ae5215ddcddcf0838b886660

Observation 48afbab3-510e-4247-8b5a-8ff59c37fd2b · outbound

This paper cites Mm-spatial: Exploring 3d spatial understanding in multimodal llms.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Mm-spatial: Exploring 3d spatial understanding in multimodal llms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.243096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.626142Z digest=sha256:9917b2f837e03ed153f01dbd22d6f31c93f8f05fcc4275ce7c8c159033eff86f

Observation 0c2694aa-3f0f-4a19-8dfa-12c24caad030 · outbound

This paper cites Robix: A Unified Model for Robot Interaction, Reasoning and Planning.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Robix: A Unified Model for Robot Interaction, Reasoning and Planning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.629999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.629999Z digest=sha256:847f8e921f849011274b48af75028495b38ffc62d1f0ca67ac7ff7abbd2d7423

Observation 536c9ea0-0a17-4a01-aa26-dab0ba904fee · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Blink: Multimodal large language models can see but not perceive

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.633875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.633875Z digest=sha256:9a8df8eed17a9efb8d7e9505bdd5f2ed0a9b177dcad9416bca238b5b94048223

Observation b2229948-db60-42f8-8924-5f2e21d61fa4 · outbound

This paper cites Gemini 3: Introducing the latest gemini ai model from google.https://blog.google/products/gemini/ gemini-3/, 2025.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Gemini 3: Introducing the latest gemini ai model from google.https://blog.google/products/gemini/ gemini-3/, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.226778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.637830Z digest=sha256:f0a1a89f3e578162f50bde7e573b70212a23a51c857320d6dc2158d83f68ea49

Observation e5cd3282-a386-4ca1-aef1-8d3c71cad95f · outbound

This paper cites GPT-4o System Card.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.641174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.641174Z digest=sha256:fc4108904f0e601908c52763e22c8d4b0f570c9c2fb2f9c7b3d0197da9334efb

Observation f7f25101-e1f7-444f-95cb-9e77f38f45b0 · outbound

This paper cites Artvip: Articulated digital assets of visual realism, modular interaction, and physical fidelity for robot learning.arXiv preprint arXiv:2506.04941, 2025.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Artvip: Articulated digital assets of visual realism, modular interaction, and physical fidelity for robot learning.arXiv preprint arXiv:2506.04941, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.645149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.645149Z digest=sha256:3ad701048f1336493ccf0864805b2daf99f88a51b67e0e6fda7b36cf9f5c4dc4

Observation 504a08fd-00e3-45b5-a17d-736f68c4a734 · outbound

This paper cites What’s “up” with vision-language models? investigating their struggle with spatial reasoning.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? What’s “up” with vision-language models? investigating their struggle with spatial reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.648647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.648647Z digest=sha256:66b73e4eb8f2397ce2c12f39afdd03988175d064d60f64c2beb283acda76a5f4

Observation a3f31ade-1f30-4e7e-97ff-e04324714436 · outbound

This paper cites Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.652690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.652690Z digest=sha256:a88d608ae00854b15e0f4e820c758361fa97e5a72a60be5e5592d86925e7b258

Observation b64c74b1-9ef2-4f84-835a-8aea39b4162d · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.655986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.655986Z digest=sha256:7b92be642afd24af97f7136448404ae2c38aff0f5df7907426d5b7cd9a8a92f3

Observation b6dab8f6-c1fb-4272-8c69-d8548ed13058 · outbound

This paper cites Sti-bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Sti-bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.205132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.659445Z digest=sha256:5ac1a1cf8bab1428664e5e8d250a34b8088cf855740e803006f1901a3a7c8238

Observation f644afb7-65c6-4da1-9893-c1236dcb0131 · outbound

This paper cites Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence.arXiv preprint arXiv:2512.10863, 2025.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence.arXiv preprint arXiv:2512.10863, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.662749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.662749Z digest=sha256:88e34e7075938c7093c2337f4fc1a97af3ecb32f5f25d99772e30805427d6402

Observation 73350104-8c20-4faa-98d6-3f227d77e263 · outbound

This paper cites Ost-bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Ost-bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.195156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.666338Z digest=sha256:9b1a314e0569fa58654d4727232813cdc2c2ea14ddd83f84f89e5079267e79bf

Observation 31c8c0ee-1531-4362-a436-6028f7be6b1f · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.670298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.670298Z digest=sha256:9e9c05af0fafdfbdd3337b6f43fd9ba0ce2be60d23269e2f59a800d80c86474d

Observation a554ac96-8048-45ec-8bfe-1dad6dd8746b · outbound

This paper cites Nvila: Efficient frontier visual language models.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Nvila: Efficient frontier visual language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.674326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.674326Z digest=sha256:3113cf7e0d9e30f7a71e3b59b721e3233aa7d4259928c69f9a3e5ebe6e5d9fc1

Observation 1be615e0-80f8-437c-9bdc-83ed3070ae71 · outbound

This paper cites 3dsrbench: A comprehensive 3d spatial reasoning benchmark.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? 3dsrbench: A comprehensive 3d spatial reasoning benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.678285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.678285Z digest=sha256:6fc71c65145c1a1d0e239a9651d7fe810aea6168fed6369b3b7446042fb71df8

Observation 8c2f3515-ed64-4a4b-afcb-da22c99c892c · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Openeqa: Embodied question answering in the era of foundation models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.167241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.682295Z digest=sha256:0f40eb5de52b786079bb1f4d7d5c971078f0e6d95b7ea8ce8d6240a8dc42eea8

Observation c4455015-2bf4-497e-b61a-b2d99d260117 · outbound

This paper cites Cosmos-reason2.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Cosmos-reason2

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.156499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.686302Z digest=sha256:ee1f679ffbf08ca724b54f1eb93b2ecc8cbdd3a142e13a8452bbd570c169a051

Observation ab821289-cf90-444a-a8b5-b00d22bd9ba8 · outbound

This paper cites Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.146707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.689975Z digest=sha256:52824a65be54580202a5c14db55a9dbe792cf3632e39a3016062ae4c48c373b1

Observation e0904a3a-a40d-4f8b-b990-0cac299f58c8 · outbound

This paper cites Seed1.8 Model Card: Towards Generalized Real-World Agency.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Seed1.8 Model Card: Towards Generalized Real-World Agency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.693233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.693233Z digest=sha256:8908f6072e6ce0c30cbe84dafe1105f5bb6edabcfabbf685c3228e989f1c1c63

Observation 29f45c7e-65c3-45db-8d64-7bef050f313b · outbound

This paper cites OpenAI GPT-5 System Card.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? OpenAI GPT-5 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.696882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.696882Z digest=sha256:22295141bff56deecc4cb26f0d1d24556f706435e01f6d1e11875c61e91a35ec

Observation f689b408-4562-4dee-b6ea-a9cf268ce6e9 · outbound

This paper cites Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.700804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.700804Z digest=sha256:195e7b5679165b31362fbe5a047ff0a609eb3f68366da3b337db7e625a8bdeb1

Observation d9576de4-a15a-465d-85f3-fd9f156fc3e9 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.136900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.704875Z digest=sha256:1ff8dea7371c2007191baaf5f69e6909ee2dbaab2a0b600a85498c25f4418091

Observation d628961d-d841-4acf-b113-9d52bf31cbbe · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.707962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.707962Z digest=sha256:1e1f5b4977aab820b96f11ae6d54bc0e43b834cfd72d28a5736fdb5dca112374

Observation 1a13155f-73e4-4fd4-b883-9ab534fe02ae · outbound

This paper cites Spatial457: A diagnostic benchmark for 6d spatial reasoning of large mutimodal models.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Spatial457: A diagnostic benchmark for 6d spatial reasoning of large mutimodal models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.126632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.711372Z digest=sha256:a082d8a9fc557451eef590c562d7afb44c001ce5ff7a834be00f1ebeca7dbe80

Observation aec5670f-4081-4fa7-a379-d4d33b1979ff · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.714534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.714534Z digest=sha256:2e0d9ff8cce156e5010c422e5fd2cf3ef9d5ac2483c6b135762a3ba61a2827ca

Observation 54fc181a-dde9-4407-9ec2-c63dd0542529 · outbound

This paper cites Cambrian-s: Towards spatial supersensing in video.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Cambrian-s: Towards spatial supersensing in video

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.110735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.717864Z digest=sha256:9c3e84ba16c695644bda8e36716be256b9cfba358c894a2b350e1e18b189ed6c

Observation 65ce0823-fb8c-4075-af19-503139ff9bc6 · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:15:18.720958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:15:18.720958Z digest=sha256:6c9d0eba51f06a755d1954bd5cf7c0a285e621a17f46530ef4d21ef9bf7a1616

Observation 1e5482ab-af6b-4d27-be5f-a376c33dfec6 · outbound

This paper cites From flatland to space: Teach- ing vision-language models to perceive and reason in 3d.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? From flatland to space: Teach- ing vision-language models to perceive and reason in 3d

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:15:19.100566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T00:15:18.724350Z digest=sha256:7fdc576363ece5ade20d27e0d2ee7b194c1f57c5496625ce6ebba258630a26c8

Pith citing papers

No inbound Pith citation observations are available.