Pith. sign in

Paper Citation Record · LEDGER

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

As of 3 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2501.01428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01428 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T13:46:28.179433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:44:27.702116Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 68282466-6117-4a8f-bee3-501a1f557853 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.929625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:29dd7a61e7302ef1ac8c8945a5ef329d256ac98a2e8b12ea43fc0a06581d557e

Observation 65e82996-56b2-4918-875e-caced0b314df · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.395130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:fb19bf83fd55a6c5e2e2875cfd598b9b120636dfbc8200e95cfd8d22e3771572

Observation c1758b5f-e01e-4635-adf9-ceae67bd6f17 · inbound

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing cites this paper.

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:58:10.236698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T04:58:10.202784Z digest=sha256:45319d8650d8a42a8355428e3a8966964e0d0e971d5a8cd02640442d7f828588

Observation b916cc03-05e3-4628-a445-d4d239b6680e · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.210840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:659b3b28fd8ef4b2006577f56660c9fa4daad6b663a1b01698ca088f5b89c055

Observation e6481afa-a945-489a-9415-830dd543b5bd · inbound

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding cites this paper.

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T13:46:28.179433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:46:28.179433Z digest=sha256:9b2ebb941f6b68c190f93b75d687f8af6f4d4e9a1ee177508ad8b798eda8e68c

Observation 48b73ff5-a861-434c-9f81-a910cb1ca327 · inbound

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments cites this paper.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:40.003799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:40.003799Z digest=sha256:30aff8056c1cecbd80c4f176786ecb09b114d2eab01ba107521f1164b63a1ec6

Observation 67e5a825-e443-4632-9d14-e9c0fcec8480 · inbound

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations cites this paper.

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:56.059324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T14:31:03.909336Z digest=sha256:63ca7aaefae5f67ce618a9ea21ffb0d96925e9e86fd9a9254a25242574eb826d

Observation 0ee15ea0-69cc-44c5-9b2c-8fcee3ae62f6 · inbound

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models cites this paper.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.926358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.926358Z digest=sha256:b4932b2297bc66e32634b54fc91b3463ee811d07471f49992d1f278c68f4a66d

Observation 4b25ab6a-9274-48e0-864f-28bfc1251243 · inbound

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding cites this paper.

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:30:22.387360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T09:30:03.668178Z digest=sha256:f60a175d8146a758a0ccbdabdae615379be2cbedad538701cc47ba6e1b951cf1

Observation fc2db593-5ae4-4870-a65a-2f0695f038cf · inbound

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs cites this paper.

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.407430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T19:45:33.950587Z digest=sha256:978fb02a1e95aa23c2973d692025e3f7922cb209cc55e927e6d599a913a7dbe4

Observation a248adb3-8099-4504-aa8b-55f56ec5bdd2 · inbound

3D-IDE: 3D Implicit Depth Emergent cites this paper.

3D-IDE: 3D Implicit Depth Emergent GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:38:11.477870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T22:34:04.833557Z digest=sha256:9524b58c621ba8fbcbaf14effd5f20914f3f1f8a54ae4cc3cbf95005c5f7686f

Observation c0a9d8a8-42c3-4b3c-b50e-a0de98bba6d2 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:43:22.885945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T22:41:09.840792Z digest=sha256:06b9d9efd623079bd6caac1d0b4b72d50d4f1927610a25fa9156eb4937a780ca

Observation e7577ef1-9c43-4f33-97a1-877f78d5de67 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T14:39:55.177552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T14:39:55.177552Z digest=sha256:f954d910b399dc8f8f75deed2cf13af00070602190852e3e89425f5f0a050464

Observation be164b81-94e3-48fe-859e-8a844c2fc348 · inbound

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs cites this paper.

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:50:49.009736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T19:32:59.381755Z digest=sha256:6eb9c31d01a6c7c6f57cd49d5b563a3d9996ba7120ab210225df5a39265660ff

Observation c5b05d49-f223-4fa4-9dd5-c4da6aee2aad · inbound

Geometry-Guided 3D Visual Token Pruning for Video-Language Models cites this paper.

Geometry-Guided 3D Visual Token Pruning for Video-Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:09.965323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T05:49:38.346274Z digest=sha256:2afc9888c9b9060c78c6011c3d2545ff1a90301ef70719b4e50006896813fde6

Observation 190ecbbf-3cc5-431f-a4d6-fb673951d3b9 · inbound

Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning cites this paper.

Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:00.789767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:29:56.514907Z digest=sha256:e899e0510e5d14c9024f87045d6ce5165254b52b74bc2a4b49f34684350f3134

Observation 5c25f57d-8586-4dea-a859-da4d1f564f55 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:36.416075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:ab9670b25dfb619a3470e8326cf34396d719eff38331715fe1a4a4b3a000c84c

Observation 733a178d-8570-40d1-876e-d01d40dfcd6a · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:23:40.994480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T19:20:04.468206Z digest=sha256:5cbf65e8806e3985665e2f33045a9a87a588679f431838b0c6355544c086ec5d

Observation fc52b842-e9ba-4cf5-89f1-e34b8519cce7 · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:51.097379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T07:54:52.926995Z digest=sha256:c375cd278767d8ebf1947162e5cbfc887af597f9e874dfd959da76b3b0f333a6

Observation d3284230-31a6-43ff-b81e-d4d6ea12c62e · inbound

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy cites this paper.

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.494005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T13:28:43.538541Z digest=sha256:6d10056f545ed9e19b982bf639a1de751b15f1da4556335c79085bcada04aaaa

Observation 10454cfd-1cdb-4d81-94e9-768de310711f · inbound

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models cites this paper.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:01.235024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:52dd611026e36cb434e9ce62a3c97938e833b393d60a96593e5dd265ab190af7

Observation 578c5a4d-d994-403e-ba8a-c0ac90ec47c1 · inbound

SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs cites this paper.

SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:03:29.805356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T13:54:04.992565Z digest=sha256:f34658f8965fb297a1bffa87b48eba39196a18e87a2ba7aaf3e074a44616c425

Observation 37d34a9c-257d-4de4-b4f5-cb1e30cd20f8 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.604835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:ae995dc70f26d405a9df8668215b575bac2a06ddc56ea1109f7dc540dbfc715e

Observation dbe0e397-02c2-4bfe-a85e-f3bc4b870a69 · inbound

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation cites this paper.

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.361485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:33:32.671550Z digest=sha256:42f66acaaab53afa844eacc382edf4cc8ab00ce137898c2b6278437fb0cb5257

Observation 00859666-c515-43b7-b10d-4520d40da0fd · inbound

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors cites this paper.

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:09.633602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T22:49:02.846428Z digest=sha256:33d43f19d4b8caaf82d2717127bf3d1066f1649c3f9156ad419a7117b30b5e81

Observation 5c320a40-97dd-43ea-91ec-3eb6229a6660 · inbound

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding cites this paper.

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:59:25.790387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T18:40:20.588652Z digest=sha256:f4c3ff8447aa932c97dd24fe2e368ee025b33197da8e749f91416cb195b40ba0

Observation 5321faac-0870-41a7-987f-5e556943f58d · inbound

SpatialSV: Internalizing Interpretable 3D Spatial Awareness in MLLMs via Task-Oriented Visual Supervision cites this paper.

SpatialSV: Internalizing Interpretable 3D Spatial Awareness in MLLMs via Task-Oriented Visual Supervision GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:30.832906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T17:59:21.032539Z digest=sha256:75c231f5ed7c5f4b4d3138eeb0648eb57ce8b250173c474225d5def41c150c62

Observation 5d5a29e3-9795-4ee7-8890-f32df3da4a9b · inbound

Agentic Collaborative Cognition for Zero-Shot 3D Understanding cites this paper.

Agentic Collaborative Cognition for Zero-Shot 3D Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.795867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T00:22:06.082183Z digest=sha256:2e79576d90ff6b6f96e7dc50040e356998f86cbbbdc5d505fe7907575287b08f

Observation 89812ce1-780f-4e00-a266-22b4f6d2de64 · inbound

Agentic Collaborative Cognition for Zero-Shot 3D Understanding cites this paper.

Agentic Collaborative Cognition for Zero-Shot 3D Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:59:52.627994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T05:37:41.407624Z digest=sha256:61b333f63d83789d44bbd3f03138941e303de5c21b74c713d2cfeb93d257093d

Observation 3c3a0780-16a3-44a3-ab22-b1bf0f143544 · inbound

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures cites this paper.

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:57.264648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T04:30:56.107507Z digest=sha256:47db649759e740e0b715bea06bbfef460566f4a7bf29924565b4bbc172fa3887

Observation 3d876ced-b9fb-4c34-9a35-ee6c351c74ff · inbound

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video cites this paper.

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.281330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T16:16:41.412451Z digest=sha256:41efc6d622a5eaf3d681fe4c627d22b9786d0238966970f8880893905fada3c0

Observation 94eb3c21-e4a1-4a59-9fde-60aeeb1b28de · inbound

Holo-Captioning: Toward the Text Equivalent of 3D Scenes cites this paper.

Holo-Captioning: Toward the Text Equivalent of 3D Scenes GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T06:12:48.722467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:12:48.722467Z digest=sha256:e2aca127717b6cbd8a8304872208af59b17381af3b9c7a394f66216ed0e36f79

Observation feaaff6a-f67b-4390-aa46-1d8ffe0c02de · inbound

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering cites this paper.

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T21:50:49.625343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:50:49.625343Z digest=sha256:6581babc2cc7c2e116580aaf7fd47f8ad1bd950f778620ed1c6b347137f0df63

Observation 8973bd45-e242-4a6f-ade2-2af26347c810 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 158

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:6456aa1450c686e2775dce4d6c5ed7d723f581080162a60e1e6ce2d399f4f16d

Observation b80da959-fcc9-43a8-a6e5-740fb519ec4f · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.703386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-08T02:44:00.608590Z digest=sha256:a5b45b940272fb59097bcb3608bc209e527d82b13d1f1d02b96ff7e4535a8f33

Observation fbb5f7de-c5d8-41ae-84f1-d333c7925267 · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:19faeee7bf5d0fb2b7d52c3a04ce64ad5e18b31b1209e121fca2145bfcae80a5

Observation f26e0aa9-9241-4119-8386-f578cd8416fa · inbound

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence cites this paper.

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:28.227722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:28.227722Z digest=sha256:fe8eda37f35fd776e02d710be7ae167732148a71ff8d309a9f196fbda4a83781

Observation d3b06a84-d677-4a0a-9a70-d483a1a880cd · inbound

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding cites this paper.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:02.432874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:02.432874Z digest=sha256:d5ad920388b37bfa501b1ffaa34e67349340ce3e862e916e0eb67596acdba42e

Observation fee6c617-29ba-4ae2-be3b-ec4fcdca95e7 · inbound

An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia cites this paper.

An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:25:57.302362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:25:57.302362Z digest=sha256:9977e2506239643150120d5fb6438d4e3e186fa8062b1ab2e16b6c39f54b3c13

Observation 4f72c9d9-a190-48ce-8cee-9d7193ae3242 · inbound

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA cites this paper.

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T07:16:36.438659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:16:36.438659Z digest=sha256:7fbcef30d322a2c178ba49f8159721348755cc2725a509906a300debf40488fd