Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2501.01428.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T13:46:28.179433Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T02:44:27.702116Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 68282466-6117-4a8f-bee3-501a1f557853 · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 65e82996-56b2-4918-875e-caced0b314df · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c1758b5f-e01e-4635-adf9-ceae67bd6f17 · inbound
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b916cc03-05e3-4628-a445-d4d239b6680e · inbound
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e6481afa-a945-489a-9415-830dd543b5bd · inbound
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48b73ff5-a861-434c-9f81-a910cb1ca327 · inbound
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e5a825-e443-4632-9d14-e9c0fcec8480 · inbound
Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0ee15ea0-69cc-44c5-9b2c-8fcee3ae62f6 · inbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b25ab6a-9274-48e0-864f-28bfc1251243 · inbound
Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fc2db593-5ae4-4870-a65a-2f0695f038cf · inbound
Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a248adb3-8099-4504-aa8b-55f56ec5bdd2 · inbound
3D-IDE: 3D Implicit Depth Emergent GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c0a9d8a8-42c3-4b3c-b50e-a0de98bba6d2 · inbound
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e7577ef1-9c43-4f33-97a1-877f78d5de67 · inbound
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be164b81-94e3-48fe-859e-8a844c2fc348 · inbound
Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c5b05d49-f223-4fa4-9dd5-c4da6aee2aad · inbound
Geometry-Guided 3D Visual Token Pruning for Video-Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 190ecbbf-3cc5-431f-a4d6-fb673951d3b9 · inbound
Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5c25f57d-8586-4dea-a859-da4d1f564f55 · inbound
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 733a178d-8570-40d1-876e-d01d40dfcd6a · inbound
Unlocking Dense Metric Depth Estimation in VLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fc52b842-e9ba-4cf5-89f1-e34b8519cce7 · inbound
Unlocking Dense Metric Depth Estimation in VLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d3284230-31a6-43ff-b81e-d4d6ea12c62e · inbound
EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 10454cfd-1cdb-4d81-94e9-768de310711f · inbound
AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 578c5a4d-d994-403e-ba8a-c0ac90ec47c1 · inbound
SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 37d34a9c-257d-4de4-b4f5-cb1e30cd20f8 · inbound
Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation dbe0e397-02c2-4bfe-a85e-f3bc4b870a69 · inbound
Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 00859666-c515-43b7-b10d-4520d40da0fd · inbound
Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5c320a40-97dd-43ea-91ec-3eb6229a6660 · inbound
Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5321faac-0870-41a7-987f-5e556943f58d · inbound
SpatialSV: Internalizing Interpretable 3D Spatial Awareness in MLLMs via Task-Oriented Visual Supervision GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5d5a29e3-9795-4ee7-8890-f32df3da4a9b · inbound
Agentic Collaborative Cognition for Zero-Shot 3D Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 89812ce1-780f-4e00-a266-22b4f6d2de64 · inbound
Agentic Collaborative Cognition for Zero-Shot 3D Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3c3a0780-16a3-44a3-ab22-b1bf0f143544 · inbound
ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3d876ced-b9fb-4c34-9a35-ee6c351c74ff · inbound
SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 94eb3c21-e4a1-4a59-9fde-60aeeb1b28de · inbound
Holo-Captioning: Toward the Text Equivalent of 3D Scenes GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feaaff6a-f67b-4390-aa46-1d8ffe0c02de · inbound
Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8973bd45-e242-4a6f-ade2-2af26347c810 · inbound
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 158
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b80da959-fcc9-43a8-a6e5-740fb519ec4f · inbound
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fbb5f7de-c5d8-41ae-84f1-d333c7925267 · inbound
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f26e0aa9-9241-4119-8386-f578cd8416fa · inbound
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b06a84-d677-4a0a-9a70-d483a1a880cd · inbound
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee6c617-29ba-4ae2-be3b-ec4fcdca95e7 · inbound
An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f72c9d9-a190-48ce-8cee-9d7193ae3242 · inbound
ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.