Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T13:49:45.095377Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 0 inbound Pith citation observations for arXiv:2605.28132.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T13:49:45.095377Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
7 of 7 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9e051b59-4c93-4ba9-a67b-aa501eed1e37 · outbound
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15ba6a80-e273-4e79-88e6-f2c44a31e0fd · outbound
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6d91e95-174f-4c4e-a0c2-7401bf783f2a · outbound
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Iggt: Instance- grounded geometry transformer for semantic 3d reconstruc- tion
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5050c484-4f37-44b6-9da9-4a3cf203184a · outbound
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1dfe407a-545d-44a0-acf5-2c08e0503a44 · outbound
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Motubrain: An Advanced World Action Model for Robot Control
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7a52f2f8-3b1c-4d11-bab1-4db4b4c3222f · outbound
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 63928e68-38f6-457d-aee4-f807976f477c · outbound
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models VGM features are spatially pooled to a fixed grid when needed: WAN/OpenSora use 15×26 , and CogVideoX/Aether use 15×22 ; VLM features keep their native visual-token grids
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.