Pith. sign in

Paper Citation Record · LEDGER

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

As of 6 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 0 inbound Pith citation observations for arXiv:2605.28132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28132 v1

Coverage vector

measured 7 of 7 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T13:49:45.095377Z

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

7 of 7 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e051b59-4c93-4ba9-a67b-aa501eed1e37 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:53:28.699603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:49:45.095377Z digest=sha256:778aa2ac8c0bc258a6f131803523bc6ded00bbfde8b3f73cbb20cb70df058b2e

Observation 15ba6a80-e273-4e79-88e6-f2c44a31e0fd · outbound

This paper cites Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization.

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:53:28.693772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:49:45.095377Z digest=sha256:973dd7bd77e907dce14306b7717656b51aaf3d306f0bbf45d77253be18fd58b3

Observation d6d91e95-174f-4c4e-a0c2-7401bf783f2a · outbound

This paper cites Iggt: Instance- grounded geometry transformer for semantic 3d reconstruc- tion.

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Iggt: Instance- grounded geometry transformer for semantic 3d reconstruc- tion

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:53:28.705111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:49:45.095377Z digest=sha256:1925d4ea74f88aae68211c3f500a4028d609af8432102dbcf5b60cf0d83c9623

Observation 5050c484-4f37-44b6-9da9-4a3cf203184a · outbound

This paper cites Pan and H.

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Pan and H

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.690788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:49:45.095377Z digest=sha256:22f0b48b8048ea7abc0e26b4a333005692ec3588544fdea13623f0521b75eb68

Observation 1dfe407a-545d-44a0-acf5-2c08e0503a44 · outbound

This paper cites Motubrain: An Advanced World Action Model for Robot Control.

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Motubrain: An Advanced World Action Model for Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:53:28.702407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:49:45.095377Z digest=sha256:caade8c8759217e66b9c1ecab99ce8a3b35b6e6dbfc219c5bec88e0589998591

Observation 7a52f2f8-3b1c-4d11-bab1-4db4b4c3222f · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:53:28.696566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:49:45.095377Z digest=sha256:c18e0341dd73a3b46ca1e69c76e5337d4c9a1f7a9a4e36a673c517050f49f793

Observation 63928e68-38f6-457d-aee4-f807976f477c · outbound

This paper cites VGM features are spatially pooled to a fixed grid when needed: WAN/OpenSora use 15×26 , and CogVideoX/Aether use 15×22 ; VLM features keep their native visual-token grids.

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models VGM features are spatially pooled to a fixed grid when needed: WAN/OpenSora use 15×26 , and CogVideoX/Aether use 15×22 ; VLM features keep their native visual-token grids

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T13:49:45.095377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:49:45.095377Z digest=sha256:29215eeb40eb6ee5f747c2dfc2d98462df97e55530fdab922cd9c15a09d0bc3b

Pith citing papers

No inbound Pith citation observations are available.