Pith. sign in

REVIEW 9 cited by

Reconstructing 4D Spatial Intelligence: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.21045 v2 pith:IB3TXDV4 submitted 2025-07-28 cs.CV

classification cs.CV
keywords levelintelligencereconstructionspatialreconstructingscenecomponentsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer vision, with broad real-world applications. These range from entertainment domains like movies, where the focus is often on reconstructing fundamental visual elements, to embodied AI, which emphasizes interaction modeling and physical realism. Fueled by rapid advances in 3D representations and deep learning architectures, the field has evolved quickly, outpacing the scope of previous surveys. Additionally, existing surveys rarely offer a comprehensive analysis of the hierarchical structure of 4D scene reconstruction. To address this gap, we present a new perspective that organizes existing methods into five progressive levels of 4D spatial intelligence: (1) Level 1 -- reconstruction of low-level 3D attributes (e.g., depth, pose, and point maps); (2) Level 2 -- reconstruction of 3D scene components (e.g., objects, humans, structures); (3) Level 3 -- reconstruction of 4D dynamic scenes; (4) Level 4 -- modeling of interactions among scene components; and (5) Level 5 -- incorporation of physical laws and constraints. We conclude the survey by discussing the key challenges at each level and highlighting promising directions for advancing toward even richer levels of 4D spatial intelligence. To track ongoing developments, we maintain an up-to-date project page: https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. One Video, One World: Turning Monocular Video into Physical 4D Scenes

    cs.CV 2026-06 unverdicted novelty 8.0 of 10

    OVOW reconstructs instance-level, simulation-ready 4D mesh scenes from monocular video via a four-stage training-free pipeline and introduces a new benchmark for structured Video-to-4D evaluation.

  2. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

    cs.CV 2026-07 accept novelty 7.0 of 10

    ACE-Data-0 is a 150-hour home HOI dataset with millisecond-synced ego/exo video, mocap body/hands, object 6-DoF, audio, and tactile signals, plus a three-level benchmark exposing large SOTA gaps.

  3. CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

    cs.CV 2026-01 unverdicted novelty 7.0 of 10

    CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.

  4. Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    HA-HOI produces physically plausible 4D HOI animations from monocular videos by anchoring object reconstruction to human motion and refining the result in a physics-based humanoid-object simulator.

  5. Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Stitch4D reconstructs coherent 4D urban scenes from sparse non-overlapping camera placements by synthesizing bridge views and enforcing inter-location spatio-temporal consistency.

  6. PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception

    cs.CV 2025-10 unverdicted novelty 6.0 of 10

    PAGE-4D is a feedforward extension of VGGT that uses a dynamics-aware aggregator and mask to disentangle pose estimation from geometry reconstruction in videos with moving objects.

  7. PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A fine-tuned VGGT with a learned dynamics mask improves camera pose, depth, and point-cloud reconstruction on dynamic-scene benchmarks over the original static-scene model.

  8. Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation

    cs.CV 2026-04 conditional novelty 5.0 of 10

    Synthesizing intermediate bridge views between sparse, non-overlapping urban cameras and jointly optimizing them stabilizes 4D reconstruction where dense-view methods collapse.

  9. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

Pith tools