OVOW reconstructs instance-level, simulation-ready 4D mesh scenes from monocular video via a four-stage training-free pipeline and introduces a new benchmark for structured Video-to-4D evaluation.
Advances in Neural Information Processing Systems37, 21875–21911 (2024)
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 10years
2026 10roles
background 1polarities
background 1representative citing papers
The subspace intervention framework reveals that pre-training objectives shape how ViTs encode geometric information in compressible low-rank subspaces, with peak precision at intermediate layers.
EgoTraj is a new open multimodal dataset of 75 long-horizon egocentric human navigation sequences in urban environments with head pose, gaze, and scene data, plus benchmarks of trajectory prediction methods.
Track2Map jointly optimizes camera poses and deformable 3D Gaussian maps online from surgical stereo video via track-anchored deformation and motion-gated pose updates.
ICDepth adapts text-to-video diffusion transformers for video depth estimation via in-context conditioning, achieving SOTA results on benchmarks with 6-13x less training data than prior generative methods.
HSDF-Lane uses a height-aligned signed distance field with differentiable rendering and lane-aware semantic positional encoding to achieve SOTA 3D lane detection and height estimation on OpenLane.
OrthoTrack is a training-free system for continuous metric 6-DoF UAV pose estimation anchored in public orthophotos and surface models, with a new MovingDrone benchmark dataset.
UniFixer is a universal reference-guided framework that fixes spatial, temporal, and backbone-related degradations in diffusion-based view synthesis via coarse-to-fine modules and achieves zero-shot SOTA results on novel view synthesis and stereo conversion.
Sphere-Depth benchmark shows substantial performance degradation in both general and spherical-aware depth estimation models under simulated camera pose variations.
FoundDP integrates DP-derived metric depth with ViT-based structural priors from monocular models, using feature alignment to mitigate defocus blur and improve depth in low-observability areas.
citing papers explorer
-
One Video, One World: Turning Monocular Video into Physical 4D Scenes
OVOW reconstructs instance-level, simulation-ready 4D mesh scenes from monocular video via a four-stage training-free pipeline and introduces a new benchmark for structured Video-to-4D evaluation.
-
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
The subspace intervention framework reveals that pre-training objectives shape how ViTs encode geometric information in compressible low-rank subspaces, with peak precision at intermediate layers.
-
EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction
EgoTraj is a new open multimodal dataset of 75 long-horizon egocentric human navigation sequences in urban environments with head pose, gaze, and scene data, plus benchmarks of trajectory prediction methods.
-
Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery
Track2Map jointly optimizes camera poses and deformable 3D Gaussian maps online from surgical stereo video via track-anchored deformation and motion-gated pose updates.
-
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
ICDepth adapts text-to-video diffusion transformers for video depth estimation via in-context conditioning, achieving SOTA results on benchmarks with 6-13x less training data than prior generative methods.
-
HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection
HSDF-Lane uses a height-aligned signed distance field with differentiable rendering and lane-aware semantic positional encoding to achieve SOTA 3D lane detection and height estimation on OpenLane.
-
OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos
OrthoTrack is a training-free system for continuous metric 6-DoF UAV pose estimation anchored in public orthophotos and surface models, with a new MovingDrone benchmark dataset.
-
UniFixer: A Universal Reference-Guided Fixer for Diffusion-Based View Synthesis
UniFixer is a universal reference-guided framework that fixes spatial, temporal, and backbone-related degradations in diffusion-based view synthesis via coarse-to-fine modules and achieves zero-shot SOTA results on novel view synthesis and stereo conversion.
-
Sphere-Depth: A Benchmark for Depth Estimation Methods with Varying Spherical Camera Orientations
Sphere-Depth benchmark shows substantial performance degradation in both general and spherical-aware depth estimation models under simulated camera pose variations.
-
FoundDP: Revisiting Weak Disparity Observability in Dual-Pixel Depth Estimation
FoundDP integrates DP-derived metric depth with ViT-based structural priors from monocular models, using feature alignment to mitigate defocus blur and improve depth in low-observability areas.