OVOW reconstructs instance-level, simulation-ready 4D mesh scenes from monocular video via a four-stage training-free pipeline and introduces a new benchmark for structured Video-to-4D evaluation.
hub Canonical reference
In: Proceedings of the Computer Vision and Pattern Recognition Conference
Canonical reference. 71% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
years
2026 19representative citing papers
MV-Forcing composes temporal and view-sequential autoregression in a single diffusion model, using a recurrent 3D reconstruction model as a geometric bridge to generate arbitrarily long, multi-view consistent videos.
Stitched Embeddings unifies 3D garment reconstruction and 2D pattern inference in a bidirectional latent space using BoxMesh as an intermediate representation.
EPO is a trackless, edge-map-alignment framework that refines pose estimates from 3D foundation models and matches or exceeds bundle-adjustment performance with substantially lower runtime and memory use.
Video diffusion models can be adapted into permutation-invariant generators for sparse novel view synthesis by treating the problem as video completion and removing temporal order cues.
EndoVGGT uses a dynamic DeGAT graph attention module to improve depth estimation and non-rigid 3D reconstruction in surgery, reporting 24.6% PSNR and 9.1% SSIM gains on SCARED with zero-shot generalization to new domains.
3AM integrates MUSt3R 3D features into SAM2 via a Feature Merger and FOV-aware sampling to deliver geometry-consistent video object segmentation from RGB alone, with large gains on wide-baseline datasets.
Coupling tetrahedra to Gaussians and pruning via a vertex-shared continuous opacity field produces unified, single-component tetrahedral meshes suitable for FEM from multi-view images.
A real multi-city, multi-kilometer surround-view driving dataset plus an urban-tailored 3DGS baseline shows that city-scale reconstruction still degrades with scale, off-trajectory views, and real-world noise.
VOCA is a causal stereo visual odometry system that achieves state-of-the-art performance on compressed streams by exploiting codec awareness.
Embody4D generates novel-view videos from monocular robot videos via a 3D-aware synthesis pipeline, confidence-aware expert modulation, and interaction-aware attention for embodied 4D world modeling.
Layer analysis of DINOv3 shows non-uniform 3D geometric knowledge concentrated in deeper layers, enabling a last-layer-centric recombination module that improves monocular depth estimation accuracy to state-of-the-art levels.
SS3D pretrains an end-to-end feed-forward 3D estimator on filtered YouTube-8M videos via SfM self-supervision, MVS filtering, and expert distillation, delivering stronger zero-shot transfer and fine-tuning than prior self-supervised baselines.
A single freehand sketch can generate a full orbit of photorealistic views in one pass, trained on a 9k synthetic sketch-to-multiview dataset with camera-aware adapters and SfM-supervised correspondences.
Scaling transformer context with sparse attention and 3D-aware block routing improves feed-forward 3D reconstruction and inverse rendering, closing much of the quality gap with dense-view optimization.
Motion-MLLM integrates IMU egomotion data into MLLMs using cascaded filtering and asymmetric fusion to ground visual content in physical trajectories for scale-aware 3D understanding, achieving competitive accuracy at higher speed.
X-Imitator is a bidirectional action-pose interaction framework for spatial-aware imitation learning that outperforms vanilla policies and explicit pose guidance on 24 simulated and 3 real-world robotic tasks.
PAD synthesizes 3D geometry in observation space via depth unprojection as anchor to eliminate pose ambiguity in image-to-3D generation.
UniGeo improves camera-controllable image editing by injecting point cloud geometry into a video diffusion model at the representation, architecture, and loss levels, achieving state-of-the-art geometric consistency on RE10K, DL3DV, and Tanks benchmarks.
citing papers explorer
-
One Video, One World: Turning Monocular Video into Physical 4D Scenes
OVOW reconstructs instance-level, simulation-ready 4D mesh scenes from monocular video via a four-stage training-free pipeline and introduces a new benchmark for structured Video-to-4D evaluation.
-
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
MV-Forcing composes temporal and view-sequential autoregression in a single diffusion model, using a recurrent 3D reconstruction model as a geometric bridge to generate arbitrarily long, multi-view consistent videos.
-
Stitched Embeddings: A Unified Latent Space for 3D Garments and 2D Patterns
Stitched Embeddings unifies 3D garment reconstruction and 2D pattern inference in a bidirectional latent space using BoxMesh as an intermediate representation.
-
EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization
EPO is a trackless, edge-map-alignment framework that refines pose estimates from 3D foundation models and matches or exceeds bundle-adjustment performance with substantially lower runtime and memory use.
-
Novel View Synthesis as Video Completion
Video diffusion models can be adapted into permutation-invariant generators for sparse novel view synthesis by treating the problem as video completion and removing temporal order cues.
-
EndoVGGT: GNN-Enhanced Depth Estimation for Surgical 3D Reconstruction
EndoVGGT uses a dynamic DeGAT graph attention module to improve depth estimation and non-rigid 3D reconstruction in surgery, reporting 24.6% PSNR and 9.1% SSIM gains on SCARED with zero-shot generalization to new domains.
-
3AM: 3egment Anything with Geometric Consistency in Videos
3AM integrates MUSt3R 3D features into SAM2 via a Feature Merger and FOV-aware sampling to deliver geometry-consistent video object segmentation from RGB alone, with large gains on wide-baseline datasets.
-
HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations
Coupling tetrahedra to Gaussians and pruning via a vertex-shared continuous opacity field produces unified, single-component tetrahedral meshes suitable for FEM from multi-view images.
-
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence
A real multi-city, multi-kilometer surround-view driving dataset plus an urban-tailored 3DGS baseline shows that city-scale reconstruction still degrades with scale, off-trajectory views, and real-world noise.
-
VOCA: Visual Odometry with Codec Awareness
VOCA is a causal stereo visual odometry system that achieves state-of-the-art performance on compressed streams by exploiting codec awareness.
-
Embody4D: A Generalist Data Engine for Embodied 4D World Modeling
Embody4D generates novel-view videos from monocular robot videos via a 3D-aware synthesis pipeline, confidence-aware expert modulation, and interaction-aware attention for embodied 4D world modeling.
-
Last-Layer-Centric Feature Recombination: Unleashing 3D Geometric Knowledge in DINOv3 for Monocular Depth Estimation
Layer analysis of DINOv3 shows non-uniform 3D geometric knowledge concentrated in deeper layers, enabling a last-layer-centric recombination module that improves monocular depth estimation accuracy to state-of-the-art levels.
-
SS3D: End2End Self-Supervised 3D from Web Videos
SS3D pretrains an end-to-end feed-forward 3D estimator on filtered YouTube-8M videos via SfM self-supervision, MVS filtering, and expert distillation, delivering stronger zero-shot transfer and fine-tuning than prior self-supervised baselines.
-
Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
A single freehand sketch can generate a full orbit of photorealistic views in one pass, trained on a 9k synthetic sketch-to-multiview dataset with camera-aware adapters and SfM-supervised correspondences.
-
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
Scaling transformer context with sparse attention and 3D-aware block routing improves feed-forward 3D reconstruction and inverse rendering, closing much of the quality gap with dense-view optimization.
-
Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding
Motion-MLLM integrates IMU egomotion data into MLLMs using cascaded filtering and asymmetric fusion to ground visual content in physical trajectories for scale-aware 3D understanding, achieving competitive accuracy at higher speed.
-
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction
X-Imitator is a bidirectional action-pose interaction framework for spatial-aware imitation learning that outperforms vanilla policies and explicit pose guidance on 24 simulated and 3 real-world robotic tasks.
-
Pose-Aware Diffusion for 3D Generation
PAD synthesizes 3D geometry in observation space via depth unprojection as anchor to eliminate pose ambiguity in image-to-3D generation.
-
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
UniGeo improves camera-controllable image editing by injecting point cloud geometry into a video diffusion model at the representation, architecture, and loss levels, achieving state-of-the-art geometric consistency on RE10K, DL3DV, and Tanks benchmarks.