REVIEW 42 cited by
ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a detailed scene from insufficient captured views is still an ill-posed optimization problem, often resulting in artifacts and distortions in unseen areas. In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction challenge as a temporal generation task. The key insight is to unleash the strong generative prior of large pre-trained video diffusion models for sparse-view reconstruction. However, 3D view consistency struggles to be accurately preserved in directly generated video frames from pre-trained models. To address this, given limited input views, the proposed ReconX first constructs a global point cloud and encodes it into a contextual space as the 3D structure condition. Guided by the condition, the video diffusion model then synthesizes video frames that are both detail-preserved and exhibit a high degree of 3D consistency, ensuring the coherence of the scene from various perspectives. Finally, we recover the 3D scene from the generated video through a confidence-aware 3D Gaussian Splatting optimization scheme. Extensive experiments on various real-world datasets show the superiority of our ReconX over state-of-the-art methods in terms of quality and generalizability.
Forward citations
Cited by 42 Pith papers
-
Wonderland: Navigating 3D Scenes from a Single Image
A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.
-
NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images
A dual-stream diffusion model that generates novel views and condition-view camera poses together, removing the need for external pose estimation in multi-view novel view synthesis.
-
Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models
Using a product-of-hyperspheres Riemannian flow matching model on VGGT's latent codes, the authors generate plausible depth, point maps, and RGB for target views from one to four unposed context images.
-
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.
-
LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis
Initializing a highway encoder-decoder NVS network from VGGT 3D-aware features yields 31.4 PSNR on RealEstate10k with real-time decoding and optional unposed inputs.
-
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
Diffusing in a unified 3D representation (geometry plus appearance and semantics) produces more cross-view-consistent 3D Gaussian scenes than 2D latent pipelines.
-
DAV-GSWT: Diffusion-Active-View Sampling for Data-Efficient Gaussian Splatting Wang Tiles
DAV-GSWT uses diffusion priors and active view sampling to synthesize high-fidelity Gaussian Splatting Wang Tiles from minimal observations while preserving visual quality and tile transitions.
-
WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
WorldPlay uses dual action representation, reconstituted context memory, and context forcing distillation to produce consistent 720p streaming video at 24 FPS for interactive world modeling.
-
PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
PanoSplatt3R adapts a perspective pretrained stereo model to unposed wide-baseline panorama reconstruction with per-head rolled rotary positional embeddings, achieving SOTA on HM3D and Replica.
-
From Virtual Games to Real-World Play
A chunk-wise video diffusion model trained on labeled game data plus unlabeled real footage transfers game-style control commands to real-world entities.
-
Emergent Temporal Correspondences from Video Diffusion Transformers
Video diffusion transformers encode temporal correspondences primarily in query-key similarities of a few specific attention layers, which can be extracted for zero-shot point tracking and used for training-free motio...
-
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.
-
Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object
Zero-P-to-3 fuses multi-view diffusion, a restoration prior, and a coarse 3D Gaussian rendering in DDIM sampling, then refines with rotated views, and reports improved invisible-region reconstruction from partial-view...
-
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
EPiC trains a 30M-parameter visibility-aware ControlNet on mask-based anchor videos from 5,000 in-the-wild videos and 500 steps, reaching SOTA camera accuracy on RealEstate10K and MiraData.
-
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
EF-VI injects temporally expanded end-frame features into a transformer-based image-to-video diffusion model, improving video inbetweening quality over direct fine-tuning and bidirectional sampling baselines.
-
SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
SpatialCrafter generates camera-controlled video from sparse input views and reconstructs a 3D Gaussian scene from the generated video latents, improving novel view synthesis.
-
Dust to Tower: Coarse-to-Fine Photo-Realistic Scene Reconstruction from Sparse Uncalibrated Images
A coarse-to-fine pipeline jointly optimizes 3D Gaussian Splatting and camera poses from sparse, uncalibrated images, using warped and inpainted pseudo-views for supervision.
-
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
StreetCrafter conditions a video diffusion model on LiDAR point cloud renderings to synthesize controllable street views, and distills it into a real-time 3D Gaussian representation.
-
SLGaussian: Fast Language Gaussian Splatting in Sparse Views
SLGaussian builds a 3D semantic field from two photos in a single forward pass, stores CLIP features in a memory bank for fast open-vocabulary queries, and reports higher IoU than LangSplat and LERF on the LERF and 3D...
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
Splatter-360: Generalizable 360$^{\circ}$ Gaussian Splatting for Wide-baseline Panoramic Images
Splatter-360 is an end-to-end generalizable 3D Gaussian splatting model that builds a spherical cost volume to improve geometry and rendering from wide-baseline panoramic images.
-
InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
A three-stage pipeline generates up to 100,000 square meters of dynamic 3D driving scenes with 200-frame videos, controlled by HD maps, bounding boxes, and text.
-
FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes
A generation-reconstruction pipeline with a diffusion enhancer trained on simulated degradations enables off-trajectory camera rendering in driving scenes.
-
MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model
A multi-view diffusion model trained on 1.6 million scenes uses warped depth-based 3D priors and a key-rescaling trick to synthesize up to 158 views in one forward pass.
-
DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes
A distributed pipeline using a pretrained feed-forward Gaussian model, global alignment, depth regularization, and distillation-based merging reconstructs sparse-view large-scale aerial scenes faster than prior methods.
-
High-Quality Exposure Correction with Diffusion-Based Image Generation Priors
DPEC fine-tunes a pretrained diffusion model for single-step exposure correction and fuses its low-frequency output into a regression network, improving perceptual metrics on LCDP, MSEC, and SICE.
-
SUMI: Scalable Unified Model for 3D Point Cloud Inference
SUMI refines coarse 3D point cloud predictions by injecting noisy geometric features into cross-attention, and reports state-of-the-art Chamfer distance on PCN, ShapeNet-55/34, and MVP.
-
PanoImager: Geometry-Guided Novel View Synthesis and Reconstruction from Sparse Panoramic Views
PanoImager is an SfM-free pipeline combining feed-forward priors, geometry-conditioned diffusion view completion, and depth-guided 3DGS optimization to reconstruct from sparse panoramic images.
-
Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation
A diffusion model generates dense 4D point trajectories from a single image, and a separate view-synthesis module renders them into novel-view videos.
-
Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling
A computational model estimates pancreatic duct pressure non-invasively from MRCP geometry, with reported agreement against ERCP pressure measurements.
-
DIP-GS: Deep Image Prior For Gaussian Splatting Sparse View Recovery
DIP-GS applies a deep image prior in a coarse-to-fine manner to enable 3D Gaussian Splatting for sparse-view reconstruction.
-
FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation
FVGen uses GAN-based adversarial distillation and softened reverse KL divergence to compress a video diffusion teacher for novel-view synthesis into a four-step student with comparable quality.
-
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
LangScene-X generates RGB, normal, and semantic videos from sparse views to reconstruct 3D language-embedded Gaussian fields that support open-ended text queries.
-
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
SceneCompleter jointly denoises RGB and depth latents, conditioned on projected depth and global scene features, yielding higher quality and more pose-consistent novel views than 2D-inpainting baselines.
-
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.
-
VRAG: Learning World Models for Interactive Video Generation
VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft a...
-
PhysAnimator: Physics-Guided Generative Cartoon Animation
PhysAnimator combines 2D deformable-body physics simulation with a sketch-guided video diffusion model to animate static anime illustrations with controllable, physically plausible motion.
-
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.
-
GBR: Generative Bundle Refinement for High-fidelity Gaussian Splatting with Enhanced Mesh Reconstruction
GBR reconstructs accurate camera poses, dense point clouds, and high-fidelity meshes from 4-6 unposed images by combining DUSt3R-based neural bundle adjustment with scale-preserving diffusion depth refinement.
-
Camera Trajectory Generation: A Comprehensive Survey of Methods, Metrics, and Future Directions
A review that organizes camera trajectory generation into representation levels, algorithm families, evaluation metrics, and datasets.
-
Reinforcement Learning: From Algorithms To Foundation Models
A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.
-
Sparse-View 3D Reconstruction: Recent Advances and Open Challenges
A comprehensive survey that organizes sparse-view 3D reconstruction methods into geometry-based, NeRF, 3DGS, and diffusion-based categories, with benchmarks and open challenges.
Discussion (0). Continue with ORCID to comment.