Pith. sign in

REVIEW 42 cited by

ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.16767 v4 pith:M6UMHDOJ submitted 2024-08-29 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords scenevideoreconstructionreconxdiffusionmodelsviewscondition
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a detailed scene from insufficient captured views is still an ill-posed optimization problem, often resulting in artifacts and distortions in unseen areas. In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction challenge as a temporal generation task. The key insight is to unleash the strong generative prior of large pre-trained video diffusion models for sparse-view reconstruction. However, 3D view consistency struggles to be accurately preserved in directly generated video frames from pre-trained models. To address this, given limited input views, the proposed ReconX first constructs a global point cloud and encodes it into a contextual space as the 3D structure condition. Guided by the condition, the video diffusion model then synthesizes video frames that are both detail-preserved and exhibit a high degree of 3D consistency, ensuring the coherence of the scene from various perspectives. Finally, we recover the 3D scene from the generated video through a confidence-aware 3D Gaussian Splatting optimization scheme. Extensive experiments on various real-world datasets show the superiority of our ReconX over state-of-the-art methods in terms of quality and generalizability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 42 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Wonderland: Navigating 3D Scenes from a Single Image

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.

  2. NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A dual-stream diffusion model that generates novel views and condition-view camera poses together, removing the need for external pose estimation in multi-view novel view synthesis.

  3. Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Using a product-of-hyperspheres Riemannian flow matching model on VGGT's latent codes, the authors generate plausible depth, point maps, and RGB for target views from one to four unposed context images.

  4. PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.

  5. LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis

    cs.CV 2026-03 accept novelty 6.0 of 10

    Initializing a highway encoder-decoder NVS network from VGGT 3D-aware features yields 31.4 PSNR on RealEstate10k with real-time decoding and optional unposed inputs.

  6. Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Diffusing in a unified 3D representation (geometry plus appearance and semantics) produces more cross-view-consistent 3D Gaussian scenes than 2D latent pipelines.

  7. DAV-GSWT: Diffusion-Active-View Sampling for Data-Efficient Gaussian Splatting Wang Tiles

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    DAV-GSWT uses diffusion priors and active view sampling to synthesize high-fidelity Gaussian Splatting Wang Tiles from minimal observations while preserving visual quality and tile transitions.

  8. WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

    cs.CV 2025-12 unverdicted novelty 6.0 of 10

    WorldPlay uses dual action representation, reconstituted context memory, and context forcing distillation to produce consistent 720p streaming video at 24 FPS for interactive world modeling.

  9. PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PanoSplatt3R adapts a perspective pretrained stereo model to unposed wide-baseline panorama reconstruction with per-head rolled rotary positional embeddings, achieving SOTA on HM3D and Replica.

  10. From Virtual Games to Real-World Play

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A chunk-wise video diffusion model trained on labeled game data plus unlabeled real footage transfers game-style control commands to real-world entities.

  11. Emergent Temporal Correspondences from Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video diffusion transformers encode temporal correspondences primarily in query-key similarities of a few specific attention layers, which can be extracted for zero-shot point tracking and used for training-free motio...

  12. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.

  13. Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Zero-P-to-3 fuses multi-view diffusion, a restoration prior, and a coarse 3D Gaussian rendering in DDIM sampling, then refines with rotated views, and reports improved invisible-region reconstruction from partial-view...

  14. EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EPiC trains a 30M-parameter visibility-aware ControlNet on mask-based anchor videos from 5,000 in-the-wild videos and 500 steps, reaching SOTA camera accuracy on RealEstate10K and MiraData.

  15. EF-VI: Enhancing End-Frame Injection for Video Inbetweening

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EF-VI injects temporally expanded end-frame features into a transformer-based image-to-video diffusion model, improving video inbetweening quality over direct fine-tuning and bidirectional sampling baselines.

  16. SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations

    cs.CV 2025-05 conditional novelty 6.0 of 10

    SpatialCrafter generates camera-controlled video from sparse input views and reconstructs a 3D Gaussian scene from the generated video latents, improving novel view synthesis.

  17. Dust to Tower: Coarse-to-Fine Photo-Realistic Scene Reconstruction from Sparse Uncalibrated Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A coarse-to-fine pipeline jointly optimizes 3D Gaussian Splatting and camera poses from sparse, uncalibrated images, using warped and inpainted pseudo-views for supervision.

  18. StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    StreetCrafter conditions a video diffusion model on LiDAR point cloud renderings to synthesize controllable street views, and distills it into a real-time 3D Gaussian representation.

  19. SLGaussian: Fast Language Gaussian Splatting in Sparse Views

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SLGaussian builds a 3D semantic field from two photos in a single forward pass, stores CLIP features in a memory bank for fast open-vocabulary queries, and reports higher IoU than LangSplat and LERF on the LERF and 3D...

  20. You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

    cs.CV 2024-12 reject novelty 6.0 of 10

    See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...

  21. Splatter-360: Generalizable 360$^{\circ}$ Gaussian Splatting for Wide-baseline Panoramic Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Splatter-360 is an end-to-end generalizable 3D Gaussian splatting model that builds a spherical cost volume to improve geometry and rendering from wide-baseline panoramic images.

  22. InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A three-stage pipeline generates up to 100,000 square meters of dynamic 3D driving scenes with 200-frame videos, controlled by HD maps, bounding boxes, and text.

  23. FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A generation-reconstruction pipeline with a diffusion enhancer trained on simulated degradations enables off-trajectory camera rendering in driving scenes.

  24. MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A multi-view diffusion model trained on 1.6 million scenes uses warped depth-based 3D priors and a key-rescaling trick to synthesize up to 158 views in one forward pass.

  25. DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A distributed pipeline using a pretrained feed-forward Gaussian model, global alignment, depth regularization, and distillation-based merging reconstructs sparse-view large-scale aerial scenes faster than prior methods.

  26. High-Quality Exposure Correction with Diffusion-Based Image Generation Priors

    cs.CV 2026-08 conditional novelty 5.0 of 10

    DPEC fine-tunes a pretrained diffusion model for single-step exposure correction and fuses its low-frequency output into a regression network, improving perceptual metrics on LCDP, MSEC, and SICE.

  27. SUMI: Scalable Unified Model for 3D Point Cloud Inference

    cs.CV 2026-08 conditional novelty 5.0 of 10

    SUMI refines coarse 3D point cloud predictions by injecting noisy geometric features into cross-attention, and reports state-of-the-art Chamfer distance on PCN, ShapeNet-55/34, and MVP.

  28. PanoImager: Geometry-Guided Novel View Synthesis and Reconstruction from Sparse Panoramic Views

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    PanoImager is an SfM-free pipeline combining feed-forward priors, geometry-conditioned diffusion view completion, and depth-guided 3DGS optimization to reconstruct from sparse panoramic images.

  29. Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A diffusion model generates dense 4D point trajectories from a single image, and a separate view-synthesis module renders them into novel-view videos.

  30. Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling

    physics.med-ph 2025-08 unverdicted novelty 5.0 of 10

    A computational model estimates pancreatic duct pressure non-invasively from MRCP geometry, with reported agreement against ERCP pressure measurements.

  31. DIP-GS: Deep Image Prior For Gaussian Splatting Sparse View Recovery

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DIP-GS applies a deep image prior in a coarse-to-fine manner to enable 3D Gaussian Splatting for sparse-view reconstruction.

  32. FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    FVGen uses GAN-based adversarial distillation and softened reverse KL divergence to compress a video diffusion teacher for novel-view synthesis into a four-step student with comparable quality.

  33. LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LangScene-X generates RGB, normal, and semantic videos from sparse views to reconstruct 3D language-embedded Gaussian fields that support open-ended text queries.

  34. SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SceneCompleter jointly denoises RGB and depth latents, conditioned on projected depth and global scene features, yielding higher quality and more pose-consistent novel views than 2D-inpainting baselines.

  35. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

  36. VRAG: Learning World Models for Interactive Video Generation

    cs.CV 2025-05 unverdicted novelty 5.0 of 10

    VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft a...

  37. PhysAnimator: Physics-Guided Generative Cartoon Animation

    cs.GR 2025-01 conditional novelty 5.0 of 10

    PhysAnimator combines 2D deformable-body physics simulation with a sketch-guided video diffusion model to animate static anime illustrations with controllable, physically plausible motion.

  38. Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.

  39. GBR: Generative Bundle Refinement for High-fidelity Gaussian Splatting with Enhanced Mesh Reconstruction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GBR reconstructs accurate camera poses, dense point clouds, and high-fidelity meshes from 4-6 unposed images by combining DUSt3R-based neural bundle adjustment with scale-preserving diffusion depth refinement.

  40. Camera Trajectory Generation: A Comprehensive Survey of Methods, Metrics, and Future Directions

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A review that organizes camera trajectory generation into representation levels, algorithm families, evaluation metrics, and datasets.

  41. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

  42. Sparse-View 3D Reconstruction: Recent Advances and Open Challenges

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A comprehensive survey that organizes sparse-view 3D reconstruction methods into geometry-based, NeRF, 3DGS, and diffusion-based categories, with benchmarks and open challenges.

Pith tools