{"work":{"id":"0a07d24b-c2f2-4a48-8f62-a71df4e53f9e","openalex_id":null,"doi":null,"arxiv_id":"2408.16767","raw_key":null,"title":"ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model","authors":null,"authors_text":"F","year":2024,"venue":"cs.CV","abstract":"Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a detailed scene from insufficient captured views is still an ill-posed optimization problem, often resulting in artifacts and distortions in unseen areas. In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction challenge as a temporal generation task. The key insight is to unleash the strong generative prior of large pre-trained video diffusion models for sparse-view reconstruction. However, 3D view consistency struggles to be accurately preserved in directly generated video frames from pre-trained models. To address this, given limited input views, the proposed ReconX first constructs a global point cloud and encodes it into a contextual space as the 3D structure condition. Guided by the condition, the video diffusion model then synthesizes video frames that are both detail-preserved and exhibit a high degree of 3D consistency, ensuring the coherence of the scene from various perspectives. Finally, we recover the 3D scene from the generated video through a confidence-aware 3D Gaussian Splatting optimization scheme. Extensive experiments on various real-world datasets show the superiority of our ReconX over state-of-the-art methods in terms of quality and generalizability.","external_url":"https://arxiv.org/abs/2408.16767","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-07T14:33:54.364325+00:00","pith_arxiv_id":"2408.16767","created_at":"2026-05-11T10:46:04.043336+00:00","updated_at":"2026-07-07T14:33:54.364325+00:00","title_quality_ok":true,"display_title":"Reconx: Reconstruct any scene from sparse views with video diffusion model","render_title":"Reconx: Reconstruct any scene from sparse views with video diffusion model"},"hub":{"state":{"work_id":"0a07d24b-c2f2-4a48-8f62-a71df4e53f9e","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":14,"external_cited_by_count":null,"distinct_field_count":2,"first_pith_cited_at":"2025-05-28T05:55:44+00:00","last_pith_cited_at":"2026-07-06T17:51:00+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-05T02:34:30.819373+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":2}],"polarity_counts":[{"context_polarity":"background","n":2}],"runs":{},"summary":{},"graph":{},"authors":[]}}