REVIEW 9 cited by
Real3D: Scaling Up Large Reconstruction Models with Real-World Images
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The default strategy for training single-view Large Reconstruction Models (LRMs) follows the fully supervised route using large-scale datasets of synthetic 3D assets or multi-view captures. Although these resources simplify the training procedure, they are hard to scale up beyond the existing datasets and they are not necessarily representative of the real distribution of object shapes. To address these limitations, in this paper, we introduce Real3D, the first LRM system that can be trained using single-view real-world images. Real3D introduces a novel self-training framework that can benefit from both the existing synthetic data and diverse single-view real images. We propose two unsupervised losses that allow us to supervise LRMs at the pixel- and semantic-level, even for training examples without ground-truth 3D or novel views. To further improve performance and scale up the image data, we develop an automatic data curation approach to collect high-quality examples from in-the-wild images. Our experiments show that Real3D consistently outperforms prior work in four diverse evaluation settings that include real and synthetic data, as well as both in-domain and out-of-domain shapes. Code and model can be found here: https://hwjiang1510.github.io/Real3D/
Forward citations
Cited by 9 Pith papers
-
4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
4D-LRM is a transformer that maps sparse posed frames scattered across time to a cloud of 4D Gaussians and renders any query view at any query time in under 1.5 seconds.
-
RayZer: A Self-supervised Large View Synthesis Model
A self-supervised transformer model predicts camera poses and scene features from unposed images and renders novel views, reaching performance on par with pose-supervised baselines.
-
F3D-Gaus: Feed-forward 3D-aware Generation on ImageNet with Cycle-Aggregative Gaussian Splatting
F3D-Gaus predicts a pixel-aligned 3D Gaussian representation from a single RGB-D image and uses cycle-aggregative self-supervision plus video-prior refinement to render consistent novel views from monocular training d...
-
Dust to Tower: Coarse-to-Fine Photo-Realistic Scene Reconstruction from Sparse Uncalibrated Images
A coarse-to-fine pipeline jointly optimizes 3D Gaussian Splatting and camera poses from sparse, uncalibrated images, using warped and inpainted pseudo-views for supervision.
-
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
A finetuned pose-conditioned diffusion model, Drive-1-to-3, synthesizes photorealistic novel views of real vehicles from a single image on Waymo and other driving datasets, beating prior methods on FID and LPIPS.
-
MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
A 700K-scene procedural, non-semantic synthetic dataset improves large reconstruction models by 1.2 to 1.8 dB PSNR when combined with real data.
-
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
LoRA3D specializes pretrained 3D foundation models to target scenes via confidence-calibrated pseudo-labels from multi-view robust optimization and LoRA fine-tuning, improving performance by up to 88%.
-
3D Arena: An Open Platform for Generative 3D Evaluation
A crowdsourced voting platform with 123,000 votes reveals that people judge AI-generated 3D assets mainly by visual appearance rather than technical quality.
-
Instructive3D: Editing Large Reconstruction Models with Text Instructions
A text-conditioned diffusion adapter operating on the triplane latents of a frozen large reconstruction model enables natural-language editing of generated 3D objects.
Discussion (0). Continue with ORCID to comment.