REVIEW 11 cited by
GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we introduce the Geometry-Aware Large Reconstruction Model (GeoLRM), an approach which can predict high-quality assets with 512k Gaussians and 21 input images in only 11 GB GPU memory. Previous works neglect the inherent sparsity of 3D structure and do not utilize explicit geometric relationships between 3D and 2D images. This limits these methods to a low-resolution representation and makes it difficult to scale up to the dense views for better quality. GeoLRM tackles these issues by incorporating a novel 3D-aware transformer structure that directly processes 3D points and uses deformable cross-attention mechanisms to effectively integrate image features into 3D representations. We implement this solution through a two-stage pipeline: initially, a lightweight proposal network generates a sparse set of 3D anchor points from the posed image inputs; subsequently, a specialized reconstruction transformer refines the geometry and retrieves textural details. Extensive experimental results demonstrate that GeoLRM significantly outperforms existing models, especially for dense view inputs. We also demonstrate the practical applicability of our model with 3D generation tasks, showcasing its versatility and potential for broader adoption in real-world applications. The project page: https://linshan-bin.github.io/GeoLRM/.
Forward citations
Cited by 11 Pith papers
-
Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction
A hybrid pixel-plus-volume Gaussian representation with triplane transformer and depth-guided training yields state-of-the-art feed-forward sparse-view reconstruction for ego-centric driving scenes.
-
Nautilus: Locality-aware Autoencoder for Scalable Mesh Generation
Nautilus introduces a locality-preserving tokenization for meshes that lets autoregressive transformers generate high-fidelity meshes with up to 5,000 faces, outperforming prior methods.
-
GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
A sparse Gaussian Transformer aligned with 2D foundation models achieves 12.27 mIoU zero-shot on Occ3D-nuScenes occupancy prediction without 3D semantic labels.
-
FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction
A feed-forward transformer that jointly predicts pixel-aligned 3D Gaussians and camera poses from uncalibrated sparse views.
-
Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D Reconstruction
Generative Densification improves feed-forward Gaussian 3D reconstruction by learning to generate fine Gaussians for detailed regions in one forward pass, and it beats baselines on object and scene datasets.
-
Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and Manipulation
Sharp-It fine-tunes a multi-view diffusion model to enhance low-quality Shap-E renderings into high-quality multi-view sets that can be reconstructed into detailed 3D assets.
-
NovelGS: Consistent Novel-view Denoising via Large Gaussian Reconstruction Model
A novel-view denoising diffusion model over pixel-aligned 3D Gaussians achieves state-of-the-art PSNR and SSIM on GSO and OmniObject3D sparse-view reconstruction.
-
Collaborative Multi-Modal Coding for High-Quality 3D Generation
TriMM fuses RGB, RGB-D, and point-cloud encoding into a shared triplane latent space and generates 3D assets from a single image with a latent diffusion model.
-
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.
-
Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction
Momentum-GS improves large-scale 3D Gaussian splatting by using a momentum teacher decoder and reconstruction-guided block weighting to boost reconstruction quality and reduce memory use.
-
DROID-Splat: Combining end-to-end SLAM with 3D Gaussian Splatting
DROID-Splat couples DROID-SLAM dense tracking with a 3D Gaussian Splatting renderer and reports state-of-the-art or near-state-of-the-art ATE and rendering scores on TUM-RGBD and Replica, with the best tracking in a s...
Discussion (0). Continue with ORCID to comment.