Pith. sign in

REVIEW 11 cited by

GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15333 v2 pith:TFLBTVVV submitted 2024-06-21 cs.CV

classification cs.CV
keywords geolrmmodelreconstructiondemonstratedensegenerationgeometry-awarehigh-quality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we introduce the Geometry-Aware Large Reconstruction Model (GeoLRM), an approach which can predict high-quality assets with 512k Gaussians and 21 input images in only 11 GB GPU memory. Previous works neglect the inherent sparsity of 3D structure and do not utilize explicit geometric relationships between 3D and 2D images. This limits these methods to a low-resolution representation and makes it difficult to scale up to the dense views for better quality. GeoLRM tackles these issues by incorporating a novel 3D-aware transformer structure that directly processes 3D points and uses deformable cross-attention mechanisms to effectively integrate image features into 3D representations. We implement this solution through a two-stage pipeline: initially, a lightweight proposal network generates a sparse set of 3D anchor points from the posed image inputs; subsequently, a specialized reconstruction transformer refines the geometry and retrieves textural details. Extensive experimental results demonstrate that GeoLRM significantly outperforms existing models, especially for dense view inputs. We also demonstrate the practical applicability of our model with 3D generation tasks, showcasing its versatility and potential for broader adoption in real-world applications. The project page: https://linshan-bin.github.io/GeoLRM/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A hybrid pixel-plus-volume Gaussian representation with triplane transformer and depth-guided training yields state-of-the-art feed-forward sparse-view reconstruction for ego-centric driving scenes.

  2. Nautilus: Locality-aware Autoencoder for Scalable Mesh Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Nautilus introduces a locality-preserving tokenization for meshes that lets autoregressive transformers generate high-fidelity meshes with up to 5,000 faces, outperforming prior methods.

  3. GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A sparse Gaussian Transformer aligned with 2D foundation models achieves 12.27 mIoU zero-shot on Occ3D-nuScenes occupancy prediction without 3D semantic labels.

  4. FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A feed-forward transformer that jointly predicts pixel-aligned 3D Gaussians and camera poses from uncalibrated sparse views.

  5. Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D Reconstruction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Generative Densification improves feed-forward Gaussian 3D reconstruction by learning to generate fine Gaussians for detailed regions in one forward pass, and it beats baselines on object and scene datasets.

  6. Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and Manipulation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Sharp-It fine-tunes a multi-view diffusion model to enhance low-quality Shap-E renderings into high-quality multi-view sets that can be reconstructed into detailed 3D assets.

  7. NovelGS: Consistent Novel-view Denoising via Large Gaussian Reconstruction Model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A novel-view denoising diffusion model over pixel-aligned 3D Gaussians achieves state-of-the-art PSNR and SSIM on GSO and OmniObject3D sparse-view reconstruction.

  8. Collaborative Multi-Modal Coding for High-Quality 3D Generation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    TriMM fuses RGB, RGB-D, and point-cloud encoding into a shared triplane latent space and generates 3D assets from a single image with a latent diffusion model.

  9. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

  10. Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Momentum-GS improves large-scale 3D Gaussian splatting by using a momentum teacher decoder and reconstruction-guided block weighting to boost reconstruction quality and reduce memory use.

  11. DROID-Splat: Combining end-to-end SLAM with 3D Gaussian Splatting

    cs.CV 2024-11 conditional novelty 4.0 of 10

    DROID-Splat couples DROID-SLAM dense tracking with a 3D Gaussian Splatting renderer and reports state-of-the-art or near-state-of-the-art ATE and rendering scores on TUM-RGBD and Replica, with the best tracking in a s...

Pith tools