REVIEW 3 cited by
PixelGaussian: Generalizable 3D Gaussian Reconstruction from Arbitrary Views
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose PixelGaussian, an efficient feed-forward framework for learning generalizable 3D Gaussian reconstruction from arbitrary views. Most existing methods rely on uniform pixel-wise Gaussian representations, which learn a fixed number of 3D Gaussians for each view and cannot generalize well to more input views. Differently, our PixelGaussian dynamically adapts both the Gaussian distribution and quantity based on geometric complexity, leading to more efficient representations and significant improvements in reconstruction quality. Specifically, we introduce a Cascade Gaussian Adapter to adjust Gaussian distribution according to local geometry complexity identified by a keypoint scorer. CGA leverages deformable attention in context-aware hypernetworks to guide Gaussian pruning and splitting, ensuring accurate representation in complex regions while reducing redundancy. Furthermore, we design a transformer-based Iterative Gaussian Refiner module that refines Gaussian representations through direct image-Gaussian interactions. Our PixelGaussian can effectively reduce Gaussian redundancy as input views increase. We conduct extensive experiments on the large-scale ACID and RealEstate10K datasets, where our method achieves state-of-the-art performance with good generalization to various numbers of views. Code: https://github.com/Barrybarry-Smith/PixelGaussian.
Forward citations
Cited by 3 Pith papers
-
SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization
A feed-forward Gaussian-splatting model that subdivides each primary Gaussian into learned sub-pixel primitives, achieving state-of-the-art high-resolution novel-view synthesis from low-resolution inputs.
-
H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
H3R combines explicit epipolar volume matching with a camera-aware transformer and a spatial-aligned SD-VAE encoder, reporting state-of-the-art sparse-view 3D reconstruction on RealEstate10K, ACID, and DTU.
-
MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models
A feed-forward architecture that reuses a frozen depth foundation model to predict 3D Gaussian primitives, improving novel view synthesis and cross-dataset generalization.
Discussion (0). Continue with ORCID to comment.