REVIEW 4 major objections 5 minor 58 references
Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that saliency-guided superpixel merging can compress feed-forward 3D Gaussian reconstructions to about one twentieth of their original primitive count while keeping novel-view quality close to the unmerged backbone, and…
desk verdict A genuinely new and well-executed compression pipeline for feed-forward 3DGS, but the LPIPS cost is understated and the depth-coherence assumption needs testing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Feature Gaussian, a single Gaussian-format token (position, scale, rotation, base color, plus a latent vector) that represents a whole superpixel group. It is produced by a Set Transformer encoder (SAB and PMA blocks) that is permutation-invariant and handles arbitrary group sizes, merged across views by an identical SAB+PMA module with zero-initialized residual heads, and expanded by a slot-based decoder into K output Gaussians. Saliency-guided BASS segmentation with Shi-Tomasi corner seeding is what makes the groups content-adaptive: dense seeds in textured regions, sparse seeds in homogeneous areas. A moment-matching teacher loss supplies initial geometry, and a diversification regularizer prevents the K>1 decoders from collapsing into identical copies.
What would settle it
Take a scene with two identical-looking flat surfaces at different depths (for example two coplanar-colored walls separated by a gap) and run the pipeline: if any superpixel crosses the depth boundary, the merged Gaussian set will smear the two surfaces and novel-view PSNR in that region should drop measurably below the unmerged backbone. A quantitative version would compare quality on the subset of rendered pixels whose superpixels contain a depth discontinuity larger than, say, ten times the local Gaussian scale; the pipeline's quality should degrade sharply there if the central claim about geometric coherence is correct.
Extended reading notes
Core claim
The central discovery is that primitive redundancy in feed-forward 3DGS is structured rather than random: per-pixel Gaussians form coherent clusters in image space that can be encoded, matched across views, and decoded back to a few Gaussians per cluster with little loss. The pipeline segments each input view with BASS superpixels seeded by the Shi-Tomasi corner response, so textured regions receive small segments and flat regions large ones; a Set Transformer encoder collapses each segment into one Feature Gaussian (mean, scale, rotation, base color, and a latent vector); a learned merger fuses Feature Gaussians from different views when their latent features and axis-aligned bounding box overlap match; a refiner updates each Feature Gaussian from its neighbors; and a level-of-detail decoder expands each Feature Gaussian into K output Gaussians, with K=1, 2, and 4 trained jointly. On the DA3 backbone, K=1 retains 4.4% of the original primitives and reaches 17.41 dB PSNR on average against 16.82 dB for the unmerged backbone, a regularizing effect the paper attributes to merging removing floaters and high-frequency noise.
Load-bearing premise
The method assumes that a superpixel in image space corresponds to a geometrically coherent cluster in 3D; if a superpixel straddles a depth discontinuity or merges visually similar but separated surfaces, the latent representation cannot recover the lost structure and compression fails.
Editorial extensions
If this is right
- At K=1 the pipeline compresses to 4.4% of the original primitive count and renders about 6x faster (365 vs 60 FPS on the online MipNeRF360 setting), with PSNR and SSIM roughly matching the unmerged backbone.
- In online reconstruction the primitive-growth slope drops 16x (0.12M vs 1.84M Gaussians per 12 views), so scenes with many views stay within memory bounds: 0.82M vs 12.9M Gaussians after 84 views.
- K and superpixel size are independent inference-time knobs that let users trade quality for speed without retraining, and the two mechanisms are complementary: K=2 at r_c=8.9% matches K=4 with larger superpixels at r_c=8.0%.
- Because the backbone stays frozen, the module can be attached to newer and stronger feed-forward predictors as they appear, and it can be combined with input-side compression methods that reduce the number of views.
- Compared with ReSplat and VolSplat, the paper reports higher quality at comparable or lower primitive counts, with the gains largest in sparse-view settings.
Reading between the lines
- The paper's results suggest that a large share of per-pixel Gaussians in feed-forward reconstructions are pure redundancy, so a future method that predicts primitives directly at a content-adaptive resolution might not need a merging step at all.
- A natural extension, not tested in the paper, is to make K adaptive per Feature Gaussian rather than uniform across the scene, allocating more output Gaussians to superpixels that retain residual detail, which could squeeze further compression at equal quality.
- The regularizing effect on large-viewpoint benchmarks hints that merging acts as a structural prior that suppresses floaters; if so, it could also improve robustness in even sparser settings than those tested, such as two-view input with a wide baseline.
- The same encode-merge-decode idea could in principle apply to other per-pixel output modalities, such as depth or normal fields, although the decoder would need a different output parameterization than Gaussian splatting.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a post-processing module for feed-forward 3D Gaussian splatting that reduces primitive count. It segments per-pixel Gaussians into superpixels guided by a saliency/corner map, encodes each group into a latent-augmented 'Feature Gaussian', matches and merges Feature Gaussians across views, refines them with self-attention, and decodes each Feature Gaussian into K output Gaussians. The pipeline is evaluated with AnySplat, DepthSplat, and Depth Anything 3 backbones on DL3DV-Bench, MipNeRF360, and Tanks & Temples, reporting relative primitive counts as low as 4.4% at K=1 with PSNR and SSIM close to the unmerged backbones but consistently higher LPIPS. It also reports an online reconstruction experiment, a superpixel proxy-quality analysis, and ablations of the refiner, cross-view merging, and superpixel sizes.
Significance. If the claims hold, the contribution is practically useful: a backbone-agnostic, inference-time quality/efficiency knob for feed-forward 3D Gaussian splatting, with a clear compression mechanism and extensive evaluation across three backbones and three benchmarks. The paper is commendable for reporting LPIPS explicitly rather than relying only on PSNR/SSIM, for ablating the refiner and cross-view merging, for including a proxy superpixel-quality analysis, and for evaluating online reconstruction. The main scientific risk is the unquantified 3D coherence of image-space superpixels; the consistent LPIPS degradation in Table 1 suggests perceptual losses that the paper does not fully characterize or localize.
major comments (4)
- [Sec. 4.1, Sec. 4.5, Tab. 1] The central grouping assumption is not verified. Superpixels are formed in image space from color and spatial compactness, so a superpixel can contain per-pixel Gaussians from two surfaces at different depths whenever the surfaces have similar appearance. A K=1 decoder cannot split such a bimodal cluster, and the refiner (Sec. 4.5) updates an existing Feature Gaussian but cannot split it into multiple surfaces. The paper does not quantify how often superpixels cross depth discontinuities, nor does it report region-level metrics at depth boundaries. The large LPIPS increases in Table 1 (e.g., DA3 K=1: 0.422 vs. 0.297 on DL3DV-Bench and 0.442 vs. 0.356 on MipNeRF360) are consistent with collapsed depth structure, and the Limitations section does not mention this failure mode. I ask the authors to add an analysis of superpixel depth coherence (e.g., fraction of superpixels with large internal depth variance, or per-region metrics at depth edges) and, if needed, an ablation with depth-aware grouping.
- [Sec. 5.2, Tab. 1] The 'largely retaining visual quality' claim is weakened by the LPIPS results. For DA3 K=1, LPIPS rises from 0.297 to 0.422 on DL3DV-Bench and from 0.356 to 0.442 on MipNeRF360, which are relative increases of roughly 40%; SSIM also drops on DL3DV-Bench from 0.620 to 0.578. Since LPIPS is a perceptual metric, the paper should either temper the wording of 'largely retaining visual quality' or provide supporting analysis (per-scene breakdown, error maps, or a perceptual metric that distinguishes acceptable texture loss from structural collapse). The current presentation understates the magnitude of the perceptual degradation.
- [Sec. 5.2, Abstract] The comparative claim of being 'better and more robust quality than achieved by previous approaches that target a reduction in primitive count' is under-tested. The experimental comparison includes only ReSplat and VolSplat; Off The Grid (Ref. [25]) and Fuse-and-Refine (Ref. [39]) are cited but not evaluated due to code unavailability, and the graph-based fusion methods FreeSplat and Gaussian Graph Network discussed in Sec. 2 are not compared. The claim should be scoped to the evaluated baselines, or the missing comparisons should be included with a clear justification for their omission.
- [Sec. 5.1, Tab. 1] The paper's emphasis on sparse-view performance is not directly supported by the main-text results. Table 1 averages over experiments with 3, 6, 9, and 12 input views, and the detailed per-view-count results are only mentioned as being in the supplementary material. Since the abstract and conclusion specifically highlight robustness 'particularly in sparse-view settings', the per-view-count table (or at least a subset, such as 3 and 6 views) should be presented in the main body, or the sparse-view claim should be softened.
minor comments (5)
- [Sec. 5.1] Please state whether the code and trained models will be released; the current text only says that certain baselines are unavailable, which limits reproducibility for reviewers and readers.
- [Fig. 1] The legend contains many markers and is hard to read at small size; consider splitting it into separate panels with clearer labels.
- [Tab. 3] The phrase 'zero shot larger SPs' should be 'zero-shot larger SPs' for consistency and clarity.
- [Sec. 4.2] The term 'Feature Gaussian' is an invented entity; please add a one-sentence summary of how it differs from an ordinary Gaussian in the matching and decoding stages, even though this is partially explained in the text.
- [Sec. 4.6] The teacher loss in Equations (3) and (4) is described as a closed-form moment-matching target; please clarify whether it is used only for the K=1 head or also for higher-K decoders during the early training phase.
Circularity Check
No significant circularity: the pipeline is a learned post-processor trained with photometric losses and evaluated on held-out benchmarks; the primitive-count reduction is an explicit design choice, not a fitted prediction.
full rationale
The derivation is self-contained. Per-pixel Gaussians from a frozen backbone are grouped by BASS superpixel segmentation; each group is encoded into a Feature Gaussian by a learned Set Transformer encoder; a learned merger fuses cross-view Feature Gaussians; a refiner updates them; and a level-of-detail decoder produces output Gaussians. All learned modules are trained end-to-end with photometric losses (MSE, SSIM, LPIPS) on rendered images, and evaluation is on held-out datasets (DL3DV-Bench, MipNeRF360, Tanks & Temples) not used for training. The moment-matching teacher loss (Eqs. 3-4) is explicitly an auxiliary, decaying regularizer ('This teacher signal decays on a schedule, allowing the learned encoder to eventually surpass the moment-matched initialization'), and the ablations show the learned pipeline outperforms the heuristic Mom. Match. baseline, so the final result is not forced to equal the teacher target. The reported reduction to 1/20th of the Gaussians is not a predicted quantity: it follows directly from the chosen superpixel granularity and K decoder head (K=1), i.e., it is a design setting, and the actual empirical claim—the rendering quality at that count—is measured against external benchmarks. The cited components (BASS, Set Transformer, 3DGS) are external methods used as building blocks, not self-citations that carry the argument. The unverified depth-boundary concern is a generalization/correctness risk, not circularity, because no equation or fitted parameter makes the quality outcome equal to an input.
Assumptions & free parameters
free parameters (6)
- tau_f (feature similarity threshold) =
not stated in paper
- tau_g (geometry IoU threshold) =
not stated in paper
- tau_a (opacity pruning threshold) =
not stated in paper
- k (kNN neighborhood size) =
not stated in paper
- superpixel seed density and size settings =
large, medium, small tested
- decoder slots K =
1, 2, 4
assumptions (4)
- domain assumption Superpixels in image space correspond to 3D-coherent groups of Gaussians
- domain assumption Shi-Tomasi corner response is a good saliency proxy for where fine Gaussian detail is needed
- domain assumption The frozen backbone's per-pixel Gaussians are accurate enough that merging only removes redundancy
- domain assumption The learned modules can be trained end-to-end with photometric losses despite random initialization
invented entities (1)
-
Feature Gaussian (FG)
Cite this review
Pith. "Pith review of Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging." pith.science (2026). https://pith.science/paper/LBULU3Q2
@misc{pith2026260810712,
author = {Pith},
title = {Pith review of: Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging},
year = {2026},
howpublished = {\url{https://pith.science/paper/LBULU3Q2}},
note = {Machine review of arXiv:2608.10712}
}
abstract
3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this context. Its feed-forward variants provide fast reconstruction from sparse input views but often produce per-pixel primitives, leading to highly redundant and thus inefficient representations. We present a structure-aware merging pipeline that takes per-pixel primitives from any feed-forward method and consolidates them into a compact, content-adaptive Gaussian set while largely retaining visual quality at just $\frac{1}{20}^\text{th}$ of the Gaussians of a per-pixel method. We group spatially coherent Gaussians of similar appearance into variable-size clusters via adaptive superpixel segmentation guided by a saliency map, which allocates fine segments to textured regions and coarse segments to homogeneous areas. We compress each cluster into a compact latent representation through a learned encoder, then match and consolidate representations across views based on geometric overlap and feature similarity via a learned merger. A level-of-detail decoder then produces the final Gaussians at a controllable resolution, enabling a flexible quality-efficiency trade-off at inference. As a post-processing module, the pipeline is backbone-agnostic, leveraging the strengths of existing feed-forward methods. This leads to better and more robust quality than achieved by previous approaches that target a reduction in primitive count, while providing a highly compact representation, that can be rendered efficiently.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[25]
Off The Grid: Detection of Prim- itives for Feed-Forward 3D Gaussian Splatting
Arthur Moreau, Richard Shaw, Michal Nazarczuk, Jisu Shin, Thomas Tanay, Zhensong Zhang, Songcen Xu, and Eduardo Pérez-Pellitero. Off The Grid: Detection of Prim- itives for Feed-Forward 3D Gaussian Splatting. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2026
work page 2026
-
[39]
Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian Splatting
Yiming Wang, Lucy Chai, Xuan Luo, Michael Niemeyer, Manuel Lagunas, Stephen Lombardi, Siyu Tang, and Tiancheng Sun. Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian Splatting. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2025
work page 2025
-
[1]
SLIC Superpixels Compared to State-of-the-Art Superpixel Meth- ods.IEEE Trans
Radhakrishna Achanta, Appu Shaji, Kevin Smith, Aurelien Lucchi, Pascal Fua, and Sabine Süsstrunk. SLIC Superpixels Compared to State-of-the-Art Superpixel Meth- ods.IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 34(11):2274– 2282, 2012. doi: 10.1109/TPAMI.2012.120
-
[2]
Tomas Akenine-Möller, Eric Haines, Naty Hoffman, Angelo Pesce, Michaël Iwanicki, and Sébastien Hillaire.Real-Time Rendering. CRC Press, 4 edition, 2018. doi: 10. 1201/b22086
work page 2018
-
[3]
MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM
Yinlong Bai, Hongxin Zhang, Sheng Zhong, Junkai Niu, Hai Li, Yijia He, and Yi Zhou. MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM. InProc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), 2025
work page 2025
-
[4]
Isabela B. Barcelos, Felipe D.C. Belem, Leonardo D.M. Joao, Zenilton K.G.D. Pa- trocínio Jr., Alexandre X. Falcao, and Silvio J.F. Guimarães. A Comprehensive Review and New Taxonomy on Superpixel Segmentation.ACM Computing Surveys, 56(8): 1–39, 2024. doi: 10.1145/3643826
-
[5]
Barron, Ben Mildenhall, Dor Verbin, Pratul P
Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hed- man. Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2022
work page 2022
-
[6]
pix- elSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Recon- struction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pix- elSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Recon- struction. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024
work page 2024
Show all 58 references
-
[7]
Fast Feedforward 3D Gaussian Splatting Compression
Yihang Chen, Qianyi Wu, Mengyao Li, Weiyao Lin, Mehrtash Harandi, and Jianfei Cai. Fast Feedforward 3D Gaussian Splatting Compression. InProc. of the Intl. Conf. on Learning Representations (ICLR), 2025
2025
-
[8]
MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images. InProc. of the Europ. Conf. on Computer Vision (ECCV), 2024. 16FAASCH, KALL, STACHNISS:...
2024
-
[9]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition ...
2021
-
[10]
Mini-Splatting: Representing Scenes with a Con- strained Number of Gaussians
Guangchi Fang and Bing Wang. Mini-Splatting: Representing Scenes with a Con- strained Number of Gaussians. InProc. of the Europ. Conf. on Computer Vision (ECCV), 2024
2024
-
[11]
AnySplat: Feed-Forward 3D Gaussian Splatting from Unconstrained Views.ACM Trans
Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu, Kerui Ren, Yichen Jin, Xudong Xu, Mulin Yu, Jiangmiao Pang, and Feng Zhao. AnySplat: Feed-Forward 3D Gaussian Splatting from Unconstrained Views.ACM Trans. on Graphics (TOG), 44(6):1–16,
-
[12]
Lau, Feng Gao, Yin Yang, and Chenfanfu Jiang
Ying Jiang, Chang Yu, Tianyi Xie, Xuan Li, Yutao Feng, Huamin Wang, Minchen Li, Henry Y .K. Lau, Feng Gao, Yin Yang, and Chenfanfu Jiang. VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual Reality. InProc. of the Intl. Conf. on Computer Graphics ...
2024
-
[13]
3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans. on Graphics (TOG), 42(4):1–14, 2023. doi: 10.1145/3592433
2023 doi
-
[14]
A Hierarchical 3D Gaussian Representation for Real- Time Rendering of Very Large Datasets.ACM Trans
Bernhard Kerbl, Andréas Meuleman, Georgios Kopanas, Michael Wimmer, Alexandre Lanvin, and George Drettakis. A Hierarchical 3D Gaussian Representation for Real- Time Rendering of Very Large Datasets.ACM Trans. on Graphics (TOG), 43(4):1–15,
-
[15]
Tanks and Temples: Benchmarking Large-Scale Scene Reconstruction.ACM Trans
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and Temples: Benchmarking Large-Scale Scene Reconstruction.ACM Trans. on Graphics (TOG), 36(4):1–13, 2017. doi: 10.1145/3072959.3073599
2017
-
[16]
Com- pact 3D Gaussian Representation for Radiance Field
Joo Chan Lee, Daniel Rho, Xiangyu Sun, Jong Hwan Ko, and Eunbyung Park. Com- pact 3D Gaussian Representation for Radiance Field. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[17]
Optimized Minimal 3D Gaussian Splatting
Joo Chan Lee, Jong Hwan Ko, and Eunbyung Park. Optimized Minimal 3D Gaussian Splatting. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2025
2025
-
[18]
Kosiorek, Seungjin Choi, and Yee Whye Teh
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam R. Kosiorek, Seungjin Choi, and Yee Whye Teh. Set Transformer: A Framework for Attention-based Permutation- Invariant Neural Networks. InProc. of the Intl. Conf. on Machine Learning (ICML), 2019
2019
-
[19]
Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang
Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y . Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth Anything 3: Recovering the Visual Space from Any Views.arXiv preprint, arXiv:2511.10647, 2025. FAASCH, KALL, STACHNISS: COMPACT FF-3DGS VIA SALIENCY -GUIDED MERGING17
2025 arXiv
-
[20]
DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-Based 3D Vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Anirud- dha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua, Tianyi Zhang, Bedrich Benes, and Aniket Bera. DL3DV-10K: A Large-Scale S...
2024
-
[21]
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization. InProc. of the Intl. Conf. on Learning Representations (ICLR), 2019
2019
-
[22]
Taming 3DGS: High-Quality Ra- diance Fields with Limited Resources
Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl, Markus Steinberger, Fran- cisco Vicente Carrasco, and Fernando De La Torre. Taming 3DGS: High-Quality Ra- diance Fields with Limited Resources. InProc. of the Intl. Conf. on Computer Graphics and Interactive Techniques (SI...
2024
-
[23]
EV olSplat: Efficient V olume-Based Gaussian Splatting for Urban View Synthesis
Sheng Miao, Jiaxin Huang, Dongfeng Bai, Xu Yan, Hongyu Zhou, Yue Wang, Bing- bing Liu, Andreas Geiger, and Yiyi Liao. EV olSplat: Efficient V olume-Based Gaussian Splatting for Urban View Synthesis. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2025
2025
-
[24]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ra- mamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. InProc. of the Europ. Conf. on Computer Vision (ECCV), 2020
2020
-
[26]
Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis
Simon Niedermayr, Josef Stumpfegger, and Rüdiger Westermann. Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[27]
Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors
Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, and Bernard Ghanem. Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors. InProc. of the ...
2024
-
[28]
SCube: Instant Large-Scale Scene Re- construction Using V oxSplats
Xuanchi Ren, Yifan Lu, Hanxue Liang, Zhangjie Wu, Huan Ling, Mike Chen, Sanja Fidler, Francis Williams, and Jiahui Huang. SCube: Instant Large-Scale Scene Re- construction Using V oxSplats. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2024
2024
-
[29]
Ze- roNVS: Zero-Shot 360-Degree View Synthesis from a Single Image
Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, and Jiajun Wu. Ze- roNVS: Zero-Shot 360-Degree View Synthesis from a Single Image. InProc. of the IEEE/CVF Conf. on Computer Vision and Pa...
2024
-
[30]
PhD thesis, Ruprecht-Karls-Universität Heidelberg, 2000
Hanno Scharr.Optimale Operatoren in der digitalen Bildverarbeitung. PhD thesis, Ruprecht-Karls-Universität Heidelberg, 2000. 18FAASCH, KALL, STACHNISS: COMPACT FF-3DGS VIA SALIENCY -GUIDED MERGING
2000
-
[31]
Normalized Cuts and Image Segmentation.IEEE Trans
Jianbo Shi and Jitendra Malik. Normalized Cuts and Image Segmentation.IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 22(8):888–905, 2000. doi: 10.1109/34.868688
-
[32]
Good Features to Track
Jianbo Shi and Carlo Tomasi. Good Features to Track. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 1994
1994
-
[33]
Splatter Image: Ultra-Fast Single-View 3D Reconstruction
Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Splatter Image: Ultra-Fast Single-View 3D Reconstruction. InProc. of the IEEE/CVF Conf. on Com- puter Vision and Pattern Recognition (CVPR), 2024
2024
-
[34]
NeuRAD: Neural Rendering for Autonomous Driving
Adam Tonderski, Carl Lindström, Georg Hess, William Ljungbergh, Lennart Svensson, and Christoffer Petersson. NeuRAD: Neural Rendering for Autonomous Driving. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[35]
Bayesian Adaptive Superpixel Segmen- tation
Roy Uziel, Meitar Ronen, and Oren Freifeld. Bayesian Adaptive Superpixel Segmen- tation. InProc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2019
2019
-
[36]
Smol-GS: Compact Rep- resentations for Abstract 3D Gaussian Splatting.arXiv preprint, arXiv:2512.00850, 2025
Haishan Wang, Mohammad Hassan Vali, and Arno Solin. Smol-GS: Compact Rep- resentations for Abstract 3D Gaussian Splatting.arXiv preprint, arXiv:2512.00850, 2025
2025 arXiv
-
[37]
Chen, Zeyu Zhang, Duochao Shi, Akide Liu, and Bohan Zhuang
Weijie Wang, Donny Y . Chen, Zeyu Zhang, Duochao Shi, Akide Liu, and Bohan Zhuang. ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2025
2025
-
[38]
Chen, and Bohan Zhuang
Weijie Wang, Yeqing Chen, Zeyu Zhang, Hengyu Liu, Haoxiao Wang, Zhiyuan Feng, Wenkang Qin, Zheng Zhu, Donny Y . Chen, and Bohan Zhuang. V olSplat: Rethinking Feed-Forward 3D Gaussian Splatting with V oxel-Aligned Prediction.arXiv preprint, arXiv:2509.19297, 2025
2025 arXiv
-
[40]
FreeSplat: Generaliz- able 3D Gaussian Splatting Towards Free-View Synthesis of Indoor Scenes
Yunsong Wang, Tianxin Huang, Hanlin Chen, and Gim Hee Lee. FreeSplat: Generaliz- able 3D Gaussian Splatting Towards Free-View Synthesis of Indoor Scenes. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2024
2024
-
[41]
Bovik, Hamid R
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image Quality Assessment: From Error Visibility to Structural Similarity.IEEE Trans. on Image Processing, 13(4):600–612, 2004. doi: 10.1109/TIP.2003.819861
2004
-
[42]
PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics
Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chenfanfu Jiang. PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[43]
ReSplat: Learning Recurrent Gaussian Splatting.arXiv preprint, arXiv:2510.08575, 2025
Haofei Xu, Daniel Barath, Andreas Geiger, and Marc Pollefeys. ReSplat: Learning Recurrent Gaussian Splatting.arXiv preprint, arXiv:2510.08575, 2025. FAASCH, KALL, STACHNISS: COMPACT FF-3DGS VIA SALIENCY -GUIDED MERGING19
2025
-
[44]
DepthSplat: Connecting Gaussian Splatting and Depth
Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger, and Marc Pollefeys. DepthSplat: Connecting Gaussian Splatting and Depth. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2025
2025
-
[45]
GRM: Large Gaussian Reconstruction Model for Effi- cient 3D Reconstruction and Generation
Yinghao Xu, Zifan Shi, Wang Yifan, Hansheng Chen, Ceyuan Yang, Sida Peng, Yujun Shen, and Gordon Wetzstein. GRM: Large Gaussian Reconstruction Model for Effi- cient 3D Reconstruction and Generation. InProc. of the Europ. Conf. on Computer Vision (ECCV), 2024
2024
-
[46]
GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting
Chi Yan, Delin Qu, Dan Xu, Bin Zhao, Zhigang Wang, Dong Wang, and Xuelong Li. GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[47]
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images.arXiv preprint, arXiv:2410.24207, 2024
Botao Ye, Sifei Liu, Haofei Xu, Xueting Li, Marc Pollefeys, Ming-Hsuan Yang, and Songyou Peng. No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images.arXiv preprint, arXiv:2410.24207, 2024
2024 arXiv
-
[48]
GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting. In Proc. of the Europ. Conf. on Computer Vision (ECCV), 2024
2024
-
[49]
Efros, Eli Shechtman, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[50]
Gaussian Graph Network: Learning Efficient and Generalizable Gaussian Representations from Multi-view Images
Shengjun Zhang, Xin Fei, Fangfu Liu, Haixu Song, and Yueqi Duan. Gaussian Graph Network: Learning Efficient and Generalizable Gaussian Representations from Multi-view Images. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2024
2024
-
[51]
Superpixels via Pseudo-Boolean Optimization
Yuhang Zhang, Richard Hartley, John Mashford, and Stewart Burn. Superpixels via Pseudo-Boolean Optimization. InProc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2011
2011
-
[52]
Content-Aware Dynamic Superpixel Segmentation
Tingyu Zhao, Bo Peng, Zhenguang Zhang, Daipeng Yang, and Xi Wu. Content-Aware Dynamic Superpixel Segmentation. InProc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025
2025
-
[53]
GaussianGrasper: 3D Language Gaus- sian Splatting for Open-V ocabulary Robotic Grasping.IEEE Robotics and Automation Letters (RA-L), 9(9):7827–7834, 2024
Yuhang Zheng, Xiangyu Chen, Yupeng Zheng, Songen Gu, Runyi Yang, Bu Jin, Pengfei Li, Chengliang Zhong, Zengmao Wang, Lina Liu, Chao Yang, Dawei Wang, Zhen Chen, Xiaoxiao Long, and Meiqing Wang. GaussianGrasper: 3D Language Gaus- sian Splatting for Open-V ocabulary Robotic Gras...
2024
-
[54]
Stereo Magnification: Learning View Synthesis Using Multiplane Images.ACM Trans
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo Magnification: Learning View Synthesis Using Multiplane Images.ACM Trans. on Graphics (TOG), 37(4):1–12, 2018. doi: 10.1145/3197517.3201323
2018
-
[55]
DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Au- tonomous Driving Scenes
Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Au- tonomous Driving Scenes. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024. 20FAA...
2024
-
[56]
EW A V olume Splatting
Matthias Zwicker, Hanspeter Pfister, Jeroen van Baar, and Markus Gross. EW A V olume Splatting. InProc. of the IEEE Visualization Conf. (VIS), 2001
2001
-
[2024]
doi: 10.1145/3658160
-
[2025]
doi: 10.1145/3763326
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.