Pith. sign in

REVIEW 4 major objections 5 minor 58 references

Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that saliency-guided superpixel merging can compress feed-forward 3D Gaussian reconstructions to about one twentieth of their original primitive count while keeping novel-view quality close to the unmerged backbone, and…

desk verdict A genuinely new and well-executed compression pipeline for feed-forward 3DGS, but the LPIPS cost is understated and the depth-coherence assumption needs testing. read the letter →

arxiv 2608.10712 v1 pith:LBULU3Q2 submitted 2026-08-11 cs.CV

classification cs.CV
keywords 3DGaussiansplattingfeed-forwardreconstructionprimitivemergingsuperpixelsegmentationsaliency-guidedseedingnovelviewsynthesislevel-of-detaildecodingsparse-view
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Feed-forward 3D Gaussian splatting produces one primitive per pixel, a highly redundant representation that slows rendering and inflates memory. The paper claims that grouping per-pixel Gaussians into saliency-guided superpixels, compressing each group into a latent Feature Gaussian, and merging matching groups across views yields a compact set with roughly 1/20th of the primitives while keeping PSNR and SSIM close to the unmerged backbone (with LPIPS degraded). The pipeline is a backbone-agnostic post-processing module, and the paper shows it holds across three different feed-forward backbones and three benchmarks, including sparse-view settings where it is more robust than competing reduced-primitive methods. If true, this makes feed-forward reconstruction practical for large-scale simulation and robotics workloads where rendering speed is the bottleneck.

What carries the argument

The load-bearing object is the Feature Gaussian, a single Gaussian-format token (position, scale, rotation, base color, plus a latent vector) that represents a whole superpixel group. It is produced by a Set Transformer encoder (SAB and PMA blocks) that is permutation-invariant and handles arbitrary group sizes, merged across views by an identical SAB+PMA module with zero-initialized residual heads, and expanded by a slot-based decoder into K output Gaussians. Saliency-guided BASS segmentation with Shi-Tomasi corner seeding is what makes the groups content-adaptive: dense seeds in textured regions, sparse seeds in homogeneous areas. A moment-matching teacher loss supplies initial geometry, and a diversification regularizer prevents the K>1 decoders from collapsing into identical copies.

What would settle it

Take a scene with two identical-looking flat surfaces at different depths (for example two coplanar-colored walls separated by a gap) and run the pipeline: if any superpixel crosses the depth boundary, the merged Gaussian set will smear the two surfaces and novel-view PSNR in that region should drop measurably below the unmerged backbone. A quantitative version would compare quality on the subset of rendered pixels whose superpixels contain a depth discontinuity larger than, say, ten times the local Gaussian scale; the pipeline's quality should degrade sharply there if the central claim about geometric coherence is correct.

Watch

Extended reading notes

Core claim

The central discovery is that primitive redundancy in feed-forward 3DGS is structured rather than random: per-pixel Gaussians form coherent clusters in image space that can be encoded, matched across views, and decoded back to a few Gaussians per cluster with little loss. The pipeline segments each input view with BASS superpixels seeded by the Shi-Tomasi corner response, so textured regions receive small segments and flat regions large ones; a Set Transformer encoder collapses each segment into one Feature Gaussian (mean, scale, rotation, base color, and a latent vector); a learned merger fuses Feature Gaussians from different views when their latent features and axis-aligned bounding box overlap match; a refiner updates each Feature Gaussian from its neighbors; and a level-of-detail decoder expands each Feature Gaussian into K output Gaussians, with K=1, 2, and 4 trained jointly. On the DA3 backbone, K=1 retains 4.4% of the original primitives and reaches 17.41 dB PSNR on average against 16.82 dB for the unmerged backbone, a regularizing effect the paper attributes to merging removing floaters and high-frequency noise.

Load-bearing premise

The method assumes that a superpixel in image space corresponds to a geometrically coherent cluster in 3D; if a superpixel straddles a depth discontinuity or merges visually similar but separated surfaces, the latent representation cannot recover the lost structure and compression fails.

Editorial extensions

If this is right

  • At K=1 the pipeline compresses to 4.4% of the original primitive count and renders about 6x faster (365 vs 60 FPS on the online MipNeRF360 setting), with PSNR and SSIM roughly matching the unmerged backbone.
  • In online reconstruction the primitive-growth slope drops 16x (0.12M vs 1.84M Gaussians per 12 views), so scenes with many views stay within memory bounds: 0.82M vs 12.9M Gaussians after 84 views.
  • K and superpixel size are independent inference-time knobs that let users trade quality for speed without retraining, and the two mechanisms are complementary: K=2 at r_c=8.9% matches K=4 with larger superpixels at r_c=8.0%.
  • Because the backbone stays frozen, the module can be attached to newer and stronger feed-forward predictors as they appear, and it can be combined with input-side compression methods that reduce the number of views.
  • Compared with ReSplat and VolSplat, the paper reports higher quality at comparable or lower primitive counts, with the gains largest in sparse-view settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's results suggest that a large share of per-pixel Gaussians in feed-forward reconstructions are pure redundancy, so a future method that predicts primitives directly at a content-adaptive resolution might not need a merging step at all.
  • A natural extension, not tested in the paper, is to make K adaptive per Feature Gaussian rather than uniform across the scene, allocating more output Gaussians to superpixels that retain residual detail, which could squeeze further compression at equal quality.
  • The regularizing effect on large-viewpoint benchmarks hints that merging acts as a structural prior that suppresses floaters; if so, it could also improve robustness in even sparser settings than those tested, such as two-view input with a wide baseline.
  • The same encode-merge-decode idea could in principle apply to other per-pixel output modalities, such as depth or normal fields, although the decoder would need a different output parameterization than Gaussian splatting.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a post-processing module for feed-forward 3D Gaussian splatting that reduces primitive count. It segments per-pixel Gaussians into superpixels guided by a saliency/corner map, encodes each group into a latent-augmented 'Feature Gaussian', matches and merges Feature Gaussians across views, refines them with self-attention, and decodes each Feature Gaussian into K output Gaussians. The pipeline is evaluated with AnySplat, DepthSplat, and Depth Anything 3 backbones on DL3DV-Bench, MipNeRF360, and Tanks & Temples, reporting relative primitive counts as low as 4.4% at K=1 with PSNR and SSIM close to the unmerged backbones but consistently higher LPIPS. It also reports an online reconstruction experiment, a superpixel proxy-quality analysis, and ablations of the refiner, cross-view merging, and superpixel sizes.

Significance. If the claims hold, the contribution is practically useful: a backbone-agnostic, inference-time quality/efficiency knob for feed-forward 3D Gaussian splatting, with a clear compression mechanism and extensive evaluation across three backbones and three benchmarks. The paper is commendable for reporting LPIPS explicitly rather than relying only on PSNR/SSIM, for ablating the refiner and cross-view merging, for including a proxy superpixel-quality analysis, and for evaluating online reconstruction. The main scientific risk is the unquantified 3D coherence of image-space superpixels; the consistent LPIPS degradation in Table 1 suggests perceptual losses that the paper does not fully characterize or localize.

major comments (4)
  1. [Sec. 4.1, Sec. 4.5, Tab. 1] The central grouping assumption is not verified. Superpixels are formed in image space from color and spatial compactness, so a superpixel can contain per-pixel Gaussians from two surfaces at different depths whenever the surfaces have similar appearance. A K=1 decoder cannot split such a bimodal cluster, and the refiner (Sec. 4.5) updates an existing Feature Gaussian but cannot split it into multiple surfaces. The paper does not quantify how often superpixels cross depth discontinuities, nor does it report region-level metrics at depth boundaries. The large LPIPS increases in Table 1 (e.g., DA3 K=1: 0.422 vs. 0.297 on DL3DV-Bench and 0.442 vs. 0.356 on MipNeRF360) are consistent with collapsed depth structure, and the Limitations section does not mention this failure mode. I ask the authors to add an analysis of superpixel depth coherence (e.g., fraction of superpixels with large internal depth variance, or per-region metrics at depth edges) and, if needed, an ablation with depth-aware grouping.
  2. [Sec. 5.2, Tab. 1] The 'largely retaining visual quality' claim is weakened by the LPIPS results. For DA3 K=1, LPIPS rises from 0.297 to 0.422 on DL3DV-Bench and from 0.356 to 0.442 on MipNeRF360, which are relative increases of roughly 40%; SSIM also drops on DL3DV-Bench from 0.620 to 0.578. Since LPIPS is a perceptual metric, the paper should either temper the wording of 'largely retaining visual quality' or provide supporting analysis (per-scene breakdown, error maps, or a perceptual metric that distinguishes acceptable texture loss from structural collapse). The current presentation understates the magnitude of the perceptual degradation.
  3. [Sec. 5.2, Abstract] The comparative claim of being 'better and more robust quality than achieved by previous approaches that target a reduction in primitive count' is under-tested. The experimental comparison includes only ReSplat and VolSplat; Off The Grid (Ref. [25]) and Fuse-and-Refine (Ref. [39]) are cited but not evaluated due to code unavailability, and the graph-based fusion methods FreeSplat and Gaussian Graph Network discussed in Sec. 2 are not compared. The claim should be scoped to the evaluated baselines, or the missing comparisons should be included with a clear justification for their omission.
  4. [Sec. 5.1, Tab. 1] The paper's emphasis on sparse-view performance is not directly supported by the main-text results. Table 1 averages over experiments with 3, 6, 9, and 12 input views, and the detailed per-view-count results are only mentioned as being in the supplementary material. Since the abstract and conclusion specifically highlight robustness 'particularly in sparse-view settings', the per-view-count table (or at least a subset, such as 3 and 6 views) should be presented in the main body, or the sparse-view claim should be softened.
minor comments (5)
  1. [Sec. 5.1] Please state whether the code and trained models will be released; the current text only says that certain baselines are unavailable, which limits reproducibility for reviewers and readers.
  2. [Fig. 1] The legend contains many markers and is hard to read at small size; consider splitting it into separate panels with clearer labels.
  3. [Tab. 3] The phrase 'zero shot larger SPs' should be 'zero-shot larger SPs' for consistency and clarity.
  4. [Sec. 4.2] The term 'Feature Gaussian' is an invented entity; please add a one-sentence summary of how it differs from an ordinary Gaussian in the matching and decoding stages, even though this is partially explained in the text.
  5. [Sec. 4.6] The teacher loss in Equations (3) and (4) is described as a closed-form moment-matching target; please clarify whether it is used only for the K=1 head or also for higher-K decoders during the early training phase.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the pipeline is a learned post-processor trained with photometric losses and evaluated on held-out benchmarks; the primitive-count reduction is an explicit design choice, not a fitted prediction.

full rationale

The derivation is self-contained. Per-pixel Gaussians from a frozen backbone are grouped by BASS superpixel segmentation; each group is encoded into a Feature Gaussian by a learned Set Transformer encoder; a learned merger fuses cross-view Feature Gaussians; a refiner updates them; and a level-of-detail decoder produces output Gaussians. All learned modules are trained end-to-end with photometric losses (MSE, SSIM, LPIPS) on rendered images, and evaluation is on held-out datasets (DL3DV-Bench, MipNeRF360, Tanks & Temples) not used for training. The moment-matching teacher loss (Eqs. 3-4) is explicitly an auxiliary, decaying regularizer ('This teacher signal decays on a schedule, allowing the learned encoder to eventually surpass the moment-matched initialization'), and the ablations show the learned pipeline outperforms the heuristic Mom. Match. baseline, so the final result is not forced to equal the teacher target. The reported reduction to 1/20th of the Gaussians is not a predicted quantity: it follows directly from the chosen superpixel granularity and K decoder head (K=1), i.e., it is a design setting, and the actual empirical claim—the rendering quality at that count—is measured against external benchmarks. The cited components (BASS, Set Transformer, 3DGS) are external methods used as building blocks, not self-citations that carry the argument. The unverified depth-boundary concern is a generalization/correctness risk, not circularity, because no equation or fitted parameter makes the quality outcome equal to an input.

Assumptions & free parameters 6 free parameters · 4 assumptions · 1 invented entities

The central claim rests on a set of hand-chosen hyperparameters for matching and pruning, plus domain assumptions about superpixel coherence in 3D and the sufficiency of the frozen backbone. The only invented entity is the Feature Gaussian, which is an internal latent without independent falsifiability.

free parameters (6)
  • tau_f (feature similarity threshold) = not stated in paper
    Controls which Feature Gaussians are matched across views; hand-chosen threshold likely tuned on validation.
  • tau_g (geometry IoU threshold) = not stated in paper
    Controls required geometric overlap for matching; hand-chosen.
  • tau_a (opacity pruning threshold) = not stated in paper
    Prunes output Gaussians with alpha below threshold at inference.
  • k (kNN neighborhood size) = not stated in paper
    Number of neighbors used in cross-view matching and refinement.
  • superpixel seed density and size settings = large, medium, small tested
    Saliency-guided seed density determines superpixel sizes, directly affecting the compression-quality trade-off.
  • decoder slots K = 1, 2, 4
    Number of output Gaussians per Feature Gaussian; an architectural choice acting as a level-of-detail knob.
assumptions (4)
  • domain assumption Superpixels in image space correspond to 3D-coherent groups of Gaussians
    Pipeline groups per-pixel Gaussians by 2D superpixel segmentation; if a superpixel straddles depth discontinuities, the compressed latent loses structure. This is assumed throughout Section 4.1.
  • domain assumption Shi-Tomasi corner response is a good saliency proxy for where fine Gaussian detail is needed
    Saliency-guided seeding allocates more segments to corners and edges; assumes visual complexity correlates with geometric detail (Section 4.1, Fig. 3).
  • domain assumption The frozen backbone's per-pixel Gaussians are accurate enough that merging only removes redundancy
    The paper states in Limitations that if initial primitives are low quality, merging may not recover a good representation.
  • domain assumption The learned modules can be trained end-to-end with photometric losses despite random initialization
    The teacher loss is introduced specifically because random initialization makes geometry random; assumes the teacher schedule enables stable training (Section 4.6).
invented entities (1)
  • Feature Gaussian (FG)
    purpose: Latent-augmented single Gaussian representing a whole superpixel group, enabling efficient matching, merging, and decoding
    Internal latent representation with no direct external validation; its quality is only measured through the final rendered images. The paper provides no falsifiable handle for the FG itself outside the pipeline.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging." pith.science (2026). https://pith.science/paper/LBULU3Q2

@misc{pith2026260810712,
  author       = {Pith},
  title        = {Pith review of: Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LBULU3Q2}},
  note         = {Machine review of arXiv:2608.10712}
}
abstract

3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this context. Its feed-forward variants provide fast reconstruction from sparse input views but often produce per-pixel primitives, leading to highly redundant and thus inefficient representations. We present a structure-aware merging pipeline that takes per-pixel primitives from any feed-forward method and consolidates them into a compact, content-adaptive Gaussian set while largely retaining visual quality at just $\frac{1}{20}^\text{th}$ of the Gaussians of a per-pixel method. We group spatially coherent Gaussians of similar appearance into variable-size clusters via adaptive superpixel segmentation guided by a saliency map, which allocates fine segments to textured regions and coarse segments to homogeneous areas. We compress each cluster into a compact latent representation through a learned encoder, then match and consolidate representations across views based on geometric overlap and feature similarity via a learned merger. A level-of-detail decoder then produces the final Gaussians at a controllable resolution, enabling a flexible quality-efficiency trade-off at inference. As a post-processing module, the pipeline is backbone-agnostic, leveraging the strengths of existing feed-forward methods. This leads to better and more robust quality than achieved by previous approaches that target a reduction in primitive count, while providing a highly compact representation, that can be rendered efficiently.

Figures

Figures reproduced from arXiv: 2608.10712 by the authors.

Figure 1
Figure 1. NVS quality vs. relative primitive count (DA3 backbone, averaged over DL3DV [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Pipeline overview. Per-pixel Gaussians G FF are grouped into saliency-guided su￾perpixels. The encoder Menc compresses each group into a single Feature Gaussian (FG), the merger Mmrg fuses FGs that overlap across views, the refiner Mref updates each FG based on its neighbors. The decoder Mdec K expands each FG into K output Gaussians, result￾ing in the output scene GK. Feed-Forward Gaussian Splatting. Feed-forward (… view at source ↗
Figure 3
Figure 3. Superpixel comparison at similar segment count: SLIC (uniform), BASS with [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Internal architecture of our modules built on the Set Transformer [ [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Qualitative novel view synthesis comparison at 12 input views (DA3 backbone). [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Gaussian count vs. integrated views in the online setting, averaged over 7 MipN [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Left: Proxy PSNR vs. superpixel count for three segmentation algorithms (100 scenes × 10 views). Each pixel is replaced by its superpixel mean colour and compared to the original image. Right: NVS PSNR vs. relative primitive count for three SP sizes (AnySplat backbone,…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 45 canonical work pages

  1. [25]

    Off The Grid: Detection of Prim- itives for Feed-Forward 3D Gaussian Splatting

    Arthur Moreau, Richard Shaw, Michal Nazarczuk, Jisu Shin, Thomas Tanay, Zhensong Zhang, Songcen Xu, and Eduardo Pérez-Pellitero. Off The Grid: Detection of Prim- itives for Feed-Forward 3D Gaussian Splatting. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2026

  2. [39]

    Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian Splatting

    Yiming Wang, Lucy Chai, Xuan Luo, Michael Niemeyer, Manuel Lagunas, Stephen Lombardi, Siyu Tang, and Tiancheng Sun. Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian Splatting. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2025

  3. [1]

    SLIC Superpixels Compared to State-of-the-Art Superpixel Meth- ods.IEEE Trans

    Radhakrishna Achanta, Appu Shaji, Kevin Smith, Aurelien Lucchi, Pascal Fua, and Sabine Süsstrunk. SLIC Superpixels Compared to State-of-the-Art Superpixel Meth- ods.IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 34(11):2274– 2282, 2012. doi: 10.1109/TPAMI.2012.120

  4. [2]

    CRC Press, 4 edition, 2018

    Tomas Akenine-Möller, Eric Haines, Naty Hoffman, Angelo Pesce, Michaël Iwanicki, and Sébastien Hillaire.Real-Time Rendering. CRC Press, 4 edition, 2018. doi: 10. 1201/b22086

  5. [3]

    MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM

    Yinlong Bai, Hongxin Zhang, Sheng Zhong, Junkai Niu, Hai Li, Yijia He, and Yi Zhou. MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM. InProc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), 2025

  6. [4]

    Barcelos, Felipe D.C

    Isabela B. Barcelos, Felipe D.C. Belem, Leonardo D.M. Joao, Zenilton K.G.D. Pa- trocínio Jr., Alexandre X. Falcao, and Silvio J.F. Guimarães. A Comprehensive Review and New Taxonomy on Superpixel Segmentation.ACM Computing Surveys, 56(8): 1–39, 2024. doi: 10.1145/3643826

  7. [5]

    Barron, Ben Mildenhall, Dor Verbin, Pratul P

    Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hed- man. Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2022

  8. [6]

    pix- elSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Recon- struction

    David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pix- elSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Recon- struction. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

Show all 58 references
  1. [7]

    Fast Feedforward 3D Gaussian Splatting Compression

    Yihang Chen, Qianyi Wu, Mengyao Li, Weiyao Lin, Mehrtash Harandi, and Jianfei Cai. Fast Feedforward 3D Gaussian Splatting Compression. InProc. of the Intl. Conf. on Learning Representations (ICLR), 2025

  2. [8]

    MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

    Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images. InProc. of the Europ. Conf. on Computer Vision (ECCV), 2024. 16FAASCH, KALL, STACHNISS:...

  3. [9]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition ...

  4. [10]

    Mini-Splatting: Representing Scenes with a Con- strained Number of Gaussians

    Guangchi Fang and Bing Wang. Mini-Splatting: Representing Scenes with a Con- strained Number of Gaussians. InProc. of the Europ. Conf. on Computer Vision (ECCV), 2024

  5. [11]

    AnySplat: Feed-Forward 3D Gaussian Splatting from Unconstrained Views.ACM Trans

    Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu, Kerui Ren, Yichen Jin, Xudong Xu, Mulin Yu, Jiangmiao Pang, and Feng Zhao. AnySplat: Feed-Forward 3D Gaussian Splatting from Unconstrained Views.ACM Trans. on Graphics (TOG), 44(6):1–16,

  6. [12]

    Lau, Feng Gao, Yin Yang, and Chenfanfu Jiang

    Ying Jiang, Chang Yu, Tianyi Xie, Xuan Li, Yutao Feng, Huamin Wang, Minchen Li, Henry Y .K. Lau, Feng Gao, Yin Yang, and Chenfanfu Jiang. VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual Reality. InProc. of the Intl. Conf. on Computer Graphics ...

  7. [13]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans. on Graphics (TOG), 42(4):1–14, 2023. doi: 10.1145/3592433

  8. [14]

    A Hierarchical 3D Gaussian Representation for Real- Time Rendering of Very Large Datasets.ACM Trans

    Bernhard Kerbl, Andréas Meuleman, Georgios Kopanas, Michael Wimmer, Alexandre Lanvin, and George Drettakis. A Hierarchical 3D Gaussian Representation for Real- Time Rendering of Very Large Datasets.ACM Trans. on Graphics (TOG), 43(4):1–15,

  9. [15]

    Tanks and Temples: Benchmarking Large-Scale Scene Reconstruction.ACM Trans

    Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and Temples: Benchmarking Large-Scale Scene Reconstruction.ACM Trans. on Graphics (TOG), 36(4):1–13, 2017. doi: 10.1145/3072959.3073599

  10. [16]

    Com- pact 3D Gaussian Representation for Radiance Field

    Joo Chan Lee, Daniel Rho, Xiangyu Sun, Jong Hwan Ko, and Eunbyung Park. Com- pact 3D Gaussian Representation for Radiance Field. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

  11. [17]

    Optimized Minimal 3D Gaussian Splatting

    Joo Chan Lee, Jong Hwan Ko, and Eunbyung Park. Optimized Minimal 3D Gaussian Splatting. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2025

  12. [18]

    Kosiorek, Seungjin Choi, and Yee Whye Teh

    Juho Lee, Yoonho Lee, Jungtaek Kim, Adam R. Kosiorek, Seungjin Choi, and Yee Whye Teh. Set Transformer: A Framework for Attention-based Permutation- Invariant Neural Networks. InProc. of the Intl. Conf. on Machine Learning (ICML), 2019

  13. [19]

    Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang

    Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y . Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth Anything 3: Recovering the Visual Space from Any Views.arXiv preprint, arXiv:2511.10647, 2025. FAASCH, KALL, STACHNISS: COMPACT FF-3DGS VIA SALIENCY -GUIDED MERGING17

  14. [20]

    DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-Based 3D Vision

    Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Anirud- dha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua, Tianyi Zhang, Bedrich Benes, and Aniket Bera. DL3DV-10K: A Large-Scale S...

  15. [21]

    Decoupled Weight Decay Regularization

    Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization. InProc. of the Intl. Conf. on Learning Representations (ICLR), 2019

  16. [22]

    Taming 3DGS: High-Quality Ra- diance Fields with Limited Resources

    Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl, Markus Steinberger, Fran- cisco Vicente Carrasco, and Fernando De La Torre. Taming 3DGS: High-Quality Ra- diance Fields with Limited Resources. InProc. of the Intl. Conf. on Computer Graphics and Interactive Techniques (SI...

  17. [23]

    EV olSplat: Efficient V olume-Based Gaussian Splatting for Urban View Synthesis

    Sheng Miao, Jiaxin Huang, Dongfeng Bai, Xu Yan, Hongyu Zhou, Yue Wang, Bing- bing Liu, Andreas Geiger, and Yiyi Liao. EV olSplat: Efficient V olume-Based Gaussian Splatting for Urban View Synthesis. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2025

  18. [24]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ra- mamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. InProc. of the Europ. Conf. on Computer Vision (ECCV), 2020

  19. [26]

    Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis

    Simon Niedermayr, Josef Stumpfegger, and Rüdiger Westermann. Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

  20. [27]

    Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

    Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, and Bernard Ghanem. Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors. InProc. of the ...

  21. [28]

    SCube: Instant Large-Scale Scene Re- construction Using V oxSplats

    Xuanchi Ren, Yifan Lu, Hanxue Liang, Zhangjie Wu, Huan Ling, Mike Chen, Sanja Fidler, Francis Williams, and Jiahui Huang. SCube: Instant Large-Scale Scene Re- construction Using V oxSplats. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2024

  22. [29]

    Ze- roNVS: Zero-Shot 360-Degree View Synthesis from a Single Image

    Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, and Jiajun Wu. Ze- roNVS: Zero-Shot 360-Degree View Synthesis from a Single Image. InProc. of the IEEE/CVF Conf. on Computer Vision and Pa...

  23. [30]

    PhD thesis, Ruprecht-Karls-Universität Heidelberg, 2000

    Hanno Scharr.Optimale Operatoren in der digitalen Bildverarbeitung. PhD thesis, Ruprecht-Karls-Universität Heidelberg, 2000. 18FAASCH, KALL, STACHNISS: COMPACT FF-3DGS VIA SALIENCY -GUIDED MERGING

  24. [31]

    Normalized Cuts and Image Segmentation.IEEE Trans

    Jianbo Shi and Jitendra Malik. Normalized Cuts and Image Segmentation.IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 22(8):888–905, 2000. doi: 10.1109/34.868688

  25. [32]

    Good Features to Track

    Jianbo Shi and Carlo Tomasi. Good Features to Track. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 1994

  26. [33]

    Splatter Image: Ultra-Fast Single-View 3D Reconstruction

    Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Splatter Image: Ultra-Fast Single-View 3D Reconstruction. InProc. of the IEEE/CVF Conf. on Com- puter Vision and Pattern Recognition (CVPR), 2024

  27. [34]

    NeuRAD: Neural Rendering for Autonomous Driving

    Adam Tonderski, Carl Lindström, Georg Hess, William Ljungbergh, Lennart Svensson, and Christoffer Petersson. NeuRAD: Neural Rendering for Autonomous Driving. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

  28. [35]

    Bayesian Adaptive Superpixel Segmen- tation

    Roy Uziel, Meitar Ronen, and Oren Freifeld. Bayesian Adaptive Superpixel Segmen- tation. InProc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2019

  29. [36]

    Smol-GS: Compact Rep- resentations for Abstract 3D Gaussian Splatting.arXiv preprint, arXiv:2512.00850, 2025

    Haishan Wang, Mohammad Hassan Vali, and Arno Solin. Smol-GS: Compact Rep- resentations for Abstract 3D Gaussian Splatting.arXiv preprint, arXiv:2512.00850, 2025

  30. [37]

    Chen, Zeyu Zhang, Duochao Shi, Akide Liu, and Bohan Zhuang

    Weijie Wang, Donny Y . Chen, Zeyu Zhang, Duochao Shi, Akide Liu, and Bohan Zhuang. ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2025

  31. [38]

    Chen, and Bohan Zhuang

    Weijie Wang, Yeqing Chen, Zeyu Zhang, Hengyu Liu, Haoxiao Wang, Zhiyuan Feng, Wenkang Qin, Zheng Zhu, Donny Y . Chen, and Bohan Zhuang. V olSplat: Rethinking Feed-Forward 3D Gaussian Splatting with V oxel-Aligned Prediction.arXiv preprint, arXiv:2509.19297, 2025

  32. [40]

    FreeSplat: Generaliz- able 3D Gaussian Splatting Towards Free-View Synthesis of Indoor Scenes

    Yunsong Wang, Tianxin Huang, Hanlin Chen, and Gim Hee Lee. FreeSplat: Generaliz- able 3D Gaussian Splatting Towards Free-View Synthesis of Indoor Scenes. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2024

  33. [41]

    Bovik, Hamid R

    Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image Quality Assessment: From Error Visibility to Structural Similarity.IEEE Trans. on Image Processing, 13(4):600–612, 2004. doi: 10.1109/TIP.2003.819861

  34. [42]

    PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics

    Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chenfanfu Jiang. PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

  35. [43]

    ReSplat: Learning Recurrent Gaussian Splatting.arXiv preprint, arXiv:2510.08575, 2025

    Haofei Xu, Daniel Barath, Andreas Geiger, and Marc Pollefeys. ReSplat: Learning Recurrent Gaussian Splatting.arXiv preprint, arXiv:2510.08575, 2025. FAASCH, KALL, STACHNISS: COMPACT FF-3DGS VIA SALIENCY -GUIDED MERGING19

  36. [44]

    DepthSplat: Connecting Gaussian Splatting and Depth

    Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger, and Marc Pollefeys. DepthSplat: Connecting Gaussian Splatting and Depth. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2025

  37. [45]

    GRM: Large Gaussian Reconstruction Model for Effi- cient 3D Reconstruction and Generation

    Yinghao Xu, Zifan Shi, Wang Yifan, Hansheng Chen, Ceyuan Yang, Sida Peng, Yujun Shen, and Gordon Wetzstein. GRM: Large Gaussian Reconstruction Model for Effi- cient 3D Reconstruction and Generation. InProc. of the Europ. Conf. on Computer Vision (ECCV), 2024

  38. [46]

    GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting

    Chi Yan, Delin Qu, Dan Xu, Bin Zhao, Zhigang Wang, Dong Wang, and Xuelong Li. GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

  39. [47]

    No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images.arXiv preprint, arXiv:2410.24207, 2024

    Botao Ye, Sifei Liu, Haofei Xu, Xueting Li, Marc Pollefeys, Ming-Hsuan Yang, and Songyou Peng. No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images.arXiv preprint, arXiv:2410.24207, 2024

  40. [48]

    GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting

    Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting. In Proc. of the Europ. Conf. on Computer Vision (ECCV), 2024

  41. [49]

    Efros, Eli Shechtman, and Oliver Wang

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2018

  42. [50]

    Gaussian Graph Network: Learning Efficient and Generalizable Gaussian Representations from Multi-view Images

    Shengjun Zhang, Xin Fei, Fangfu Liu, Haixu Song, and Yueqi Duan. Gaussian Graph Network: Learning Efficient and Generalizable Gaussian Representations from Multi-view Images. InProc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2024

  43. [51]

    Superpixels via Pseudo-Boolean Optimization

    Yuhang Zhang, Richard Hartley, John Mashford, and Stewart Burn. Superpixels via Pseudo-Boolean Optimization. InProc. of the IEEE/CVF Intl. Conf. on Computer Vision (ICCV), 2011

  44. [52]

    Content-Aware Dynamic Superpixel Segmentation

    Tingyu Zhao, Bo Peng, Zhenguang Zhang, Daipeng Yang, and Xi Wu. Content-Aware Dynamic Superpixel Segmentation. InProc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025

  45. [53]

    GaussianGrasper: 3D Language Gaus- sian Splatting for Open-V ocabulary Robotic Grasping.IEEE Robotics and Automation Letters (RA-L), 9(9):7827–7834, 2024

    Yuhang Zheng, Xiangyu Chen, Yupeng Zheng, Songen Gu, Runyi Yang, Bu Jin, Pengfei Li, Chengliang Zhong, Zengmao Wang, Lina Liu, Chao Yang, Dawei Wang, Zhen Chen, Xiaoxiao Long, and Meiqing Wang. GaussianGrasper: 3D Language Gaus- sian Splatting for Open-V ocabulary Robotic Gras...

  46. [54]

    Stereo Magnification: Learning View Synthesis Using Multiplane Images.ACM Trans

    Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo Magnification: Learning View Synthesis Using Multiplane Images.ACM Trans. on Graphics (TOG), 37(4):1–12, 2018. doi: 10.1145/3197517.3201323

  47. [55]

    DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Au- tonomous Driving Scenes

    Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Au- tonomous Driving Scenes. InProc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2024. 20FAA...

  48. [56]

    EW A V olume Splatting

    Matthias Zwicker, Hanspeter Pfister, Jeroen van Baar, and Markus Gross. EW A V olume Splatting. InProc. of the IEEE Visualization Conf. (VIS), 2001

  49. [2024]

    doi: 10.1145/3658160

  50. [2025]

    doi: 10.1145/3763326

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.