REVIEW 2 major objections 8 minor 52 references
GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing
T0 review · 2 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A 28nm 3D Gaussian Splatting accelerator called GCC reports a 5.24x area-normalized speedup and 3.35x energy-efficiency gain over GSCore by rendering Gaussian-wise and skipping unnecessary preprocessing.
desk verdict Genuinely new 3DGS dataflow with an honest GPU comparison, but the headline speedup rests on an unaudited GSCore baseline simulator. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the four-stage Gaussian-wise pipeline, whose scheduling unit is the depth group: Stage I bins Gaussians by view-space depth, Stage II projects the mean and reconstructs a 2D inverse covariance matrix, Stage III evaluates spherical-harmonic colors and sorts within each group, and Stage IV computes $\alpha$ and blends front-to-back. Two mechanisms carry the efficiency story. Cross-stage conditional processing interleaves these stages and, because blending tracks accumulated transmittance $T$, lets the controller skip every deeper group once $T$ passes the early-termination threshold, so preprocessing and rendering are never performed for Gaussians that cannot affect the image. Gaussian-wise rendering loads each Gaussian's 59 floating-point parameters exactly once per frame instead of once per overlapped tile. A third mechanism, the $\alpha$-based Gaussian boundary identifier, uses a breadth-first traversal from the projected center and exploits the convexity of the Gaussian's elliptical footprint to mark whole pixel blocks as pruned, giving a tighter per-Gaussian region than axis-aligned or oriented bounding boxes. The named identity is the $\omega$-$\sigma$ law, $r = \lceil \sqrt{2\ln(255\omega)\max(\lambda_1,\lambda_2)} \rceil$, which replaces the fixed $3\sigma$ envelope with an opacity-aware radius.
What would settle it
A decisive check is to re-run the six-scene benchmark with GCC's released RTL or cycle-accurate simulator and an independently built GSCore simulator using the same LPDDR4-3200 memory model: if GCC does not reach 667 FPS on Lego, or if the GSCore reproduction deviates more than 3% from the paper's reported cycles, the central speedup claim is falsified.
Extended reading notes
Core claim
The paper's central claim is that the standard decoupled preprocess-then-render, tile-by-tile dataflow is the root cause of wasted work in 3DGS accelerators, and that reordering work around individual Gaussians removes most of it. GCC computes each Gaussian's view-space depth, groups Gaussians into depth-ordered bins, and then, for each Gaussian in that order, performs projection, spherical-harmonic color evaluation, $\alpha$ computation, and blending before moving on; accumulated transmittance lets later groups be skipped entirely. To keep per-Gaussian work small, the $\alpha$-based boundary identifier visits only pixels where a Gaussian's $\alpha$ contribution exceeds the perceptual threshold $1/255$, starting from the projected center and pruning whole directions once the elliptical boundary fails. The authors report that across six indoor, outdoor, and synthetic scenes, this design reaches a geometric-mean 5.24x area-normalized speedup and a 3.35x area-normalized energy-efficiency improvement over GSCore, with peak throughput of 667 FPS on Lego, PSNR within 0.1 dB of the GPU reference, and a 2.71 mm$^2$, 790 mW implementation in 28nm.
Load-bearing premise
The claim rests on the assumption that the authors' simulator faithfully reproduces both their own chip design and the GSCore accelerator it is compared with, within the reported 3% deviation of the baseline; the 5.24x and 3.35x numbers are computed against that reproduction, and no code or raw measurements are released to check it independently.
Editorial extensions
If this is right
- Across the six benchmark scenes, GCC reports a geometric-mean 5.24x area-normalized speedup and 3.35x area-normalized energy efficiency over GSCore, with per-scene speedups from 4.27x on Playroom to 6.22x on Lego.
- DRAM traffic falls by more than 50%: Gaussian-wise loading eliminates repeated fetches and conditional processing stops useless Gaussians before they are read, and above about 220 GB/s of DRAM bandwidth GCC becomes compute-bound while GSCore stays memory-bound.
- Rendering quality is preserved: PSNR stays within 0.1 dB of the GPU reference and LPIPS is identical to GSCore's, so the pruning decisions do not cost noticeable fidelity.
- A 128x128 sub-view compatibility mode lets large scenes render with negligible redundant Gaussian processing, which is what allows the design to operate with 190 KB of on-chip SRAM on edge-class platforms.
- Ablations attribute the gains to both mechanisms: Gaussian-wise rendering dominates on compact scenes like Palace, while cross-stage conditional processing contributes more on large, sparse scenes like Drjohnson.
Reading between the lines
- My inference: the 5.24x and 3.35x ratios measure GCC against a reproduced GSCore simulator, not against GSCore silicon; a third-party re-implementation of the baseline is the decisive external check, because a pessimistic baseline model would inflate the ratios.
- My inference: the paper's own GPU experiments show the GCC dataflow slows down on GPUs because deterministic blending requires costly atomic updates; that makes the dataflow's benefit specific to architectures with small on-chip storage and ordered pipelines rather than a general-purpose algorithmic improvement.
- My inference: the alpha-based boundary identifier and the opacity-aware radius are transferable ideas for any opacity-weighted splatting or particle renderer, though the gains will depend on how elliptical and how transparent the primitives are.
- My inference: because GCC is already compute-bound past about 220 GB/s of DRAM bandwidth, pairing it with faster LPDDR5-class memory would buy little; further gains would have to come from cheaper spherical-harmonic evaluation or from compressing the 48 SH coefficients that dominate each Gaussian's footprint.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GCC, a 28nm 3D Gaussian Splatting inference accelerator built around two dataflow innovations: Gaussian-wise rendering, which processes each Gaussian completely before moving to the next, and cross-stage conditional processing, which interleaves preprocessing and rendering so that Gaussians that would be killed by alpha-blending early termination are never projected or color-evaluated. The authors additionally introduce an alpha-based boundary identification method that replaces static 3-sigma bounding boxes with opacity-aware elliptical footprints. The claims are supported by a SystemVerilog RTL implementation synthesized at 1GHz, a cycle-accurate Python simulator, and a re-implemented GSCore baseline simulator. The headline results are an average 5.24x area-normalized speedup and 3.35x area-normalized energy efficiency over GSCore across six scenes, with PSNR within 0.1 dB and identical LPIPS, plus a peak throughput of 667 FPS on the Lego scene.
Significance. If the results are reproducible, this is a significant architecture contribution. The Gaussian-wise dataflow is a principled departure from the tile-centric pipeline used by prior 3DGS accelerators, and the paper gives a detailed module-level design, area/power breakdown, ablation study, and a GPU implementation study showing why the dataflow does not translate to GPUs. The quality-preservation check (PSNR within 0.1 dB, identical LPIPS in Table 2) is appropriate and directly addresses the main accuracy risk. The paper also provides sensitivity analysis for image buffer size, PE array size, and DRAM bandwidth. The main weakness is verification: the comparative speedup and energy numbers rest on two simulators whose fidelity is asserted but not demonstrated in the paper, and no artifacts are released. The mathematical derivation for the opacity-aware footprint also contains an overclaim that should be corrected.
major comments (2)
- [§5.1, §5.2] The central performance claims (5.24x speedup and 3.35x energy efficiency in Figure 10) are computed against a GSCore simulator whose fidelity is justified only by the sentence 'we also develop a simulator for GSCore based on the architectural details provided in its paper with less than 3% performance deviation.' No per-scene comparison to GSCore's published cycles or energy, no DRAM/bank/burst modeling details, and no artifacts are provided. Since the two designs differ in dataflow, memory scheduling, and buffer arbitration, small modeling decisions can change a speedup ratio by a factor, not by a few percent. The same applies to the GCC simulator, whose 'validated with the HDL design at the cycle level' claim is not accompanied by any per-module validation table. Please provide a detailed validation appendix: per-module cycle-count comparisons against RTL, per-scene GSCore published-vs-reproduced cycle and energy numbers, and explicit DRAM and memory-system modeling assumptions; releasing the simulators or a reproducible artifact would resolve this concern most directly.
- [§3 Stage II, Eq. (8)] Equation (8) is not universally tighter than the 3-sigma envelope of Eq. (6). For omega = 1, the radius is sqrt(2 ln 255) * sqrt(lambda_max) ≈ 3.33 * sqrt(lambda_max), which is larger than 3 * sqrt(lambda_max). The opacity-aware radius is tighter only when 2 ln(255 omega) ≤ 9, i.e., omega ≤ e^{4.5}/255 ≈ 0.353. The text states without qualification that 'compared to the static 3σ rule, this opacity-aware culling removes more redundant Gaussians' and calls Eq. (8) a 'tighter radius estimate.' This overclaim affects the interpretation of the pixel-count reductions in Table 1. Please qualify the statement with the opacity regime, or report the opacity distribution and show that the average/rendered-pixel count decreases; otherwise the reduction cannot be attributed to Eq. (8) for all Gaussians.
minor comments (8)
- [§3 Stage IV] The text says 'once the transmittance surpasses a predefined threshold, subsequent Gaussians are skipped'; since transmittance decreases as Gaussians are composited, this should read 'falls below' the threshold.
- [Algorithm 1] The predicate E(p) used in Algorithm 1 is never defined; state explicitly that E(p) holds when the alpha value from Eq. (9) is at least 1/255.
- [§4.4] The claim that EXP inputs above 0 are saturated to alpha = 1 conflicts with the min(0.99, ...) cap in Eq. (9): for omega = 1 and d = 0 the exponent is exactly 0, and the correct alpha is 0.99, not 1. Clarify the LUT boundary behavior and quantify the resulting error.
- [Figure 13] The axis labels render as 'FPS/mm/uni00B2' and 'mJ/mm/uni00B2'; the superscripts need to be fixed.
- [§2.2, Challenge 1] The phrase '81.4% (48 out of 59) of the SH coefficients remain unused before alpha-blending begins' is confusing, since SH coefficients are used to compute RGB before blending; what is meant is that these 48 parameters are unnecessarily loaded and processed for Gaussians that are later discarded by early termination.
- [§4.6 and Abstract] The abstract states that Gaussian-wise rendering eliminates duplicated Gaussian loading, but in Compatibility Mode Gaussians overlapping sub-view boundaries are processed more than once; Section 4.6 acknowledges this, so the abstract/contribution wording should include the caveat.
- [Table 3 and §5.2] The throughput numbers, including the 667 FPS peak on Lego, should state the image resolution and camera configuration; otherwise the comparison with GSCore's Lego throughput is not reproducible.
- [§5.1] The phrase 'same configuration settings as reported' should enumerate the actual settings (DRAM frequency, tile size, data widths, clock, and buffer sizes) used for both simulators.
Circularity Check
No significant circularity: the performance and energy-efficiency claims are simulation outputs evaluated against an external baseline, not results assumed by the model; the omega-sigma law and alpha-boundary identification are algebraic and algorithmic consequences of the externally defined 3DGS alpha threshold.
full rationale
The paper's central claims are architectural and empirical rather than derived from a model that assumes its own conclusions. The 5.24x area-normalized speedup and 3.35x energy-efficiency figures come from a cycle-accurate Python simulator and synthesized RTL, compared against a re-implemented GSCore baseline; nothing in the paper fits a parameter to the reported speedup and then presents that same speedup as a prediction. The omega-sigma law (Equations 7 and 8) is a direct rearrangement of the externally defined alpha formula (Equation 3) together with the 1/255 alpha cutoff inherited from the original 3DGS algorithm, so it is an algebraic consequence rather than a self-referential definition. The alpha-based Gaussian Boundary Identification uses the same externally given threshold to define a Gaussian's effective pixel region; the Table 1 comparison to AABB and OBB regions is an empirical measurement of scene statistics, not a tautology. The paper includes a same-group citation (GSArch, reference [11]), but it appears only in related work and does not carry any load-bearing argument. The main weaknesses are reproducibility concerns: the assertions that the simulators were validated at the cycle level with the HDL design and that the GSCore reproduction has less than 3% performance deviation are not backed by released artifacts, and Section 6 concedes that a GPU implementation of the GCC dataflow reaches only 6-20 FPS on Jetson AGX Xavier. These are correctness and verifiability limitations, not circular reasoning. Because no derived quantity reduces by construction to an input or to a self-citation chain, the circularity score is minimal.
Assumptions & free parameters
free parameters (5)
- Z-axis frustum pivot =
0.2
- Depth subgroup size N =
256
- Compatibility sub-view size =
128x128
- Alpha/blending PE array dimension n =
8
- EXP LUT segment count =
16
assumptions (4)
- domain assumption 3D Gaussian Splatting rendering equations and their numerical thresholds (alpha cutoff 1/255, transmittance threshold 0.0001)
- standard math Convexity of the alpha superlevel set for a single Gaussian
- ad hoc to paper The cycle-accurate simulator matches the RTL at cycle level for every module
- ad hoc to paper The GSCore baseline simulator reproduces GSCore's reported performance within 3%
Cite this review
Pith. "Pith review of GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing." pith.science (2026). https://pith.science/paper/WWYOMM35
@misc{pith2026250715300,
author = {Pith},
title = {Pith review of: GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing},
year = {2026},
howpublished = {\url{https://pith.science/paper/WWYOMM35}},
note = {Machine review of arXiv:2507.15300}
}
read the original abstract
3D Gaussian Splatting (3DGS) has emerged as a leading neural rendering technique for high-fidelity view synthesis, prompting the development of dedicated 3DGS accelerators for resource-constrained platforms. The conventional decoupled preprocessing-rendering dataflow in existing accelerators has two major limitations: 1) a significant portion of preprocessed Gaussians are not used in rendering, and 2) the same Gaussian gets repeatedly loaded across different tile renderings, resulting in substantial computational and data movement overhead. To address these issues, we propose GCC, a novel accelerator designed for fast and energy-efficient 3DGS inference. GCC introduces a novel dataflow featuring: 1) \textit{cross-stage conditional processing}, which interleaves preprocessing and rendering to dynamically skip unnecessary Gaussian preprocessing; and 2) \textit{Gaussian-wise rendering}, ensuring that all rendering operations for a given Gaussian are completed before moving to the next, thereby eliminating duplicated Gaussian loading. We also propose an alpha-based boundary identification method to derive compact and accurate Gaussian regions, thereby reducing rendering costs. We implement our GCC accelerator in 28nm technology. Extensive experiments demonstrate that GCC significantly outperforms the state-of-the-art 3DGS inference accelerator, GSCore, in both performance and energy efficiency.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Yanqi Bao, Tianyu Ding, Jing Huo, Yaoli Liu, Yuxin Li, Wenbin Li, Yang Gao, and Jiebo Luo. 2025. 3d gaussian splatting: Survey, technologies, challenges, and opportunities. IEEE Transactions on Circuits and Systems for Video Technology (2025)
work page 2025
-
[2]
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. 2022. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5470–5479
work page 2022
-
[3]
Zhiqin Chen, Thomas Funkhouser, Peter Hedman, and Andrea Tagliasacchi. 2023. Mobilenerf: Exploiting the polygon rasterization pipeline for efficient neural field rendering on mobile architectures. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16569–16578
work page 2023
-
[4]
Anurag Dalal, Daniel Hagen, Kjell G Robbersmyr, and Kristian Muri Knausgård
-
[5]
Ben Fei, Jingyi Xu, Rui Zhang, Qingyuan Zhou, Weidong Yang, and Ying He. 2024. 3d gaussian splatting as new era: A survey. IEEE Transactions on Visualization and Computer Graphics (2024)
work page 2024
-
[6]
Guofeng Feng, Siyan Chen, Rong Fu, Zimu Liao, Yi Wang, Tao Liu, Boni Hu, Linning Xu, Zhilin Pei, Hengjie Li, et al . 2025. Flashgs: Efficient 3d gaussian splatting for large-scale and high-resolution rendering. In Proceedings of the Computer Vision and Pattern Recognition Conference . 26652–26662
work page 2025
-
[7]
Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. 2022. Plenoxels: Radiance fields without neural net- works. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5501–5510
work page 2022
-
[8]
Andrew S Glassner. 1989. An introduction to ray tracing . Morgan Kaufmann
work page 1989
Show all 52 references
-
[9]
Donghyeon Han, Junha Ryu, Sangyeob Kim, Sangjin Kim, Jongjun Park, and Hoi-Jun Yoo. 2023. MetaVRain: A mobile neural 3-D rendering processor with bundle-frame-familiarity-based NeRF acceleration and hybrid DNN computing. IEEE Journal of Solid-State Circuits 59, 1 (2023), 65–78
2023
-
[10]
Alex Hanson, Allen Tu, Geng Lin, Vasu Singla, Matthias Zwicker, and Tom Goldstein. 2025. Speedy-splat: Fast 3d gaussian splatting with sparse pixels and sparse primitives. In Proceedings of the Computer Vision and Pattern Recognition Conference. 21537–21546
2025
-
[11]
Houshu He, Gang Li, Fangxin Liu, Li Jiang, Xiaoyao Liang, and Zhuoran Song
-
[12]
Peter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, and Gabriel Brostow. 2018. Deep blending for free-viewpoint image-based ren- dering. ACM Transactions on Graphics (ToG) 37, 6 (2018), 1–15
2018
-
[13]
Xiaotong Huang, He Zhu, Zihan Liu, Weikai Lin, Xiaohong Liu, Zhezhi He, Jingwen Leng, Minyi Guo, and Yu Feng. 2025. Seele: A unified acceleration framework for real-time gaussian splatting. arXiv preprint arXiv:2503.05168 (2025)
2025
-
[14]
Ying Jiang, Chang Yu, Tianyi Xie, Xuan Li, Yutao Feng, Huamin Wang, Minchen Li, Henry Lau, Feng Gao, Yin Yang, et al. 2024. Vr-gs: A physical dynamics-aware interactive gaussian splatting system in virtual reality. In ACM SIGGRAPH 2024 Conference Papers. 1–1
2024
-
[15]
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis
-
[16]
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. 2017. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (ToG) 36, 4 (2017), 1–13
2017
-
[17]
Junseo Lee, Kwanseok Choi, Jungi Lee, Seokwon Lee, Joonho Whangbo, and Jaewoong Sim. 2023. Neurex: A case for neural rendering acceleration. In Pro- ceedings of the 50th Annual International Symposium on Computer Architecture . 1–13
2023
-
[18]
Junseo Lee, Jaisung Kim, Junyong Park, and Jaewoong Sim. 2025. VR-Pipe: Streamlining Hardware Graphics Pipeline for Volume Rendering. arXiv preprint arXiv:2502.17078 (2025)
2025 arXiv
-
[19]
Junseo Lee, Seokwon Lee, Jungi Lee, Junyong Park, and Jaewoong Sim. 2024. Gscore: Efficient radiance field rendering via architectural support for 3d gaussian splatting. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages ...
2024
-
[20]
Marc Levoy. 1988. Display of surfaces from volume data. IEEE Computer graphics and Applications 8, 3 (1988), 29–37
1988
-
[21]
Chaojian Li, Sixu Li, Yang Zhao, Wenbo Zhu, and Yingyan Lin. 2022. Rt-nerf: Real-time on-device neural radiance fields towards immersive ar/vr rendering. In Proceedings of the 41st IEEE/ACM International Conference on Computer-Aided Design. 1–9
2022
-
[22]
Brockman, and Norman P
Sheng Li, Ke Chen, Jung Ho Ahn, Jay B. Brockman, and Norman P. Jouppi
-
[23]
Sixu Li, Chaojian Li, Wenbo Zhu, Boyang Yu, Yang Zhao, Cheng Wan, Haoran You, Huihong Shi, and Yingyan Lin. 2023. Instant-3d: Instant neural radiance field training towards on-device ar/vr 3d reconstruction. In Proceedings of the 50th Annual International Symposium on Computer...
2023
-
[24]
Sixu Li, Yang Zhao, Chaojian Li, Bowei Guo, Jingqun Zhang, Wenbo Zhu, Zhifan Ye, Cheng Wan, and Yingyan Celine Lin. 2024. Fusion-3D: Integrated Acceleration for Instant 3D Reconstruction and Real-Time Rendering. In 2024 57th IEEE/ACM International Symposium on Microarchitectur...
2024
-
[25]
Stefan Mach, Fabian Schuiki, Florian Zaruba, and Luca Benini. 2020. FPnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 29, 4 (2020), 774–787
2020
-
[26]
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 1 (2021), 99–106
2021
-
[27]
Nicolas Moenne-Loccoz, Ashkan Mirzaei, Or Perel, Riccardo de Lutio, Janick Martinez Esturo, Gavriel State, Sanja Fidler, Nicholas Sharp, and Zan Gojcic. 2024. 3D Gaussian Ray Tracing: Fast Tracing of Particle Scenes. ACM Transactions on Graphics (TOG) 43, 6 (2024), 1–19
2024
-
[28]
Muhammad Husnain Mubarik, Ramakrishna Kanungo, Tobias Zirr, and Rakesh Kumar. 2023. Hardware acceleration of neural graphics. In Proceedings of the 50th Annual International Symposium on Computer Architecture . 1–12
2023
-
[29]
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. In- stant neural graphics primitives with a multiresolution hash encoding. ACM transactions on graphics (TOG) 41, 4 (2022), 1–15
2022
-
[30]
NVIDIA Corporation. 2018. NVIDIA Jetson Xavier. https://www.nvidia.com/en- us/autonomous-machines/embedded-systems/jetson-xavier-series/
2018
-
[31]
NVIDIA Corporation. 2020. GeForce RTX 3090 Family. https://www.nvidia.com/ en-us/geforce/graphics-cards/30-series/rtx-3090/
2020
-
[32]
NVIDIA Corporation. 2020. NVIDIA RTX A6000 Graphics Card. https://www. nvidia.com/en-us/products/workstations/rtx-a6000/
2020
-
[33]
Gwangtae Park, Seokchan Song, Haoyang Sang, Dongseok Im, Donghyeon Han, Sangyeob Kim, Hongseok Lee, and Hoi-Jun Yoo. 2024. 20.8 Space-Mate: A 303.5 mW Real-Time Sparse Mixture-of-Experts-Based NeRF-SLAM Processor for Mo- bile Spatial Computing. In 2024 IEEE International Solid...
2024
-
[34]
Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, and Hoi-Jun Yoo. 2024. 20.7 NeuGPU: A 18.5 mJ/Iter neural-graphics processing unit for instant-modeling and real-time rendering with segmented-hashing arch...
2024
-
[35]
Kotaro Shimamura, Ayumi Ohno, and Shinya Takamaeda-Yamazaki. 2025. Ex- ploring the Versal AI Engine for 3D Gaussian Splatting. arXiv preprint arXiv:2502.11782 (2025)
2025 arXiv
-
[36]
Peter-Pike Sloan, Jan Kautz, and John Snyder. 2023. Precomputed radiance transfer for real-time rendering in dynamic, low-frequency lighting environments. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2 . 339–348
2023
-
[37]
Xinzhe Wang, Ran Yi, and Lizhuang Ma. 2024. AdR-Gaussian: Accelerating Gaussian Splatting with Adaptive Radius. In SIGGRAPH Asia 2024 Conference Papers. 1–10
2024
-
[38]
Xiaobao Wei, Peng Chen, Ming Lu, Hui Chen, and Feng Tian. 2025. Graphavatar: Compact head avatars with gnn-generated 3d gaussians. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 8295–8303
2025
-
[39]
Xiaobao Wei, Qingpo Wuwu, Zhongyu Zhao, Zhuangzhe Wu, Nan Huang, Ming Lu, Ningning Ma, and Shanghang Zhang. 2024. Emd: Explicit motion modeling for high-quality street gaussian splatting. arXiv preprint arXiv:2411.15582 (2024)
2024 arXiv
-
[40]
Lizhou Wu, Haozhe Zhu, Siqi He, Jiapei Zheng, Chixiao Chen, and Xiaoyang Zeng. 2024. GauSPU: 3D Gaussian Splatting Processor for Real-Time SLAM Systems. In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1562–1573
2024
-
[41]
Tong Wu, Yu-Jie Yuan, Ling-Xiao Zhang, Jie Yang, Yan-Pei Cao, Ling-Qi Yan, and Lin Gao. 2024. Recent advances in 3d gaussian splatting. Computational Visual Media 10, 4 (2024), 613–642
2024
-
[42]
Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. 2022. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5438–5448
2022
-
[43]
Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa
-
[44]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang
-
[45]
Yang Katie Zhao, Shang Wu, Jingqun Zhang, Sixu Li, Chaojian Li, and Yingyan Ce- line Lin. 2023. Instant-nerf: Instant on-device neural radiance field training via algorithm-accelerator co-designed near-memory processing. In 2023 60th ACM/IEEE Design Automation Conference (DAC)...
2023
-
[46]
Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. 2002. EWA splatting. IEEE Transactions on Visualization and Computer Graphics 8, 3 (2002), 223–238. 14
2002
-
[2011]
In ICCAD: International Conference on Computer-Aided Design
CACTI-P: Architecture-level modeling for SRAM-based structures with advanced leakage reduction techniques. In ICCAD: International Conference on Computer-Aided Design. 694–701
-
[2018]
In Proceedings of the IEEE conference on computer vision and pattern recognition
The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition . 586–595. 13 MICRO 2025, October 18–22, 2025, Seoul, Korea Minnan Pei, et al
2025
-
[2021]
In Proceedings of the IEEE/CVF international conference on computer vision
Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE/CVF international conference on computer vision . 5752–5761
-
[2023]
ACM Trans
3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph. 42, 4 (2023), 139–1
2023
-
[2024]
IEEE Access (2024)
Gaussian splatting: 3d reconstruction and novel view synthesis, a review. IEEE Access (2024)
2024
-
[2025]
In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA)
GSArch: Breaking Memory Barriers in 3D Gaussian Splatting Training via Architectural Support. In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). 366–379. https://doi.org/10.1109/HPCA61900. 2025.00037
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.