REVIEW 4 major objections 6 minor 57 references
A learned budget-aware policy can pick far fewer Gaussian anchors and still beat fixed FPS in 4D streaming.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 23:16 UTC pith:QOKYYXSB
load-bearing objection The arXiv title/abstract claim a null sampler result; the body claims EGS@256 beats IGS@8192 by ~0.5 dB and ~1.3× speed—until that is reconciled the central number is not checkable. the 4 major comments →
Does it matter which Gaussians you pick in 4D Gaussian streaming?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that Efficient Gaussian Streaming (EGS), a plug-in reinforcement-learned sampler that jointly selects an anchor budget and an informative subset, improves the quality–efficiency trade-off of anchor-driven 4D Gaussian streaming relative to fixed farthest-point sampling at 8192 anchors, including roughly half a decibel PSNR gains and about 1.3× lower latency at 256 anchors on unseen data.
What carries the argument
The budgeted anchor action (budget κ plus same-size subset Ω): candidates are scored by a point-MLP plus small Transformer, then trained with a short FPS-imitation warm-start followed by a contextual-bandit reward that penalizes excess anchors, excess time, and PSNR shortfall while rewarding quality gains over FPS teacher targets.
Load-bearing premise
The claim depends on a frame-wise reward built from farthest-point-sampling teacher targets, hand-tuned weights, and a short imitation warm-start actually producing a general rule for informative anchors rather than a fit to the training interface and scenes.
What would settle it
Freeze the released checkpoint and, under the same fast-rendering protocol with no online refinement, measure whether 256-anchor EGS still beats FPS at 8192 in both PSNR and time per frame on a held-out multi-view dynamic dataset outside the paper’s N3DV and MeetingRoom splits; a clear loss on quality or speed falsifies the claimed general trade-off.
If this is right
- The default 8192-anchor FPS budget is over-provisioned for many scenes if selection is content-aware.
- Anchor selection can be swapped as a plug-in without retraining the streaming reconstruction backbone.
- Same-budget ablations imply gains come from subset quality, not only from shrinking the budget.
- Learned selection pays off most in the low-budget regime where sampling overhead is offset by fewer anchors.
- A shared budget–sampler sweep can turn “which Gaussians to pick” into a comparable subproblem rather than a fixed default.
Where Pith is reading between the lines
- If refinement dominates runtime, further latency wins may come more from lighter refinement than from smarter anchor choice alone.
- The same budgeted set-selection pattern may transfer to other online pipelines that drive a scene from sparse control points.
- Tying the reward to FPS teachers may cap how far the policy can move beyond spatial-coverage heuristics.
- Releasing the full sampler sweep would let later work treat random, uniform, and learned policies under one protocol.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents Efficient Gaussian Streaming (EGS), a plug-in budget-aware anchor sampler for IGS-style 4D Gaussian streaming. A point-MLP + Transformer policy jointly chooses a discrete budget κ and a subset of anchors; training uses one-epoch FPS imitation (SFT) followed by a frame-wise contextual-bandit REINFORCE objective whose reward trades sparsity, runtime, and PSNR against FPS teacher targets (Eqs. 3–5). With the IGS backbone frozen, the paper reports that low-budget EGS (256–1024 anchors) improves PSNR by roughly +0.5 dB and is 1.2–1.35× faster than IGS@8192 on held-out N3DV and MeetingRoom scenes in fast mode, with same-budget RL-vs-FPS wins in 14/15 pairs (Tab. 5) and competitive HQ refinement at reduced budgets (Tab. 6).
Significance. If the positive low-budget results hold under a single, reproducible protocol, the work is a useful systems contribution: it shows that anchor-driven 4DGS streaming can operate far below the conventional 8192-FPS budget without quality collapse, and that subset choice (not only κ) can matter under tight constraints. The frozen-backbone plug-in design, same-budget FPS ablation, dual-protocol (fast/HQ) reporting, and promised code/checkpoints are concrete strengths. The absolute quality gains are modest, and the method is tightly coupled to the IGS candidate-pool interface, so impact is primarily engineering/protocol rather than a new representation principle.
major comments (4)
- Title/abstract vs body claim conflict: the arXiv title and front-matter abstract assert a largely null result (sampler choice has no measurable effect at deployment budgets; random/uniform@4096 matches FPS@8192; learned policy is mixed and not a stable cross-dataset rule), while the body abstract, Fig. 2, and Tabs. 1–6 assert the opposite strongest claim (EGS@256 beats IGS@8192 by +0.52–0.61 dB and 1.29–1.35× speed on unseen data). These cannot both be true under one protocol. The manuscript must present a single reconciled claim set, matching experiments (including the random/uniform/opacity-scale sweep named only in the null abstract), and consistent numbers before the central result is checkable.
- §4 / Tabs. 1–5: all headline ΔPSNR and speedup figures are point estimates with no multi-seed means, standard errors, or paired significance tests. Given reported gains of ~0.4–0.6 dB and the abstract’s own language of “within measurement error,” error bars (or at least 3-seed replications of the frozen-checkpoint evaluation) are load-bearing for deciding whether the sampler effect is real or noise—especially for the same-budget 14/15 win claim in Tab. 5.
- §3.3, Eqs. (3)–(4): the reward and SFT warm-start are deliberately FPS-anchored (ψ_tgt mixes FPS@8k and FPS@16k; positives are FPS-overlap labels; time is normalized to FPS@8k). Combined with a hand-tuned {λκ, λt, λv, λg, δ, η} and a fixed discrete budget set B, this leaves open whether the policy learns a transferable “informative subset” rule or a specialization to the IGS candidate pool and N3DV training scenes. A minimal stress test—e.g., training without the FPS teacher mix, or evaluating the released policy under a non-IGS backbone with the same descriptor—should be reported, or the claim scope narrowed to “IGS-compatible budgeted selection.”
- Tab. 6 (HQ) vs Tab. 1 (fast): on MeetingRoom HQ, IGS@8192 remains best (29.55 PSNR) while EGS low-budget variants trail by ~0.6–0.8 dB; on N3DV HQ the ranking flips. The paper’s primary claim is framed as a consistent quality–efficiency improvement, but refinement-mode results are mixed and protocol-dependent. Either HQ should be clearly demoted to secondary context with that caveat in the abstract, or the conditions under which low-κ selection helps vs hurts under refinement need an explicit analysis.
minor comments (6)
- Fig. 2 caption claims “32× fewer anchors on unseen data” and large margins vs 3DGStream/StreamRF; those external methods are correctly marked as non-plug-in context in §4.3—keep that caveat in the figure caption so the Pareto plot is not over-read as a same-interface comparison.
- §3.2, Eq. (2): descriptor ξ_m uses only position, opacity, log-scale, and distance-to-center. A one-sentence justification for omitting color/SH or motion cues would help readers assess whether the policy can respond to appearance-critical anchors.
- §4.2: report the exact values of η, δ, and the λ weights used for the main checkpoint (currently only described qualitatively), so the reward in Eq. (4) is reproducible from the text alone.
- Table 2: at budgets ≥3072, ΔP_IGS turns negative; a brief discussion of why the learned policy underperforms FPS when the budget is no longer tight would strengthen the “low-budget Pareto” narrative.
- Limitations correctly note frozen-backbone and IGS-interface coupling; consider also stating that the bandit is frame-wise (no multi-frame credit assignment), matching the axiom used in training.
- Minor polish: “Imitation warm-start” / “Sthochastic” typo in Fig. 3; ensure PSNR/DSSIM/LPIPS aggregation (per-frame vs per-scene mean) is stated once in §4.3.
Circularity Check
No load-bearing circular derivation: reported gains are external GT metrics, not algebraically forced by the FPS teacher used only as a training reference.
specific steps
-
other
[Sec. 3.3 Training Objective, Eqs. (3)–(4)]
"We define a teacher target by mixing two FPS references: an 8k reference (used for time normalization and compatibility with IGS) and a stronger 16k teacher... The reward trades off sparsity, runtime, and quality: ρt = −λκ κt/κmax − λt max(0, tRL/tref − 1) − λv max(0, ψtgt − ψRL − δ) + λg max(0, ψRL − ψtgt)"
FPS is used both as SFT imitation target and as the quality/time reference inside the RL reward. This anchors the learning loop to FPS, but does not make reported test PSNR equal to the reward by construction: ψ is PSNR vs external GT, evaluation is on held-out/unseen scenes with deterministic top-κ selection, and the policy can under- or over-perform FPS. Mild training-loop dependence only; not a definitional reduction of the central empirical claim.
full rationale
The paper’s central claims are empirical quality–speed comparisons (PSNR/DSSIM/LPIPS vs ground-truth images, and wall-clock time) of a learned anchor sampler against IGS/FPS and other streaming baselines on held-out N3DV and MeetingRoom. Training does use FPS as an imitation warm-start and as a teacher target inside the contextual-bandit reward (Eq. 3–4: ψ_tgt mixes FPS@8k and FPS@16k; time is normalized to FPS@8k), but that is a standard baseline-relative training signal, not a definitional identity of the reported test metrics. Test PSNR is measured against external ground truth under frozen IGS reconstruction; same-budget ablations (Tab. 5) show the policy can beat FPS-matched subsets, and high-budget frontier points can be worse than FPS (Tab. 2 negative ΔP), so outcomes are not forced by construction. There is no self-citation uniqueness theorem, no ansatz smuggled from the authors’ prior work, and no renaming of a known closed-form result. The packaging tension between the null-finding abstract and the positive body claims is a claim-consistency issue, not circularity of the derivation chain. Score 1 only for the mild, non-load-bearing FPS teacher anchoring in the learning loop.
Axiom & Free-Parameter Ledger
free parameters (5)
- Reward weights {λκ, λt, λv, λg} and PSNR tolerance δ
- Teacher mix η in ψ_tgt = η ψ_ref + (1−η) ψ_tea
- Discrete budget set B (256–8192) and κ_max=8192
- M_max=16384 candidate pool via voxel-hash downsampling
- Optimizer and architecture hyperparameters (lr 1e-4, d=128, 4 heads, 2 layers, dropout 0.1, EMA baseline μ, entropy β)
axioms (4)
- domain assumption A frozen IGS anchor-graph + linear-blend-skinning renderer is a valid fixed backbone so that only the sampler changes quality and runtime.
- domain assumption PSNR (with DSSIM/LPIPS) is an adequate primary quality signal for the bandit reward and for claiming reconstruction improvement.
- ad hoc to paper Frame-wise contextual bandit (no explicit long-horizon temporal credit) is sufficient for streaming anchor selection.
- domain assumption N3DV training scenes plus MeetingRoom transfer tests represent the deployment distribution for multi-view dynamic capture.
invented entities (2)
-
EGS budgeted sampler policy (point-MLP + Transformer selection/budget heads)
no independent evidence
-
Composite sparsity–time–PSNR reward ρ_t with FPS teachers
no independent evidence
Cite this review
Pith. "Pith review of Does it matter which Gaussians you pick in 4D Gaussian streaming?." pith.science (2026). https://pith.science/paper/QOKYYXSB
@misc{pith2026260317227,
author = {Pith},
title = {Pith review of: Does it matter which Gaussians you pick in 4D Gaussian streaming?},
year = {2026},
howpublished = {\url{https://pith.science/paper/QOKYYXSB}},
note = {Machine review of arXiv:2603.17227}
}
read the original abstract
Anchor-driven 4D Gaussian streaming methods such as Instant Gaussian Stream (IGS) update a dynamic scene each frame from a compact set of Gaussian anchors, chosen by default with Farthest Point Sampling (FPS) at a fixed budget of $8{,}192$. Because these anchors act as control points that drive the whole scene through linear blend skinning, the rule used to choose them ought to affect reconstruction quality. We test this by holding the IGS pipeline fixed and changing only the sampler, comparing FPS, random, uniform, an opacity-scale heuristic, and a learned policy across budgets and refinement settings on N3DV and MeetingRoom. At deployment budgets the sampler has no measurable effect: a cheap random or uniform sampler at $4{,}096$ anchors matches FPS@8192 within measurement error, the default budget is over-provisioned, and the result holds on a second backbone (3DGStream). The learned policy is mixed rather than consistently better: it can improve the N3DV validation set at tight budgets, but does not give a stable cross-dataset rule, and selection is never the bottleneck because refinement dominates runtime. We will release our full sweep and evaluation protocol as a sampler benchmark.
Figures
Reference graph
Works this paper leans on
-
[1]
Per-gaussian embedding- based deformation for deformable 3d gaussian splatting
Jeongmin Bae, Seoha Kim, Youngsik Yun, Hahyun Lee, Gun Bang, and Youngjung Uh. Per-gaussian embedding- based deformation for deformable 3d gaussian splatting. In European Conference on Computer Vision, pages 321–335. Springer, 2024. 4
2024
-
[2]
Hexplane: A fast representa- tion for dynamic scenes.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 130–141, 2023
Ang Cao and Justin Johnson. Hexplane: A fast representa- tion for dynamic scenes.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 130–141, 2023. 2
2023
-
[3]
Junli Deng, Ping Shi, Qipei Li, and Jinyang Guo. Dynas- plat: Dynamic-static gaussian splatting with hierarchical mo- tion decomposition for scene reconstruction.arXiv preprint arXiv:2506.09836, 2025. 2
Pith/arXiv arXiv 2025
-
[4]
Learning to sam- ple
Oren Dovrat, Itai Lang, and Shai Avidan. Learning to sam- ple. InCVPR, 2019. 2
2019
-
[5]
4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes
Yuanxing Duan, Fangyin Wei, Qiyu Dai, Yuhang He, Wen- zheng Chen, and Baoquan Chen. 4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes. InACM SIGGRAPH 2024 Conference Papers, pages 1–11,
2024
-
[6]
Reconstructing 3d human pose by watching humans in the mirror
Qi Fang, Qing Shuai, Junting Dong, Hujun Bao, and Xiaowei Zhou. Reconstructing 3d human pose by watching humans in the mirror. InCVPR, 2021. 2
2021
-
[7]
Chaoran Feng, Zhenyu Tang, Wangbo Yu, Yatian Pang, Yian Zhao, Jianbin Zhao, Li Yuan, and Yonghong Tian. E-4dgs: High-fidelity dynamic reconstruction from the multi-view event cameras.arXiv preprint arXiv:2508.09912, 2025. 2
Pith/arXiv arXiv 2025
-
[8]
K-planes: Explicit radiance fields in space, time, and appearance
Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pages 12479–12488, June 2023. 2
2023
-
[9]
Gonzalez
Teofilo F. Gonzalez. Clustering to minimize the maxi- mum intercluster distance.Theoretical Computer Science, 38:293–306, 1985. 3
1985
-
[10]
Pup 3d-gs: Principled uncertainty pruning for 3d gaussian splatting
Alex Hanson, Allen Tu, Vasu Singla, Mayuka Jayawardhana, Matthias Zwicker, and Tom Goldstein. Pup 3d-gs: Principled uncertainty pruning for 3d gaussian splatting. InProceedings of the Computer Vision and Pattern Recognition Conference (CVPR), pages 5949–5958, June 2025. 2
2025
-
[11]
Amc: Automl for model compression and accel- eration on mobile devices
Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. Amc: Automl for model compression and accel- eration on mobile devices. InECCV, 2018. 2
2018
-
[12]
Yanqin Jiang, Li Zhang, Jin Gao, Weiming Hu, and Yao Yao. Consistent4d: Consistent 360° dynamic object gener- ation from monocular video.ArXiv, abs/2311.02848, 2023. 2
Pith/arXiv arXiv 2023
-
[13]
3d gaussian splatting for real-time radiance field rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM Trans. Graph., 42(4):139–1,
-
[14]
Physgaia: A physics-aware dataset of multi-body interactions for dynamic novel view synthesis
Mijeong Kim, Gunhee Kim, Jungyoon Choi, Wonjae Roh, and Bohyung Han. Physgaia: A physics-aware dataset of multi-body interactions for dynamic novel view synthesis. arXiv preprint arXiv:2506.02794, 2025. 2
Pith/arXiv arXiv 2025
-
[15]
Nersemble: Multi-view radi- ance field reconstruction of human heads.ACM Transactions on Graphics (TOG), 42(4):1–14, 2023
Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. Nersemble: Multi-view radi- ance field reconstruction of human heads.ACM Transactions on Graphics (TOG), 42(4):1–14, 2023. 2
2023
-
[16]
Samplenet: Differ- entiable point cloud sampling
Itai Lang, Asaf Manor, and Shai Avidan. Samplenet: Differ- entiable point cloud sampling. InCVPR, 2020. 2
2020
-
[17]
The epoch-greedy algo- rithm for multi-armed bandits with side information
John Langford and Tong Zhang. The epoch-greedy algo- rithm for multi-armed bandits with side information. In Advances in Neural Information Processing Systems 20 (NeurIPS 2007), 2007. 2
2007
-
[18]
Set transformer: A frame- work for attention-based permutation-invariant neural net- works
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Se- ungjin Choi, and Yee Whye Teh. Set transformer: A frame- work for attention-based permutation-invariant neural net- works. InInternational Conference on Machine Learning (ICML), 2019. 2, 3
2019
-
[19]
Fully explicit dynamic gaussian splat- ting.Advances in Neural Information Processing Systems, 37:5384–5409, 2024
Junoh Lee, ChangYeon Won, Hyunjun Jung, Inhwan Bae, and Hae-Gon Jeon. Fully explicit dynamic gaussian splat- ting.Advances in Neural Information Processing Systems, 37:5384–5409, 2024. 4
2024
-
[20]
Gifstream: 4d gaussian-based immersive video with feature stream
Hao Li, Sicheng Li, Xiang Gao, Abudouaihati Batuer, Lu Yu, and Yiyi Liao. Gifstream: 4d gaussian-based immersive video with feature stream. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 21761– 21770, 2025. 2
2025
-
[21]
Schapire
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire. A contextual-bandit approach to personalized news article recommendation. InProceedings of the 19th International Conference on World Wide Web (WWW), pages 661–670,
-
[22]
Streaming radiance fields for 3d video synthesis
Lingzhi Li, Zhen Shen, Zhongshu Wang, Li Shen, and Ping Tan. Streaming radiance fields for 3d video synthesis. InAd- vances in Neural Information Processing Systems (NeurIPS),
-
[23]
Mtgs: Multi-traversal gaussian splatting.arXiv preprint arXiv:2503.12552, 2025
Tianyu Li, Yihang Qiu, Zhenhua Wu, Carl Lind- str¨om, Peng Su, Matthias Nießner, and Hongyang Li. Mtgs: Multi-traversal gaussian splatting.arXiv preprint arXiv:2503.12552, 2025. 2
Pith/arXiv arXiv 2025
-
[24]
Neural 3d video synthesis from multi-view video
Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. Neural 3d video synthesis from multi-view video. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), pages 5521–5531,
-
[25]
Spacetime gaus- sian feature splatting for real-time dynamic view synthesis
Zhan Li, Zhang Chen, Zhong Li, and Yi Xu. Spacetime gaus- sian feature splatting for real-time dynamic view synthesis. 9 InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8508–8520, June 2024. 2
2024
-
[26]
Neural scene flow fields for space-time view synthesis of dy- namic scenes.2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6494–6504,
Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dy- namic scenes.2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6494–6504,
2021
-
[27]
Decoupled weight de- cay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight de- cay regularization. InInternational Conference on Learning Representations (ICLR), 2019. 4
2019
-
[28]
Diva- 360: The dynamic visual dataset for immersive neural fields
Cheng-You Lu, Peisen Zhou, Angela Xing, Chandradeep Pokhariya, Arnab Dey, Ishaan Nikhil Shah, Rugved Mavidi- palli, Dylan Hu, Andrew I Comport, Kefan Chen, et al. Diva- 360: The dynamic visual dataset for immersive neural fields. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22466–22476, 2024. 2
2024
-
[29]
Nerf: Representing scenes as neural radiance fields for view syn- thesis.Communications of the ACM, 65(1):99–106, 2021
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis.Communications of the ACM, 65(1):99–106, 2021. 1, 2
2021
-
[30]
Real-time 3d reconstruction at scale using voxel hashing.ACM Transactions on Graphics, 32(6):169:1–169:11, 2013
Matthias Nießner, Michael Zollh ¨ofer, Shahram Izadi, and Marc Stamminger. Real-time 3d reconstruction at scale using voxel hashing.ACM Transactions on Graphics, 32(6):169:1–169:11, 2013. 3
2013
-
[31]
Seungjun Oh, Younggeun Lee, Hyejin Jeon, and Eunbyung Park. Hybrid 3d-4d gaussian splatting for fast dynamic scene representation.arXiv preprint arXiv:2505.13215, 2025. 2, 4
Pith/arXiv arXiv 2025
-
[32]
Nerfies: Deformable neural radiance fields
Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. InProceedings of the IEEE/CVF international conference on computer vision, pages 5865–5874, 2021. 2
2021
-
[33]
Sinha, Peter Hedman, Jonathan T
Keunhong Park, U. Sinha, Peter Hedman, Jonathan T. Bar- ron, Sofien Bouaziz, Dan B. Goldman, Ricardo Martin- Brualla, and Steven M. Seitz. Hypernerf.ACM Transactions on Graphics (TOG), 40:1 – 12, 2021. 2
2021
-
[34]
Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans
Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. InCVPR,
-
[35]
D-NeRF: Neural Radiance Fields for Dynamic Scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-NeRF: Neural Radiance Fields for Dynamic Scenes. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
2020
-
[36]
Qi, Li Yi, Hao Su, and Leonidas J
Charles R. Qi, Li Yi, Hao Su, and Leonidas J. Guibas. Point- net++: Deep hierarchical feature learning on point sets in a metric space. InAdvances in Neural Information Processing Systems (NeurIPS), 2017. 2, 3
2017
-
[37]
Octree-gs: Towards consistent real-time rendering with lod-structured 3d gaussians.IEEE transactions on pattern analysis and machine intelligence, PP, 2024
Kerui Ren, Lihan Jiang, Tao Lu, Mulin Yu, Linning Xu, Zhangkai Ni, and Bo Dai. Octree-gs: Towards consistent real-time rendering with lod-structured 3d gaussians.IEEE transactions on pattern analysis and machine intelligence, PP, 2024. 2
2024
-
[38]
Dataset and pipeline for multi-view light-field video
Neus Sabater, Guillaume Boisson, Benoit Vandame, Paul Kerbiriou, Frederic Babon, Matthieu Hog, Remy Gendrot, Tristan Langlois, Olivier Bureller, Arno Schubert, and Va- lerie Alli ´e. Dataset and pipeline for multi-view light-field video. In2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1743– 1753, 2017. 2
2017
-
[39]
Replay: Multi- modal multi-view acted videos for casual holography
Roman Shapovalov, Yanir Kleiman, Ignacio Rocco, David Novotny, Andrea Vedaldi, Changan Chen, Filippos Kokki- nos, Ben Graham, and Natalia Neverova. Replay: Multi- modal multi-view acted videos for casual holography. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20338–20348, 2023. 2
2023
-
[40]
3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free- viewpoint videos
Jiakai Sun, Han Jiao, Guangyuan Li, Zhanjie Zhang, Lei Zhao, and Wei Xing. 3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free- viewpoint videos. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 20675–20685, 2024. 2, 4, 5, 8
2024
-
[41]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto.Reinforcement Learning: An Introduction. MIT Press, 2 edition, 2018. 2, 4
2018
-
[42]
Allen Tu, Haiyang Ying, Alex Hanson, Yonghan Lee, Tom Goldstein, and Matthias Zwicker. Speedy deformable 3d gaussian splatting: Fast rendering and compression of dy- namic scenes.arXiv preprint arXiv:2506.07917, 2025. 2
Pith/arXiv arXiv 2025
-
[43]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems (NeurIPS), 2017. 2, 3
2017
-
[44]
Neural residual radiance fields for streamably free-viewpoint videos
Liao Wang, Qiang Hu, Qihan He, Ziyu Wang, Jingyi Yu, Tinne Tuytelaars, Lan Xu, and Minye Wu. Neural residual radiance fields for streamably free-viewpoint videos. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 76–87, June 2023. 2
2023
-
[45]
Gonzalez
Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E. Gonzalez. Skipnet: Learning dynamic routing in convolutional networks. InECCV, 2018. 2
2018
-
[46]
Freetimegs: Free gaussian primitives at anytime any- where for dynamic scene reconstruction
Yifan Wang, Peishan Yang, Zhen Xu, Jiaming Sun, Zhan- hua Zhang, Yong Chen, Hujun Bao, Sida Peng, and Xiaowei Zhou. Freetimegs: Free gaussian primitives at anytime any- where for dynamic scene reconstruction. InCVPR, 2025. 2
2025
-
[47]
Bovik, Hamid R
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: From error visibility to structural similarity.IEEE Transactions on Image Pro- cessing, 13(4):600–612, 2004. 4
2004
-
[48]
Williams
Ronald J. Williams. Simple statistical gradient-following al- gorithms for connectionist reinforcement learning.Machine Learning, 8(3–4):229–256, 1992. 2, 4
1992
-
[49]
4d gaussian splatting for real-time dynamic scene render- ing
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene render- ing. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 20310– 20320, June 2024. 2
2024
-
[50]
Davis, Kristen Grauman, and Rogerio Feris
Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar, Steven Rennie, Larry S. Davis, Kristen Grauman, and Rogerio Feris. Blockdrop: Dynamic inference paths in residual networks. InCVPR, 2018. 2 10
2018
-
[51]
Instant gaussian stream: Fast and generalizable streaming of dy- namic scene reconstruction via gaussian splatting
Jinbo Yan, Rui Peng, Zhiyan Wang, Luyang Tang, Jiayu Yang, Jie Liang, Jiahao Wu, and Ronggang Wang. Instant gaussian stream: Fast and generalizable streaming of dy- namic scene reconstruction via gaussian splatting. InPro- ceedings of the Computer Vision and Pattern Recognition Conference, pages 16520–16531, 2025. 2, 3, 4, 5, 8
2025
-
[52]
Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction
Ziyi Yang, Xinyu Gao, Wenming Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction. 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 20331–20341, 2023. 2
2024
-
[53]
Wei Yao, Shuzhao Xie, Letian Li, Weixiang Zhang, Zhixin Lai, Shiqi Dai, Ke Zhang, and Zhi Wang. Sd-gs: Structured deformable 3d gaussians for efficient dynamic scene recon- struction.arXiv preprint arXiv:2507.07465, 2025. 4
Pith/arXiv arXiv 2025
-
[54]
Splat4d: Diffusion-enhanced 4d gaussian splatting for tem- porally and spatially consistent content creation
Minghao Yin, Yukang Cao, Songyou Peng, and Kai Han. Splat4d: Diffusion-enhanced 4d gaussian splatting for tem- porally and spatially consistent content creation. InProceed- ings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, pages 1–10, 2025. 2
2025
-
[55]
Y . Yuan et al. Efficient differentiable hardware rasterization for 3d gaussian splatting.arXiv preprint arXiv:2505.18764,
-
[56]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 4
2018
-
[57]
Pku- dymvhumans: A multi-view video benchmark for high- fidelity dynamic human modeling
Xiaoyun Zheng, Liwei Liao, Xufeng Li, Jianbo Jiao, Rongjie Wang, Feng Gao, Shiqi Wang, and Ronggang Wang. Pku- dymvhumans: A multi-view video benchmark for high- fidelity dynamic human modeling. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22530–22540, 2024. 2 11
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.