Pith. sign in

REVIEW 4 major objections 6 minor 57 references

A learned budget-aware policy can pick far fewer Gaussian anchors and still beat fixed FPS in 4D streaming.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 23:16 UTC pith:QOKYYXSB

load-bearing objection The arXiv title/abstract claim a null sampler result; the body claims EGS@256 beats IGS@8192 by ~0.5 dB and ~1.3× speed—until that is reconciled the central number is not checkable. the 4 major comments →

arxiv 2603.17227 v3 pith:QOKYYXSB submitted 2026-03-18 cs.CV

Does it matter which Gaussians you pick in 4D Gaussian streaming?

classification cs.CV
keywords 4D Gaussian SplattingGaussian streaminganchor selectionbudgeted samplingreinforcement learningcontextual banditdynamic novel view synthesisfree-viewpoint video
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Streaming dynamic scenes with Gaussian splatting usually freezes a large, fixed set of anchors chosen by farthest-point sampling. This paper argues that the choice of those anchors can be turned into a budgeted decision problem: a learned policy picks both how many anchors to keep and which ones, while the rest of the reconstruction pipeline stays frozen. Trained first by imitating farthest-point sampling and then with a contextual-bandit reward that trades sparsity, runtime, and image quality, the policy is reported to match or improve reconstruction at budgets as low as 256 anchors—32 times fewer than the usual 8192—while cutting per-frame time on unseen multi-view scenes. A reader who cares about free-viewpoint video under strict compute limits cares because the same streaming backbone can be made leaner without redesigning the renderer or the skinning graph, if the right control points can be selected on the fly.

Core claim

The paper claims that Efficient Gaussian Streaming (EGS), a plug-in reinforcement-learned sampler that jointly selects an anchor budget and an informative subset, improves the quality–efficiency trade-off of anchor-driven 4D Gaussian streaming relative to fixed farthest-point sampling at 8192 anchors, including roughly half a decibel PSNR gains and about 1.3× lower latency at 256 anchors on unseen data.

What carries the argument

The budgeted anchor action (budget κ plus same-size subset Ω): candidates are scored by a point-MLP plus small Transformer, then trained with a short FPS-imitation warm-start followed by a contextual-bandit reward that penalizes excess anchors, excess time, and PSNR shortfall while rewarding quality gains over FPS teacher targets.

Load-bearing premise

The claim depends on a frame-wise reward built from farthest-point-sampling teacher targets, hand-tuned weights, and a short imitation warm-start actually producing a general rule for informative anchors rather than a fit to the training interface and scenes.

What would settle it

Freeze the released checkpoint and, under the same fast-rendering protocol with no online refinement, measure whether 256-anchor EGS still beats FPS at 8192 in both PSNR and time per frame on a held-out multi-view dynamic dataset outside the paper’s N3DV and MeetingRoom splits; a clear loss on quality or speed falsifies the claimed general trade-off.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • The default 8192-anchor FPS budget is over-provisioned for many scenes if selection is content-aware.
  • Anchor selection can be swapped as a plug-in without retraining the streaming reconstruction backbone.
  • Same-budget ablations imply gains come from subset quality, not only from shrinking the budget.
  • Learned selection pays off most in the low-budget regime where sampling overhead is offset by fewer anchors.
  • A shared budget–sampler sweep can turn “which Gaussians to pick” into a comparable subproblem rather than a fixed default.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If refinement dominates runtime, further latency wins may come more from lighter refinement than from smarter anchor choice alone.
  • The same budgeted set-selection pattern may transfer to other online pipelines that drive a scene from sparse control points.
  • Tying the reward to FPS teachers may cap how far the policy can move beyond spatial-coverage heuristics.
  • Releasing the full sampler sweep would let later work treat random, uniform, and learned policies under one protocol.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript presents Efficient Gaussian Streaming (EGS), a plug-in budget-aware anchor sampler for IGS-style 4D Gaussian streaming. A point-MLP + Transformer policy jointly chooses a discrete budget κ and a subset of anchors; training uses one-epoch FPS imitation (SFT) followed by a frame-wise contextual-bandit REINFORCE objective whose reward trades sparsity, runtime, and PSNR against FPS teacher targets (Eqs. 3–5). With the IGS backbone frozen, the paper reports that low-budget EGS (256–1024 anchors) improves PSNR by roughly +0.5 dB and is 1.2–1.35× faster than IGS@8192 on held-out N3DV and MeetingRoom scenes in fast mode, with same-budget RL-vs-FPS wins in 14/15 pairs (Tab. 5) and competitive HQ refinement at reduced budgets (Tab. 6).

Significance. If the positive low-budget results hold under a single, reproducible protocol, the work is a useful systems contribution: it shows that anchor-driven 4DGS streaming can operate far below the conventional 8192-FPS budget without quality collapse, and that subset choice (not only κ) can matter under tight constraints. The frozen-backbone plug-in design, same-budget FPS ablation, dual-protocol (fast/HQ) reporting, and promised code/checkpoints are concrete strengths. The absolute quality gains are modest, and the method is tightly coupled to the IGS candidate-pool interface, so impact is primarily engineering/protocol rather than a new representation principle.

major comments (4)
  1. Title/abstract vs body claim conflict: the arXiv title and front-matter abstract assert a largely null result (sampler choice has no measurable effect at deployment budgets; random/uniform@4096 matches FPS@8192; learned policy is mixed and not a stable cross-dataset rule), while the body abstract, Fig. 2, and Tabs. 1–6 assert the opposite strongest claim (EGS@256 beats IGS@8192 by +0.52–0.61 dB and 1.29–1.35× speed on unseen data). These cannot both be true under one protocol. The manuscript must present a single reconciled claim set, matching experiments (including the random/uniform/opacity-scale sweep named only in the null abstract), and consistent numbers before the central result is checkable.
  2. §4 / Tabs. 1–5: all headline ΔPSNR and speedup figures are point estimates with no multi-seed means, standard errors, or paired significance tests. Given reported gains of ~0.4–0.6 dB and the abstract’s own language of “within measurement error,” error bars (or at least 3-seed replications of the frozen-checkpoint evaluation) are load-bearing for deciding whether the sampler effect is real or noise—especially for the same-budget 14/15 win claim in Tab. 5.
  3. §3.3, Eqs. (3)–(4): the reward and SFT warm-start are deliberately FPS-anchored (ψ_tgt mixes FPS@8k and FPS@16k; positives are FPS-overlap labels; time is normalized to FPS@8k). Combined with a hand-tuned {λκ, λt, λv, λg, δ, η} and a fixed discrete budget set B, this leaves open whether the policy learns a transferable “informative subset” rule or a specialization to the IGS candidate pool and N3DV training scenes. A minimal stress test—e.g., training without the FPS teacher mix, or evaluating the released policy under a non-IGS backbone with the same descriptor—should be reported, or the claim scope narrowed to “IGS-compatible budgeted selection.”
  4. Tab. 6 (HQ) vs Tab. 1 (fast): on MeetingRoom HQ, IGS@8192 remains best (29.55 PSNR) while EGS low-budget variants trail by ~0.6–0.8 dB; on N3DV HQ the ranking flips. The paper’s primary claim is framed as a consistent quality–efficiency improvement, but refinement-mode results are mixed and protocol-dependent. Either HQ should be clearly demoted to secondary context with that caveat in the abstract, or the conditions under which low-κ selection helps vs hurts under refinement need an explicit analysis.
minor comments (6)
  1. Fig. 2 caption claims “32× fewer anchors on unseen data” and large margins vs 3DGStream/StreamRF; those external methods are correctly marked as non-plug-in context in §4.3—keep that caveat in the figure caption so the Pareto plot is not over-read as a same-interface comparison.
  2. §3.2, Eq. (2): descriptor ξ_m uses only position, opacity, log-scale, and distance-to-center. A one-sentence justification for omitting color/SH or motion cues would help readers assess whether the policy can respond to appearance-critical anchors.
  3. §4.2: report the exact values of η, δ, and the λ weights used for the main checkpoint (currently only described qualitatively), so the reward in Eq. (4) is reproducible from the text alone.
  4. Table 2: at budgets ≥3072, ΔP_IGS turns negative; a brief discussion of why the learned policy underperforms FPS when the budget is no longer tight would strengthen the “low-budget Pareto” narrative.
  5. Limitations correctly note frozen-backbone and IGS-interface coupling; consider also stating that the bandit is frame-wise (no multi-frame credit assignment), matching the axiom used in training.
  6. Minor polish: “Imitation warm-start” / “Sthochastic” typo in Fig. 3; ensure PSNR/DSSIM/LPIPS aggregation (per-frame vs per-scene mean) is stated once in §4.3.

Circularity Check

1 steps flagged

No load-bearing circular derivation: reported gains are external GT metrics, not algebraically forced by the FPS teacher used only as a training reference.

specific steps
  1. other [Sec. 3.3 Training Objective, Eqs. (3)–(4)]
    "We define a teacher target by mixing two FPS references: an 8k reference (used for time normalization and compatibility with IGS) and a stronger 16k teacher... The reward trades off sparsity, runtime, and quality: ρt = −λκ κt/κmax − λt max(0, tRL/tref − 1) − λv max(0, ψtgt − ψRL − δ) + λg max(0, ψRL − ψtgt)"

    FPS is used both as SFT imitation target and as the quality/time reference inside the RL reward. This anchors the learning loop to FPS, but does not make reported test PSNR equal to the reward by construction: ψ is PSNR vs external GT, evaluation is on held-out/unseen scenes with deterministic top-κ selection, and the policy can under- or over-perform FPS. Mild training-loop dependence only; not a definitional reduction of the central empirical claim.

full rationale

The paper’s central claims are empirical quality–speed comparisons (PSNR/DSSIM/LPIPS vs ground-truth images, and wall-clock time) of a learned anchor sampler against IGS/FPS and other streaming baselines on held-out N3DV and MeetingRoom. Training does use FPS as an imitation warm-start and as a teacher target inside the contextual-bandit reward (Eq. 3–4: ψ_tgt mixes FPS@8k and FPS@16k; time is normalized to FPS@8k), but that is a standard baseline-relative training signal, not a definitional identity of the reported test metrics. Test PSNR is measured against external ground truth under frozen IGS reconstruction; same-budget ablations (Tab. 5) show the policy can beat FPS-matched subsets, and high-budget frontier points can be worse than FPS (Tab. 2 negative ΔP), so outcomes are not forced by construction. There is no self-citation uniqueness theorem, no ansatz smuggled from the authors’ prior work, and no renaming of a known closed-form result. The packaging tension between the null-finding abstract and the positive body claims is a claim-consistency issue, not circularity of the derivation chain. Score 1 only for the mild, non-load-bearing FPS teacher anchoring in the learning loop.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 2 invented entities

The central empirical claim rests on the frozen IGS streaming interface, a hand-designed multi-term reward with several scalar weights, discrete budget buckets, and a candidate-pool construction—not on a parameter-free derivation. No new physical entities are postulated; the invented pieces are the policy architecture and reward. Free parameters and domain assumptions dominate what is not paid for by prior IGS/3DGS literature.

free parameters (5)
  • Reward weights {λκ, λt, λv, λg} and PSNR tolerance δ
    Hand-chosen scalars in Eq. 4 that define the sparsity/time/quality trade-off the policy optimizes; paper says training is 'insensitive to moderate weight changes' but does not report a full sensitivity grid for the final claims.
  • Teacher mix η in ψ_tgt = η ψ_ref + (1−η) ψ_tea
    Controls how much the reward targets FPS@8k vs a stronger FPS@16k teacher; fitted/chosen for training stability.
  • Discrete budget set B (256–8192) and κ_max=8192
    Action space is restricted to fixed buckets aligned with the IGS default; the operating points reported (tiny/small/med) are chosen from this set.
  • M_max=16384 candidate pool via voxel-hash downsampling
    Caps and reshapes the set the policy can select from; changes which Gaussians are even visible to the sampler.
  • Optimizer and architecture hyperparameters (lr 1e-4, d=128, 4 heads, 2 layers, dropout 0.1, EMA baseline μ, entropy β)
    Standard but claim-relevant free choices that affect whether the policy converges to the reported low-budget regime.
axioms (4)
  • domain assumption A frozen IGS anchor-graph + linear-blend-skinning renderer is a valid fixed backbone so that only the sampler changes quality and runtime.
    Stated throughout §3 and evaluation protocol; all gains are conditional on this interface remaining optimal when anchors change.
  • domain assumption PSNR (with DSSIM/LPIPS) is an adequate primary quality signal for the bandit reward and for claiming reconstruction improvement.
    Reward Eq. 4 and all main tables use PSNR differences vs FPS teachers / IGS.
  • ad hoc to paper Frame-wise contextual bandit (no explicit long-horizon temporal credit) is sufficient for streaming anchor selection.
    §3.3 formulates per-frame REINFORCE; limitations section admits more explicit temporal modeling is future work.
  • domain assumption N3DV training scenes plus MeetingRoom transfer tests represent the deployment distribution for multi-view dynamic capture.
    §4.1 splits; generalization claims rest on this dataset pair only.
invented entities (2)
  • EGS budgeted sampler policy (point-MLP + Transformer selection/budget heads) no independent evidence
    purpose: Map candidate Gaussian descriptors to a joint budget and subset action under discrete constraints.
    New module relative to FPS; independent evidence is only the paper’s own ablations, not an external measurement.
  • Composite sparsity–time–PSNR reward ρ_t with FPS teachers no independent evidence
    purpose: Provide a scalar training signal that encourages fewer anchors, lower latency, and higher PSNR than FPS references.
    Hand-designed objective; not derived from a unique optimality principle outside this paper.

pith-pipeline@v1.1.0-grok45 · 20341 in / 3757 out tokens · 66417 ms · 2026-07-13T23:16:21.018648+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Does it matter which Gaussians you pick in 4D Gaussian streaming?." pith.science (2026). https://pith.science/paper/QOKYYXSB

@misc{pith2026260317227,
  author       = {Pith},
  title        = {Pith review of: Does it matter which Gaussians you pick in 4D Gaussian streaming?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QOKYYXSB}},
  note         = {Machine review of arXiv:2603.17227}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Anchor-driven 4D Gaussian streaming methods such as Instant Gaussian Stream (IGS) update a dynamic scene each frame from a compact set of Gaussian anchors, chosen by default with Farthest Point Sampling (FPS) at a fixed budget of $8{,}192$. Because these anchors act as control points that drive the whole scene through linear blend skinning, the rule used to choose them ought to affect reconstruction quality. We test this by holding the IGS pipeline fixed and changing only the sampler, comparing FPS, random, uniform, an opacity-scale heuristic, and a learned policy across budgets and refinement settings on N3DV and MeetingRoom. At deployment budgets the sampler has no measurable effect: a cheap random or uniform sampler at $4{,}096$ anchors matches FPS@8192 within measurement error, the default budget is over-provisioned, and the result holds on a second backbone (3DGStream). The learned policy is mixed rather than consistently better: it can improve the N3DV validation set at tight budgets, but does not give a stable cross-dataset rule, and selection is never the bottleneck because refinement dominates runtime. We will release our full sweep and evaluation protocol as a sampler benchmark.

Figures

Figures reproduced from arXiv: 2603.17227 by Ashim Dahal, Nick Rahimi, Rabab Abdelfattah.

Figure 2
Figure 2. Figure 2: Fast Rendering Mode. Our RL policy selects anchors that match or improve quality with 32× fewer anchors on unseen data. Left: qualitative comparison (Ours, IGS, 3DGStream). Right: N3DV PSNR vs time/frame trade-off. Abstract Dynamic scene reconstruction with Gaussian Splatting has enabled efficient streaming for real-time rendering and free-viewpoint video. However, most pipelines rely on fixed anchor selec… view at source ↗
Figure 3
Figure 3. Figure 3: Method overview of EGS. Our adaptive sampler selects a budgeted set of anchor Gaussians (stochastic during training; deter￾ministic at inference) using point-MLP embeddings and a lightweight Transformer [18, 43], then feeds the selected anchors to the frozen IGS pipeline for anchor-graph construction and rasterization [51]. The training objective trades off sparsity, runtime, and PSNR against FPS-based tea… view at source ↗
Figure 4
Figure 4. Figure 4: Training dynamics and analysis. Across training, the sampler learns stable budget control and converges under constraint-aware rewards; final checkpoints show positive ∆PSNR vs IGS@8192 on seen and unseen splits. Training curves use stochastic actions and may differ from deterministic inference. dates. In parallel, the on-policy ∆PSNR signal shown in Fig. 4c initially improves as exploration discovers bett… view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison on N3DV and MeetingRoom scenes. We compare our method against IGS, 3DGStream, and ground truth across representative scenes. f122 Trimming | t~4.0s +0.00s f130 +0.27s f138 +0.53s f146 +0.80s f152 Trimming | t~5.0s f160 f168 f176 f182 Trimming | t~6.0s f190 f198 f206 [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Motion progression on MeetingRoom:Trimming. Re￾constructions across frames demonstrating temporal consistency [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

57 extracted references · 9 linked inside Pith

  1. [1]

    Per-gaussian embedding- based deformation for deformable 3d gaussian splatting

    Jeongmin Bae, Seoha Kim, Youngsik Yun, Hahyun Lee, Gun Bang, and Youngjung Uh. Per-gaussian embedding- based deformation for deformable 3d gaussian splatting. In European Conference on Computer Vision, pages 321–335. Springer, 2024. 4

  2. [2]

    Hexplane: A fast representa- tion for dynamic scenes.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 130–141, 2023

    Ang Cao and Justin Johnson. Hexplane: A fast representa- tion for dynamic scenes.2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 130–141, 2023. 2

  3. [3]

    Dynas- plat: Dynamic-static gaussian splatting with hierarchical mo- tion decomposition for scene reconstruction.arXiv preprint arXiv:2506.09836, 2025

    Junli Deng, Ping Shi, Qipei Li, and Jinyang Guo. Dynas- plat: Dynamic-static gaussian splatting with hierarchical mo- tion decomposition for scene reconstruction.arXiv preprint arXiv:2506.09836, 2025. 2

  4. [4]

    Learning to sam- ple

    Oren Dovrat, Itai Lang, and Shai Avidan. Learning to sam- ple. InCVPR, 2019. 2

  5. [5]

    4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes

    Yuanxing Duan, Fangyin Wei, Qiyu Dai, Yuhang He, Wen- zheng Chen, and Baoquan Chen. 4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes. InACM SIGGRAPH 2024 Conference Papers, pages 1–11,

  6. [6]

    Reconstructing 3d human pose by watching humans in the mirror

    Qi Fang, Qing Shuai, Junting Dong, Hujun Bao, and Xiaowei Zhou. Reconstructing 3d human pose by watching humans in the mirror. InCVPR, 2021. 2

  7. [7]

    E-4dgs: High-fidelity dynamic reconstruction from the multi-view event cameras.arXiv preprint arXiv:2508.09912, 2025

    Chaoran Feng, Zhenyu Tang, Wangbo Yu, Yatian Pang, Yian Zhao, Jianbin Zhao, Li Yuan, and Yonghong Tian. E-4dgs: High-fidelity dynamic reconstruction from the multi-view event cameras.arXiv preprint arXiv:2508.09912, 2025. 2

  8. [8]

    K-planes: Explicit radiance fields in space, time, and appearance

    Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pages 12479–12488, June 2023. 2

  9. [9]

    Gonzalez

    Teofilo F. Gonzalez. Clustering to minimize the maxi- mum intercluster distance.Theoretical Computer Science, 38:293–306, 1985. 3

  10. [10]

    Pup 3d-gs: Principled uncertainty pruning for 3d gaussian splatting

    Alex Hanson, Allen Tu, Vasu Singla, Mayuka Jayawardhana, Matthias Zwicker, and Tom Goldstein. Pup 3d-gs: Principled uncertainty pruning for 3d gaussian splatting. InProceedings of the Computer Vision and Pattern Recognition Conference (CVPR), pages 5949–5958, June 2025. 2

  11. [11]

    Amc: Automl for model compression and accel- eration on mobile devices

    Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. Amc: Automl for model compression and accel- eration on mobile devices. InECCV, 2018. 2

  12. [12]

    Consistent4d: Consistent 360° dynamic object gener- ation from monocular video.ArXiv, abs/2311.02848, 2023

    Yanqin Jiang, Li Zhang, Jin Gao, Weiming Hu, and Yao Yao. Consistent4d: Consistent 360° dynamic object gener- ation from monocular video.ArXiv, abs/2311.02848, 2023. 2

  13. [13]

    3d gaussian splatting for real-time radiance field rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM Trans. Graph., 42(4):139–1,

  14. [14]

    Physgaia: A physics-aware dataset of multi-body interactions for dynamic novel view synthesis

    Mijeong Kim, Gunhee Kim, Jungyoon Choi, Wonjae Roh, and Bohyung Han. Physgaia: A physics-aware dataset of multi-body interactions for dynamic novel view synthesis. arXiv preprint arXiv:2506.02794, 2025. 2

  15. [15]

    Nersemble: Multi-view radi- ance field reconstruction of human heads.ACM Transactions on Graphics (TOG), 42(4):1–14, 2023

    Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. Nersemble: Multi-view radi- ance field reconstruction of human heads.ACM Transactions on Graphics (TOG), 42(4):1–14, 2023. 2

  16. [16]

    Samplenet: Differ- entiable point cloud sampling

    Itai Lang, Asaf Manor, and Shai Avidan. Samplenet: Differ- entiable point cloud sampling. InCVPR, 2020. 2

  17. [17]

    The epoch-greedy algo- rithm for multi-armed bandits with side information

    John Langford and Tong Zhang. The epoch-greedy algo- rithm for multi-armed bandits with side information. In Advances in Neural Information Processing Systems 20 (NeurIPS 2007), 2007. 2

  18. [18]

    Set transformer: A frame- work for attention-based permutation-invariant neural net- works

    Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Se- ungjin Choi, and Yee Whye Teh. Set transformer: A frame- work for attention-based permutation-invariant neural net- works. InInternational Conference on Machine Learning (ICML), 2019. 2, 3

  19. [19]

    Fully explicit dynamic gaussian splat- ting.Advances in Neural Information Processing Systems, 37:5384–5409, 2024

    Junoh Lee, ChangYeon Won, Hyunjun Jung, Inhwan Bae, and Hae-Gon Jeon. Fully explicit dynamic gaussian splat- ting.Advances in Neural Information Processing Systems, 37:5384–5409, 2024. 4

  20. [20]

    Gifstream: 4d gaussian-based immersive video with feature stream

    Hao Li, Sicheng Li, Xiang Gao, Abudouaihati Batuer, Lu Yu, and Yiyi Liao. Gifstream: 4d gaussian-based immersive video with feature stream. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 21761– 21770, 2025. 2

  21. [21]

    Schapire

    Lihong Li, Wei Chu, John Langford, and Robert E. Schapire. A contextual-bandit approach to personalized news article recommendation. InProceedings of the 19th International Conference on World Wide Web (WWW), pages 661–670,

  22. [22]

    Streaming radiance fields for 3d video synthesis

    Lingzhi Li, Zhen Shen, Zhongshu Wang, Li Shen, and Ping Tan. Streaming radiance fields for 3d video synthesis. InAd- vances in Neural Information Processing Systems (NeurIPS),

  23. [23]

    Mtgs: Multi-traversal gaussian splatting.arXiv preprint arXiv:2503.12552, 2025

    Tianyu Li, Yihang Qiu, Zhenhua Wu, Carl Lind- str¨om, Peng Su, Matthias Nießner, and Hongyang Li. Mtgs: Multi-traversal gaussian splatting.arXiv preprint arXiv:2503.12552, 2025. 2

  24. [24]

    Neural 3d video synthesis from multi-view video

    Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. Neural 3d video synthesis from multi-view video. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), pages 5521–5531,

  25. [25]

    Spacetime gaus- sian feature splatting for real-time dynamic view synthesis

    Zhan Li, Zhang Chen, Zhong Li, and Yi Xu. Spacetime gaus- sian feature splatting for real-time dynamic view synthesis. 9 InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8508–8520, June 2024. 2

  26. [26]

    Neural scene flow fields for space-time view synthesis of dy- namic scenes.2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6494–6504,

    Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dy- namic scenes.2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6494–6504,

  27. [27]

    Decoupled weight de- cay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight de- cay regularization. InInternational Conference on Learning Representations (ICLR), 2019. 4

  28. [28]

    Diva- 360: The dynamic visual dataset for immersive neural fields

    Cheng-You Lu, Peisen Zhou, Angela Xing, Chandradeep Pokhariya, Arnab Dey, Ishaan Nikhil Shah, Rugved Mavidi- palli, Dylan Hu, Andrew I Comport, Kefan Chen, et al. Diva- 360: The dynamic visual dataset for immersive neural fields. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22466–22476, 2024. 2

  29. [29]

    Nerf: Representing scenes as neural radiance fields for view syn- thesis.Communications of the ACM, 65(1):99–106, 2021

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis.Communications of the ACM, 65(1):99–106, 2021. 1, 2

  30. [30]

    Real-time 3d reconstruction at scale using voxel hashing.ACM Transactions on Graphics, 32(6):169:1–169:11, 2013

    Matthias Nießner, Michael Zollh ¨ofer, Shahram Izadi, and Marc Stamminger. Real-time 3d reconstruction at scale using voxel hashing.ACM Transactions on Graphics, 32(6):169:1–169:11, 2013. 3

  31. [31]

    Hybrid 3d-4d gaussian splatting for fast dynamic scene representation.arXiv preprint arXiv:2505.13215, 2025

    Seungjun Oh, Younggeun Lee, Hyejin Jeon, and Eunbyung Park. Hybrid 3d-4d gaussian splatting for fast dynamic scene representation.arXiv preprint arXiv:2505.13215, 2025. 2, 4

  32. [32]

    Nerfies: Deformable neural radiance fields

    Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. InProceedings of the IEEE/CVF international conference on computer vision, pages 5865–5874, 2021. 2

  33. [33]

    Sinha, Peter Hedman, Jonathan T

    Keunhong Park, U. Sinha, Peter Hedman, Jonathan T. Bar- ron, Sofien Bouaziz, Dan B. Goldman, Ricardo Martin- Brualla, and Steven M. Seitz. Hypernerf.ACM Transactions on Graphics (TOG), 40:1 – 12, 2021. 2

  34. [34]

    Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans

    Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. InCVPR,

  35. [35]

    D-NeRF: Neural Radiance Fields for Dynamic Scenes

    Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-NeRF: Neural Radiance Fields for Dynamic Scenes. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2

  36. [36]

    Qi, Li Yi, Hao Su, and Leonidas J

    Charles R. Qi, Li Yi, Hao Su, and Leonidas J. Guibas. Point- net++: Deep hierarchical feature learning on point sets in a metric space. InAdvances in Neural Information Processing Systems (NeurIPS), 2017. 2, 3

  37. [37]

    Octree-gs: Towards consistent real-time rendering with lod-structured 3d gaussians.IEEE transactions on pattern analysis and machine intelligence, PP, 2024

    Kerui Ren, Lihan Jiang, Tao Lu, Mulin Yu, Linning Xu, Zhangkai Ni, and Bo Dai. Octree-gs: Towards consistent real-time rendering with lod-structured 3d gaussians.IEEE transactions on pattern analysis and machine intelligence, PP, 2024. 2

  38. [38]

    Dataset and pipeline for multi-view light-field video

    Neus Sabater, Guillaume Boisson, Benoit Vandame, Paul Kerbiriou, Frederic Babon, Matthieu Hog, Remy Gendrot, Tristan Langlois, Olivier Bureller, Arno Schubert, and Va- lerie Alli ´e. Dataset and pipeline for multi-view light-field video. In2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1743– 1753, 2017. 2

  39. [39]

    Replay: Multi- modal multi-view acted videos for casual holography

    Roman Shapovalov, Yanir Kleiman, Ignacio Rocco, David Novotny, Andrea Vedaldi, Changan Chen, Filippos Kokki- nos, Ben Graham, and Natalia Neverova. Replay: Multi- modal multi-view acted videos for casual holography. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20338–20348, 2023. 2

  40. [40]

    3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free- viewpoint videos

    Jiakai Sun, Han Jiao, Guangyuan Li, Zhanjie Zhang, Lei Zhao, and Wei Xing. 3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free- viewpoint videos. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 20675–20685, 2024. 2, 4, 5, 8

  41. [41]

    Sutton and Andrew G

    Richard S. Sutton and Andrew G. Barto.Reinforcement Learning: An Introduction. MIT Press, 2 edition, 2018. 2, 4

  42. [42]

    Speedy deformable 3d gaussian splatting: Fast rendering and compression of dy- namic scenes.arXiv preprint arXiv:2506.07917, 2025

    Allen Tu, Haiyang Ying, Alex Hanson, Yonghan Lee, Tom Goldstein, and Matthias Zwicker. Speedy deformable 3d gaussian splatting: Fast rendering and compression of dy- namic scenes.arXiv preprint arXiv:2506.07917, 2025. 2

  43. [43]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems (NeurIPS), 2017. 2, 3

  44. [44]

    Neural residual radiance fields for streamably free-viewpoint videos

    Liao Wang, Qiang Hu, Qihan He, Ziyu Wang, Jingyi Yu, Tinne Tuytelaars, Lan Xu, and Minye Wu. Neural residual radiance fields for streamably free-viewpoint videos. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 76–87, June 2023. 2

  45. [45]

    Gonzalez

    Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E. Gonzalez. Skipnet: Learning dynamic routing in convolutional networks. InECCV, 2018. 2

  46. [46]

    Freetimegs: Free gaussian primitives at anytime any- where for dynamic scene reconstruction

    Yifan Wang, Peishan Yang, Zhen Xu, Jiaming Sun, Zhan- hua Zhang, Yong Chen, Hujun Bao, Sida Peng, and Xiaowei Zhou. Freetimegs: Free gaussian primitives at anytime any- where for dynamic scene reconstruction. InCVPR, 2025. 2

  47. [47]

    Bovik, Hamid R

    Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: From error visibility to structural similarity.IEEE Transactions on Image Pro- cessing, 13(4):600–612, 2004. 4

  48. [48]

    Williams

    Ronald J. Williams. Simple statistical gradient-following al- gorithms for connectionist reinforcement learning.Machine Learning, 8(3–4):229–256, 1992. 2, 4

  49. [49]

    4d gaussian splatting for real-time dynamic scene render- ing

    Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene render- ing. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 20310– 20320, June 2024. 2

  50. [50]

    Davis, Kristen Grauman, and Rogerio Feris

    Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar, Steven Rennie, Larry S. Davis, Kristen Grauman, and Rogerio Feris. Blockdrop: Dynamic inference paths in residual networks. InCVPR, 2018. 2 10

  51. [51]

    Instant gaussian stream: Fast and generalizable streaming of dy- namic scene reconstruction via gaussian splatting

    Jinbo Yan, Rui Peng, Zhiyan Wang, Luyang Tang, Jiayu Yang, Jie Liang, Jiahao Wu, and Ronggang Wang. Instant gaussian stream: Fast and generalizable streaming of dy- namic scene reconstruction via gaussian splatting. InPro- ceedings of the Computer Vision and Pattern Recognition Conference, pages 16520–16531, 2025. 2, 3, 4, 5, 8

  52. [52]

    Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction

    Ziyi Yang, Xinyu Gao, Wenming Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction. 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 20331–20341, 2023. 2

  53. [53]

    Sd-gs: Structured deformable 3d gaussians for efficient dynamic scene recon- struction.arXiv preprint arXiv:2507.07465, 2025

    Wei Yao, Shuzhao Xie, Letian Li, Weixiang Zhang, Zhixin Lai, Shiqi Dai, Ke Zhang, and Zhi Wang. Sd-gs: Structured deformable 3d gaussians for efficient dynamic scene recon- struction.arXiv preprint arXiv:2507.07465, 2025. 4

  54. [54]

    Splat4d: Diffusion-enhanced 4d gaussian splatting for tem- porally and spatially consistent content creation

    Minghao Yin, Yukang Cao, Songyou Peng, and Kai Han. Splat4d: Diffusion-enhanced 4d gaussian splatting for tem- porally and spatially consistent content creation. InProceed- ings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, pages 1–10, 2025. 2

  55. [55]

    Yuan et al

    Y . Yuan et al. Efficient differentiable hardware rasterization for 3d gaussian splatting.arXiv preprint arXiv:2505.18764,

  56. [56]

    Efros, Eli Shecht- man, and Oliver Wang

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 4

  57. [57]

    Pku- dymvhumans: A multi-view video benchmark for high- fidelity dynamic human modeling

    Xiaoyun Zheng, Liwei Liao, Xufeng Li, Jianbo Jiao, Rongjie Wang, Feng Gao, Shiqi Wang, and Ronggang Wang. Pku- dymvhumans: A multi-view video benchmark for high- fidelity dynamic human modeling. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22530–22540, 2024. 2 11