Pith. sign in

REVIEW 4 minor 103 references

Two different MoE integration strategies for multi-deformation Gaussian models outperform any single deformation prior on dynamic scenes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 15:36 UTC pith:CYE7VFYL

load-bearing objection Solid design-space paper: two concrete MoE integration strategies for dynamic 3DGS, with real gains and honest trade-offs, not a universal claim.

arxiv 2607.08250 v2 pith:CYE7VFYL submitted 2026-07-09 cs.CV

On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting

classification cs.CV
keywords dynamic Gaussian splattingmixture of expertsdeformation modelingnovel view synthesis3D reconstructionMoDEMoE-GS
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Real-world dynamic scenes contain heterogeneous motion that no single deformation model captures well: different methods win on different scenes, patches, and timestamps because each deformation prior favors only certain motion regimes. The paper treats multi-deformation modeling as a Mixture-of-Experts problem and shows two concrete ways to combine specialized experts inside a unified 3D Gaussian representation. MoDE jointly optimizes multiple canonical deformation experts that share the same Gaussian primitives and are gated by a continuous temporal spline, adding almost no extra stages. MoE-GS instead trains heterogeneous experts independently (including non-canonical ones) and later blends their rendered images with a volume-aware pixel router. Together the two designs map the trade-offs among reconstruction fidelity, training stability, and compute, giving practitioners clear alternatives rather than a single universal recipe.

Core claim

Performance gaps among existing dynamic Gaussian Splatting methods arise from reliance on one deformation prior, not from lack of capacity. Combining multiple specialized deformation experts under either joint canonical optimization (MoDE) or decoupled optimization plus volume-aware routing (MoE-GS) systematically improves novel-view quality by exploiting complementary motion behaviors across space and time.

What carries the argument

Two integration constraints that decide when experts interact: MoDE (shared canonical Gaussians + spline temporal gating + baseline-only gradient flow) versus MoE-GS (independent experts + volume-aware pixel router that splat per-Gaussian routing weights then blends images).

Load-bearing premise

The chosen experts keep complementary, non-interfering motion priors under the proposed gating so the mixture reliably beats the strongest single expert instead of averaging or destabilizing.

What would settle it

Find a dynamic scene (or large ROI set) in which one fixed deformation model already wins every spatial region and every timestamp; on that data both MoDE and MoE-GS must then match or fall below that single expert’s PSNR rather than improve it.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • When a shared canonical space exists, MoDE gives multi-deformation modeling with only modest extra training time and direct 3D Gaussian output.
  • When experts are heterogeneous or already trained, MoE-GS yields larger PSNR gains by full specialization, at the price of multiple training runs and a routing stage.
  • Gate-aware pruning and distillation recover real-time speed while retaining most of the mixture’s quality.
  • Image-space routing can still be lifted to a coherent post-hoc 3D Gaussian model whose multi-view depth consistency matches or exceeds single experts.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same joint-versus-decoupled design choice likely applies to other explicit dynamic representations (meshes, particles, hash grids) whose motion models also carry conflicting inductive biases.
  • An online residual-driven expert pool that can add or drop deformation modules mid-training would reduce the need for hand-selected candidate sets.
  • Volume-aware routing may serve as a general post-hoc calibration layer for any ensemble of 3D renderers that lack direct primitive correspondence.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper studies multi-deformation modeling for dynamic 3D Gaussian Splatting under two integration constraints, framed as Mixture-of-Experts. MoDE jointly optimizes multiple canonical deformation experts (HexPlane, hash-grid, per-Gaussian embeddings) on a shared canonical Gaussian representation with spline-based temporal Top-K gating and asymmetric gradient flow. MoE-GS independently trains heterogeneous experts (including non-canonical polynomial and keyframe-interpolation models plus static experts) then combines them via a volume-aware pixel router that lifts per-Gaussian routing weights; efficiency is recovered by single-pass multi-expert rendering, gate-aware pruning, and distillation. Extensive experiments on N3V, Technicolor, HyperNeRF, PanopticSports and D-NeRF, plus ablations of routers, pruning, distillation, expert candidates, run-to-run stability and multi-view depth consistency, demonstrate that both strategies improve robustness over single-deformation baselines while exposing complementary trade-offs in 3D fidelity, training stability and cost (Tables I–XI, Figs. 5–11).

Significance. If the reported gains and trade-offs hold, the work supplies a clear design-space map for multi-deformation dynamic Gaussian representations rather than a single new SOTA method. The explicit contrast between joint canonical composition (MoDE) and decoupled expert-plus-router composition (MoE-GS), the geometry-aware lifting analysis, the efficiency mechanisms, and the released code (https://github.com/cvsp-lab/MoE-GS-studio) make the contribution reusable for future dynamic GS research. The empirical breadth—multiple datasets, large-motion benchmarks, repeated-run stability, and static-region degradation analysis—raises the bar for claims about deformation priors in this literature.

minor comments (4)
  1. [Sec. IV-B3 / Fig. 6] Table II and Fig. 6: the static-region degradation for E-D3DGS-based MoDE is shown for one illustrative case; a short quantitative summary of static vs. dynamic ROI PSNR across all six N3V scenes would make the trade-off fully transparent.
  2. [Sec. III-C2] Eqs. (12)–(15) and Fig. 4: the residual MLP Φ that refines the splatted routing features is described only at a high level; a one-sentence statement of its layer count / channel width would aid exact re-implementation.
  3. [Appendix D3 / Fig. 11] Appendix D3: the Multi-view Depth Consistency formula is clear, yet the precise set of viewpoint pairs and the depth-map resolution used for the curves in Fig. 11 are not stated; adding them would strengthen reproducibility of the geometry claim.
  4. [Throughout] A few typographical inconsistencies remain (e.g., “V olume-aware”, occasional missing spaces after citations). A final proof-reading pass would polish the manuscript.

Circularity Check

1 steps flagged

No significant circularity: empirical design-space comparison of two MoE integration strategies, with results measured on held-out views against external baselines.

specific steps
  1. self citation load bearing [Sec. II-C Related Works (Mixture of Experts) and abstract/intro framing of MoE-GS]
    "Inspired by this perspective, MoE-GS [78] applies mixture-of-experts to Dynamic Gaussian Splatting through rendering-level expert routing. ... Building upon this line of research, the present work introduces MoDE ..."

    The paper cites the authors’ own prior MoE-GS work as the foundation for one of the two integration strategies under study. This is ordinary incremental research and is fully disclosed; it is not load-bearing for any derivation, uniqueness claim, or numerical result in the present paper, which re-implements, re-evaluates, and contrasts MoE-GS with the new MoDE formulation on independent benchmarks.

full rationale

This is a standard computer-vision systems/engineering paper. Its central claims are comparative empirical results (PSNR/SSIM/LPIPS tables, qualitative figures, efficiency ablations, multi-view depth consistency) obtained by training and evaluating MoDE and MoE-GS on public dynamic-scene benchmarks (N3V, Technicolor, HyperNeRF, PanopticSports, D-NeRF). Routing/gating weights are optimized from data under standard L1+SSIM losses; they are not defined to equal any target metric. The only self-citation is the authors’ prior MoE-GS conference paper, which is openly disclosed as the starting point for one of the two integration strategies and is not used as a uniqueness theorem, ansatz, or definitional premise that forces the new results. No equation reduces a claimed prediction to a fitted input by construction, no uniqueness is imported from the authors’ own prior work, and no known empirical pattern is merely renamed. The paper is therefore self-contained against external benchmarks; the minor self-citation does not raise the score above 1.

Axiom & Free-Parameter Ledger

4 free parameters · 3 axioms · 3 invented entities

Empirical systems paper. The load-bearing content is architectural design plus measured performance; free parameters are the usual training and architectural hyperparameters. Domain assumptions are standard 3DGS and dynamic-deformation premises. Invented entities are the two MoE formulations and the volume-aware router.

free parameters (4)
  • number of experts N and Top-K gating
    Chosen by hand (N=2/3/4 fixed combinations; Top-K sparse softmax); performance depends on the selection.
  • spline control points Nw and warm-up iterations
    Architectural and schedule choices that control temporal gating smoothness and early expert balance.
  • router learning rates and pruning threshold τ
    Hand-tuned (e.g., 0.05/0.5 for MLP vs. per-Gaussian weights; Ei < τ for pruning); directly affect final quality/efficiency trade-off.
  • distillation balance λ
    Weight between GT and MoE pseudo-supervision in the KD loss; fitted for the reported distilled-expert gains.
axioms (3)
  • domain assumption Standard 3D Gaussian Splatting rasterization and optimization (Kerbl et al.) correctly approximate the radiance field for novel-view synthesis.
    All experts and both MoE variants rest on this representation; invoked throughout Sec. III.
  • domain assumption Distinct deformation formulations (HexPlane, per-Gaussian embedding, polynomial, keyframe interpolation) induce complementary motion priors that can be usefully specialized.
    Core motivation (Fig. 1 and trajectory analysis); without it the mixture has no advantage.
  • ad hoc to paper Image-space blending of independently trained experts can be made geometry-aware via per-Gaussian routing weights that are liftable back to 3D.
    Assumed by the volume-aware pixel router and the post-hoc fusion / MDC evaluation (Sec. III-C3, Appendix D).
invented entities (3)
  • Mixture of Deformation Experts (MoDE) no independent evidence
    purpose: Joint multi-expert deformation composition on a shared canonical Gaussian set with spline gating and asymmetric gradient flow.
    New architectural construct introduced in Sec. III-B; no independent external evidence beyond the paper’s own experiments.
  • Volume-aware Pixel Router (and associated lifting) no independent evidence
    purpose: Spatially/temporally adaptive blending of heterogeneous expert renders while retaining Gaussian-level structural cues.
    Central mechanism of MoE-GS (Sec. III-C2); evaluated only inside this paper.
  • Gate-aware Gaussian pruning + single-pass multi-expert rendering no independent evidence
    purpose: Make multi-expert inference practical by removing low-influence Gaussians and eliminating redundant rasterization.
    Efficiency inventions specific to MoE-GS (Sec. III-C3).

pith-pipeline@v1.1.0-grok45 · 40280 in / 3023 out tokens · 31958 ms · 2026-07-14T15:36:17.107229+00:00 · methodology

0 comments
read the original abstract

Dynamic scene reconstruction remains challenging due to the heterogeneous and spatially varying nature of real-world motion. Although recent 3D Gaussian Splatting methods have introduced diverse deformation formulations for dynamic novel view synthesis, each method typically relies on a single deformation model within its representation, which limits robustness across diverse dynamic scenarios. In this work, we study a fundamental problem-multi-deformation modeling for dynamic 3D Gaussian representations-under two distinct integration constraints that differ in when and how multiple deformation experts interact during training. From a Mixture-of-Experts (MoE) perspective, we view multi-deformation modeling as the problem of combining multiple specialized deformation models within a unified 3D representation. We first introduce Mixture of Deformation Experts (MoDE), which integrates multiple deformation experts directly into the deformable Gaussian Splatting pipeline through joint optimization. In MoDE, experts operate on a shared canonical Gaussian representation, enabling multi-deformation modeling without introducing additional training stages or modifying the original optimization schedule. In contrast, we further present Mixture of Experts for Dynamic Gaussian Splatting (MoE-GS) under a different integration constraint, where deformation experts are optimized independently and combined through a separate routing stage. As a result, expert interaction occurs over non-canonical Gaussian representations after individual optimization. Together, these two approaches provide alternative strategies for multi-deformation modeling, clarifying how integration constraints shape the design and behavior of deformation experts in dynamic 3D Gaussian representations. Our code is available at: https://github.com/cvsp-lab/MoE-GS-studio.

Figures

Figures reproduced from arXiv: 2607.08250 by Hyeongju Mun, In-Hwan Jin, Joonsoo Kim, Kugjin Yun, Kyeongbo Kong.

Figure 1
Figure 1. Figure 1: Limitations of existing dynamic Gaussian splatting methods. (a) Scene-level: No single method consistently dominates [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of Mixture of Deformation Experts (MoDE). MoDE augments a canonical Gaussian deformation model [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the MoE-GS framework. In Stage 1 (Expert Training), each expert is independently trained to reconstruct [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Comparison of Router Architectures. The Pixel Router (top-left) assigns weights purely at the pixel level, ignoring [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: MoDE Qualitative Results Comparison of MoDE with baseline dynamic Gaussian splatting methods on Neural 3D Video dataset [84]. corresponding routing weights G′ k provide confidence estimates for each expert. The distillation loss is L KD k = λL(G ′ k IEk , G′ k IGT ) + (1 − λ)L((1 − G ′ k )IEk , (1 − G ′ k )IMoE). (18) where L combines L1 and SSIM losses, and λ balances ground-truth vs. MoE supervision. Thi… view at source ↗
Figure 6
Figure 6. Figure 6: Trade-off analysis of the E-D3DGS-based MoDE variant. Dynamic-region gains (blue) and static-region degra￾dation (red). use each method in its original single-deformation configura￾tion as a baseline reference. MoDE variants are constructed by augmenting each baseline with one additional canonical deformation expert operating on the same set of canonical Gaussians. All deformation experts are trained joint… view at source ↗
Figure 7
Figure 7. Figure 7: N3V Qualitative Results Comparison of our MoE-GS with other dynamic Gaussian splatting methods on Neural 3D Video dataset [84]. Blue backgrounds highlight the method that produces the most visually accurate result among the baselines for each region. Coffee_Martini Ex4DGS (Inter.) 4DGaussians (Hex.) E-D3DGS (Per.) STG (Poly.) [PITH_FULL_IMAGE:figures/full_fig_p013_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Expert-specific motion patterns. Representative motion trajectories produced by different dynamic Gaussian Splatting experts. These observations suggest that different deformation formulations induce distinct motion priors, which in turn lead to different expert specialization behaviors in MoE-GS. This complementary behavior motivates the expert composition strategy adopted in MoE-GS. Furthermore, addition… view at source ↗
Figure 9
Figure 9. Figure 9 [PITH_FULL_IMAGE:figures/full_fig_p014_9.png] view at source ↗
Figure 10
Figure 10. Figure 10 [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Multi-view Depth Consistency on the N3V dataset [84]. Comparison of our MoE-GS with other dynamic Gaussian [PITH_FULL_IMAGE:figures/full_fig_p015_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Architectural details of the Volume-aware Pixel Router. [PITH_FULL_IMAGE:figures/full_fig_p019_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Additional N3V Qualitative Results. Comparison of our MoE-GS with other dynamic Gaussian splatting methods on the Neural 3D Video dataset [84] [PITH_FULL_IMAGE:figures/full_fig_p024_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Additional Qualitative Results on the Technicolor Dataset [85]. Visual comparison of our MoE-GS with other dynamic Gaussian splatting methods [PITH_FULL_IMAGE:figures/full_fig_p025_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Visual comparison between retrained and distilled expert models on the Technicolor dataset [85]. [PITH_FULL_IMAGE:figures/full_fig_p026_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Motion Trajectory 1 Comparison across dynamic Gaussian Splatting methods. Cook_Spinach Ex4DGS (Inter.) 4DGaussians (Hex.) E-D3DGS (Per.) STG (Poly.) [PITH_FULL_IMAGE:figures/full_fig_p028_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Motion Trajectory 2 Comparison across dynamic Gaussian Splatting methods. Flame_Salmon Ex4DGS (Inter.) 4DGaussians (Hex.) E-D3DGS (Per.) STG (Poly.) [PITH_FULL_IMAGE:figures/full_fig_p028_17.png] view at source ↗
Figure 18
Figure 18. Figure 18 [PITH_FULL_IMAGE:figures/full_fig_p028_18.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

103 extracted references · 6 linked inside Pith

  1. [1]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,”Communications of the ACM, vol. 65, pp. 99–106, 2021

  2. [2]

    3d gaussian splatting for real-time radiance field rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,”ACM Transactions on Graphics, vol. 42, pp. 139–1, 2023

  3. [3]

    4d gaussian splatting for real-time dynamic scene rendering,

    G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 310–20 320

  4. [4]

    Spacetime gaussian feature splatting for real-time dynamic view synthesis,

    Z. Li, Z. Chen, Z. Li, and Y . Xu, “Spacetime gaussian feature splatting for real-time dynamic view synthesis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 8508–8520

  5. [5]

    Per-gaussian embedding-based deformation for deformable 3d gaussian splatting,

    J. Bae, S. Kim, Y . Yun, H. Lee, G. Bang, and Y . Uh, “Per-gaussian embedding-based deformation for deformable 3d gaussian splatting,” in European Conference on Computer Vision. Springer, 2024, pp. 321–335

  6. [6]

    Fully explicit dynamic gaussian splatting,

    J. Lee, C. Won, H. Jung, I. Bae, and H.-G. Jeon, “Fully explicit dynamic gaussian splatting,”Advances in Neural Information Processing Systems, vol. 37, pp. 5384–5409, 2024

  7. [7]

    Grid4d: 4d decomposed hash encoding for high-fidelity dynamic gaussian splatting,

    J. Xu, Z. Fan, J. Yang, and J. Xie, “Grid4d: 4d decomposed hash encoding for high-fidelity dynamic gaussian splatting,” inAdvances in Neural Information Processing Systems, vol. 37, 2024, pp. 123 787–123 811

  8. [8]

    The plenoptic function and the elements of early vision,

    J. R. Bergen and E. H. Adelson, “The plenoptic function and the elements of early vision,”Computational models of visual processing, vol. 1, no. 8, p. 3, 1991

  9. [9]

    Light field rendering,

    M. Levoy and P. Hanrahan, “Light field rendering,” inSeminal Graphics Papers: Pushing the Boundaries, Volume 2, 2023, pp. 441–452

  10. [10]

    Ms-nerf: Multi- space neural radiance fields,

    Z.-X. Yin, P.-Y . Jiao, J. Qiu, M.-M. Cheng, and B. Ren, “Ms-nerf: Multi- space neural radiance fields,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  11. [11]

    Ref-nerf: Structured view-dependent appearance for neural radiance fields,

    D. Verbin, P. Hedman, B. Mildenhall, T. Zickler, J. T. Barron, and P. P. Srinivasan, “Ref-nerf: Structured view-dependent appearance for neural radiance fields,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 9426–9437, 2024

  12. [12]

    Forward flow for novel view synthesis of dynamic scenes,

    X. Guo, J. Sun, Y . Dai, G. Chen, X. Ye, X. Tan, E. Ding, Y . Zhang, and J. Wang, “Forward flow for novel view synthesis of dynamic scenes,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 16 022–16 033

  13. [13]

    Devrf: Fast deformable voxel radiance fields for dynamic scenes,

    J.-W. Liu, Y .-P. Cao, W. Mao, W. Zhang, D. J. Zhang, J. Keppo, Y . Shan, X. Qie, and M. Z. Shou, “Devrf: Fast deformable voxel radiance fields for dynamic scenes,”Advances in Neural Information Processing Systems, vol. 35, pp. 36 762–36 775, 2022

  14. [14]

    Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields,

    K. Park, U. Sinha, P. Hedman, J. T. Barron, S. Bouaziz, D. B. Goldman, R. Martin-Brualla, and S. M. Seitz, “Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields,” in SIGGRAPH Asia, 2021, pp. 1–12

  15. [15]

    D- nerf: Neural radiance fields for dynamic scenes,

    A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno-Noguer, “D- nerf: Neural radiance fields for dynamic scenes,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 10 318–10 327

  16. [16]

    Nerfplayer: A streamable dynamic scene representation with decomposed neural radiance fields,

    L. Song, A. Chen, Z. Li, Z. Chen, L. Chen, J. Yuan, Y . Xu, and A. Geiger, “Nerfplayer: A streamable dynamic scene representation with decomposed neural radiance fields,”IEEE Transactions on Visualization and Computer Graphics, vol. 29, no. 5, pp. 2732–2742, 2023

  17. [17]

    Neural trajectory fields for dynamic novel view synthesis,

    C. Wang, B. Eckart, S. Lucey, and O. Gallo, “Neural trajectory fields for dynamic novel view synthesis,”arXiv preprint arXiv:2105.05994, 2021

  18. [18]

    Nerfies: Deformable neural radiance fields,

    K. Park, U. Sinha, J. T. Barron, S. Bouaziz, D. B. Goldman, S. M. Seitz, and R. Martin-Brualla, “Nerfies: Deformable neural radiance fields,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5865–5874

  19. [19]

    Deformgs: Scene flow in highly deformable scenes for deformable object manipulation,

    B. P. Duisterhof, M. Zhao, Y . Yao, J.-W. Liu, J. Seidenschwarz, M. Z. Shou, D. Ramanan, S. Song, S. Birchfield, B. Wenet al., “Deformgs: Scene flow in highly deformable scenes for deformable object manipulation,” inInternational Workshop on the Algorithmic Foundations of Robotics. Springer, 2024, pp. 263–282

  20. [20]

    Hexplane: A fast representation for dynamic scenes,

    A. Cao and J. Johnson, “Hexplane: A fast representation for dynamic scenes,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 130–141

  21. [21]

    K-planes: Explicit radiance fields in space, time, and appearance,

    S. Fridovich-Keil, G. Meanti, F. R. Warburg, B. Recht, and A. Kanazawa, “K-planes: Explicit radiance fields in space, time, and appearance,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12 479–12 488

  22. [22]

    High-fidelity and real-time novel view synthesis for dynamic scenes,

    H. Lin, S. Peng, Z. Xu, T. Xie, X. He, H. Bao, and X. Zhou, “High-fidelity and real-time novel view synthesis for dynamic scenes,” inSIGGRAPH Asia 2023 Conference Papers, 2023, pp. 1–9

  23. [23]

    Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering,

    R. Shao, Z. Zheng, H. Tu, B. Liu, H. Zhang, and Y . Liu, “Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 16 632–16 642

  24. [24]

    Masked space-time hash encoding for efficient dynamic scene reconstruction,

    F. Wang, Z. Chen, G. Wang, Y . Song, and H. Liu, “Masked space-time hash encoding for efficient dynamic scene reconstruction,”Advances in neural information processing systems, vol. 36, pp. 70 497–70 510, 2023

  25. [25]

    Neural residual radiance fields for streamably free-viewpoint videos,

    L. Wang, Q. Hu, Q. He, Z. Wang, J. Yu, T. Tuytelaars, L. Xu, and M. Wu, “Neural residual radiance fields for streamably free-viewpoint videos,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 76–87

  26. [26]

    Hac++: Towards 100x compression of 3d gaussian splatting,

    Y . Chen, Q. Wu, W. Lin, M. Harandi, and J. Cai, “Hac++: Towards 100x compression of 3d gaussian splatting,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  27. [27]

    Z-splat: Z-axis gaussian splatting for camera-sonar fusion,

    Z. Qu, O. Vengurlekar, M. Qadri, K. Zhang, M. Kaess, C. Metzler, S. Jayasuriya, and A. Pediredla, “Z-splat: Z-axis gaussian splatting for camera-sonar fusion,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  28. [28]

    3dgstream: On- the-fly training of 3d gaussians for efficient streaming of photo-realistic free-viewpoint videos,

    J. Sun, H. Jiao, G. Li, Z. Zhang, L. Zhao, and W. Xing, “3dgstream: On- the-fly training of 3d gaussians for efficient streaming of photo-realistic free-viewpoint videos,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 675–20 685

  29. [29]

    Dynamics-aware gaussian splatting streaming towards fast on-the-fly 4d reconstruction,

    Z. Liu, Y . Hu, X. Zhang, R. Song, J. Shao, Z. Lin, and J. Zhang, “Dynamics-aware gaussian splatting streaming towards fast on-the-fly 4d reconstruction,”IEEE Transactions on Visualization and Computer Graphics, 2026

  30. [30]

    4d gaussian splatting with scale- aware residual field and adaptive optimization for real-time rendering of temporally complex dynamic scenes,

    J. Yan, R. Peng, L. Tang, and R. Wang, “4d gaussian splatting with scale- aware residual field and adaptive optimization for real-time rendering of temporally complex dynamic scenes,” inProceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 7871–7880. JIN et al.: ON THE DESIGN OF MIXTURE-OF-EXPERTS FOR DYNAMIC GAUSSIAN SPLATTING 17

  31. [31]

    Swings: Sliding window gaussian splatting for volumetric video streaming with arbitrary length,

    B. Liu and S. Banerjee, “Swings: Sliding window gaussian splatting for volumetric video streaming with arbitrary length,”arXiv preprint arXiv:2409.07759, vol. 2409, pp. 1–12, 2024

  32. [32]

    4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes,

    Y . Duan, F. Wei, Q. Dai, Y . He, W. Chen, and B. Chen, “4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes,” in ACM SIGGRAPH 2024 Conference Papers, 2024, pp. 1–11

  33. [33]

    Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle,

    Y . Lin, Z. Dai, S. Zhu, and Y . Yao, “Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 136–21 145

  34. [34]

    Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis,

    J. Luiten, G. Kopanas, B. Leibe, and D. Ramanan, “Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis,” in2024 International Conference on 3D Vision (3DV). IEEE, 2024, pp. 800–809

  35. [35]

    Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,

    Z. Yang, X. Gao, W. Zhou, S. Jiao, Y . Zhang, and X. Jin, “Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 331–20 341

  36. [36]

    Dynmf: Neural motion factorization for real-time dynamic view synthesis with 3d gaussian splatting,

    A. Kratimenos, J. Lei, and K. Daniilidis, “Dynmf: Neural motion factorization for real-time dynamic view synthesis with 3d gaussian splatting,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 252–269

  37. [37]

    Gaufre: Gaussian deformation fields for real-time dynamic novel view synthesis,

    Y . Liang, N. Khan, Z. Li, T. Nguyen-Phuoc, D. Lanman, J. Tompkin, and L. Xiao, “Gaufre: Gaussian deformation fields for real-time dynamic novel view synthesis,” in2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025, pp. 2642–2652

  38. [38]

    Sc-gs: Sparse-controlled gaussian splatting for editable dynamic scenes,

    Y .-H. Huang, Y .-T. Sun, Z. Yang, X. Lyu, Y .-P. Cao, and X. Qi, “Sc-gs: Sparse-controlled gaussian splatting for editable dynamic scenes,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 4220–4230

  39. [39]

    Cogs: Controllable gaussian splatting,

    H. Yu, J. Julin, Z. ´A. Milacski, K. Niinuma, and L. A. Jeni, “Cogs: Controllable gaussian splatting,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 624–21 633

  40. [40]

    Dash: 4d hash encoding with self-supervised decomposition for real-time dynamic scene rendering,

    J. Chen, Z. Hu, P. Wu, H. Zhu, H. Li, and X. Sun, “Dash: 4d hash encoding with self-supervised decomposition for real-time dynamic scene rendering,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 26 349–26 359

  41. [41]

    Localdygs: Multi-view global dynamic scene modeling via adaptive local implicit feature decoupling,

    J. Wu, R. Peng, J. Jiao, J. Yang, L. Tang, K. Xiong, J. Liang, J. Yan, R. Liu, and R. Wang, “Localdygs: Multi-view global dynamic scene modeling via adaptive local implicit feature decoupling,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 9519–9529

  42. [42]

    Haif-gs: Hierarchical and induced flow-guided gaussian splatting for dynamic scene,

    J. Chen, Z. Li, Y . Cai, H. Jiang, C. Qian, J. Kang, S. Gao, H. Zhao, T. Mao, and Y . Zhang, “Haif-gs: Hierarchical and induced flow-guided gaussian splatting for dynamic scene,”Advances in Neural Information Processing Systems, vol. 38, pp. 125 539–125 563, 2026

  43. [43]

    Timeformer: Capturing temporal relationships of deformable 3d gaussians for robust re- construction,

    D. Jiang, Z. Hou, Z. Ke, X. Yang, X. Zhou, and T. Qiu, “Timeformer: Capturing temporal relationships of deformable 3d gaussians for robust re- construction,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 8721–8732

  44. [44]

    Freetimegs: Free gaussian primitives at anytime anywhere for dynamic scene reconstruction,

    Y . Wang, P. Yang, Z. Xu, J. Sun, Z. Zhang, Y . Chen, H. Bao, S. Peng, and X. Zhou, “Freetimegs: Free gaussian primitives at anytime anywhere for dynamic scene reconstruction,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 750–21 760

  45. [45]

    7dgs: Unified spatial-temporal-angular gaussian splatting,

    Z. Gao, B. Planche, M. Zheng, A. Choudhuri, T. Chen, and Z. Wu, “7dgs: Unified spatial-temporal-angular gaussian splatting,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 26 316–26 325

  46. [46]

    Modgs: Dynamic gaussian splatting from casually-captured monocular videos with depth priors,

    Q. Liu, Y . Liu, J. Wang, X. Lyu, P. Wang, W. Wang, and J. Hou, “Modgs: Dynamic gaussian splatting from casually-captured monocular videos with depth priors,” inInternational Conference on Learning Representations, vol. 2025, 2025, pp. 97 048–97 074

  47. [47]

    Real-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting,

    Z. Yang, H. Yang, Z. Pan, and L. Zhang, “Real-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting,” inInternational Conference on Learning Representations, vol. 2024, 2024, pp. 9142–9159

  48. [48]

    4d gaussian splatting: Modeling dynamic scenes with native 4d primitives,

    Z. Yang, Z. Pan, X. Zhu, L. Zhang, J. Feng, Y .-G. Jiang, and P. H. Torr, “4d gaussian splatting: Modeling dynamic scenes with native 4d primitives,”arXiv preprint arXiv:2412.20720, 2024

  49. [49]

    Mega: Memory-efficient 4d gaussian splatting for dynamic scenes,

    X. Zhang, Z. Liu, Y . Zhang, X. Ge, D. He, T. Xu, Y . Wang, Z. Lin, S. Yan, and J. Zhang, “Mega: Memory-efficient 4d gaussian splatting for dynamic scenes,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 27 828–27 838

  50. [50]

    4d scaffold gaussian splatting with dynamic-aware anchor growing for efficient and high-fidelity dynamic scene reconstruction,

    W. O. Cho, I. Cho, S. Kim, J. Bae, Y . Uh, and S. J. Kim, “4d scaffold gaussian splatting with dynamic-aware anchor growing for efficient and high-fidelity dynamic scene reconstruction,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 5, 2026, pp. 3363–3371

  51. [51]

    Shape of motion: 4d reconstruction from a single video,

    Q. Wang, V . Ye, H. Gao, W. Zeng, J. Austin, Z. Li, and A. Kanazawa, “Shape of motion: 4d reconstruction from a single video,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 9660–9672

  52. [52]

    Slowfast networks for video recognition,

    C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 6202–6211

  53. [53]

    R. H. Bartels, J. C. Beatty, and B. A. Barsky,An introduction to splines for use in computer graphics and geometric modeling. Morgan Kaufmann, 1995

  54. [54]

    Animating rotation with quaternion curves,

    K. Shoemake, “Animating rotation with quaternion curves,” inProceed- ings of the 12th annual conference on Computer graphics and interactive techniques, 1985, pp. 245–254

  55. [55]

    Simple and scalable predictive uncertainty estimation using deep ensembles,

    B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,”Advances in neural information processing systems, vol. 30, 2017

  56. [56]

    Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” inInternational Conference on Learning Representations, 2017, pp. 1–14

  57. [57]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

    W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”Journal of Machine Learning Research, vol. 23, no. 120, pp. 1–39, 2022

  58. [58]

    Mod-squad: Designing mixtures of experts as modular multi-task learners,

    Z. Chen, Y . Shen, M. Ding, Z. Chen, H. Zhao, E. G. Learned-Miller, and C. Gan, “Mod-squad: Designing mixtures of experts as modular multi-task learners,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 11 828–11 837

  59. [59]

    Base layers: Simplifying training of large, sparse models,

    M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer, “Base layers: Simplifying training of large, sparse models,” inInternational Conference on Machine Learning. PMLR, 2021, pp. 6265–6274

  60. [60]

    Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning,

    H. Hazimeh, Z. Zhao, A. Chowdhery, M. Sathiamoorthy, Y . Chen, R. Mazumder, L. Hong, and E. Chi, “Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning,”Advances in Neural Information Processing Systems, vol. 34, pp. 29 335–29 347, 2021

  61. [61]

    Modeling task relationships in multi-task learning with multi-gate mixture-of-experts,

    J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi, “Modeling task relationships in multi-task learning with multi-gate mixture-of-experts,” inProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2018, pp. 1930–1939

  62. [62]

    On the representation collapse of sparse mixture of experts,

    Z. Chi, L. Dong, S. Huang, D. Dai, S. Ma, B. Patra, S. Singhal, P. Bajaj, X. Song, X.-L. Maoet al., “On the representation collapse of sparse mixture of experts,”Advances in Neural Information Processing Systems, vol. 35, pp. 34 600–34 613, 2022

  63. [63]

    Fastmoe: A fast mixture-of-expert training system,

    J. He, J. Qiu, A. Zeng, Z. Yang, J. Zhai, and J. Tang, “Fastmoe: A fast mixture-of-expert training system,”arXiv preprint arXiv:2103.13262, 2021

  64. [64]

    Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,

    S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y . Aminabadi, A. A. Awan, J. Rasley, and Y . He, “Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,” inInternational conference on machine learning. PMLR, 2022, pp. 18 332–18 346

  65. [65]

    Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,

    D. Yu, L. Shen, H. Hao, W. Gong, H. Wu, J. Bian, L. Dai, and H. Xiong, “Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,”IEEE Transactions on Services Computing, vol. 17, no. 5, pp. 2626–2639, 2024

  66. [66]

    Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,

    J. He, J. Zhai, T. Antunes, H. Wang, F. Luo, S. Shi, and Q. Li, “Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,” inProceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, 2022, pp. 120–134

  67. [67]

    A hybrid tensor-expert-data parallelism approach to optimize mixture-of- experts training,

    S. Singh, O. Ruwase, A. A. Awan, S. Rajbhandari, Y . He, and A. Bhatele, “A hybrid tensor-expert-data parallelism approach to optimize mixture-of- experts training,” inProceedings of the 37th International Conference on Supercomputing, 2023, pp. 203–214

  68. [68]

    Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

    X. Nie, X. Miao, Z. Wang, Z. Yang, J. Xue, L. Ma, G. Cao, and B. Cui, “Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,”Proceedings of the ACM on Management of Data, vol. 1, no. 1, pp. 1–19, 2023

  69. [69]

    {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,

    M. Zhai, J. He, Z. Ma, Z. Zong, R. Zhang, and J. Zhai, “ {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,” in2023 USENIX Annual Technical Conference (USENIX ATC 23), 2023, pp. 961–975

  70. [70]

    Gshard: Scaling giant models with conditional computation and automatic sharding,

    D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,” inInternational Conference on Learning Representations, 2021, pp. 1–14

  71. [71]

    Uni-moe: Scaling unified multimodal llms with mixture of experts,

    Y . Li, S. Jiang, B. Hu, L. Wang, W. Zhong, W. Luo, L. Ma, and M. Zhang, “Uni-moe: Scaling unified multimodal llms with mixture of experts,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. 18 IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE

  72. [72]

    Moe-adapters++: Towards more efficient continual learning of vision-language models via dynamic mixture-of-experts adapters,

    J. Yu, Z. Huang, Y . Zhuge, L. Zhang, P. Hu, D. Wang, H. Lu, and Y . He, “Moe-adapters++: Towards more efficient continual learning of vision-language models via dynamic mixture-of-experts adapters,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  73. [73]

    Boosting continual learning of vision-language models via mixture-of-experts adapters,

    J. Yu, Y . Zhuge, L. Zhang, P. Hu, D. Wang, H. Lu, and Y . He, “Boosting continual learning of vision-language models via mixture-of-experts adapters,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 23 219–23 230

  74. [74]

    Efficient face forgery detection with mixture of experts,

    Y . Kong, X. Lu, J. Shen, L. Liu, and B. Chen, “Efficient face forgery detection with mixture of experts,” inEuropean Conference on Computer Vision, 2022, pp. 1–12

  75. [75]

    Moead: A parameter-efficient model for multi-class anomaly detection,

    S. Meng, W. Meng, Q. Zhou, S. Li, W. Hou, and S. He, “Moead: A parameter-efficient model for multi-class anomaly detection,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 345–361

  76. [76]

    Learning heterogeneous mixture of scene experts for large-scale neural radiance fields,

    Z. Mi, P. Yin, X. Xiao, and D. Xu, “Learning heterogeneous mixture of scene experts for large-scale neural radiance fields,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  77. [77]

    Mocae: Mixture of calibrated experts significantly improves object detection,

    K. Oksuz, S. Kalkan, and E. Akbas, “Mocae: Mixture of calibrated experts significantly improves object detection,”arXiv preprint arXiv:2309.14976, 2023

  78. [78]

    Moe-gs: Mixture of experts for dynamic gaussian splatting,

    I.-H. Jin, H. Mun, J. Kim, K. Yun, and K. Kong, “Moe-gs: Mixture of experts for dynamic gaussian splatting,” inInternational Conference on Learning Representations, 2026

  79. [79]

    3d gaussian splatting as markov chain monte carlo,

    S. Kheradmand, D. Rebain, G. Sharma, W. Sun, Y .-C. Tseng, H. Isack, A. Kar, A. Tagliasacchi, and K. M. Yi, “3d gaussian splatting as markov chain monte carlo,”Advances in Neural Information Processing Systems, vol. 37, pp. 80 965–80 986, 2024

  80. [80]

    Distilling the knowledge in a neural network,

    G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,”arXiv preprint arXiv:1503.02531, vol. 1503, pp. 1–9, 2015

Showing first 80 references.