Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Advances in Radiance Field for Dynamic Scene: From Neural Field to Gaussian Field

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This survey of over 200 papers argues that every method for reconstructing moving scenes fits one framework: a static reference space plus a motion model, with the number of reference frames setting the trade-off between detail and cost.

desk verdict A useful, broad survey whose unified-framework claim omits factorization (one of its own categories); worth refereeing after fixes. read the letter →

arxiv 2505.10049 v2 pith:5QBILHJ4 submitted 2025-05-15 cs.CV

classification cs.CV
keywords motionrepresentationdynamicscenesneuralradiancefield3Dgaussiansplattingdeformationsceneflownovelviewsynthesis4Dreconstruction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey of more than 200 papers tries to establish that the literature on reconstructing moving scenes with radiance fields is one coherent design space rather than an unstructured collection of approaches. Its central claim is that every method, from implicit neural radiance fields to explicit Gaussian primitives, can be read as a static reference space plus a motion model, with approaches differing mainly in how many reference frames anchor the motion. The paper organizes motion into four types (rigid, articulated, non-rigid, hybrid) and five representation paradigms, and argues that the shift to explicit Gaussian primitives is what unlocked real-time rendering, dense tracking, and editing. If the framework holds, it gives newcomers a single map of a fast-moving field and shows experienced researchers where design combinations remain unfilled.

What carries the argument

The organizing mechanism is the reference-frame spectrum introduced in the paper's unified framework, formalized by the point-motion map $x_t = T_\theta(x_{t-1}; \pi(t))$, which sends any 3D point to its position at a later time through a transformation conditioned on a temporal code. The spectrum spans two base representations: NeRF, the implicit neural field mapping position and view direction to color and density, and 3D Gaussian Splatting (3DGS), explicit collections of anisotropic Gaussian primitives rendered by splatting. The framework classifies every method by how many static reference frames anchor this transformation - one canonical space, multiple keyframes, consecutive-frame flow, per-frame 4D spacetime, or per-point trajectories - and by a motion-type taxonomy (rigid, articulated, non-rigid, hybrid). The survey's tables then classify each of the 200+ papers along these axes together with the auxiliary information (depth, segmentation, optical flow, data-driven priors) and regularizers (smoothness, rigidity, volume preservation) they employ.

What would settle it

Take a random sample of roughly twenty recent dynamic-scene papers and check each against the taxonomy: if a well-known method resists placement in one of the four motion types and one of the five representation paradigms, or if a table entry contradicts what the method's own paper claims, the unified framework would be incomplete rather than universal.

Watch

Extended reading notes

Core claim

The paper's central organizational claim is that any dynamic scene can be conceptualized as a static reference space coupled with an appropriate motion representation, and that all reconstruction methods differ in how many reference frames they use. A single canonical space with a learned deformation field covers rigid, articulated, and mildly non-rigid scenes; several keyframes serve as local references when one global space fails; reducing the window to two frames yields frame-to-frame flow fields; letting each frame be its own reference gives full 4D spacetime optimization, where the fourth dimension is time; and at the finest granularity, per-point tracking builds continuous trajectories across the whole sequence. The survey claims that finer granularity captures more temporal detail at higher computational cost, and that hybrid methods - structured coarse motion plus a neural residual - tend to win on real scenes with mixed motion. Alongside this spectrum, the paper maps the field's trajectory from implicit MLP fields to explicit 3D Gaussian primitives and credits that shift with enabling real-time rendering, dense long-term tracking, and object-level editing.

Load-bearing premise

The survey's map of the field depends on its corpus of over 200 papers being complete and representative and on its taxonomy being applied consistently, but it reports no auditable search protocol or inclusion criteria.

Editorial extensions

If this is right

  • A reader encountering a new dynamic-scene method can place it on the reference-frame spectrum and immediately read off the expected trade-off between temporal detail and computational cost.
  • Hybrid representations that layer a structured coarse motion - rigid or articulated - over a neural residual field are presented as the most effective pattern for real scenes with mixed motion, because they keep interpretability while capturing fine deformation.
  • The shift from implicit MLP radiance fields to explicit Gaussian primitives is framed as the enabler of real-time rendering, dense long-term tracking, and part-level editing, at the price of higher memory use.
  • Auxiliary information such as depth, segmentation, optical flow, and data-driven priors is treated as a load-bearing component of monocular reconstruction, without which the motion and appearance ambiguities are underdetermined.
  • The open challenges the survey names - editing, scalability to long videos and large spaces, reconstruction by generation, and LLM-guided semantic priors - are the directions where it expects the field's next advances.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The spectrum predicts that an adaptive method choosing its reference-frame granularity region by region - canonical where motion is small, flow where it is large - should beat any fixed-paradigm method; the paper argues for the spectrum's usefulness but does not test this construction.
  • The framework makes canonical-space and flow-field methods endpoints of one continuum rather than rivals; a published method that interpolates between them as sequence length grows would directly test the framework's predictive value.
  • By the paper's own logic, the analogue of the SMPL template for arbitrary object categories is a foundation-model semantic field: data-driven semantic features should supply the priors that hand-built kinematic trees supplied for humans, enabling template-free articulated reconstruction at scale.
  • The taxonomy yields a quantitative corollary the survey does not state: hybrid-classified methods in the tables should outperform single-paradigm methods on benchmarks with mixed rigid and non-rigid motion, which a reader could verify by aggregating the reported metrics of the cited papers.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey reviews over 200 papers on dynamic scene representation and reconstruction using neural radiance fields and 3D Gaussian splatting. It proposes a taxonomy of motion types (rigid, articulated, non-rigid, hybrid) and of motion representation paradigms (4D spacetime, canonical space with deformation fields, flow fields, point tracking, and factorization), and organizes reconstruction methods into tables by motion type. It further discusses auxiliary information, regularization, and future trends, and claims to provide a unified representational framework (Sec. 4.6, Fig. 5) that relates methods by the number of reference frames used. The paper also maintains an active online repository of methods and implementations.

Significance. If the survey's claims hold, it would serve as a valuable entry point and reference for researchers in dynamic scene reconstruction, bridging implicit neural fields and explicit Gaussian primitives. The paper's strengths include a broad corpus of methods, a structured taxonomy of motion types and representation paradigms, useful summary tables (Tables 1–4), and an explicit discussion of regularization and auxiliary information. However, the central claim of a "unified representational framework" is weakened by the omission of one of the paper's own categories (factorization) from that framework, and the lack of a systematic search protocol reduces confidence in the survey's completeness and in its "definitive reference" claim. The paper's equations (1)–(20) are standard formulations from the cited literature, and the taxonomy is internally consistent apart from the noted gap.

major comments (3)
  1. [Sec. 4.6, Fig. 5, Sec. 3.2.5, Table 3] The unified framework proposed in Sec. 4.6 and illustrated in Fig. 5 organizes methods along a spectrum defined by the number of reference frames (single canonical space, multiple keyframes, two-frame flow fields, per-frame 4D spacetime, per-point tracking). This framework omits factorization, which is one of the five motion-representation paradigms defined in Sec. 3.2.5 and which occupies a substantial block in Table 3 (e.g., Hexplane, K-Planes, 4K4D, NPGs, DynMF). Consequently, the claim in the Abstract and Sec. 1.3 that the survey organizes diverse methodological approaches under a unified representational framework is internally incomplete: a major category of the surveyed literature is not placed on the proposed organizing axis. The authors should either extend the framework to include factorization (e.g., as a complementary dimension concerned with how the motion field is decomposed rather than how many reference frames are used) or explicitly explain how factorization relates to the reference-frame spectrum.
  2. [Abstract, Sec. 1.3] The paper claims to provide a "systematic analysis of over 200 papers" and to be a "definitive reference," but it reports no systematic search protocol, no inclusion or exclusion criteria, no database or time-window specification, and no audit trail linking each table entry to search decisions. Without such methodology, the completeness and representativeness of the corpus cannot be assessed, and the assertion of a definitive reference is not supported. The authors should describe their literature collection process in the paper or in a supplementary document, including how papers were selected and how the tables were populated.
  3. [Sec. 3.1.4, Eq. (9)] Equation (9), which formalizes the decomposition of hybrid motion into a coarse global transformation and a fine non-rigid residual, is corrupted in the manuscript: it contains a long run of repeated tokens (e.g., "rl", "mo", "r") that makes the formula unreadable. Because hybrid motion is a core concept in the taxonomy and the equation is meant to provide the mathematical basis for the discussion that follows, this corruption is a load-bearing presentation error and must be fixed.
minor comments (5)
  1. [Sec. 3.3 heading] The heading "Disscussion" should be "Discussion".
  2. [Tables 1–4] The column header "Auxilary" is misspelled; it should be "Auxiliary".
  3. [Author affiliation, Abstract footnote] The author affiliation contains "T ao" where "Tao" is intended.
  4. [Sec. 2.1.1] The term "Lidar" is inconsistently capitalized; use "LiDAR" throughout.
  5. [Sec. 2.1.2 and elsewhere] The phrase "casual video capture" is used where "casually captured video" or "casual capture" would be clearer; the meaning is understandable but the phrasing is informal for a survey.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey proposes an organizational taxonomy but derives no result from its own inputs.

full rationale

This paper is an expository survey, not a derivation: it makes no numerical or theoretical prediction that could reduce to a fitted parameter or to an assumed ansatz. Its central claims are the scope of the surveyed corpus (over 200 papers) and the usefulness of a proposed taxonomy. I checked each circularity pattern: there is no self-definitional equation, no fitted input relabeled as a prediction, no load-bearing self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation. The closest potential issue is that the 'unified representational framework' of Sec. 4.6 and Fig. 5 is the authors' own construction; but a taxonomy is an organizing premise, not an input-output cycle, so it cannot be circular in the sense of deriving its conclusion from its own definition. The one internal tension visible in the text, namely that factorization methods (Sec. 3.2.5, Table 3) are not assigned a position on the reference-frame spectrum discussed in Sec. 4.6, is an organizational omission or incompleteness, not a circular derivation. Self-citations are incidental (e.g., [215] as an optical-flow prior) and do not carry the paper's argument. I therefore find no specific circular step and assign score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities appear because the paper is a survey. The load-bearing assumptions are about corpus completeness, the validity of the authors' taxonomy, and the accuracy of the table entries. These assumptions are plausible but not demonstrated with a search protocol, benchmark, or public audit trail.

assumptions (3)
  • domain assumption The surveyed 200+ papers are representative and correctly categorized
    The survey's conclusions and 'definitive reference' claim rest on coverage; no search protocol or audit data is provided.
  • ad hoc to paper The proposed taxonomy of motion types and representation paradigms is a valid organizing scheme
    This is the authors' own framework, stated in Sections 1 and 4, and is not derived from prior work or validated against alternative taxonomies.
  • domain assumption Prior work is accurately represented by the summary tables
    Tables 1-4 assign each method to a motion type, representation style, and auxiliary information set; these assignments are not independently verified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Advances in Radiance Field for Dynamic Scene: From Neural Field to Gaussian Field." pith.science (2026). https://pith.science/paper/5QBILHJ4

@misc{pith2026250510049,
  author       = {Pith},
  title        = {Pith review of: Advances in Radiance Field for Dynamic Scene: From Neural Field to Gaussian Field},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5QBILHJ4}},
  note         = {Machine review of arXiv:2505.10049}
}
read the original abstract

Dynamic scene representation and reconstruction have undergone transformative advances in recent years, catalyzed by breakthroughs in neural radiance fields and 3D Gaussian splatting techniques. While initially developed for static environments, these methodologies have rapidly evolved to address the complexities inherent in 4D dynamic scenes through an expansive body of research. Coupled with innovations in differentiable volumetric rendering, these approaches have significantly enhanced the quality of motion representation and dynamic scene reconstruction, thereby garnering substantial attention from the computer vision and graphics communities. This survey presents a systematic analysis of over 200 papers focused on dynamic scene representation using radiance field, spanning the spectrum from implicit neural representations to explicit Gaussian primitives. We categorize and evaluate these works through multiple critical lenses: motion representation paradigms, reconstruction techniques for varied scene dynamics, auxiliary information integration strategies, and regularization approaches that ensure temporal consistency and physical plausibility. We organize diverse methodological approaches under a unified representational framework, concluding with a critical examination of persistent challenges and promising research directions. By providing this comprehensive overview, we aim to establish a definitive reference for researchers entering this rapidly evolving field while offering experienced practitioners a systematic understanding of both conceptual principles and practical frontiers in dynamic scene reconstruction.

Figures

Figures reproduced from arXiv: 2505.10049 by the authors.

Figure 1
Figure 1. Survey at A Glance. (a) Introduction and Foundation. We trace the evolution from static to dynamic scene representation, highlighting the challenges of jointly modeling motion, geometry, and appearance using radiance fields. (b) Motion Representation. We categorize motion patterns and their representation paradigms, examining how they enable complex motion modeling while addressing inherent limitations. (c) Scene Re… view at source ↗
Figure 2
Figure 2. Roadmap of Dynamic Scenes in Radiance Fields. This chronological timeline illustrates the evolution of the field, organizing works into methodological clusters based on their representation paradigms. The representative or first work within each cluster appears in black with accompanying paradigm illustrations, while the dates of remaining works may vary within clusters. Seminal contributions that significantly adva… view at source ↗
Figure 3
Figure 3. A 2D illustration of various motion types. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Illustration of typical motion representation methods. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: We propose a unified framework to encapsulate [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Relative motion speed between camera and objects [PITH_FULL_IMAGE:figures/full_fig_p021_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

Reference graph

Works this paper leans on

235 extracted references · 69 canonical work pages · cited by 1 Pith paper

  1. [1]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Commun. ACM, vol. 65, no. 1, 2021

  2. [2]

    3d gaussian splatting for real-time radiance field ren- dering,

    B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3d gaussian splatting for real-time radiance field ren- dering,” ACM TOG, vol. 42, no. 4, 2023

  3. [3]

    Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision,

    M. Niemeyer, L. Mescheder, M. Oechsle, and A. Geiger, “Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision,” in CVPR, 2020

  4. [4]

    Instant neural graphics primitives with a multiresolution hash encoding,

    T. Müller, A. Evans, C. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,” ACM TOG, vol. 41, no. 4, 2022

  5. [5]

    Tensorf: Tensorial radiance fields,

    A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su, “Tensorf: Tensorial radiance fields,” in ECCV, Springer, 2022

  6. [6]

    Mip-nerf: A mul- tiscale representation for anti-aliasing neural radiance fields,

    J. T. Barron, B. Mildenhall, M. Tancik, P . Hedman, R. Martin-Brualla, and P . P . Srinivasan, “Mip-nerf: A mul- tiscale representation for anti-aliasing neural radiance fields,” in ICCV, 2021

  7. [7]

    Mip- splatting: Alias-free 3d gaussian splatting,

    Z. Yu, A. Chen, B. Huang, T. Sattler, and A. Geiger, “Mip- splatting: Alias-free 3d gaussian splatting,” in CVPR, 2024

  8. [8]

    D-nerf: Neural radiance fields for dynamic scenes,

    A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno- Noguer, “D-nerf: Neural radiance fields for dynamic scenes,” in CVPR, 2021

Show all 235 references
  1. [9]

    Nerfies: Deformable neural radiance fields,

    K. Park, U. Sinha, J. T. Barron, S. Bouaziz, D. B. Goldman, S. M. Seitz, and R. Martin-Brualla, “Nerfies: Deformable neural radiance fields,” in ICCV, 2021

  2. [10]

    Hy- pernerf: a higher-dimensional representation for topolog- ically varying neural radiance fields,

    K. Park, U. Sinha, P . Hedman, J. T. Barron, S. Bouaziz, D. B. Goldman, R. Martin-Brualla, and S. M. Seitz, “Hy- pernerf: a higher-dimensional representation for topolog- ically varying neural radiance fields,” ACM TOG, vol. 40, no. 6, 2021

  3. [11]

    4d gaussian splatting for real-time dynamic scene rendering,

    G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” in CVPR, 2024

  4. [12]

    Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,

    Z. Yang, X. Gao, W. Zhou, S. Jiao, Y. Zhang, and X. Jin, “Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction,” in CVPR, 2024

  5. [13]

    Neural scene flow fields for space-time view synthesis of dynamic scenes,

    Z. Li, S. Niklaus, N. Snavely, and O. Wang, “Neural scene flow fields for space-time view synthesis of dynamic scenes,” in CVPR, 2021

  6. [14]

    K-planes: Explicit radiance fields in space, time, and appearance,

    S. Fridovich-Keil, G. Meanti, F. R. Warburg, B. Recht, and A. Kanazawa, “K-planes: Explicit radiance fields in space, time, and appearance,” in CVPR, 2023

  7. [15]

    Hexplane: A fast representation for dynamic scenes,

    A. Cao and J. Johnson, “Hexplane: A fast representation for dynamic scenes,” in CVPR, 2023

  8. [16]

    Robust dynamic radiance fields,

    Y.-L. Liu, C. Gao, A. Meuleman, H.-Y. Tseng, A. Saraf, C. Kim, Y.-Y. Chuang, J. Kopf, and J.-B. Huang, “Robust dynamic radiance fields,” in CVPR, 2023

  9. [17]

    Dy- namic 3d gaussians: Tracking by persistent dynamic view synthesis,

    J. Luiten, G. Kopanas, B. Leibe, and D. Ramanan, “Dy- namic 3d gaussians: Tracking by persistent dynamic view synthesis,” in International Conference on 3D Vision (3DV) , 2024

  10. [18]

    Tracking everything ev- erywhere all at once,

    Q. Wang, Y.-Y. Chang, R. Cai, Z. Li, B. Hariharan, A. Holynski, and N. Snavely, “Tracking everything ev- erywhere all at once,” in ICCV, 2023

  11. [19]

    Flow supervision for deformable nerf,

    C. Wang, L. E. MacDonald, L. A. Jeni, and S. Lucey, “Flow supervision for deformable nerf,” in CVPR, 2023

  12. [20]

    Dynibar: Neural dynamic image-based rendering,

    Z. Li, Q. Wang, F. Cole, R. Tucker, and N. Snavely, “Dynibar: Neural dynamic image-based rendering,” in CVPR, 2023

  13. [21]

    Panoptic nerf: 3d-to-2d label transfer for panoptic urban scene segmentation,

    X. Fu, S. Zhang, T. Chen, Y. Lu, L. Zhu, X. Zhou, A. Geiger, and Y. Liao, “Panoptic nerf: 3d-to-2d label transfer for panoptic urban scene segmentation,” in 3DV, 2022

  14. [22]

    Panoptic neural fields: A semantic object-aware neural scene representation,

    A. Kundu, K. Genova, X. Yin, A. Fathi, C. Pantofaru, L. J. Guibas, A. Tagliasacchi, F. Dellaert, and T. Funkhouser, “Panoptic neural fields: A semantic object-aware neural scene representation,” in CVPR, 2022

  15. [23]

    Neu- ral trajectory fields for dynamic novel view synthesis,

    C. Wang, B. Eckart, S. Lucey, and O. Gallo, “Neu- ral trajectory fields for dynamic novel view synthesis,” ArXiv:2105.05994, 2021

  16. [24]

    Optical models for direct volume rendering,

    N. Max, “Optical models for direct volume rendering,” TVCG, vol. 1, no. 2, 1995

  17. [25]

    Neural scene graphs for dynamic scenes,

    J. Ost, F. Mannan, N. Thuerey, J. Knodt, and F. Heide, “Neural scene graphs for dynamic scenes,” in CVPR, 2021

  18. [26]

    Space-time neural irradiance fields for free-viewpoint video,

    W. Xian, J.-B. Huang, J. Kopf, and C. Kim, “Space-time neural irradiance fields for free-viewpoint video,” in CVPR, 2021

  19. [27]

    Real-time pho- torealistic dynamic scene representation and rendering with 4d gaussian splatting,

    Z. Yang, H. Yang, Z. Pan, and L. Zhang, “Real-time pho- torealistic dynamic scene representation and rendering with 4d gaussian splatting,” in The Twelfth ICLR

  20. [28]

    Drivinggaussian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes,

    X. Zhou, Z. Lin, X. Shan, Y. Wang, D. Sun, and M.-H. Yang, “Drivinggaussian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes,” in CVPR, 2024

  21. [29]

    Ewa volume splatting,

    M. Zwicker, H. Pfister, J. Van Baar, and M. Gross, “Ewa volume splatting,” in IEEE Visualization 2001, IEEE, 2001

  22. [30]

    Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,

    S. Peng, Y. Zhang, Y. Xu, Q. Wang, Q. Shuai, H. Bao, and X. Zhou, “Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans,” in CVPR, 2021

  23. [31]

    Hugs: Human gaussian splats,

    M. Kocabas, J.-H. R. Chang, J. Gabriel, O. Tuzel, and A. Ranjan, “Hugs: Human gaussian splats,” in CVPR, 2024

  24. [32]

    Star: Self- supervised tracking and reconstruction of rigid objects in motion with neural rendering,

    W. Yuan, Z. Lv, T. Schmidt, and S. Lovegrove, “Star: Self- supervised tracking and reconstruction of rigid objects in motion with neural rendering,” in CVPR, 2021

  25. [33]

    Smpl: A skinned multi-person linear model,

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “Smpl: A skinned multi-person linear model,” ACM TOG, vol. 34, no. 6, 2015

  26. [34]

    Embodied hands: Modeling and capturing hands and bodies together,

    J. Romero, D. Tzionas, and M. J. Black, “Embodied hands: Modeling and capturing hands and bodies together,” ACM TOG, vol. 36, no. 6, 2017

  27. [35]

    Humannerf: Free- viewpoint rendering of moving people from monocular video,

    C.-Y. Weng, B. Curless, P . P . Srinivasan, J. T. Bar- ron, and I. Kemelmacher-Shlizerman, “Humannerf: Free- viewpoint rendering of moving people from monocular video,” in CVPR, 2022

  28. [36]

    Neural actor: Neural free-view synthesis 16 of human actors with pose control,

    L. Liu, M. Habermann, V . Rudnev, K. Sarkar, J. Gu, and C. Theobalt, “Neural actor: Neural free-view synthesis 16 of human actors with pose control,” ACM TOG , vol. 40, no. 6, 2021

  29. [37]

    Dynamic view synthesis from dynamic monocular video,

    C. Gao, A. Saraf, J. Kopf, and J.-B. Huang, “Dynamic view synthesis from dynamic monocular video,” inICCV, 2021

  30. [38]

    Neural radiance flow for 4d view synthesis and video processing,

    Y. Du, Y. Zhang, H.-X. Yu, J. B. Tenenbaum, and J. Wu, “Neural radiance flow for 4d view synthesis and video processing,” in ICCV, 2021

  31. [39]

    Fast dynamic radiance fields with time-aware neural voxels,

    J. Fang, T. Yi, X. Wang, L. Xie, X. Zhang, W. Liu, M. Nießner, and Q. Tian, “Fast dynamic radiance fields with time-aware neural voxels,” inSIGGRAPH Asia, 2022

  32. [40]

    Nerfplayer: A streamable dynamic scene representation with decomposed neural radiance fields,

    L. Song, A. Chen, Z. Li, Z. Chen, L. Chen, J. Yuan, Y. Xu, and A. Geiger, “Nerfplayer: A streamable dynamic scene representation with decomposed neural radiance fields,” TVCG, vol. 29, no. 5, 2023

  33. [41]

    Neural 3d video synthesis from multi-view video,

    T. Li, M. Slavcheva, M. Zollhoefer, S. Green, C. Lassner, C. Kim, T. Schmidt, S. Lovegrove, M. Goesele, R. New- combe, et al., “Neural 3d video synthesis from multi-view video,” in CVPR, 2022

  34. [42]

    Tensor4d: Efficient neural 4d decomposition for high- fidelity dynamic reconstruction and rendering,

    R. Shao, Z. Zheng, H. Tu, B. Liu, H. Zhang, and Y. Liu, “Tensor4d: Efficient neural 4d decomposition for high- fidelity dynamic reconstruction and rendering,” inCVPR, 2023

  35. [43]

    Gaufre: Gaussian deformation fields for real-time dynamic novel view synthesis,

    Y. Liang, N. Khan, Z. Li, T. Nguyen-Phuoc, D. Lanman, J. Tompkin, and L. Xiao, “Gaufre: Gaussian deformation fields for real-time dynamic novel view synthesis,” in 2025 IEEE/CVF Winter Conference on Applications of Com- puter Vision (WACV), IEEE, 2025

  36. [44]

    3d geometry-aware deformable gaussian splatting for dynamic view synthesis,

    Z. Lu, X. Guo, L. Hui, T. Chen, M. Yang, X. Tang, F. Zhu, and Y. Dai, “3d geometry-aware deformable gaussian splatting for dynamic view synthesis,” in CVPR, 2024

  37. [45]

    Dynamic gaussian marbles for novel view synthesis of casual monocular videos,

    C. Stearns, A. Harley, M. Uy, F. Dubost, F. Tombari, G. Wetzstein, and L. Guibas, “Dynamic gaussian marbles for novel view synthesis of casual monocular videos,” in SIGGRAPH Asia, 2024

  38. [46]

    Periodic vibration gaussian: Dynamic urban scene reconstruction and real-time rendering,

    Y. Chen, C. Gu, J. Jiang, X. Zhu, and L. Zhang, “Periodic vibration gaussian: Dynamic urban scene reconstruction and real-time rendering,” ArXiv:2311.18561, 2023

  39. [47]

    Deformgs: Scene flow in highly deformable scenes for deformable object manipulation,

    B. P . Duisterhof, Z. Mandi, Y. Yao, J.-W. Liu, J. Seiden- schwarz, M. Z. Shou, D. Ramanan, S. Song, S. Birch- field, B. Wen, et al. , “Deformgs: Scene flow in highly deformable scenes for deformable object manipulation,” ArXiv:2312.00583, 2023

  40. [48]

    Gart: Gaussian articulated template models,

    J. Lei, Y. Wang, G. Pavlakos, L. Liu, and K. Daniilidis, “Gart: Gaussian articulated template models,” in CVPR, 2024

  41. [49]

    3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting,

    Z. Qian, S. Wang, M. Mihajlovic, A. Geiger, and S. Tang, “3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting,” in CVPR, 2024

  42. [50]

    Hugs: Holistic urban 3d scene understanding via gaussian splatting,

    H. Zhou, J. Shao, L. Xu, D. Bai, W. Qiu, B. Liu, Y. Wang, A. Geiger, and Y. Liao, “Hugs: Holistic urban 3d scene understanding via gaussian splatting,” in CVPR, 2024

  43. [51]

    Dynamic 3d gaussian fields for urban areas,

    T. Fischer, J. Kulhanek, S. R. Bulò, L. Porzi, M. Pollefeys, and P . Kontschieder, “Dynamic 3d gaussian fields for urban areas,” in NeurIPS, 2024

  44. [52]

    Street gaussians: Mod- eling dynamic urban scenes with gaussian splatting,

    Y. Yan, H. Lin, C. Zhou, W. Wang, H. Sun, K. Zhan, X. Lang, X. Zhou, and S. Peng, “Street gaussians: Mod- eling dynamic urban scenes with gaussian splatting,” in ECCV, 2024

  45. [53]

    A Brief Review on Differentiable Rendering: Recent Advances and Challenges,

    R. Gao and Y. Qi, “A Brief Review on Differentiable Rendering: Recent Advances and Challenges,” Electron- ics, vol. 13, no. 17, 2024

  46. [54]

    Neural fields in visual computing and beyond,

    Y. Xie, T. Takikawa, S. Saito, O. Litany, S. Yan, N. Khan, F. Tombari, J. Tompkin, V . Sitzmann, and S. Sridhar, “Neural fields in visual computing and beyond,” in Com- put. Graph. Forum , vol. 41, 2022

  47. [55]

    Nerf: Neu- ral radiance field in 3d vision, a comprehensive review,

    K. Gao, Y. Gao, H. He, D. Lu, L. Xu, and J. Li, “Nerf: Neu- ral radiance field in 3d vision, a comprehensive review,” ArXiv:221000379, 2022

  48. [56]

    A survey on 3d gaussian splat- ting,

    G. Chen and W. Wang, “A survey on 3d gaussian splat- ting,” CoRR, 2024

  49. [57]

    Recent advances in 3d gaussian splatting,

    T. Wu, Y.-J. Yuan, L.-X. Zhang, J. Yang, Y.-P . Cao, L.- Q. Yan, and L. Gao, “Recent advances in 3d gaussian splatting,” Comput. Vis. Media, vol. 10, no. 4, 2024

  50. [58]

    3d gaussian splatting: Survey, technologies, challenges, and opportunities,

    Y. Bao, T. Ding, J. Huo, Y. Liu, Y. Li, W. Li, Y. Gao, and J. Luo, “3d gaussian splatting: Survey, technologies, challenges, and opportunities,” IEEE T rans. Circuits Syst. Video T echnol., 2025

  51. [59]

    How nerfs and 3d gaussian splatting are reshaping slam: A survey,

    F. Tosi, Y. Zhang, Z. Gong, E. Sandström, S. Mattoccia, M. R. Oswald, and M. Poggi, “How nerfs and 3d gaussian splatting are reshaping slam: A survey,”ArXiv:240213255, vol. 4, 2024

  52. [60]

    Neural radiance fields in the industrial and robotics domain: Applications, research opportunities and use cases,

    E. Šlapak, E. Pardo, M. Dopiriak, T. Maksymyuk, and J. Gazda, “Neural radiance fields in the industrial and robotics domain: Applications, research opportunities and use cases,” Robot. Comput.-Integr. Manuf. , vol. 90, 2024

  53. [61]

    NeRF in robotics: A survey,

    G. Wang, L. Pan, S. Peng, S. Liu, C. Xu, Y. Miao, W. Zhan, M. Tomizuka, M. Pollefeys, and H. Wang, “NeRF in robotics: A survey,” ArXiv:240501333, 2024

  54. [62]

    Human 3d avatar modeling with implicit neural representation: A brief survey,

    M. Sun, D. Yang, D. Kou, Y. Jiang, W. Shan, Z. Yan, and L. Zhang, “Human 3d avatar modeling with implicit neural representation: A brief survey,” in ISCPS, 2022

  55. [63]

    Benchmarking neural radiance fields for autonomous robots: An overview,

    Y. Ming, X. Yang, W. Wang, Z. Chen, J. Feng, Y. Xing, and G. Zhang, “Benchmarking neural radiance fields for autonomous robots: An overview,” EAAI, vol. 140, 2025

  56. [64]

    Recent Trends in 3D Reconstruc- tion of General Non-Rigid Scenes,

    R. Yunus, J. E. Lenssen, M. Niemeyer, Y. Liao, C. Rupprecht, C. Theobalt, G. Pons-Moll, J.-B. Huang, V . Golyanik, and E. Ilg, “Recent Trends in 3D Reconstruc- tion of General Non-Rigid Scenes,” 2024

  57. [65]

    State of the Art in Dense Monocular Non-Rigid 3D Reconstruction,

    E. Tretschk, N. Kairanda, M. BR, R. Dabral, A. Ko- rtylewski, B. Egger, M. Habermann, P . Fua, C. Theobalt, and V . Golyanik, “State of the Art in Dense Monocular Non-Rigid 3D Reconstruction,” in CGF, vol. 42, 2023

  58. [66]

    3d human avatar reconstruction with neural fields: A recent survey,

    M. Gu, J. Li, Y. Wu, H. Luo, J. Zheng, and X. Bai, “3d human avatar reconstruction with neural fields: A recent survey,” Image and Vision Computing , vol. 154, 2025

  59. [67]

    Survey on modeling of human-made articulated objects,

    J. Liu, M. Savva, and A. Mahdavi-Amiri, “Survey on modeling of human-made articulated objects,” in Com- puter Graphics Forum , Wiley Online Library, 2024

  60. [68]

    Monocular dynamic view synthesis: A reality check,

    H. Gao, R. Li, S. Tulsiani, B. Russell, and A. Kanazawa, “Monocular dynamic view synthesis: A reality check,” NeurIPS, vol. 35, 2022

  61. [69]

    Ar- ticulated and elastic non-rigid motion: A review,

    J. K. Aggarwal, Q. Cai, W. Liao, and B. Sabata, “Ar- ticulated and elastic non-rigid motion: A review,” in Proceedings of 1994 IEEE Workshop on Motion of Non-rigid and Articulated Objects , IEEE, 1994

  62. [70]

    Neural Radiance Field in Autonomous Driving: A Survey,

    L. He, L. Li, W. Sun, Z. Han, Y. Liu, S. Zheng, J. Wang, and K. Li, “Neural Radiance Field in Autonomous Driving: A Survey,” ArXiv:240413816, 2024

  63. [71]

    Suds: Scalable urban dynamic scenes,

    H. Turki, J. Y. Zhang, F. Ferroni, and D. Ramanan, “Suds: Scalable urban dynamic scenes,” in CVPR, 2023

  64. [72]

    OmniRe: Omni Urban Scene Reconstruction,

    Z. Chen, J. Yang, J. Huang, R. de Lutio, J. M. Esturo, B. Ivanovic, O. Litany, Z. Gojcic, S. Fidler, M. Pavone, L. Song, and Y. Wang, “OmniRe: Omni Urban Scene Reconstruction,” 2024

  65. [73]

    The space of human body shapes: reconstruction and parameterization from range scans,

    B. Allen, B. Curless, and Z. Popovi´ c, “The space of human body shapes: reconstruction and parameterization from range scans,” ACM TOG, vol. 22, no. 3, 2003

  66. [74]

    Expressive body capture: 3d hands, face, and body from a single image,

    G. Pavlakos, V . Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, “Expressive body capture: 3d hands, face, and body from a single image,” in CVPR, 2019

  67. [75]

    Npc: Neural point characters from video,

    S.-Y. Su, T. Bagautdinov, and H. Rhodin, “Npc: Neural point characters from video,” in ICCV, 2023

  68. [76]

    Tava: Template-free animatable volu- metric actors,

    R. Li, J. Tanke, M. Vo, M. Zollhofer, J. Gall, A. Kanazawa, and C. Lassner, “Tava: Template-free animatable volu- metric actors,” in ECCV, 2022

  69. [77]

    Rec-mv: Reconstructing 3d dynamic cloth from monoc- 17 ular videos,

    L. Qiu, G. Chen, J. Zhou, M. Xu, J. Wang, and X. Han, “Rec-mv: Reconstructing 3d dynamic cloth from monoc- 17 ular videos,” in CVPR, 2023

  70. [78]

    Selfrecon: Self reconstruction your digital avatar from monocular video,

    B. Jiang, Y. Hong, H. Bao, and J. Zhang, “Selfrecon: Self reconstruction your digital avatar from monocular video,” in CVPR, 2022

  71. [79]

    " the plenoptic function and the elements of early vision

    E. ADELSON, “" the plenoptic function and the elements of early vision", computational models of visual,”Process- ing, Chap. 1, Edited by M. Landy and JA Movshon , 1991

  72. [80]

    Temporal interpolation is all you need for dynamic neural radiance fields,

    S. Park, M. Son, S. Jang, Y. C. Ahn, J.-Y. Kim, and N. Kang, “Temporal interpolation is all you need for dynamic neural radiance fields,” in CVPR, 2023

  73. [81]

    Neural volumes: Learning dynamic renderable volumes from images,

    S. Lombardi, T. Simon, J. Saragih, G. Schwartz, A. Lehrmann, and Y. Sheikh, “Neural volumes: Learning dynamic renderable volumes from images,” TOG, vol. 38, no. 4, 2019

  74. [82]

    D^ 2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video,

    T. Wu, F. Zhong, A. Tagliasacchi, F. Cole, and C. Oztireli, “D^ 2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video,” NeurIPS, vol. 35, 2022

  75. [83]

    MoDGS: Dynamic Gaussian Splat- ting from Causually-captured Monocular Videos,

    Q. Liu, Y. Liu, J. Wang, X. Lv, P . Wang, W. Wang, and J. Hou, “MoDGS: Dynamic Gaussian Splat- ting from Causually-captured Monocular Videos,” ArXiv:240600434, 2024

  76. [84]

    Neural surface reconstruction of dynamic scenes with monocular rgb-d camera,

    H. Cai, W. Feng, X. Feng, Y. Wang, and J. Zhang, “Neural surface reconstruction of dynamic scenes with monocular rgb-d camera,” NeurIPS, vol. 35, 2022

  77. [85]

    Occupancy flow: 4d reconstruction by learning particle dynamics,

    M. Niemeyer, L. Mescheder, M. Oechsle, and A. Geiger, “Occupancy flow: 4d reconstruction by learning particle dynamics,” in ICCV, 2019

  78. [86]

    Orb: An efficient alternative to sift or surf,

    E. Rublee, V . Rabaud, K. Konolige, and G. Bradski, “Orb: An efficient alternative to sift or surf,” inICCV, Ieee, 2011

  79. [87]

    An iterative image registra- tion technique with an application to stereo vision,

    B. D. Lucas and T. Kanade, “An iterative image registra- tion technique with an application to stereo vision,” in IJCAI, vol. 2, 1981

  80. [88]

    Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,

    D. Sun, X. Yang, M.-Y. Liu, and J. Kautz, “Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,” in CVPR, 2018

  81. [89]

    Raft: Recurrent all-pairs field trans- forms for optical flow,

    Z. Teed and J. Deng, “Raft: Recurrent all-pairs field trans- forms for optical flow,” in ECCV, Springer, 2020

  82. [90]

    Particle video: Long-range motion estimation using point trajectories,

    P . Sand and S. Teller, “Particle video: Long-range motion estimation using point trajectories,” IJCV, vol. 80, 2008

  83. [91]

    Tap-vid: A benchmark for tracking any point in a video,

    C. Doersch, A. Gupta, L. Markeeva, A. Recasens, L. Smaira, Y. Aytar, J. Carreira, A. Zisserman, and Y. Yang, “Tap-vid: A benchmark for tracking any point in a video,” NeurIPS, vol. 35, 2022

  84. [92]

    Particle video revisited: Tracking through occlusions using point trajectories,

    A. W. Harley, Z. Fang, and K. Fragkiadaki, “Particle video revisited: Tracking through occlusions using point trajectories,” in ECCV, Springer, 2022

  85. [93]

    Efficient geometry-aware 3d generative adversar- ial networks,

    E. R. Chan, C. Z. Lin, M. A. Chan, K. Nagano, B. Pan, S. De Mello, O. Gallo, L. J. Guibas, J. Tremblay, S. Khamis, et al. , “Efficient geometry-aware 3d generative adversar- ial networks,” in CVPR, 2022

  86. [94]

    Mars: An instance- aware, modular and realistic simulator for autonomous driving,

    Z. Wu, T. Liu, L. Luo, Z. Zhong, J. Chen, H. Xiao, C. Hou, H. Lou, Y. Chen, R. Yang, et al. , “Mars: An instance- aware, modular and realistic simulator for autonomous driving,” in CAAI, 2023

  87. [95]

    Unisim: A neural closed- loop sensor simulator,

    Z. Yang, Y. Chen, J. Wang, S. Manivasagam, W.-C. Ma, A. J. Yang, and R. Urtasun, “Unisim: A neural closed- loop sensor simulator,” in CVPR, 2023

  88. [96]

    S-nerf: Neural radiance fields for street views,

    Z. Xie, J. Zhang, W. Li, F. Zhang, and L. Zhang, “S-nerf: Neural radiance fields for street views,” in ICLR, 2023

  89. [97]

    Multi-level neural scene graphs for dynamic urban environments,

    T. Fischer, L. Porzi, S. R. Bulo, M. Pollefeys, and P . Kontschieder, “Multi-level neural scene graphs for dynamic urban environments,” in CVPR, 2024

  90. [98]

    Neurad: Neural render- ing for autonomous driving,

    A. Tonderski, C. Lindström, G. Hess, W. Ljungbergh, L. Svensson, and C. Petersson, “Neurad: Neural render- ing for autonomous driving,” in CVPR, 2024

  91. [99]

    Autosplat: Constrained gaussian splatting for autonomous driving scene reconstruction,

    M. Khan, H. Fazlali, D. Sharma, T. Cao, D. Bai, Y. Ren, and B. Liu, “Autosplat: Constrained gaussian splatting for autonomous driving scene reconstruction,” CoRR, 2024

  92. [100]

    Prosgnerf: Progressive dynamic neural scene graph with frequency modulated auto-encoder in urban scenes,

    T. Deng, S. Liu, X. Wang, Y. Liu, D. Wang, and W. Chen, “Prosgnerf: Progressive dynamic neural scene graph with frequency modulated auto-encoder in urban scenes,” ArXiv:231209076, 2023

  93. [101]

    A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose,

    S.-Y. Su, F. Yu, M. Zollhöfer, and H. Rhodin, “A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose,” NeurIPS, vol. 34, 2021

  94. [102]

    Neural articulated radiance field,

    A. Noguchi, X. Sun, S. Lin, and T. Harada, “Neural articulated radiance field,” in ICCV, 2021

  95. [103]

    Animatable neural radiance fields for mod- eling dynamic human bodies,

    S. Peng, J. Dong, Q. Wang, S. Zhang, Q. Shuai, X. Zhou, and H. Bao, “Animatable neural radiance fields for mod- eling dynamic human bodies,” in ICCV, 2021

  96. [104]

    Neural human performer: Learning generalizable radiance fields for human performance rendering,

    Y. Kwon, D. Kim, D. Ceylan, and H. Fuchs, “Neural human performer: Learning generalizable radiance fields for human performance rendering,” NeurIPS, vol. 34, 2021

  97. [105]

    Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition,

    C. Guo, T. Jiang, X. Chen, J. Song, and O. Hilliges, “Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition,” in CVPR, 2023

  98. [106]

    Mono- human: Animatable human neural field from monocular video,

    Z. Yu, W. Cheng, X. Liu, W. Wu, and K.-Y. Lin, “Mono- human: Animatable human neural field from monocular video,” in CVPR, 2023

  99. [107]

    Expressive whole- body 3D gaussian avatar,

    G. Moon, T. Shiratori, and S. Saito, “Expressive whole- body 3D gaussian avatar,” in ECCV, 2024

  100. [108]

    Gauhuman: Articulated gaus- sian splatting from monocular human videos,

    S. Hu, T. Hu, and Z. Liu, “Gauhuman: Articulated gaus- sian splatting from monocular human videos,” in CVPR, 2024

  101. [109]

    Animatable gaus- sians: Learning pose-dependent gaussian maps for high- fidelity human avatar modeling,

    Z. Li, Z. Zheng, L. Wang, and Y. Liu, “Animatable gaus- sians: Learning pose-dependent gaussian maps for high- fidelity human avatar modeling,” in CVPR, 2024

  102. [110]

    Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians,

    L. Hu, H. Zhang, Y. Zhang, B. Zhou, B. Liu, S. Zhang, and L. Nie, “Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians,” in CVPR, 2024

  103. [111]

    Ash: Animatable gaussian splats for efficient and photoreal human rendering,

    H. Pang, H. Zhu, A. Kortylewski, C. Theobalt, and M. Habermann, “Ash: Animatable gaussian splats for efficient and photoreal human rendering,” inCVPR, 2024

  104. [112]

    Moda: Modeling deformable 3d objects from casual videos,

    C. Song, J. Wei, T. Chen, Y. Chen, C.-S. Foo, F. Liu, and G. Lin, “Moda: Modeling deformable 3d objects from casual videos,” IJCV, 2024

  105. [113]

    Lisa: Learning implicit shape and appearance of hands,

    E. Corona, T. Hodan, M. Vo, F. Moreno-Noguer, C. Sweeney, R. Newcombe, and L. Ma, “Lisa: Learning implicit shape and appearance of hands,” in CVPR, 2022

  106. [114]

    Hand avatar: Free- pose hand animation and rendering from monocular video,

    X. Chen, B. Wang, and H.-Y. Shum, “Hand avatar: Free- pose hand animation and rendering from monocular video,” in CVPR, 2023

  107. [115]

    Livehand: Real-time and photorealistic neural hand rendering,

    A. Mundra, J. Wang, M. Habermann, C. Theobalt, M. El- gharib, et al. , “Livehand: Real-time and photorealistic neural hand rendering,” in ICCV, 2023

  108. [116]

    Gaussian- Hand: Real-time 3D gaussian rendering for hand avatar animation,

    L. Zhao, X. Lu, R. Fan, S. K. Im, and L. Wang, “Gaussian- Hand: Real-time 3D gaussian rendering for hand avatar animation,” TVCG, 2024

  109. [117]

    Manus: Markerless grasp capture using articulated 3d gaussians,

    C. Pokhariya, I. N. Shah, A. Xing, Z. Li, K. Chen, A. Sharma, and S. Sridhar, “Manus: Markerless grasp capture using articulated 3d gaussians,” in CVPR, 2024

  110. [118]

    Artemis: articulated neural pets with appearance and motion synthesis,

    H. Luo, T. Xu, Y. Jiang, C. Zhou, Q. Qiu, Y. Zhang, W. Yang, L. Xu, and J. Yu, “Artemis: articulated neural pets with appearance and motion synthesis,” ACM TOG, vol. 41, no. 4, 2022

  111. [119]

    Banmo: Building animatable 3d neural mod- els from many casual videos,

    G. Yang, M. Vo, N. Neverova, D. Ramanan, A. Vedaldi, and H. Joo, “Banmo: Building animatable 3d neural mod- els from many casual videos,” in CVPR, 2022

  112. [120]

    Magicpony: Learning articulated 3d animals in the wild,

    S. Wu, R. Li, T. Jakab, C. Rupprecht, and A. Vedaldi, “Magicpony: Learning articulated 3d animals in the wild,” in CVPR, 2023

  113. [121]

    Common pets in 3d: Dynamic new-view synthesis of real-life de- formable categories,

    S. Sinha, R. Shapovalov, J. Reizenstein, I. Rocco, N. Neverova, A. Vedaldi, and D. Novotny, “Common pets in 3d: Dynamic new-view synthesis of real-life de- formable categories,” in CVPR, 2023

  114. [122]

    Animal 18 avatars: Reconstructing animatable 3D animals from ca- sual videos,

    R. Sabathier, N. J. Mitra, and D. Novotny, “Animal 18 avatars: Reconstructing animatable 3D animals from ca- sual videos,” in ECCV, 2024

  115. [123]

    Cla- nerf: Category-level articulated neural radiance field,

    W.-C. Tseng, H.-J. Liao, L. Yen-Chen, and M. Sun, “Cla- nerf: Category-level articulated neural radiance field,” in ICRA, 2022

  116. [124]

    Paris: Part- level reconstruction and motion analysis for articulated objects,

    J. Liu, A. Mahdavi-Amiri, and M. Savva, “Paris: Part- level reconstruction and motion analysis for articulated objects,” in ICCV, 2023

  117. [125]

    Leia: Latent view- invariant embeddings for implicit 3d articulation,

    A. Swaminathan, A. Gupta, K. Gupta, S. R. Maiya, V . Agarwal, and A. Shrivastava, “Leia: Latent view- invariant embeddings for implicit 3d articulation,” in ECCV, 2024

  118. [126]

    Reacto: Reconstructing articulated objects from a single video,

    C. Song, J. Wei, C. S. Foo, G. Lin, and F. Liu, “Reacto: Reconstructing articulated objects from a single video,” in CVPR, 2024

  119. [127]

    Artgs: Building interactable replicas of complex articulated ob- jects via gaussian splatting,

    Y. Liu, B. Jia, R. Lu, J. Ni, S.-C. Zhu, and S. Huang, “Artgs: Building interactable replicas of complex articulated ob- jects via gaussian splatting,” ArXiv:2502.19459, 2025

  120. [128]

    Deep learning for 3D human pose estimation and mesh recovery: A survey,

    Y. Liu, C. Qiu, and Z. Zhang, “Deep learning for 3D human pose estimation and mesh recovery: A survey,” Neurocomputing, 2024

  121. [129]

    Deepsdf: Learning continuous signed dis- tance functions for shape representation,

    J. J. Park, P . Florence, J. Straub, R. Newcombe, and S. Lovegrove, “Deepsdf: Learning continuous signed dis- tance functions for shape representation,” in CVPR, 2019

  122. [130]

    Occupancy networks: Learning 3d recon- struction in function space,

    L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger, “Occupancy networks: Learning 3d recon- struction in function space,” in CVPR, 2019

  123. [131]

    Neuman: Neural human radiance field from a single video,

    W. Jiang, K. M. Yi, G. Samei, O. Tuzel, and A. Ranjan, “Neuman: Neural human radiance field from a single video,” in ECCV, Springer, 2022

  124. [132]

    Structured local radiance fields for human avatar modeling,

    Z. Zheng, H. Huang, T. Yu, H. Zhang, Y. Guo, and Y. Liu, “Structured local radiance fields for human avatar modeling,” in CVPR, 2022

  125. [133]

    Instant-NVR: Instant neural volumetric render- ing for human-object interactions from monocular RGBD stream,

    Y. Jiang, K. Yao, Z. Su, Z. Shen, H. Luo, and L. Xu, “Instant-NVR: Instant neural volumetric render- ing for human-object interactions from monocular RGBD stream,” in CVPR, 2023

  126. [134]

    Instantavatar: Learning avatars from monocular video in 60 seconds,

    T. Jiang, X. Chen, J. Song, and O. Hilliges, “Instantavatar: Learning avatars from monocular video in 60 seconds,” in CVPR, 2023

  127. [135]

    Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes,

    X. Chen, Y. Zheng, M. J. Black, O. Hilliges, and A. Geiger, “Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes,” in ICCV, 2021

  128. [136]

    Fast-SNARF: A fast deformer for articulated neural fields,

    X. Chen, T. Jiang, J. Song, M. Rietmann, A. Geiger, M. J. Black, and O. Hilliges, “Fast-SNARF: A fast deformer for articulated neural fields,” TP AMI, vol. 45, no. 10, 2023

  129. [137]

    Pina: Learning a personalized implicit neu- ral avatar from a single rgb-d video sequence,

    Z. Dong, C. Guo, J. Song, X. Chen, A. Geiger, and O. Hilliges, “Pina: Learning a personalized implicit neu- ral avatar from a single rgb-d video sequence,” in CVPR, 2022

  130. [138]

    X-avatar: Expressive human avatars,

    K. Shen, C. Guo, M. Kaufmann, J. J. Zarate, J. Valentin, J. Song, and O. Hilliges, “X-avatar: Expressive human avatars,” in CVPR, 2023

  131. [139]

    Generalizable neural performer: Learning robust radiance fields for human novel view synthesis,

    W. Cheng, S. Xu, J. Piao, C. Qian, W. Wu, K.-Y. Lin, and H. Li, “Generalizable neural performer: Learning robust radiance fields for human novel view synthesis,” ArXiv:220411798, 2022

  132. [140]

    Pixel-aligned volumetric avatars,

    A. Raj, M. Zollhofer, T. Simon, J. Saragih, S. Saito, J. Hays, and S. Lombardi, “Pixel-aligned volumetric avatars,” in CVPR, 2021

  133. [141]

    pixelnerf: Neural radiance fields from one or few images,

    A. Yu, V . Ye, M. Tancik, and A. Kanazawa, “pixelnerf: Neural radiance fields from one or few images,” inCVPR, 2021

  134. [142]

    4k4d: Real-time 4d view synthesis at 4k resolution,

    Z. Xu, S. Peng, H. Lin, G. He, J. Sun, Y. Shen, H. Bao, and X. Zhou, “4k4d: Real-time 4d view synthesis at 4k resolution,” in CVPR, 2024

  135. [143]

    Real-time deep dynamic charac- ters,

    M. Habermann, L. Liu, W. Xu, M. Zollhoefer, G. Pons- Moll, and C. Theobalt, “Real-time deep dynamic charac- ters,” ACM TOG, vol. 40, no. 4, 2021

  136. [144]

    Animatable 3D Gaus- sians for modeling dynamic humans,

    Y. Xu, K. Ye, T. Shao, and Y. Weng, “Animatable 3D Gaus- sians for modeling dynamic humans,” Front. Comput. Sci., vol. 19, no. 9, 2025

  137. [145]

    Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos,

    R. Jena, G. S. Iyer, S. Choudhary, B. Smith, P . Chaudhari, and J. Gee, “Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos,” ArXiv:231110812, 2023

  138. [146]

    Splattingavatar: Realistic real-time human avatars with mesh-embedded gaussian splat- ting,

    Z. Shao, Z. Wang, Z. Li, D. Wang, X. Lin, Y. Zhang, M. Fan, and Z. Wang, “Splattingavatar: Realistic real-time human avatars with mesh-embedded gaussian splat- ting,” in CVPR, 2024

  139. [147]

    Gomavatar: Efficient animatable human modeling from monocular video using gaussians-on-mesh,

    J. Wen, X. Zhao, Z. Ren, A. G. Schwing, and S. Wang, “Gomavatar: Efficient animatable human modeling from monocular video using gaussians-on-mesh,” in CVPR, 2024

  140. [148]

    Bags: Building animatable gaussian splatting from a monocular video with diffusion priors,

    T. Zhang, Q. Gao, W. Li, L. Liu, and B. Chen, “Bags: Building animatable gaussian splatting from a monocular video with diffusion priors,” CoRR, 2024

  141. [149]

    Learning neural volumetric representations of dynamic humans in minutes,

    C. Geng, S. Peng, Z. Xu, H. Bao, and X. Zhou, “Learning neural volumetric representations of dynamic humans in minutes,” in CVPR, 2023

  142. [150]

    Nasa neural artic- ulated shape approximation,

    B. Deng, J. P . Lewis, T. Jeruzalski, G. Pons-Moll, G. Hin- ton, M. Norouzi, and A. Tagliasacchi, “Nasa neural artic- ulated shape approximation,” in ECCV, Springer, 2020

  143. [151]

    Lasr: Learn- ing articulated shape reconstruction from a monocular video,

    G. Yang, D. Sun, V . Jampani, D. Vlasic, F. Cole, H. Chang, D. Ramanan, W. T. Freeman, and C. Liu, “Lasr: Learn- ing articulated shape reconstruction from a monocular video,” in CVPR, 2021

  144. [152]

    Viser: Video-specific surface embeddings for articulated 3d shape reconstruction,

    G. Yang, D. Sun, V . Jampani, D. Vlasic, F. Cole, C. Liu, and D. Ramanan, “Viser: Video-specific surface embeddings for articulated 3d shape reconstruction,” NeurIPS, vol. 34, 2021

  145. [153]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V . Khalidov, P . Fernandez, D. Haziza, F. Massa, A. El- Nouby, et al. , “Dinov2: Learning robust visual features without supervision,” T ransactions on Machine Learning Research Journal, 2024

  146. [154]

    Continuous surface embed- dings,

    N. Neverova, D. Novotny, M. Szafraniec, V . Khalidov, P . Labatut, and A. Vedaldi, “Continuous surface embed- dings,” NeurIPS, vol. 33, 2020

  147. [155]

    Ohta: One-shot hand avatar via data-driven implicit priors,

    X. Zheng, C. Wen, Z. Su, Z. Xu, Z. Li, Y. Zhao, and Z. Xue, “Ohta: One-shot hand avatar via data-driven implicit priors,” in CVPR, 2024

  148. [156]

    HandNeRF: Neural radiance fields for animatable interacting hands,

    Z. Guo, W. Zhou, M. Wang, L. Li, and H. Li, “HandNeRF: Neural radiance fields for animatable interacting hands,” in CVPR, 2023

  149. [157]

    Html: A parametric hand texture model for 3d hand reconstruction and personalization,

    N. Qian, J. Wang, F. Mueller, F. Bernard, V . Golyanik, and C. Theobalt, “Html: A parametric hand texture model for 3d hand reconstruction and personalization,” in ECCV, 2020

  150. [158]

    Harp: Personalized hand reconstruction from a monoc- ular rgb video,

    K. Karunratanakul, S. Prokudin, O. Hilliges, and S. Tang, “Harp: Personalized hand reconstruction from a monoc- ular rgb video,” in CVPR, 2023

  151. [159]

    URhand: Universal relightable hands,

    Z. Chen, G. Moon, K. Guo, C. Cao, S. Pidhorskyi, T. Si- mon, R. Joshi, Y. Dong, Y. Xu, B. Pires, et al. , “URhand: Universal relightable hands,” in CVPR, 2024

  152. [160]

    Relightablehands: Efficient neural relighting of articu- lated hand models,

    S. Iwase, S. Saito, T. Simon, S. Lombardi, T. Bagautdinov, R. Joshi, F. Prada, T. Shiratori, Y. Sheikh, and J. Saragih, “Relightablehands: Efficient neural relighting of articu- lated hand models,” in CVPR, 2023

  153. [161]

    HandRT: Simultaneous hand shape and appearance reconstruction with pose tracking from monocular RGB-d video,

    P . Kalshetti and P . Chaudhuri, “HandRT: Simultaneous hand shape and appearance reconstruction with pose tracking from monocular RGB-d video,” TP AMI, 2025

  154. [162]

    Im2hands: Learning attentive implicit representation of interacting two-hand shapes,

    J. Lee, M. Sung, H. Choi, and T.-K. Kim, “Im2hands: Learning attentive implicit representation of interacting two-hand shapes,” in CVPR, 2023

  155. [163]

    What’s in your hands? 3d reconstruction of generic objects in hands,

    Y. Ye, A. Gupta, and S. Tulsiani, “What’s in your hands? 3d reconstruction of generic objects in hands,” in CVPR, 2022

  156. [164]

    Consistent 3d hand reconstruction in video via 19 self-supervised learning,

    Z. Tu, Z. Huang, Y. Chen, D. Kang, L. Bao, B. Yang, and J. Yuan, “Consistent 3d hand reconstruction in video via 19 self-supervised learning,” TP AMI, vol. 45, no. 8, 2023

  157. [165]

    Nimble: A non-rigid hand model with bones and muscles,

    Y. Li, L. Zhang, Z. Qiu, Y. Jiang, N. Li, Y. Ma, Y. Zhang, L. Xu, and J. Yu, “Nimble: A non-rigid hand model with bones and muscles,” ACM TOG, vol. 41, no. 4, 2022

  158. [166]

    Rep- resenting volumetric videos as dynamic mlp maps,

    S. Peng, Y. Yan, Q. Shuai, H. Bao, and X. Zhou, “Rep- resenting volumetric videos as dynamic mlp maps,” in CVPR, 2023

  159. [167]

    Spacetime gaussian feature splatting for real-time dynamic view synthesis,

    Z. Li, Z. Chen, Z. Li, and Y. Xu, “Spacetime gaussian feature splatting for real-time dynamic view synthesis,” in CVPR, 2024

  160. [168]

    Gflow: Recovering 4d world from monocular video,

    S. Wang, X. Yang, Q. Shen, Z. Jiang, and X. Wang, “Gflow: Recovering 4d world from monocular video,” CoRR, 2024

  161. [169]

    Neural deformable voxel grid for fast optimization of dynamic view synthesis,

    X. Guo, G. Chen, Y. Dai, X. Ye, J. Sun, X. Tan, and E. Ding, “Neural deformable voxel grid for fast optimization of dynamic view synthesis,” in ACCV, 2022

  162. [170]

    Hyperreel: High-fidelity 6-dof video with ray-conditioned sampling,

    B. Attal, J.-B. Huang, C. Richardt, M. Zollhoefer, J. Kopf, M. O’Toole, and C. Kim, “Hyperreel: High-fidelity 6-dof video with ray-conditioned sampling,” in CVPR, 2023

  163. [171]

    Mononerf: Learning a gener- alizable dynamic radiance field from monocular videos,

    F. Tian, S. Du, and Y. Duan, “Mononerf: Learning a gener- alizable dynamic radiance field from monocular videos,” in ICCV, 2023

  164. [172]

    Dynpoint: Dynamic neural point for view synthesis,

    K. Zhou, J.-X. Zhong, S. Shin, K. Lu, Y. Yang, A. Markham, and N. Trigoni, “Dynpoint: Dynamic neural point for view synthesis,” NeurIPS, vol. 36, 2024

  165. [173]

    Neural parametric gaussians for monocular non-rigid object reconstruction,

    D. Das, C. Wewer, R. Yunus, E. Ilg, and J. E. Lenssen, “Neural parametric gaussians for monocular non-rigid object reconstruction,” in CVPR, 2024

  166. [174]

    Fourier plenoctrees for dynamic radiance field rendering in real-time,

    L. Wang, J. Zhang, X. Liu, F. Zhao, Y. Zhang, Y. Zhang, M. Wu, J. Yu, and L. Xu, “Fourier plenoctrees for dynamic radiance field rendering in real-time,” in CVPR, 2022

  167. [175]

    Dynmf: Neural motion factorization for real-time dynamic view synthe- sis with 3d gaussian splatting,

    A. Kratimenos, J. Lei, and K. Daniilidis, “Dynmf: Neural motion factorization for real-time dynamic view synthe- sis with 3d gaussian splatting,” in ECCV, Springer, 2024

  168. [176]

    3d menagerie: Modeling the 3d shape and pose of animals,

    S. Zuffi, A. Kanazawa, D. W. Jacobs, and M. J. Black, “3d menagerie: Modeling the 3d shape and pose of animals,” in CVPR, 2017

  169. [177]

    Who left the dogs out? 3d animal recon- struction with expectation maximization in the loop,

    B. Biggs, O. Boyne, J. Charles, A. Fitzgibbon, and R. Cipolla, “Who left the dogs out? 3d animal recon- struction with expectation maximization in the loop,” in ECCV, 2020

  170. [178]

    Recon- structing animatable categories from videos,

    G. Yang, C. Wang, N. D. Reddy, and D. Ramanan, “Recon- structing animatable categories from videos,” in CVPR, 2023

  171. [179]

    Lassie: Learning articulated shapes from sparse image ensemble via 3d part discovery,

    C.-H. Yao, W.-C. Hung, Y. Li, M. Rubinstein, M.-H. Yang, and V . Jampani, “Lassie: Learning articulated shapes from sparse image ensemble via 3d part discovery,” NeurIPS, vol. 35, 2022

  172. [180]

    Learning the 3d fauna of the web,

    Z. Li, D. Litvak, R. Li, Y. Zhang, T. Jakab, C. Rupprecht, S. Wu, A. Vedaldi, and J. Wu, “Learning the 3d fauna of the web,” in CVPR, 2024

  173. [181]

    Self-supervised neural articulated shape and appearance models,

    F. Wei, R. Chabra, L. Ma, C. Lassner, M. Zoll- höfer, S. Rusinkiewicz, C. Sweeney, R. Newcombe, and M. Slavcheva, “Self-supervised neural articulated shape and appearance models,” in CVPR, 2022

  174. [182]

    Masked space-time hash encoding for efficient dynamic scene reconstruction,

    F. Wang, Z. Chen, G. Wang, Y. Song, and H. Liu, “Masked space-time hash encoding for efficient dynamic scene reconstruction,” NeurIPS, vol. 36, 2024

  175. [183]

    Dˆ 2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video,

    T. Wu, F. Zhong, A. Tagliasacchi, F. Cole, and C. Oztireli, “Dˆ 2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video,” NeurIPS, vol. 35, 2022

  176. [184]

    Mixed neural voxels for fast multi-view video synthesis,

    F. Wang, S. Tan, X. Li, Z. Tian, Y. Song, and H. Liu, “Mixed neural voxels for fast multi-view video synthesis,” in ICCV, 2023

  177. [185]

    V4d: Voxel for 4d novel view synthesis,

    W. Gan, H. Xu, Y. Huang, S. Chen, and N. Yokoya, “V4d: Voxel for 4d novel view synthesis,” TVCG, 2023

  178. [186]

    Sc-gs: Sparse-controlled gaussian splatting for editable dynamic scenes,

    Y.-H. Huang, Y.-T. Sun, Z. Yang, X. Lyu, Y.-P . Cao, and X. Qi, “Sc-gs: Sparse-controlled gaussian splatting for editable dynamic scenes,” in CVPR, 2024

  179. [187]

    Devrf: Fast deformable voxel radiance fields for dynamic scenes,

    J.-W. Liu, Y.-P . Cao, W. Mao, W. Zhang, D. J. Zhang, J. Keppo, Y. Shan, X. Qie, and M. Z. Shou, “Devrf: Fast deformable voxel radiance fields for dynamic scenes,” NeurIPS, vol. 35, 2022

  180. [188]

    Emernerf: Emergent spatial-temporal scene decomposition via self- supervision,

    J. Yang, B. Ivanovic, O. Litany, X. Weng, S. W. Kim, B. Li, T. Che, D. Xu, S. Fidler, M. Pavone, et al. , “Emernerf: Emergent spatial-temporal scene decomposition via self- supervision,” CoRR, 2023

  181. [189]

    NVFi: Neural velocity fields for 3D physics learning from dynamic videos,

    J. Li, Z. Song, and B. Yang, “NVFi: Neural velocity fields for 3D physics learning from dynamic videos,” NeurIPS, vol. 36, 2024

  182. [190]

    Hosnerf: Dynamic human- object-scene neural radiance fields from a single video,

    J.-W. Liu, Y.-P . Cao, T. Yang, Z. Xu, J. Keppo, Y. Shan, X. Qie, and M. Z. Shou, “Hosnerf: Dynamic human- object-scene neural radiance fields from a single video,” in ICCV, 2023

  183. [191]

    Fast High Dynamic Range Radiance Fields for Dynamic Scenes,

    G. Wu, T. Yi, J. Fang, W. Liu, and X. Wang, “Fast High Dynamic Range Radiance Fields for Dynamic Scenes,” in 3DV, 2024

  184. [192]

    High-fidelity and real-time novel view synthesis for dynamic scenes,

    H. Lin, S. Peng, Z. Xu, T. Xie, X. He, H. Bao, and X. Zhou, “High-fidelity and real-time novel view synthesis for dynamic scenes,” in SIGGRAPH Asia, 2023

  185. [193]

    Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sam- pling,

    X. Liu, Y.-W. Tai, C.-K. Tang, P . Miraldo, S. Lohit, and M. Chatterjee, “Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sam- pling,” in CVPR, 2024

  186. [194]

    BLiRF: Bandlimited Radiance Fields for Dynamic Scene Modeling,

    S. Ramasinghe, V . Shevchenko, G. Avraham, and A. Van Den Hengel, “BLiRF: Bandlimited Radiance Fields for Dynamic Scene Modeling,” in AAAI, vol. 38, 2024

  187. [195]

    Shape of Motion: 4D Reconstruction from a Single Video,

    Q. Wang, V . Ye, H. Gao, J. Austin, Z. Li, and A. Kanazawa, “Shape of Motion: 4D Reconstruction from a Single Video,” ArXiv:240713764, 2024

  188. [196]

    Forward flow for novel view synthesis of dynamic scenes,

    X. Guo, J. Sun, Y. Dai, G. Chen, X. Ye, X. Tan, E. Ding, Y. Zhang, and J. Wang, “Forward flow for novel view synthesis of dynamic scenes,” in ICCV, 2023

  189. [197]

    Learning dynamic view synthesis with few rgbd cameras,

    S. Wang, Y. Kwon, Y. Shen, Q. Zhang, A. State, J.-B. Huang, and H. Fuchs, “Learning dynamic view synthesis with few rgbd cameras,” ArXiv:220410477, 2022

  190. [198]

    Towards robust monocular depth estima- tion: Mixing datasets for zero-shot cross-dataset transfer,

    R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V . Koltun, “Towards robust monocular depth estima- tion: Mixing datasets for zero-shot cross-dataset transfer,” TP AMI, vol. 44, no. 3, 2020

  191. [199]

    Depth anything: Unleashing the power of large-scale unlabeled data,

    L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” in CVPR, 2024

  192. [200]

    Depth- supervised nerf: Fewer views and faster training for free,

    K. Deng, A. Liu, J.-Y. Zhu, and D. Ramanan, “Depth- supervised nerf: Fewer views and faster training for free,” in CVPR, 2022

  193. [201]

    Dynamic point fields,

    S. Prokudin, Q. Ma, M. Raafat, J. Valentin, and S. Tang, “Dynamic point fields,” in ICCV, 2023

  194. [202]

    Nerf-ds: Neural radiance fields for dynamic specular objects,

    Z. Yan, C. Li, and G. H. Lee, “Nerf-ds: Neural radiance fields for dynamic specular objects,” in CVPR, 2023

  195. [203]

    Econ: Explicit clothed humans optimized via normal integration,

    Y. Xiu, J. Yang, X. Cao, D. Tzionas, and M. J. Black, “Econ: Explicit clothed humans optimized via normal integration,” in CVPR, 2023

  196. [204]

    Icon: Implicit clothed humans obtained from normals,

    Y. Xiu, J. Yang, D. Tzionas, and M. J. Black, “Icon: Implicit clothed humans obtained from normals,” in CVPR, IEEE, 2022

  197. [205]

    Normal- gan: Learning detailed 3d human from a single rgb-d image,

    L. Wang, X. Zhao, T. Yu, S. Wang, and Y. Liu, “Normal- gan: Learning detailed 3d human from a single rgb-d image,” in ECCV, 2020

  198. [206]

    Enhancing neural ra- diance fields with depth and normal completion priors from sparse views,

    J. Guo, H. Chou, and N. Ding, “Enhancing neural ra- diance fields with depth and normal completion priors from sparse views,” ArXiv:2407.05666, 2024

  199. [207]

    Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization,

    J. Li, J. Zhang, X. Bai, J. Zheng, X. Ning, J. Zhou, and L. Gu, “Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization,” in CVPR, 2024

  200. [208]

    Dn-splatter: Depth and normal priors for gaussian splatting and meshing,

    M. Turkulainen, X. Ren, I. Melekhov, O. Seiskari, E. Rahtu, and J. Kannala, “Dn-splatter: Depth and normal priors for gaussian splatting and meshing,” in WACV, 20 IEEE, 2025

  201. [209]

    Metric3d: Towards zero-shot metric 3d prediction from a single image,

    W. Yin, C. Zhang, H. Chen, Z. Cai, G. Yu, K. Wang, X. Chen, and C. Shen, “Metric3d: Towards zero-shot metric 3d prediction from a single image,” in ICCV, 2023

  202. [210]

    Learning multi-object dynamics with composi- tional neural radiance fields,

    D. Driess, Z. Huang, Y. Li, R. Tedrake, and M. Tou- ssaint, “Learning multi-object dynamics with composi- tional neural radiance fields,” in CoRL, PMLR, 2023

  203. [211]

    Drivable 3d gaussian avatars,

    W. Zielonka, T. Bagautdinov, S. Saito, M. Zollhöfer, J. Thies, and J. Romero, “Drivable 3d gaussian avatars,” ArXiv:231108581, 2023

  204. [212]

    Sg-gs: Photo-realistic animatable human avatars with semantically-guided gaussian splatting,

    H. Zhao, C. Yang, H. Wang, X. Zhao, and W. Shen, “Sg-gs: Photo-realistic animatable human avatars with semantically-guided gaussian splatting,” ArXiv:2408.09665, 2024

  205. [213]

    Semantic attention flow fields for monocular dynamic scene decomposition,

    Y. Liang, E. Laidlaw, A. Meyerowitz, S. Sridhar, and J. Tompkin, “Semantic attention flow fields for monocular dynamic scene decomposition,” in ICCV, 2023

  206. [214]

    Drivable volumetric avatars using texel-aligned features,

    E. Remelli, T. Bagautdinov, S. Saito, C. Wu, T. Simon, S.-E. Wei, K. Guo, Z. Cao, F. Prada, J. Saragih, et al., “Drivable volumetric avatars using texel-aligned features,” in ACM SIGGRAPH, 2022

  207. [215]

    Unifying flow, stereo and depth estimation,

    H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,” TP AMI, vol. 45, no. 11, 2023

  208. [216]

    S3 gaussian: Self-supervised street gaussians for autonomous driv- ing,

    N. Huang, X. Wei, W. Zheng, P . An, M. Lu, W. Zhan, M. Tomizuka, K. Keutzer, and S. Zhang, “ S3 gaussian: Self-supervised street gaussians for autonomous driv- ing,” CoRR, 2024

  209. [217]

    Stylegaussian: Instant 3d style transfer with gaussian splatting,

    K. Liu, F. Zhan, M. Xu, C. Theobalt, L. Shao, and S. Lu, “Stylegaussian: Instant 3d style transfer with gaussian splatting,” in SIGGRAPH Asia 2024 T echnical Communi- cations, 2024

  210. [218]

    Dynvideo-e: Harness- ing dynamic nerf for large-scale motion-and view-change human-centric video editing,

    J.-W. Liu, Y.-P . Cao, J. Z. Wu, W. Mao, Y. Gu, R. Zhao, J. Keppo, Y. Shan, and M. Z. Shou, “Dynvideo-e: Harness- ing dynamic nerf for large-scale motion-and view-change human-centric video editing,” in CVPR, 2024

  211. [219]

    Compact 3D Gaussian Splatting for Static and Dynamic Radiance Fields,

    J. C. Lee, D. Rho, X. Sun, J. H. Ko, and E. Park, “Compact 3D Gaussian Splatting for Static and Dynamic Radiance Fields,” ArXiv:240803822, 2024

  212. [220]

    Rodus: Robust decomposition of static and dynamic elements in urban scenes,

    L. Roldão, N. Piasco, M. Bennehar, D. Tsishkou, et al. , “Rodus: Robust decomposition of static and dynamic elements in urban scenes,” CoRR, 2024

  213. [221]

    Block-nerf: Scalable large scene neural view synthesis,

    M. Tancik, V . Casser, X. Yan, S. Pradhan, B. Milden- hall, P . P . Srinivasan, J. T. Barron, and H. Kretzschmar, “Block-nerf: Scalable large scene neural view synthesis,” in CVPR, 2022

  214. [222]

    Vastgaussian: Vast 3d gaussians for large scene reconstruction,

    J. Lin, Z. Li, X. Tang, J. Liu, S. Liu, J. Liu, Y. Lu, X. Wu, S. Xu, Y. Yan, et al. , “Vastgaussian: Vast 3d gaussians for large scene reconstruction,” in CVPR, 2024

  215. [223]

    Grid-guided neural radiance fields for large urban scenes,

    L. Xu, Y. Xiangli, S. Peng, X. Pan, N. Zhao, C. Theobalt, B. Dai, and D. Lin, “Grid-guided neural radiance fields for large urban scenes,” in CVPR, 2023

  216. [224]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P . Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR, 2022

  217. [225]

    Sdxl: Im- proving latent diffusion models for high-resolution image synthesis,

    D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dock- horn, J. Müller, J. Penna, and R. Rombach, “Sdxl: Im- proving latent diffusion models for high-resolution image synthesis,” in ICLR

  218. [226]

    Magic3d: High-resolution text-to-3d content creation,

    C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y. Liu, and T.-Y. Lin, “Magic3d: High-resolution text-to-3d content creation,” in CVPR, 2023

  219. [227]

    Dreambooth3d: Subject-driven text-to- 3d generation,

    A. Raj, S. Kaza, B. Poole, M. Niemeyer, N. Ruiz, B. Mildenhall, S. Zada, K. Aberman, M. Rubinstein, J. Barron, et al. , “Dreambooth3d: Subject-driven text-to- 3d generation,” in ICCV, 2023

  220. [228]

    Sora: A review on background, technology, limitations, and opportunities of large vision models,

    Y. Liu, K. Zhang, Y. Li, Z. Yan, C. Gao, R. Chen, Z. Yuan, Y. Huang, H. Sun, J. Gao, et al. , “Sora: A review on background, technology, limitations, and opportunities of large vision models,” CoRR, 2024

  221. [229]

    A survey on evaluation of large language models,

    Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al. , “A survey on evaluation of large language models,” ACM T rans. Intell. Syst. T echnol., vol. 15, no. 3, 2024

  222. [230]

    Text-to-4d dynamic scene generation,

    U. Singer, S. Sheynin, A. Polyak, O. Ashual, I. Makarov, F. Kokkinos, N. Goyal, A. Vedaldi, D. Parikh, J. Johnson, et al. , “Text-to-4d dynamic scene generation,” in ICML, PMLR, 2023

  223. [231]

    Avatarclip: zero-shot text-driven generation and anima- tion of 3d avatars,

    F. Hong, M. Zhang, L. Pan, Z. Cai, L. Yang, and Z. Liu, “Avatarclip: zero-shot text-driven generation and anima- tion of 3d avatars,” TOG, vol. 41, no. 4, 2022

  224. [232]

    Align your gaussians: Text-to-4d with dynamic 3d gaus- sians and composed diffusion models,

    H. Ling, S. W. Kim, A. Torralba, S. Fidler, and K. Kreis, “Align your gaussians: Text-to-4d with dynamic 3d gaus- sians and composed diffusion models,” in CVPR, 2024

  225. [233]

    Langsplat: 3d language gaussian splatting,

    M. Qin, W. Li, J. Zhou, H. Wang, and H. Pfister, “Langsplat: 3d language gaussian splatting,” in CVPR, 2024

  226. [234]

    Raysplats: Ray tracing based gaussian splatting,

    K. Byrski, M. Mazur, J. Tabor, T. Dziarmaga, M. K ˛ adziołka, D. Baran, and P . Spurek, “Raysplats: Ray tracing based gaussian splatting,” ArXiv:2501.19196, 2025

  227. [235]

    3d gaussian ray tracing: Fast tracing of particle scenes,

    N. Moenne-Loccoz, A. Mirzaei, O. Perel, R. de Lutio, J. Martinez Esturo, G. State, S. Fidler, N. Sharp, and Z. Gojcic, “3d gaussian ray tracing: Fast tracing of particle scenes,” ACM TOG, vol. 43, no. 6, 2024. APPENDIX .1 More Detailed Discussion about Capture Setting Dynamic ...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.