Pith. sign in

REVIEW 3 major objections 6 minor 3 cited by

Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals

T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This survey is the first to organize the last five years of deep-learning-based 3D animal shape and motion reconstruction across explicit, template-based, implicit, and Gaussian-splatting representations, and to show that every current meth

desk verdict Useful, well-organized survey of 3D animal reconstruction; the Section IX comparison is confounded because the 'template-free' methods are generic object models, so its main trade-off conclusion doesn't hold as stated. read the letter →

arxiv 2508.16062 v1 pith:Z2MVGUYT submitted 2025-08-22 cs.CV

classification cs.CV
keywords 3DreconstructionanimalshapeandposeneuralradiancefieldsGaussiansplattingstatisticalmodelsself-supervisedlearningtemplatedeformationsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a survey claiming to be the first to cover the full spread of recent deep-learning methods for reconstructing 3D animal shape, pose, and motion from ordinary RGB images and video, going beyond parametric models such as SMAL. It organizes roughly five years of work along three axes: input modality, how the 3D geometry and motion are represented, and how training is supervised. Its central finding, drawn from a six-method comparison on four real-world scenes, is a trade-off: template-based methods are fast and anatomically plausible but species-limited and miss fine detail, while template-free methods capture detail and generalize but are slow and can produce anatomical artifacts. The authors argue that this trade-off, together with scarce 3D ground truth, is what currently limits the field.

What carries the argument

The organizing device is a taxonomy in which every method is viewed as learning a function f(s,t) that maps an input point and time to output geometry and appearance, conditioned on observations. The survey's comparison then pins the field's current frontier to the tension between a fixed template/skeleton prior and free-form deformation: the template restricts the solution space to plausible shapes, while free-form methods pay for flexibility with artifacts and runtime.

What would settle it

A fair test would run a larger, pre-registered set of methods across many species and score anatomical validity and surface detail separately; if template-free methods no longer show extra limbs and template-based methods capture fur, the trade-off claim would collapse. Alternatively, finding a published 3D animal reconstruction method whose input, representation, or supervision falls outside the survey's taxonomy would falsify the coverage claim.

Watch

Extended reading notes

Core claim

The paper's central claim is that the recent surge in 3D animal reconstruction can be coherently organized as a choice of representation: explicit surfaces and point maps, deformable templates and statistical models, neural implicit fields (SDF/NeRF), and 3D Gaussian splatting. Each representation inherits a specific trade-off among speed, geometric detail, anatomical plausibility, and cross-species generalization. The paper further claims that no current method resolves this trade-off: template-free methods such as generalist diffusion-based generators produce fine detail but sometimes add or lose limbs, while template-based methods such as SMAL-derived models stay anatomically safe but can

Load-bearing premise

The comparison's general conclusions rest on six hand-picked methods; if those six are not representative of their representation families, the claimed template trade-off is not established.

Editorial extensions

If this is right

  • If the taxonomy holds, future work can be positioned by choosing an input modality, a representation, and a supervision level, and the field's missing combinations become visible as research opportunities.
  • Template-based reconstruction will remain the practical choice for well-studied species such as dogs and horses, while template-free methods will be preferred when cross-species generalization matters more than anatomical guarantees.
  • Self-supervision from 2D keypoints, silhouettes, and perceptual losses is now the dominant training regime, so progress will likely depend on synthetic data and diffusion-generated multi-view supervision rather than on collecting new 3D scans.
  • Gaussian-splatting methods are the newest branch and combine explicit rendering efficiency with implicit quality, indicating a near-term direction for animatable animal avatars.
  • Metric-scale reconstruction, multi-animal scenes with occlusions, and biologically accurate fine details remain named open challenges that current representations do not yet solve.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the same template-versus-freeform trade-off probably applies beyond quadrupeds, so species-specific priors and generalist models are likely to coexist rather than converge.
  • Editorial extension: because several compared methods use diffusion-generated novel views, a testable extension is to measure reconstruction quality as a function of viewpoint deviation from the input; if quality degrades mainly when the animal is not in a neutral pose, part of what is called 3D learning may actually be symmetry.
  • Editorial extension: the field would benefit from a standardized benchmark with explicit inclusion criteria and per-species coverage, so claims about template trade-offs can be measured instead of illustrated on a few scenes.
  • Editorial extension: the common failure on fur and fine surface detail suggests a targeted test: feed methods high-resolution close-ups of the same species and score surface-normal and detail metrics separately from global shape metrics.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper surveys recent deep learning methods for reconstructing the 3D shape, pose, and motion of animals from RGB images and/or videos. It proposes a general problem formulation, then organizes the literature by input modality, by shape representation (explicit pointmaps/neural surfaces, template deformation and parametric models, implicit SDF/NeRF representations, and Gaussian splatting), and by training supervision. It also reviews recent generation methods, tabulates datasets and losses, presents a qualitative comparison of six methods on four real-world scenarios, and closes with open challenges. The authors claim this is the first comprehensive survey covering animal reconstruction beyond parametric models.

Significance. The survey is timely and well-structured: it covers a rapidly growing body of work from roughly 2020–2025, includes a useful dataset table (Table I) and a method taxonomy (Table II), and gives a readable account of the main representation families and supervision strategies. The survey also acknowledges some limitations explicitly, such as the non-comparable timing of One2345++. If the comparative claims in Section IX were properly supported, this would be a valuable reference for researchers entering the field. As it stands, the central comparative conclusion—that template-free methods excel at fine detail but produce anatomical artifacts while template-based methods are fast but limited—is weakened by a confounded experimental design, so the survey needs substantial revision before it can be recommended for publication.

major comments (3)
  1. [Section IX, Figures 13-14] The headline trade-off between 'template-based' and 'template-free' methods is not supported by the selected methods. The template-free set—Zero123 [23], One2345++ [24], and Hunyuan3D-2 [127]—consists of generic single-image-to-3D generation/reconstruction models with no animal-specific priors, no articulated skeleton estimation, and no animal-specific training data. The template-based set—3D-Fauna [5], SAOR [56], AniMer [78]—consists of animal-specific methods. The observed differences in speed, fine-detail fidelity, and anatomical artifacts (e.g., Hunyuan3D-2's extra limbs on the giraffe) are therefore confounded: they may reflect domain-specific training and task scope rather than the template versus template-free representation distinction. The conclusion that template-free approaches 'excel at capturing fine-grained details' while 'occasionally produce anatomical artifacts' is not e
  2. [Section IX, Table II and method selection] The comparison lacks explicit inclusion criteria for the six selected methods and reports no quantitative metrics. The text says the methods were chosen to 'span the different representations,' but no rationale is given for why these six and not others within the same categories. Moreover, One2345++ was evaluated through its demonstration website, and the paper itself states that its processing time 'cannot be directly compared.' Even with this caveat, the qualitative conclusions are used to support general claims about reconstruction speed and quality. Please specify the selection criteria, report quantitative results (e.g., chamfer distance, F-score, keypoint reprojection error) on a common benchmark such as Animal3D [120] or a fixed set of test images, and separate timing from accuracy when drawing conclusions.
  3. [Section IX and Section I (terminology)] The term 'template-free' is used in Section IX without aligning it with the survey's own taxonomy. Sections III–VI classify methods by representation (explicit, template deformation, implicit, Gaussian splatting), and none of those sections defines 'template-free' as a category. Zero123 and Hunyuan3D-2 are not 'template-free' in the sense of TAGA [105] or BANMo [6]; they are general feedforward or optimization-based generators. This terminological imprecision contributes to the confound described above. Please define 'template-free' operationally and ensure the comparison set matches that definition.
minor comments (6)
  1. [References [121] and [124]] References [121] and [124] are duplicate entries for the same DigiDogs paper (Shooter et al., WACV 2024). One should be removed and the in-text citations renumbered.
  2. [Section IV-D] The method name 'MagicPonny' appears where the cited work is 'MagicPony' (Wu et al., 2022). Please correct the typo.
  3. [Figure 1] The timeline contains the typo 'CASANeruIPS2022'; it should read 'NeurIPS2022'.
  4. [Section IV-C2] The subsection heading 'Leaning-based methods' should be 'Learning-based methods.'
  5. [Table II] The row for Zeng et al. [126] (STAG4D) lists it as a 3DGS animal reconstruction method, but STAG4D is a general generative 4D Gaussian method. If it is included because it can be applied to animals, clarify the criterion; otherwise remove it from the animal-method table.
  6. [Section V-B] The reference to 'CADEX [93]' should be 'CaDeX' to match the cited paper's title and the rest of the text.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a descriptive survey; its Section IX comparison is a qualitative evaluation with acknowledged limitations, not a derivation that reduces to its inputs.

full rationale

This is a survey paper, not a derivation or prediction paper. It organizes existing methods by input modality, representation, supervision, and datasets, and it reports qualitative comparisons. The formal equations it presents (e.g., Eq. 1 for the reconstruction function, Eq. 6 for template deformation, Eq. 12 for Gaussian covariance, and the loss definitions in Section VIII) are standard definitions drawn from the cited literature or generic formulations, not fitted parameters renamed as predictions. No central claim is obtained by a chain of reasoning whose conclusion is equivalent to its premises. The only evaluative component is Section IX's qualitative comparison of six methods. The paper explicitly acknowledges a limitation there: One2345++ 'does not have publicly available source code, and thus its results were obtained via the official demonstration website. Consequently, the processing time for this method cannot be directly compared with the others.' The qualitative conclusions about template-based methods being fast but species-limited and template-free methods capturing fine detail but occasionally producing artifacts are empirical observations based on Figures 13 and 14, not consequences of the definitions of 'template-based' and 'template-free.' The concern that the selected template-free methods are generic object reconstruction models rather than animal-specific methods is a representativeness/validity critique of the comparison, but it is not circular: the conclusions do not reduce to the selection criteria by construction. Regarding self-citations: the authors (notably H. Laga) appear in references [8], [9], [52], and [71], but these are used as general background pointers to prior surveys and statistical shape models, or as one entry in the method taxonomy (Nizamani et al. [52]). None is invoked as a uniqueness theorem, none is load-bearing for the paper's taxonomy or comparative claims, and the paper's positioning as 'the first comprehensive survey... beyond parametric models' is assessed against the cited prior surveys [13], [14], not justified by a self-citation chain. Overall, there is no circular step requiring a score above 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a survey, the paper introduces no free parameters and no invented entities. It rests on the assumption that its taxonomy is a meaningful way to organize the field and that the selected comparison methods are representative. Standard definitions are taken from the cited literature.

assumptions (3)
  • domain assumption The field can be meaningfully organized by input modality, shape representation, and training supervision (Section I.B).
    The entire taxonomy rests on this organizing principle; if it fails to capture important distinctions, the survey structure would mislead readers.
  • domain assumption The six methods selected in Section IX are representative of template-based and template-free approaches.
    The qualitative comparison is used to draw general conclusions, but the selection is a judgment call and One2345++ was run via a demo website.
  • standard math Standard mathematical definitions from prior literature (SDF, NeRF, Gaussian splatting) are accepted as given.
    The survey relies on published formulations without re-deriving them.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals." pith.science (2026). https://pith.science/paper/Z2MVGUYT

@misc{pith2026250816062,
  author       = {Pith},
  title        = {Pith review of: Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z2MVGUYT}},
  note         = {Machine review of arXiv:2508.16062}
}
read the original abstract

Reconstructing the 3D geometry, pose, and motion of animals is a long-standing problem, which has a wide range of applications, from biology, livestock management, and animal conservation and welfare to content creation in digital entertainment and Virtual/Augmented Reality (VR/AR). Traditionally, 3D models of real animals are obtained using 3D scanners. These, however, are intrusive, often prohibitively expensive, and difficult to deploy in the natural environment of the animals. In recent years, we have seen a significant surge in deep learning-based techniques that enable the 3D reconstruction, in a non-intrusive manner, of the shape and motion of dynamic objects just from their RGB image and/or video observations. Several papers have explored their application and extension to various types of animals. This paper surveys the latest developments in this emerging and growing field of research. It categorizes and discusses the state-of-the-art methods based on their input modalities, the way the 3D geometry and motion of animals are represented, the type of reconstruction techniques they use, and the training mechanisms they adopt. It also analyzes the performance of some key methods, discusses their strengths and limitations, and identifies current challenges and directions for future research.

Figures

Figures reproduced from arXiv: 2508.16062 by the authors.

Figure 1
Figure 1. Representative recent papers that tackled the problem of 3D recon [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the different challenges in reconstructing the 3D shape, pose, and motion of animals in-the-wild. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Input modalities used for 3D and 4D reconstruction of animals. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Input encoding using various types of encoders. Here, we show [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Various explicit representations used for deep learning-based 3D reconstruction of animals. [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Animal reconstruction using point map-based representations [43]. Image adapted from [43]. Polygonal meshes and point clouds are irregular data structures, and thus, they cannot be directly processed with convolutional operations. On the other hand, depth maps ( [PITH…
Figure 7
Figure 7. Figure 7: Neural architectures for neural surface-based 3D animal representation and reconstruction. [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: (a) and (b) show deformation-based representation of animals. (c) shows that the PCA spaces of SMAL [53], which combines multiple quadruped species, and hSMAL [25] and VAREN [54], which are specifically designed for horses. For each example, we show the mean shape in t…
Figure 9
Figure 9. Figure 9: Illustration of the difference between SMAL model [53] and models [PITH_FULL_IMAGE:figures/full_fig_p009_9.png]
Figure 10
Figure 10. Figure 10: Neural architectures for signed distance fields representation. [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Gaussian Avatars of dogs. Image adapted from GART [98]. [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: Comparison of Gaussian Avatar-based 3D reconstruction and rendering of dogs. Images adapted from GART [98] and DogRecon [104]. [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13: Qualitative comparison of geometric detail in 3D animal reconstruction, including Zero123 [23] (2023), SAOR [56] (2024), 3D Fauna [5] (2024), [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]
Figure 14
Figure 14. Figure 14: The same results as in Figure 13, rendered from a different viewpoint (back view). [PITH_FULL_IMAGE:figures/full_fig_p021_14.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    A new pipeline using canonical LoRAs for view synthesis, deformable 3D Gaussian splatting anchored on D-SMAL, and generative repair to produce animatable 3D dogs from single wild images without 3D supervision.

  2. CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CORGI reconstructs high-fidelity, animatable 3D dogs from a single in-the-wild image via canonical orbital generation, deformable 3DGS anchored to D-SMAL, and self-supervised generative repair, without 3D supervision.

  3. PRIMA: Boosting Animal Mesh Recovery with Biological Priors and Test-Time Adaptation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    PRIMA boosts 3D quadruped mesh recovery by injecting BioCLIP biological priors and using test-time adaptation with 2D constraints to build the Quadruped3D pseudo-3D dataset and reach SOTA on imbalanced animal benchmarks.

Reference graph

Works this paper leans on

127 extracted references · 68 canonical work pages · cited by 2 Pith papers

  1. [23]

    Zero-1-to-3: Zero-shot One Image to 3D Object,

    R. Liu, R. Wu, B. V . Hoorick, P. Tokmakov, S. Zakharov, and C. V ondrick, “Zero-1-to-3: Zero-shot One Image to 3D Object,” 2023

  2. [24]

    One-2-3-45++: Fast Single Image to 3D Objects XXX 23 with Consistent Multi-View Generation and 3D Diffusion,

    M. Liu, R. Shi, L. Chen, Z. Zhang, C. Xu, X. Wei, H. Chen, C. Zeng, J. Gu, and H. Su, “One-2-3-45++: Fast Single Image to 3D Objects XXX 23 with Consistent Multi-View Generation and 3D Diffusion,” ArXiv, vol. abs/2311.07885, 2023

  3. [127]

    Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation,

    T. H. Team, “Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation,” 2025

  4. [5]

    Learning the 3D Fauna of the Web

    Z. Li, D. Litvak, R. Li, Y . Zhang, T. Jakab, C. Rupprecht, S. Wu, A. Vedaldi, and J. Wu, “Learning the 3D Fauna of the Web,”IEEE/CVF CVPR, vol. abs/2401.02400, 2024

  5. [56]

    SAOR: Single-View Articulated Object Reconstruction,

    M. Aygun and O. M. Aodha, “SAOR: Single-View Articulated Object Reconstruction,” IEEE/CVF CVPR, 2024

  6. [78]

    Animer: Animal pose and shape estimation using family aware transformer,

    J. Lyu, T. Zhu, Y . Gu, L. Lin, P. Cheng, Y . Liu, X. Tang, and L. An, “Animer: Animal pose and shape estimation using family aware transformer,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 17 486–17 496

  7. [120]

    Animal3d: A comprehensive dataset of 3d animal pose and shape,

    J. Xu, Y . Zhang, J.-X. Peng, W. Ma, A. Jesslen, P. Ji, Q. Hu, J. Zhang, Q. Liu, J. Wang, W. Ji, C. Wang, X. Yuan, P. Kaushik, G. Zhang, J. Liu, Y . Xie, Y . Cui, A. L. Yuille, and A. Kortylewski, “Animal3d: A comprehensive dataset of 3d animal pose and shape,” IEEE/CVF ICCV, pp. 9065–9075, 2023

  8. [105]

    TAGA: Self- supervised Learning for Template-free Animatable Gaussian Articu- lated Model,

    Z. Zhai, G. Chen, W. Wang, D. Zheng, and J. Xiao, “TAGA: Self- supervised Learning for Template-free Animatable Gaussian Articu- lated Model,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 159–21 169

  9. [6]

    BANMo: Building Animatable 3D Neural Models from Many Casual Videos,

    G. Yang, M. V o, N. Neverova, D. Ramanan, A. Vedaldi, and H. Joo, “BANMo: Building Animatable 3D Neural Models from Many Casual Videos,” IEEE/CVF CVPR, pp. 2853–2863, 2022

Show all 127 references
  1. [1]

    Common Pets in 3D: Dynamic New- View Synthesis of Real-Life Deformable Categories,

    S. Sinha, R. Shapovalov, J. Reizenstein, I. Rocco, N. Neverova, A. Vedaldi, and D. Novotn ´y, “Common Pets in 3D: Dynamic New- View Synthesis of Real-Life Deformable Categories,” IEEE/CVF CVPR, pp. 4881–4891, 2023

  2. [2]

    Learning 3D Deformation of Animals from 2D Images,

    A. Kanazawa, S. Z. Kovalsky, R. Basri, and D. W. Jacobs, “Learning 3D Deformation of Animals from 2D Images,” Computer Graphics Forum, vol. 35, 2015

  3. [3]

    Reconstruct- ing Animatable Categories from Videos,

    G. Yang, C. Wang, D. R. Narapureddy, and D. Ramanan, “Reconstruct- ing Animatable Categories from Videos,” 2023 IEEE/CVF CVPR , pp. 16 995–17 005, 2023

  4. [4]

    Virtual Pets: Animatable Animal Generation in 3D Scenes,

    Y .-C. Cheng, C. H. Lin, C. Wang, Y . Kant, S. Tulyakov, A. Schwing, L. Gui, and H.-Y . Lee, “Virtual Pets: Animatable Animal Generation in 3D Scenes,” ArXiv, vol. abs/2312.14154, 2023

  5. [7]

    Viser: Video-specific surface embeddings for articulated 3D shape reconstruction,

    G. Yang, D. Sun, V . Jampani, D. Vlasic, F. Cole, C. Liu, and D. Ramanan, “Viser: Video-specific surface embeddings for articulated 3D shape reconstruction,” Advances in Neural Information Processing Systems, vol. 34, pp. 19 326–19 338, 2021

  6. [8]

    Image-based 3d object re- construction: State-of-the-art and trends in the deep learning era,

    X.-F. Han, H. Laga, and M. Bennamoun, “Image-based 3d object re- construction: State-of-the-art and trends in the deep learning era,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 5, pp. 1578–1604, 2021

  7. [9]

    A survey on deep learning techniques for stereo-based depth estimation,

    H. Laga, L. V . Jospin, F. Boussaid, and M. Bennamoun, “A survey on deep learning techniques for stereo-based depth estimation,” IEEE transactions on pattern analysis and machine intelligence , vol. 44, no. 4, pp. 1738–1764, 2022

  8. [10]

    Single-view 3d reconstruction: A survey of deep learning methods,

    G. Fahim, K. Amin, and S. Zarif, “Single-view 3d reconstruction: A survey of deep learning methods,” Comput. Graph., vol. 94, pp. 164– 190, 2021

  9. [11]

    3d reconstruction using deep learning: a survey,

    Y . Jin, D. Jiang, and M. Cai, “3d reconstruction using deep learning: a survey,” Commun. Inf. Syst. , vol. 20, pp. 389–413, 2020

  10. [12]

    Gaussian splatting: 3d reconstruction and novel view synthesis, a review,

    A. Dalal, D. Hagen, K. G. Robbersmyr, and K. M. Knausg ˚ard, “Gaussian splatting: 3d reconstruction and novel view synthesis, a review,” ArXiv, vol. abs/2405.03417, 2024

  11. [13]

    State of the Art in Dense Monocular Non-Rigid 3D Reconstruction,

    E. Tretschk, N. Kairanda, R. MallikarjunB., R. Dabral, A. Kortylewski, B. Egger, M. Habermann, P. Fua, C. Theobalt, and V . Golyanik, “State of the Art in Dense Monocular Non-Rigid 3D Reconstruction,” Computer Graphics Forum, vol. 42, 2022

  12. [14]

    Recent trends in 3d reconstruction of general non-rigid scenes,

    R. Yunus, J. E. Lenssen, M. Niemeyer, Y . Liao, C. Rupprecht, C. Theobalt, G. Pons-Moll, J.-B. Huang, V . Golyanik, and E. Ilg, “Recent trends in 3d reconstruction of general non-rigid scenes,” Computer Graphics Forum, p. e15062, 2024

  13. [15]

    A morphable model for the synthesis of 3d faces,

    V . Blanz and T. Vetter, “A morphable model for the synthesis of 3d faces,” Proceedings of the 26th annual conference on Computer graphics and interactive techniques , pp. 187–194, 1999

  14. [16]

    3D morphable face models—past, present, and future,

    B. Egger, W. A. Smith, A. Tewari, S. Wuhrer, M. Zollhoefer, T. Beeler, F. Bernard, T. Bolkart, A. Kortylewski, S. Romdhani et al. , “3D morphable face models—past, present, and future,” ACM TOG, vol. 39, no. 5, pp. 1–38, 2020

  15. [17]

    SMPL: A skinned multi-person linear model,

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “SMPL: A skinned multi-person linear model,” ACM Trans. Graphics (Proc. SIGGRAPH Asia), vol. 34, no. 6, pp. 248:1–248:16, Oct. 2015

  16. [18]

    Depth anything: Unleashing the power of large-scale unlabeled data,

    L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” ArXiv, vol. abs/2401.10891, 2024

  17. [19]

    LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part Discovery,

    C. Yao, W.-C. Hung, Y . Li, M. Rubinstein, M. Yang, and V . Jampani, “LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part Discovery,” Advances in Neural Information Processing Systems, vol. abs/2207.03434, 2022

  18. [20]

    Artic3D: Learning robust articulated 3Dshapes from noisy web image collections,

    C.-H. Yao, A. Raj, W.-C. Hung, M. Rubinstein, Y . Li, M.-H. Yang, and V . Jampani, “Artic3D: Learning robust articulated 3Dshapes from noisy web image collections,” Advances in Neural Information Processing Systems, vol. 36, 2024

  19. [21]

    Hi-LASSIE: High-Fidelity Articulated Shape and Skeleton Discovery from Sparse Image Ensemble,

    C. Yao, W.-C. Hung, Y . Li, M. Rubinstein, M.-H. Yang, and V . Jampani, “Hi-LASSIE: High-Fidelity Articulated Shape and Skeleton Discovery from Sparse Image Ensemble,” IEEE/CVF CVPR , pp. 4853–4862, 2022

  20. [22]

    One-2-3-45: Any single image to 3D mesh in 45 seconds without per-shape optimization,

    M. Liu, C. Xu, H. Jin, L. Chen, M. Varma T, Z. Xu, and H. Su, “One-2-3-45: Any single image to 3D mesh in 45 seconds without per-shape optimization,” Advances in Neural Information Processing Systems, vol. 36, 2024

  21. [25]

    hSMAL: Detailed Horse Shape and Pose Reconstruction for Motion Pattern Recognition,

    C. Li, N. Ghorbani, S. Broom ´e, M. Rashid, M. J. Black, E. Hernlund, H. Kjellstr ¨om, and S. Zuffi, “hSMAL: Detailed Horse Shape and Pose Reconstruction for Motion Pattern Recognition,” 2021

  22. [26]

    Emerging properties in self-supervised vision transformers,

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” Proceedings of the IEEE/CVF international conference on computer vision, pp. 9650–9660, 2021

  23. [27]

    DINOv2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khali- dov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al., “DINOv2: Learning robust visual features without supervision,” arXiv preprint arXiv:2304.07193, 2023

  24. [28]

    4DCom- plete: Non-Rigid Motion Estimation Beyond the Observable Surface,

    Y . Li, H. Takehara, T. Taketomi, B. Zheng, and M. Nießner, “4DCom- plete: Non-Rigid Motion Estimation Beyond the Observable Surface,” IEEE/CVF ICCV, pp. 12 686–12 696, 2021

  25. [29]

    DOVE: Learning deformable 3d objects by watching videos,

    S. Wu, T. Jakab, C. Rupprecht, and A. Vedaldi, “DOVE: Learning deformable 3d objects by watching videos,” IJCV, 2023

  26. [30]

    Learning category-specific mesh reconstruction from image collections,

    A. Kanazawa, S. Tulsiani, A. A. Efros, and J. Malik, “Learning category-specific mesh reconstruction from image collections,” ECCV, pp. 371–386, 2018

  27. [31]

    Articulation- Aware Canonical Surface Mapping,

    N. Kulkarni, A. K. Gupta, D. F. Fouhey, and S. Tulsiani, “Articulation- Aware Canonical Surface Mapping,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 449–458, 2020

  28. [32]

    Implicit Mesh Reconstruc- tion from Unannotated Image Collections,

    S. Tulsiani, N. Kulkarni, and A. K. Gupta, “Implicit Mesh Reconstruc- tion from Unannotated Image Collections,” ArXiv, vol. abs/2007.08504, 2020

  29. [33]

    Self-supervised single-view 3D reconstruction via semantic consistency,

    X. Li, S. Liu, K. Kim, S. De Mello, V . Jampani, M.-H. Yang, and J. Kautz, “Self-supervised single-view 3D reconstruction via semantic consistency,” ECCV, pp. 677–693, 2020

  30. [34]

    Casa: Category-agnostic skeletal animal reconstruction,

    Y . Wu, Z. Chen, S. Liu, Z. Ren, and S. Wang, “Casa: Category-agnostic skeletal animal reconstruction,” Advances in Neural Information Pro- cessing Systems, vol. 35, pp. 28 559–28 574, 2022

  31. [35]

    BARC: Learning to Regress 3D Dog Shape from Images by Exploiting Breed Infor- mation,

    N. Rueegg, S. Zuffi, K. Schindler, and M. J. Black, “BARC: Learning to Regress 3D Dog Shape from Images by Exploiting Breed Infor- mation,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3866–3874, 2022

  32. [36]

    Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images “In the Wild

    S. Zuffi, A. Kanazawa, T. Berger-Wolf, and M. J. Black, “Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images “In the Wild”,” IEEE/CVF ICCV, pp. 5358–5367, 2019

  33. [37]

    Who left the dogs out? 3d animal reconstruction with expectation maxi- mization in the loop,

    B. Biggs, O. Boyne, J. Charles, A. Fitzgibbon, and R. Cipolla, “Who left the dogs out? 3d animal reconstruction with expectation maxi- mization in the loop,” in European Conference on Computer Vision . Springer, 2020, pp. 195–211

  34. [38]

    LEP- ARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction,

    D. Liu, A. Stathopoulos, Q. Zhangli, Y . Gao, and D. Metaxas, “LEP- ARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction,” Advances in Neural Information Processing Systems , vol. 36, 2024

  35. [39]

    MagicPony: Learning Articulated 3D Animals in the Wild,

    S. Wu, R. Li, T. Jakab, C. Rupprecht, and A. Vedaldi, “MagicPony: Learning Articulated 3D Animals in the Wild,” IEEE CVPR, pp. 8792– 8802, 2022

  36. [40]

    Pulsar: Efficient sphere-based neural rendering,

    C. Lassner and M. Zollh ¨ofer, “Pulsar: Efficient sphere-based neural rendering,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1440–1449, 2021

  37. [41]

    Surfels: surface elements as rendering primitives,

    H. Pfister, M. Zwicker, J. van Baar, and M. H. Gross, “Surfels: surface elements as rendering primitives,” Proceedings of the 27th annual conference on Computer graphics and interactive techniques , 2000

  38. [42]

    Deepsur- fels: Learning online appearance fusion,

    M. Mihajlovic, S. Weder, M. Pollefeys, and M. R. Oswald, “Deepsur- fels: Learning online appearance fusion,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 14 519– 14 530, 2020

  39. [43]

    DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruc- tion,

    B. Kaye, T. Jakab, S. Wu, C. Rupprecht, and A. Vedaldi, “DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruc- tion,” IEEE/CVF CVPR, 2025

  40. [44]

    DUSt3R: Geometric 3d vision made easy,

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud, “DUSt3R: Geometric 3d vision made easy,” in IEEE/CVF CVPR , 2024, pp. 20 697–20 709

  41. [45]

    Grounding image matching in 3D with MASt3R,

    V . Leroy, Y . Cabon, and J. Revaud, “Grounding image matching in 3D with MASt3R,” in European Conference on Computer Vision . Springer, 2024, pp. 71–91

  42. [46]

    MONSt3R: A simple approach for estimating geometry in the presence of motion,

    J. Zhang, C. Herrmann, J. Hur, V . Jampani, T. Darrell, F. Cole, D. Sun, and M.-H. Yang, “MONSt3R: A simple approach for estimating geometry in the presence of motion,” ICLR, 2025

  43. [47]

    Dynamicfusion: Recon- struction and tracking of non-rigid scenes in real-time,

    R. A. Newcombe, D. Fox, and S. M. Seitz, “Dynamicfusion: Recon- struction and tracking of non-rigid scenes in real-time,” in IEEE/CVF CVPR, 2015, pp. 343–352

  44. [48]

    A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence,

    J. Zhang, C. Herrmann, J. Hur, L. Polania Cabrera, V . Jampani, D. Sun, and M.-H. Yang, “A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence,” Advances in Neural Information Processing Systems , vol. 36, pp. 45 533–45 547, 2023

  45. [49]

    NeRS: Neural reflectance surfaces for sparse-view 3D reconstruction in the wild,

    J. Zhang, G. Yang, S. Tulsiani, and D. Ramanan, “NeRS: Neural reflectance surfaces for sparse-view 3D reconstruction in the wild,” Advances in Neural Information Processing Systems , vol. 34, pp. 29 835–29 847, 2021

  46. [50]

    Neural Surface Maps,

    L. Morreale, N. Aigerman, V . G. Kim, and N. J. Mitra, “Neural Surface Maps,” IEEE CVPR, pp. 4639–4648, 2021

  47. [51]

    Neural Geometry Processing via Spherical Neural Surfaces,

    R. Williamson and N. J. Mitra, “Neural Geometry Processing via Spherical Neural Surfaces,” in Computer Graphics Forum . Wiley Online Library, 2024, p. e70021

  48. [52]

    Dynamic Neural Surfaces for Elastic 4D Shape Repre- sentation and Analysis,

    A. Nizamani, H. Laga, G. Wang, F. Boussaid, M. Bennamoun, and A. Srivastava, “Dynamic Neural Surfaces for Elastic 4D Shape Repre- sentation and Analysis,” IEEE CVPR, 2025

  49. [53]

    3D Menagerie: Modeling the 3D Shape and Pose of Animals,

    S. Zuffi, A. Kanazawa, D. W. Jacobs, and M. J. Black, “3D Menagerie: Modeling the 3D Shape and Pose of Animals,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 5524–5532, 2017

  50. [54]

    Varen: Very accurate and realistic equine network,

    S. Zuffi, Y . Mellbin, C. Li, M. Hoeschle, H. Kjellstr ¨om, S. Polikovsky, E. Hernlund, and M. J. Black, “Varen: Very accurate and realistic equine network,” in IEEE/CVF CVPR, 2024, pp. 5374–5383

  51. [55]

    LASR: Learning Articulated Shape Reconstruction from a Monocular Video,

    G. Yang, D. Sun, V . Jampani, D. Vlasic, F. Cole, H. Chang, D. Ra- manan, W. T. Freeman, and C. Liu, “LASR: Learning Articulated Shape Reconstruction from a Monocular Video,”IEE/CVF CVPR, pp. 15 975– 15 984, 2021

  52. [57]

    3D Bird Reconstruction: a Dataset, Model, and Shape Recovery from a Single View,

    M. Badger, Y . Wang, A. Modh, A. Perkes, N. Kolotouros, B. Pfrommer, M. F. Schmidt, and K. Daniilidis, “3D Bird Reconstruction: a Dataset, Model, and Shape Recovery from a Single View,” ECCV, vol. 12363, pp. 1–17, 2020

  53. [58]

    Learning monocular 3D reconstruction of articulated categories from motion,

    F. Kokkinos and I. Kokkinos, “Learning monocular 3D reconstruction of articulated categories from motion,” IEEE/CVF CVPR , pp. 1737– 1746, 2021

  54. [59]

    Laplacian surface editing,

    O. Sorkine, D. Cohen-Or, Y . Lipman, M. Alexa, C. R ¨ossl, and H.- P. Seidel, “Laplacian surface editing,” in Proceedings of the 2004 Eurographics/ACM SIGGRAPH symposium on Geometry processing , 2004, pp. 175–184

  55. [60]

    As-rigid-as-possible surface modeling,

    O. Sorkine and M. Alexa, “As-rigid-as-possible surface modeling,” Eurographics Association, p. 109–116, 2007

  56. [61]

    Online adaptation for consistent mesh reconstruction in the wild,

    X. Li, S. Liu, S. De Mello, K. Kim, X. Wang, M.-H. Yang, and J. Kautz, “Online adaptation for consistent mesh reconstruction in the wild,” Advances in Neural Information Processing Systems , vol. 33, pp. 15 009–15 019, 2020

  57. [62]

    Active shape models- their training and application,

    T. Cootes, C. Taylor, D. Cooper, and J. Graham, “Active shape models- their training and application,” CVIU, vol. 61, no. 1, pp. 38 – 59, 1995

  58. [63]

    Active appearance models,

    T. F. Cootes, G. J. Edwards, and C. J. Taylor, “Active appearance models,” Computer Vision—ECCV’98: 5th European Conference on Computer Vision Freiburg, Germany, June 2–6, 1998 Proceedings, Volume II 5, pp. 484–498, 1998

  59. [64]

    Active appearance models,

    T. Cootes, G. Edwards, and C. Taylor, “Active appearance models,” IEEE PAMI, vol. 23, no. 6, pp. 681–685, 2001

  60. [65]

    The space of human body shapes: reconstruction and parameterization from range scans,

    B. Allen, B. Curless, and Z. Popovi ´c, “The space of human body shapes: reconstruction and parameterization from range scans,” ACM TOG, vol. 22, no. 3, pp. 587–594, 2003

  61. [66]

    Expressive Body Capture: 3D Hands, Face, and Body From a Single Image,

    G. Pavlakos, V . Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black, “Expressive Body Capture: 3D Hands, Face, and Body From a Single Image,” IEEE CVPR, June 2019

  62. [67]

    BITE: Beyond priors for improved three-D dog pose estimation,

    N. R ¨uegg, S. Tripathi, K. Schindler, M. J. Black, and S. Zuffi, “BITE: Beyond priors for improved three-D dog pose estimation,” in IEEE/CVF CVPR, 2023, pp. 8867–8876

  63. [68]

    Distilling Neural Fields for Real- Time Articulated Shape Reconstruction,

    J. Tan, G. Yang, and D. Ramanan, “Distilling Neural Fields for Real- Time Articulated Shape Reconstruction,” IEEE/CVF CVPR, pp. 4692– 4701, 2023

  64. [69]

    Model- based metric 3d shape and motion reconstruction of wild bottlenose dolphins in drone-shot videos,

    D. Baieri, R. Cicciarella, M. Kr ¨utzen, E. Rodol`a, and S. Zuffi, “Model- based metric 3d shape and motion reconstruction of wild bottlenose dolphins in drone-shot videos,” arXiv preprint arXiv:2504.15782, 2025

  65. [70]

    Lions and Tigers and Bears: Capturing Non-Rigid, 3D, Articulated Shape from Images,

    S. Zuffi, A. Kanazawa, and M. J. Black, “Lions and Tigers and Bears: Capturing Non-Rigid, 3D, Articulated Shape from Images,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018

  66. [71]

    A survey on nonrigid 3D shape analysis,

    H. Laga, “A survey on nonrigid 3D shape analysis,” Academic Press Library in Signal Processing, Volume 6 , pp. 261–304, 2018. XXX 24

  67. [72]

    Coarse-to-fine Animal Pose and Shape Estima- tion,

    C. Li and G. H. Lee, “Coarse-to-fine Animal Pose and Shape Estima- tion,” Advances in Neural Information Processing Systems, vol. 34, pp. 11 757–11 768, 2021

  68. [73]

    Pose space deformation: a uni- fied approach to shape interpolation and skeleton-driven deformation,

    J. P. Lewis, M. Cordner, and N. Fong, “Pose space deformation: a uni- fied approach to shape interpolation and skeleton-driven deformation,” in Seminal Graphics Papers: Pushing the Boundaries, Volume 2, 2023, pp. 811–818

  69. [74]

    Skeleton-free pose transfer for stylized 3d characters,

    Z. Liao, J. Yang, J. Saito, G. Pons-Moll, and Y . Zhou, “Skeleton-free pose transfer for stylized 3d characters,” in European Conference on Computer Vision. Springer, 2022, pp. 640–656

  70. [75]

    Deepface: Closing the gap to human-level performance in face verification,

    Y . Taigman, M. Yang, M. Ranzato, and L. Wolf, “Deepface: Closing the gap to human-level performance in face verification,” in IEEE/CVF CVPR, 2014, pp. 1701–1708

  71. [76]

    Facenet: A unified embedding for face recognition and clustering,

    F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in IEEE/CVF CVPR , 2015, pp. 815–823

  72. [77]

    Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image,

    F. Bogo, A. Kanazawa, C. Lassner, P. Gehler, J. Romero, and M. J. Black, “Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image,” ECCV, pp. 561–578, 2016

  73. [79]

    Hu- mans in 4d: Reconstructing and tracking humans with transformers,

    S. Goel, G. Pavlakos, J. Rajasegaran, A. Kanazawa, and J. Malik, “Hu- mans in 4d: Reconstructing and tracking humans with transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 14 783–14 794

  74. [80]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017

  75. [81]

    NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” in ECCV. Springer, 2020, pp. 405–421

  76. [82]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021

  77. [83]

    Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,

    P. Wang, L. Liu, Y . Liu, C. Theobalt, T. Komura, and W. Wang, “Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,” NeurIPS, 2021

  78. [84]

    Reconstruction and representation of 3d objects with radial basis functions,

    J. C. Carr, R. K. Beatson, J. B. Cherrie, T. J. Mitchell, W. R. Fright, B. C. McCallum, and T. R. Evans, “Reconstruction and representation of 3d objects with radial basis functions,” in Proceedings of the 28th annual conference on Computer graphics and interactive techniques ...

  79. [85]

    Surface reconstruction based on compactly supported radial basis functions,

    N. Kojekine, V . Savchenko, and I. Hagiwara, “Surface reconstruction based on compactly supported radial basis functions,” in Geometric modeling: techniques, applications, systems and tools. Springer, 2004, pp. 217–231

  80. [86]

    Multi- level partition of unity implicits,

    Y . Ohtake, A. Belyaev, M. Alexa, G. Turk, and H.-P. Seidel, “Multi- level partition of unity implicits,” inAcm Siggraph 2005 Courses, 2005, pp. 173–es

  81. [87]

    Affine transformations of 3D objects represented with neural networks,

    E. Piperakis and I. Kumazawa, “Affine transformations of 3D objects represented with neural networks,” in Proceedings Third International Conference on 3-D Digital Imaging and Modeling . IEEE, 2001, pp. 213–223

  82. [88]

    3d object & light source representation with multi layer feed forward networks,

    E. Piperakis, I. Kumazawa, and R. Piperakis, “3d object & light source representation with multi layer feed forward networks,” Neural, Parallel & Scientific Computations , vol. 9, no. 2, pp. 161–173, 2001

  83. [89]

    Occupancy networks: Learning 3d reconstruction in func- tion space,

    L. M. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger, “Occupancy networks: Learning 3d reconstruction in func- tion space,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4455–4465, 2018

  84. [90]

    Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape syn- thesis,

    T. Shen, J. Gao, K. Yin, M.-Y . Liu, and S. Fidler, “Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape syn- thesis,” Advances in Neural Information Processing Systems , vol. 34, pp. 6087–6101, 2021

  85. [91]

    Deepsdf: Learning continuous signed distance functions for shape representation,

    J. J. Park, P. R. Florence, J. Straub, R. A. Newcombe, and S. Lovegrove, “Deepsdf: Learning continuous signed distance functions for shape representation,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 165–174, 2019

  86. [92]

    Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video,

    Y . Jiang, L. Zhang, J. Gao, W. Hu, and Y . Yao, “Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video,” The Twelfth International Conference on Learning Representations , 2024

  87. [93]

    CaDeX: Learning Canonical Deformation Coordinate Space for Dynamic Surface Representation via Neural Homeomorphism,

    J. Lei and K. Daniilidis, “CaDeX: Learning Canonical Deformation Coordinate Space for Dynamic Surface Representation via Neural Homeomorphism,” IEEE/CVF CVPR, pp. 6614–6624, 2022

  88. [94]

    Density estimation using real nvp,

    L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Density estimation using real nvp,” arXiv preprint arXiv:1605.08803 , 2016

  89. [95]

    Nice: Non-linear independent components estimation,

    L. Dinh, D. Krueger, and Y . Bengio, “Nice: Non-linear independent components estimation,” arXiv preprint arXiv:1410.8516 , 2014

  90. [96]

    K-planes: Explicit radiance fields in space, time, and appearance,

    S. Fridovich-Keil, G. Meanti, F. R. Warburg, B. Recht, and A. Kanazawa, “K-planes: Explicit radiance fields in space, time, and appearance,” in IEEE/CVF CVPR, 2023, pp. 12 479–12 488

  91. [97]

    3D gaussian splatting for real-time radiance field rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3D gaussian splatting for real-time radiance field rendering,” ACM Trans. Graph. , vol. 42, no. 4, pp. 139–1, 2023

  92. [98]

    GART: Gaussian Articulated Template Models,

    J. Lei, Y . Wang, G. Pavlakos, L. Liu, and K. Daniilidis, “GART: Gaussian Articulated Template Models,” in IEEE/CVF CVPR , 2024, pp. 19 876–19 887

  93. [99]

    Ewa splatting,

    M. Zwicker, H. Pfister, J. Van Baar, and M. Gross, “Ewa splatting,” IEEE Transactions on Visualization and Computer Graphics , vol. 8, no. 3, pp. 223–238, 2002

  94. [100]

    Real-time photorealistic dy- namic scene representation and rendering with 4d gaussian splatting,

    Z. Yang, H. Yang, Z. Pan, and L. Zhang, “Real-time photorealistic dy- namic scene representation and rendering with 4d gaussian splatting,” arXiv preprint arXiv:2310.10642 , 2023

  95. [101]

    Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis,

    J. Luiten, G. Kopanas, B. Leibe, and D. Ramanan, “Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis,” in 2024 International Conference on 3D Vision (3DV) . IEEE, 2024, pp. 800– 809

  96. [102]

    4d gaussian splatting for real-time dynamic scene rendering,

    G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2024, pp. 20 310–20 320

  97. [103]

    Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruc- tion,

    Z. Yang, X. Gao, W. Zhou, S. Jiao, Y . Zhang, and X. Jin, “Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruc- tion,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2024, pp. 20 331–20 341

  98. [104]

    DogRecon: Canine Prior- Guided Animatable 3D Gaussian Dog Reconstruction From A Single Image,

    G. Cho, C. Kang, D. Soon, and K. Joo, “DogRecon: Canine Prior- Guided Animatable 3D Gaussian Dog Reconstruction From A Single Image,” International Journal of Computer Vision , pp. 1–15, 2025

  99. [106]

    OmniMotionGPT: Animal Motion Generation with Limited Data,

    Z. Yang, M. Zhou, M. Shan, B. Wen, Z. Xuan, M. Hill, J. Bai, G.-J. Qi, and Y . Wang, “OmniMotionGPT: Animal Motion Generation with Limited Data,” IEEE/CVF CVPR, pp. 1249–1259, 2024

  100. [107]

    AniMo: Species-Aware Model for Text-Driven Animal Motion Generation,

    X. Wang, K. Ruan, X. Zhang, and G. Wang, “AniMo: Species-Aware Model for Text-Driven Animal Motion Generation,” in IEEE/CVF CVPR, 2025, pp. 1929–1939

  101. [108]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PmLR, 2021, pp. 8748–8763

  102. [109]

    A. M. Bronstein, M. M. Bronstein, and R. Kimmel, Numerical geom- etry of non-rigid shapes . Springer Science & Business Media, 2008

  103. [110]

    Articulated motion discovery using pairs of trajectories,

    L. Del Pero, S. Ricco, R. Sukthankar, and V . Ferrari, “Articulated motion discovery using pairs of trajectories,” in IEEE/CVF CVPR (CVPR), 2015

  104. [111]

    Cross- domain adaptation for animal pose estimation,

    J. Cao, H. Tang, H.-S. Fang, X. Shen, C. Lu, and Y .-W. Tai, “Cross- domain adaptation for animal pose estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 9498–9507

  105. [112]

    Rgbd-dog: Predicting canine pose from rgbd sensors,

    S. Kearney, W. Li, M. Parsons, K. I. Kim, and D. Cosker, “Rgbd-dog: Predicting canine pose from rgbd sensors,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020

  106. [113]

    Ap-10k: A benchmark for animal pose estimation in the wild,

    H. Yu, Y . Xu, J. Zhang, W. Zhao, Z. Guan, and D. Tao, “Ap-10k: A benchmark for animal pose estimation in the wild,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , 2021

  107. [114]

    Pretraining boosts out-of-domain robustness for pose estimation,

    A. Mathis, T. Biasi, S. Schneider, M. Yuksekgonul, B. Rogers, M. Bethge, and M. W. Mathis, “Pretraining boosts out-of-domain robustness for pose estimation,” inProceedings of the IEEE/CVF winter conference on applications of computer vision , 2021, pp. 1859–1868

  108. [115]

    Sydog: A synthetic dog dataset for improved 2d pose estimation,

    M. Shooter, C. Malleson, and A. Hilton, “Sydog: A synthetic dog dataset for improved 2d pose estimation,” arXiv preprint arXiv:2108.00249, 2021

  109. [116]

    Acinoset: a 3d pose estimation dataset and baseline models for cheetahs in the wild,

    D. Joska, L. Clark, N. Muramatsu, R. Jericevich, F. Nicolls, A. Mathis, M. W. Mathis, and A. Patel, “Acinoset: a 3d pose estimation dataset and baseline models for cheetahs in the wild,” in 2021 IEEE international XXX 25 conference on robotics and automation (ICRA) . IEEE, 202...

  110. [117]

    Apt-36k: A large-scale benchmark for animal pose estimation and tracking,

    Y . Yang, J. Yang, Y . Xu, J. Zhang, L. Lan, and D. Tao, “Apt-36k: A large-scale benchmark for animal pose estimation and tracking,” Advances in Neural Information Processing Systems , vol. 35, pp. 17 301–17 313, 2022

  111. [118]

    Animal kingdom: A large and diverse dataset for animal behavior understand- ing,

    X. L. Ng, K. E. Ong, Q. Zheng, Y . Ni, S. Y . Yeo, and J. Liu, “Animal kingdom: A large and diverse dataset for animal behavior understand- ing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 19 023–19 034

  112. [119]

    Artemis: articulated neural pets with appearance and motion synthesis,

    H. Luo, T. Xu, Y . Jiang, C. Zhou, Q. Qiu, Y . Zhang, W. Yang, L. Xu, and J. Yu, “Artemis: articulated neural pets with appearance and motion synthesis,” ACM Transactions on Graphics (TOG) , vol. 41, no. 4, pp. 1–19, 2022

  113. [122]

    Sydog-video: A synthetic dog video dataset for temporal pose estimation,

    M. Shooter, C. Malleson, and A. Hilton, “Sydog-video: A synthetic dog video dataset for temporal pose estimation,” International Journal of Computer Vision , vol. 132, no. 6, pp. 1986–2002, 2024

  114. [123]

    Digital Life 3D,

    Digital Life Project, “Digital Life 3D,” https://digitallife3d.org/

  115. [124]

    Digidogs: Single-view 3d pose estimation of dogs using synthetic training data,

    M. Shooter, C. Malleson, and A. Hilton, “Digidogs: Single-view 3d pose estimation of dogs using synthetic training data,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 80–89

  116. [125]

    Creatures great and smal: Recovering the shape and motion of animals from video,

    B. Biggs, T. Roddick, A. Fitzgibbon, and R. Cipolla, “Creatures great and smal: Recovering the shape and motion of animals from video,” Asian Conference on Computer Vision , pp. 3–19, 2019

  117. [126]

    STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians,

    Y . Zeng, Y . Jiang, S. Zhu, Y . Lu, Y . Lin, H. Zhu, W. Hu, X. Cao, and Y . Yao, “STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians,” ECCV, pp. 163–179, 2024

  118. [128]

    Reconstructing Animals and the Wild,

    P. Kulits, M. J. Black, and S. Zuffi, “Reconstructing Animals and the Wild,” in IEEE/CVF CVPR, 2025, pp. 16 565–16 577

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.