REVIEW 3 major objections 5 minor 1 cited by
AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Family-aware Transformer beats previous animal 3D reconstruction
desk verdict A solid, useful step for multi-species animal mesh recovery, but the synthetic label fidelity and missing uncertainty reporting keep it at conditional rather than accept. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is a family-aware Transformer: a ViT backbone with a learnable class token, followed by a SMAL Transformer decoder that regresses SMAL parameters through MLPs, together with an animal-family supervised contrastive loss (Lcon) applied to the class token. The class token acts as a compact representation of the animal family, and the contrastive loss makes the token's embedding cluster animals of the same family while separating different families, which improves the network's ability to predict family-appropriate body shapes. The other key mechanism is the CtrlAni3D generation pipeline, which turns SMAL mesh labels into realistic images via ControlNet conditioned on rendered depth and mask maps, then filters them with SAM2-based mask IoU and manual review.
What would settle it
Render a random sample of CtrlAni3D images, then run an independent three-dimensional fitting optimizer that does not use the dataset's own labels to estimate SMAL parameters from each image, and compare those fits to the dataset labels; systematic divergence on pose or shape, beyond the expected fitting noise, would indicate the labels are biased by two-dimensional-only filtering.
Extended reading notes
Core claim
The paper's central claim is that a single network can reconstruct the 3D pose and shape of diverse quadruped species from one RGB image with greater accuracy than existing methods, provided the network is built on a Transformer backbone, a supervised contrastive loss on animal family labels, and a sufficiently large and diverse training set. Following the recipe of HMR2.0, the authors connect a ViT encoder to a Transformer decoder to regress SMAL shape, pose, and camera parameters directly rather than as residuals. The family-aware contrastive loss operates on a class token that interacts with image tokens, pulling same-family animals together and separating different families, which the paper argues is essential because animal shape distributions differ across families. The paper further claims that the new CtrlAni3D synthetic dataset, produced by prompting ControlNet with SMAL-rendered depth and mask maps, improves in-the-wild generalization compared to training without it and compared to traditional CG-rendered synthetic data.
Load-bearing premise
The whole approach assumes that the images in CtrlAni3D genuinely match the 3D pose and shape of their SMAL labels, but the automated filter only checks 2D mask similarity, so images could lie about 3D geometry and still pass.
Editorial extensions
If this is right
- Applying the same scaling recipe used for human mesh recovery to animals yields consistent gains across 3D and 2D benchmarks, not just on the training distribution.
- A diffusion-based conditional image generator can be used as a data engine for parametric mesh labels, replacing expensive manual 3D annotation and providing pixel-aligned supervision at scale.
- Animal-family supervised contrastive learning makes the model more shape-aware for families that are underrepresented in the dataset, such as boars and cats.
- Training on synthetic images generated from SMAL improves accuracy on real in-the-wild images enough to make synthetic data a practical component of animal mesh recovery.
Reading between the lines
- Because the family-aware contrastive loss clusters same-family shapes, AniMer should be directly applicable to rare or unseen species within the families it has seen, potentially improving few-shot adaptation to new quadruped species if the SMAL shape space covers them.
- The CtrlAni3D pipeline is not tied to SMAL: it could be reused to generate pixel-aligned training data for any parametric mesh model, including new animal models with more expressive shape spaces, as long as depth and mask renders can be produced.
- If the paper's reported gains on Animal Kingdom hold in independent testing, a practical near-term consequence is that behavioral biologists could use off-the-shelf AniMer reconstructions for quantitative movement analysis in the wild, without per-species fine-tuning.
- One implication the paper does not chase is whether the family contrastive token can be swapped for a continuous shape embedding to interpolate between families rather than classify them; that would be a natural test of how much of the gain comes from the contrastive objective versus the Transformer capacity.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents AniMer, a Transformer-based architecture for estimating SMAL pose and shape parameters of quadruped animals from a single image. The method uses a ViT encoder with a family-aware supervised contrastive loss on a class token, followed by a Transformer decoder that directly regresses SMAL parameters. The authors also introduce CtrlAni3D, a synthetic dataset of roughly 10k images generated with ControlNet conditioned on rendered SMAL depth and mask maps, with semi-automated filtering based on SAM2 silhouette IoU and manual review. Training aggregates existing 2D and 3D animal datasets plus CtrlAni3D (41.3k images total). The paper reports improvements over HMR, WLDO, and HMR2.0 on Animal3D, on a CtrlAni3D test split, and on the out-of-distribution Animal Kingdom benchmark, along with ablations for the dataset and the contrastive loss.
Significance. If the claims hold, the paper makes a meaningful step for animal mesh recovery: it demonstrates that a high-capacity Transformer backbone together with a large, multi-species training set yields substantial gains over prior CNN-based methods, and it contributes a scalable synthetic-data generation pipeline that is likely useful beyond this specific task. The CtrlAni3D dataset itself is a potentially valuable community resource, and the family-aware contrastive design is a sensible way to exploit taxonomic structure. The evidence is partly strong: the improvement on real Animal3D when CtrlAni3D is added to training (Supplementary Table 7) provides an independent signal that the synthetic data helps genuine 3D accuracy. However, the quantitative support for the headline 'state-of-the-art' claim is weakened by the lack of repeated runs/error bars, the uncertainty about the HMR2.0 baseline protocol, and the absence of a direct 3D-consistency validation of the CtrlAni3D labels.
major comments (3)
- [Supplementary Sec. D and Table 1] The supplement states that AniMer's performance varies substantially across settings: PA-MPJPE on Animal3D ranges from 87 to 78 mm and PCK@0.15 on Animal Kingdom ranges from 0.5 to 0.6. All tables in the main text report a single rune without error bars. Given that the claimed improvement over HMR2.0 on Animal3D PA-MPJPE is 94.1 vs. 80.4 (Table 1), the reported variation of up to 9 mm is comparable in magnitude to that difference, and the ablation differences in Tables 2 and 4 (e.g., 82.7 vs. 82.9 AUC) are well within the observed variation. Please report results over multiple seeds (at least 3) with mean and standard deviation, and indicate whether the improvements over baselines are statistically significant.
- [Sec. 4, CtrlAni3D pipeline] The quality-control pipeline verifies only 2D consistency: SAM2 mask IoU > 0.95 compares the generated foreground silhouette with the rendered SMAL mask, and manual review checks visual plausibility. This does not confirm that the generated image actually depicts the conditioning SMAL mesh's 3D pose and shape; a hallucinated image can match the 2D silhouette while differing in limb foreshortening, joint angles, or body proportions. Because the CtrlAni3D test split is produced by the same pipeline, the PA-MPJPE/PA-MPVPE numbers on CtrlAni3D in Table 1 are not independent evidence of 3D accuracy. Moreover, if the training labels are biased, the model could learn that bias and still show gains on Animal3D. The authors should provide a direct 3D validation of CtrlAni3D labels, for example by manually fitting SMAL to a random subset of generated images, or by comparing CtrlAni3D-trained models against an independently annotated real dataset, and report the resulting label noise level.
- [Sec. 5.2, Table 1, comparison protocol] The text says 'all above methods' (HMR, WLDO, AniMer-a, AniMer-b) are retrained on the full aggregated dataset with the same losses and two-stage strategy, but HMR2.0 is introduced with the phrase 'we compare to HMR2.0 to highlight our special design choices.' It is not specified whether HMR2.0 was also retrained on the same data or used with its original pretrained weights. Since HMR2.0 is designed for SMPL (humans), a fair comparison requires adapting its output head to SMAL and retraining it on the same aggregated data. If the reported HMR2.0 numbers come from an unmodified model, the comparison is confounded and the headline claim 'outperforms HMR2.0' is not supported. Please clarify the exact HMR2.0 training protocol, including data, loss functions, and number of epochs, and either retrain it under identical conditions or label the comparison as 'off-the-shelf'.
minor comments (5)
- [Sec. 3.2] The sentence 'the weights of ViT encoder are pretrained using Xu et al. [47]' refers to ViTPose++, but [47] is a pose-estimation work, not a general ViT pretraining method. Please specify that the backbone is initialized with ViTPose++ pretrained weights, or use a more direct citation for the ViT backbone.
- [Sec. 4] The rendering position is described as 'uniformly sampled between [−0.5, −0.5, 4] and [0.5, 0.5, 8]'; it would be clearer to state that these are world-coordinate ranges and give the exact sampling distribution.
- [Sec. 5.2, Table 1] The row for HMR2.0 on Animal3D reports PA-MPJPE 94.1 and PA-MPVPE 98.5, but on CtrlAni3D the same model reports 60.9 and 66.4; it would be helpful to explain this large discrepancy, especially if the model is trained on the same aggregated data.
- [Sec. 5.3, Table 2] The 'others' column aggregates Animal Pose, APT-36K, AwA-Pose, Stanford Extra, and Zebra synthetic, but the exact composition is not given in the table or the text. Please define this set explicitly, since the ablation conclusions depend on it.
- [Abstract] The term 'family aware' is used without a hyphen in the abstract and title; please use a consistent hyphenation ('family-aware') throughout.
Circularity Check
No significant circularity: AniMer's results rest on external benchmarks and standard supervised evaluation, with the CtrlAni3D self-benchmark being a data-validity limitation rather than a definitional reduction.
full rationale
The paper's central claim is empirical: a Transformer-based network trained on aggregated real and synthetic quadruped data outperforms baselines on Animal3D, CtrlAni3D, and Animal Kingdom. No equation defines a predicted quantity in terms of its own input. CtrlAni3D is a synthetic dataset whose images are generated by ControlNet conditioned on rendered SMAL depth and mask maps, then filtered by SAM2 IoU and manual review; evaluating on a held-out portion of the same synthetic distribution is a standard synthetic-data benchmark, not a circular derivation, because the metric compares predicted SMAL parameters to the conditioning SMAL parameters and all compared methods are evaluated under identical conditions. The concern that 2D silhouette filtering does not certify 3D pose/shape alignment is a legitimate validity caveat, but it is not circularity. Independent external evidence is provided by Animal3D (Supplementary Table 7 shows PA-MPJPE improving from 82.6 to 80.4 when CtrlAni3D is added) and by the out-of-distribution Animal Kingdom benchmark, which is unseen during training. The only overlapping-author citation, reference [2] on pig motion capture, appears in related work and is not load-bearing for the architecture, dataset pipeline, or results. Design choices such as the ViT backbone and Transformer decoder are attributed to external prior work [10, 47], and the family contrastive loss is a standard supervised contrastive formulation [15]. No circular step can be exhibited from the paper's own equations or self-citation chain.
Assumptions & free parameters
free parameters (6)
- Loss weights in Eq. 2 =
lambda_3D=0.05, lambda_2D=0.01, lambda_prior=0.001, lambda_adv=0.0005, lambda_con=0.0005
- Sub-loss weights in Eq. 3 and 5 =
lambda_beta=0.01, lambda_theta=0.2, lambda_beta_prior=0.5
- Temperature tau in contrastive loss
- Training epochs =
500 then 700
- Initial learning rate =
1.25e-6
- Dataset sampling weights =
Animal3D=1, CtrlAni3D=0.5, Animal Pose=0.15, AwA-Pose=0.15, Zebra Synthetic=0.05, Stanford Extra=0.15, APT-36K=0.15
assumptions (4)
- domain assumption The SMAL model space and its Gaussian priors are adequate to represent the shape and pose of all test quadrupeds.
- domain assumption A weak-perspective camera model with fixed focal length 1000 is sufficient for projection.
- domain assumption ControlNet conditioned on rendered SMAL mask and depth maps produces images whose 3D structure is aligned with the SMAL mesh.
- domain assumption The 10 species can be reliably grouped into 5 families for contrastive learning.
Cite this review
Pith. "Pith review of AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer." pith.science (2026). https://pith.science/paper/6HHADCJB
@misc{pith2026241200837,
author = {Pith},
title = {Pith review of: AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer},
year = {2026},
howpublished = {\url{https://pith.science/paper/6HHADCJB}},
note = {Machine review of arXiv:2412.00837}
}
read the original abstract
Quantitative analysis of animal behavior and biomechanics requires accurate animal pose and shape estimation across species, and is important for animal welfare and biological research. However, the small network capacity of previous methods and limited multi-species dataset leave this problem underexplored. To this end, this paper presents AniMer to estimate animal pose and shape using family aware Transformer, enhancing the reconstruction accuracy of diverse quadrupedal families. A key insight of AniMer is its integration of a high-capacity Transformer-based backbone and an animal family supervised contrastive learning scheme, unifying the discriminative understanding of various quadrupedal shapes within a single framework. For effective training, we aggregate most available open-sourced quadrupedal datasets, either with 3D or 2D labels. To improve the diversity of 3D labeled data, we introduce CtrlAni3D, a novel large-scale synthetic dataset created through a new diffusion-based conditional image generation pipeline. CtrlAni3D consists of about 10k images with pixel-aligned SMAL labels. In total, we obtain 41.3k annotated images for training and validation. Consequently, the combination of a family aware Transformer network and an expansive dataset enables AniMer to outperform existing methods not only on 3D datasets like Animal3D and CtrlAni3D, but also on out-of-distribution Animal Kingdom dataset. Ablation studies further demonstrate the effectiveness of our network design and CtrlAni3D in enhancing the performance of AniMer for in-the-wild applications. The project page of AniMer is https://luoxue-star.github.io/AniMer_project_page/.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos
MorphGS retargets motion from a monocular video onto a rigged 3D character by optimizing target morphology and pose with image-space losses, without 3D source reconstruction or parametric templates.
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 5
arXiv 2023
-
[2]
Three-dimensional surface motion capture of multiple freely moving pigs using mammal
Liang An, Jilong Ren, Tao Yu, Tang Hai, Yichang Jia, and Yebin Liu. Three-dimensional surface motion capture of multiple freely moving pigs using mammal. Nature Commu- nications, 14(1):7727, 2023. 2
work page 2023
-
[3]
A novel dataset for keypoint detection of quadruped animals from images
Prianka Banik, Lin Li, and Xishuang Dong. A novel dataset for keypoint detection of quadruped animals from images. arXiv preprint arXiv:2108.13958, 2021. 6, 1
arXiv 2021
-
[4]
Creatures great and smal: Recovering the shape and motion of animals from video
Benjamin Biggs, Thomas Roddick, Andrew Fitzgibbon, and Roberto Cipolla. Creatures great and smal: Recovering the shape and motion of animals from video. InComputer Vision– ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2–6, 2018, Revised Selected Pa- pers, Part V 14, pages 3–19. Springer, 2019. 1
work page 2018
-
[5]
Who left the dogs out?: 3D animal reconstruction with expectation maximization in the loop
Benjamin Biggs, Oliver Boyne, James Charles, Andrew Fitzgibbon, and Roberto Cipolla. Who left the dogs out?: 3D animal reconstruction with expectation maximization in the loop. In ECCV, 2020. 1, 2, 3, 5, 6, 7, 4
work page 2020
-
[6]
Smpler-x: Scaling up expressive human pose and shape estimation
Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qing- ping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, et al. Smpler-x: Scaling up expressive human pose and shape estimation. Advances in Neural Infor- mation Processing Systems, 36, 2024. 1, 3
work page 2024
-
[7]
Cross-domain adaptation for animal pose estimation
Jinkun Cao, Hongyang Tang, Hao-Shu Fang, Xiaoyong Shen, Cewu Lu, and Yu-Wing Tai. Cross-domain adaptation for animal pose estimation. In Proceedings of the IEEE/CVF international conference on computer vision , pages 9498– 9507, 2019. 3, 6, 1
work page 2019
-
[8]
3d-lfm: Lifting foundation model
Mosam Dabhi, L ´aszl´o A Jeni, and Simon Lucey. 3d-lfm: Lifting foundation model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10466–10475, 2024. 2
work page 2024
Show all 57 references
-
[9]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 3, 4
2010 arXiv
-
[10]
Humans in 4d: Recon- structing and tracking humans with transformers
Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4d: Recon- structing and tracking humans with transformers. In Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pages 14783–14794, 2023. 1, 2, 3, 4, 6, 7
2023
-
[11]
Farm3d: Learning articulated 3d animals by distilling 2d diffusion
Tomas Jakab, Ruining Li, Shangzhe Wu, Christian Rupprecht, and Andrea Vedaldi. Farm3d: Learning articulated 3d animals by distilling 2d diffusion. In 2024 International Conference on 3D Vision (3DV), pages 852–861. IEEE, 2024. 2
2024
-
[12]
Spac-net: synthetic pose- aware animal controlnet for enhanced pose estimation
Le Jiang and Sarah Ostadabbas. Spac-net: synthetic pose- aware animal controlnet for enhanced pose estimation. arXiv preprint arXiv:2305.17845, 2023. 3
2023 arXiv
-
[13]
End-to-end recovery of human shape and pose
Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7122–7131, 2018. 1, 2, 5, 6, 7, 4
2018
-
[14]
Rgbd-dog: Predicting canine pose from rgbd sensors
Sinead Kearney, Wenbin Li, Martin Parsons, Kwang In Kim, and Darren Cosker. Rgbd-dog: Predicting canine pose from rgbd sensors. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 3
2020
-
[15]
Supervised contrastive learning
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. Advances in neural information processing systems, 33:18661–18673,
-
[16]
Pare: Part attention regressor for 3d human body estimation
Muhammed Kocabas, Chun-Hao P Huang, Otmar Hilliges, and Michael J Black. Pare: Part attention regressor for 3d human body estimation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11127– 11137, 2021. 3
2021
-
[17]
Coarse-to-fine animal pose and shape estimation
Chen Li and Gim Hee Lee. Coarse-to-fine animal pose and shape estimation. Advances in Neural Information Processing Systems, 34:11757–11768, 2021. 1, 3
2021
-
[18]
hsmal: Detailed horse shape and pose recon- struction for motion pattern recognition
Ci Li, Nima Ghorbani, Sofia Broom ´e, Maheen Rashid, Michael J Black, Elin Hernlund, Hedvig Kjellstr ¨om, and Silvia Zuffi. hsmal: Detailed horse shape and pose recon- struction for motion pattern recognition. arXiv preprint arXiv:2106.10102, 2021. 1, 3
2021 arXiv
-
[19]
The poses for equine research dataset (pferd)
Ci Li, Ylva Mellbin, Johanna Krogager, Senya Polikovsky, Martin Holmberg, Nima Ghorbani, Michael J Black, Hedvig Kjellstr¨om, Silvia Zuffi, and Elin Hernlund. The poses for equine research dataset (pferd). Scientific Data, 11(1):497,
-
[20]
Learning the 3d fauna of the web
Zizhang Li, Dor Litvak, Ruining Li, Yunzhi Zhang, Tomas Jakab, Christian Rupprecht, Shangzhe Wu, Andrea Vedaldi, and Jiajun Wu. Learning the 3d fauna of the web. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9752–9762, 2024. 2
2024
-
[21]
4dhands: Reconstructing interactive hands in 4d with transformers
Dixuan Lin, Yuxiang Zhang, Mengcheng Li, Yebin Liu, Wei Jing, Qi Yan, Qianying Wang, and Hongwen Zhang. 4dhands: Reconstructing interactive hands in 4d with transformers. arXiv preprint arXiv:2405.20330, 2024. 3
2024 arXiv
-
[22]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceeding...
2014
-
[23]
Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIG- GRAPH Asia), 34(6):248:1–248:16, 2015. 1
2015
-
[24]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In The Seventh International Conference on Learning Representations, 2019. 6
2019
-
[25]
Generating images with 3d annotations using diffusion models
Wufei Ma, Qihao Liu, Jiahao Wang, Angtian Wang, Xiaoding Yuan, Yi Zhang, Zihao Xiao, Guofeng Zhang, Beijia Lu, Rux- iao Duan, et al. Generating images with 3d annotations using diffusion models. In The Twelfth International Conference on Learning Representations, 2023. 3
2023
-
[26]
Deeplabcut: markerless pose estimation of user-defined body parts with deep learning
Alexander Mathis, Pranav Mamidanna, Kevin M Cury, Taiga Abe, Venkatesh N Murthy, Mackenzie Weygandt Mathis, and Matthias Bethge. Deeplabcut: markerless pose estimation of user-defined body parts with deep learning. Nature neuro- science, 21(9):1281–1289, 2018. 2
2018
-
[27]
Learning from synthetic animals
Jiteng Mu, Weichao Qiu, Gregory D Hager, and Alan L Yuille. Learning from synthetic animals. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12386–12395, 2020. 3
2020
-
[28]
Continu- ous surface embeddings
Natalia Neverova, David Novotny, Marc Szafraniec, Vasil Khalidov, Patrick Labatut, and Andrea Vedaldi. Continu- ous surface embeddings. Advances in Neural Information Processing Systems, 33:17258–17270, 2020. 2
2020
-
[29]
Animal kingdom: A large and diverse dataset for animal behavior understanding
Xun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni, Si Yong Yeo, and Jun Liu. Animal kingdom: A large and diverse dataset for animal behavior understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19023–19034, 2022. 2, 6, 1
2022
-
[30]
Generative zoo
Tomasz Niewiadomski, Anastasios Yiannakidis, Hanz Cuevas-Velasquez, Soubhik Sanyal, Michael J Black, Sil- via Zuffi, and Peter Kulits. Generative zoo. arXiv preprint arXiv:2412.08101, 2024. 3
2024 arXiv
-
[31]
Reconstruct- ing hands in 3d with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstruct- ing hands in 3d with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9826–9836, 2024. 1, 3, 4
2024
-
[32]
Sleap: A deep learning system for multi-animal pose tracking
Talmo D Pereira, Nathaniel Tabris, Arie Matsliah, David M Turner, Junyu Li, Shruthi Ravindranath, Eleni S Papadoyan- nis, Edna Normand, David S Deutsch, Z Yan Wang, et al. Sleap: A deep learning system for multi-animal pose tracking. Nature methods, 19(4):486–495, 2022. 2
2022
-
[33]
replicant: a pipeline for generating an- notated images of animals in complex environments using unreal engine
Fabian Plum, Ren´e Bulla, Hendrik K Beck, Natalie Imirzian, and David Labonte. replicant: a pipeline for generating an- notated images of animals in complex environments using unreal engine. Nature Communications, 14(1):7195, 2023. 3
2023
-
[34]
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Doll ´ar, and Christoph Feicht...
2024 arXiv
-
[35]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 3
2022
-
[36]
Barc: Learning to regress 3d dog shape from im- ages by exploiting breed information
Nadine Rueegg, Silvia Zuffi, Konrad Schindler, and Michael J Black. Barc: Learning to regress 3d dog shape from im- ages by exploiting breed information. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3876–3884, 2022. 1, 3
2022
-
[37]
Bite: Beyond priors for improved three-d dog pose estimation
Nadine R ¨uegg, Shashank Tripathi, Konrad Schindler, Michael J Black, and Silvia Zuffi. Bite: Beyond priors for improved three-d dog pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8867–8876, 2023. 3, 5
2023
-
[38]
An- imal avatars: Reconstructing animatable 3D animals from casual videos
Remy Sabathier, Niloy Jyoti Mitra, and David Novotny. An- imal avatars: Reconstructing animatable 3D animals from casual videos. ArXiv, abs/2403.17103, 2024. 1, 3
2024 arXiv
-
[39]
Global-to-local modeling for video-based 3d human pose and shape estimation
Xiaolong Shen, Zongxin Yang, Xiaohan Wang, Jianxin Ma, Chang Zhou, and Yi Yang. Global-to-local modeling for video-based 3d human pose and shape estimation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8887–8896, 2023. 3
2023
-
[40]
Wham: Reconstructing world-grounded humans with accu- rate 3d motion
Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J Black. Wham: Reconstructing world-grounded humans with accu- rate 3d motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2070– 2080, 2024. 3
2024
-
[41]
Digi- dogs: Single-view 3d pose estimation of dogs using synthetic training data
Moira Shooter, Charles Malleson, and Adrian Hilton. Digi- dogs: Single-view 3d pose estimation of dogs using synthetic training data. In Proceedings of the IEEE/CVF Winter Confer- ence on Applications of Computer Vision, pages 80–89, 2024. 3
2024
-
[42]
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11),
-
[43]
Attention is all you need
A Vaswani. Attention is all you need. Advances in Neural Information Processing Systems, 2017. 3
2017
-
[44]
Encoder-decoder with multi-level atten- tion for 3d human shape and pose estimation
Ziniu Wan, Zhengjia Li, Maoqing Tian, Jianbo Liu, Shuai Yi, and Hongsheng Li. Encoder-decoder with multi-level atten- tion for 3d human shape and pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 13033–13042, 2021. 3
2021
-
[45]
Birds of a feather: Capturing avian shape models from images
Yufu Wang, Nikos Kolotouros, Kostas Daniilidis, and Marc Badger. Birds of a feather: Capturing avian shape models from images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14739– 14749, 2021. 2
2021
-
[46]
Animal3d: A comprehensive dataset of 3d ani- mal pose and shape
Jiacong Xu, Yi Zhang, Jiawei Peng, Wufei Ma, Artur Jesslen, Pengliang Ji, Qixin Hu, Jiehua Zhang, Qihao Liu, Jiahao Wang, et al. Animal3d: A comprehensive dataset of 3d ani- mal pose and shape. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages...
-
[47]
Vitpose++: Vision transformer for generic body pose estima- tion
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. Vitpose++: Vision transformer for generic body pose estima- tion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. 1, 3
2023
-
[48]
Banmo: Building animat- able 3d neural models from many casual videos
Gengshan Yang, Minh V o, Natalia Neverova, Deva Ramanan, Andrea Vedaldi, and Hanbyul Joo. Banmo: Building animat- able 3d neural models from many casual videos. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2863–2873, 2022. 2
2022
-
[49]
Apt-36k: A large-scale benchmark for animal pose estimation and tracking
Yuxiang Yang, Junjie Yang, Yufei Xu, Jing Zhang, Long Lan, and Dacheng Tao. Apt-36k: A large-scale benchmark for animal pose estimation and tracking. Advances in Neural Information Processing Systems, 35:17301–17313, 2022. 6, 1
2022
-
[50]
Lassie: Learning articulated shapes from sparse image ensemble via 3d part discovery.Advances in Neural Information Processing Systems, 35:15296–15308, 2022
Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Ru- binstein, Ming-Hsuan Yang, and Varun Jampani. Lassie: Learning articulated shapes from sparse image ensemble via 3d part discovery.Advances in Neural Information Processing Systems, 35:15296–15308, 2022. 2
2022
-
[51]
Superanimal pretrained pose estimation models for behavioral analysis
Shaokai Ye, Anastasiia Filippova, Jessy Lauer, Steffen Schnei- der, Maxime Vidal, Tian Qiu, Alexander Mathis, and Macken- zie Weygandt Mathis. Superanimal pretrained pose estimation models for behavioral analysis. Nature Communications, 15 (1):5165, 2024. 2
2024
-
[52]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 2, 3, 5
2023
-
[53]
3d menagerie: Modeling the 3d shape and pose of animals
Silvia Zuffi, Angjoo Kanazawa, David W Jacobs, and Michael J Black. 3d menagerie: Modeling the 3d shape and pose of animals. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6365–6373,
-
[54]
Lions and tigers and bears: Capturing non-rigid, 3d, articulated shape from images
Silvia Zuffi, Angjoo Kanazawa, and Michael J Black. Lions and tigers and bears: Capturing non-rigid, 3d, articulated shape from images. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 3955– 3963, 2018. 2, 6
2018
-
[55]
Three-d safari: Learning to estimate ze- bra pose, shape, and texture from images” in the wild”
Silvia Zuffi, Angjoo Kanazawa, Tanya Berger-Wolf, and Michael J Black. Three-d safari: Learning to estimate ze- bra pose, shape, and texture from images” in the wild”. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5359–5368, 2019. 1, 6
2019
-
[56]
Varen: Very accurate and realistic equine network
Silvia Zuffi, Ylva Mellbin, Ci Li, Markus Hoeschle, Hedvig Kjellstr¨om, Senya Polikovsky, Elin Hernlund, and Michael J Black. Varen: Very accurate and realistic equine network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5374–538...
2024
-
[57]
Failure cases
includes a diverse range of animal species. We only use 8 major animal classes of pose estimation dataset to evaluate our method. Animal3D dataset. Animal3D dataset [46] contains a total of 3379 images, which are classified into 40 classes. Each image is annotated with SMAL [ ...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.