REVIEW 4 major objections 5 minor 1 cited by
Generative Zoo
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that synthetic training data for 3D animal pose and shape can be produced by a conditional image-generation model — rendering depth and edge maps of a parametric mesh into photorealistic images — and that a regressor…
desk verdict Fresh, useful idea built on an unverified control-fidelity premise, with the paper's own Table 2 contradicting its ablation narrative; still deserves serious refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrier of the argument is a pair of control signals — a depth map and a Canny-edge map — rendered from the posed SMAL mesh and fed, together with a text prompt, into a conditional diffusion model through a control network. These signals are what keep the generated image aligned with the sampled pose and shape parameters; the ablation shows that depth alone gives realism without alignment, Canny edges alone give alignment without realism, and the two combined at reduced strength strike the balance. Around this core, the pipeline adds: shape sampling from a multivariate Gaussian fit to text-embedding descriptors decoded by a flow-based generator; pose sampling from a large collection of plausible dog poses extracted by an optimization-based estimator; and prompt assembly that combines a vision-language caption of the rendered orientation with an LLM-composed scene and camera description.
What would settle it
Render one fixed set of SMAL pose and shape parameters through the pipeline and through a deterministic rasterizer, train identical regressors on both image sets, and compare on Animal3D; if the rasterizer-trained model matches or beats the diffusion-trained one, then photorealism is not what drives the result. An even more direct check is to reconstruct a generated image with an independent multi-view system and measure per-joint error against the sampled parameters: if that error is as large as the depth-only versus Canny-only gap in the paper's ablation, the labels are too noisy to validate the transfer claim.
Extended reading notes
Core claim
The discovery is that a diffusion-based image generator can serve as the renderer in a synthetic-data pipeline for parametric 3D animal estimation. Conditioned on depth and edge maps rendered from a SMAL mesh plus a text prompt, the generator produces photorealistic images whose associated pose and shape parameters are known exactly because the control signals come from those parameters. Training a vision-transformer regressor solely on one million such images transfers to real photographs and sets a new state of the art on the Animal3D benchmark: S-MPJPE drops from 374.9 mm (best baseline) to 160.1 mm, and PCK@0.5 rises from 85.6 to 97.0. The authors also show that the gain concentrates in the scale-sensitive metric, and they argue that Animal3D's pseudo-labeled ground truth contains physically implausible shapes that cap the PA-MPJPE improvement.
Load-bearing premise
The load-bearing premise is that the images produced from the depth and edge maps genuinely share the 3D pose and shape of the SMAL mesh those maps were rendered from, closely enough that a regressor trained on the images can learn a mapping that transfers to real photographs.
Editorial extensions
If this is right
- A regressor trained entirely on these generated images can beat models trained on pseudo-labeled real images, suggesting that label fidelity can matter more than photorealism at the data-source level.
- Adding a new species to a synthetic dataset no longer requires artist-built 3D assets; a text prompt and a sampled body model suffice.
- The same conditioning strategy could produce training data for any parametric model with a renderable mesh, including humans, hands, and other articulated objects.
- Dataset statistics can be re-balanced by resampling taxa and appearance descriptors, giving practitioners control over class proportions without new assets.
Reading between the lines
- We infer that the large S-MPJPE improvement over baselines likely reflects, in part, systematic implausibility in the pseudo-labeled Animal3D ground truth rather than purely better 3D reasoning; the paper's own perceptual study and upper-bound discussion gesture at this without quantifying it.
- We infer that the depth-plus-Canny compromise leaves a clear failure mode: when the diffusion model refuses a rare species, the control maps still enforce geometry but the image no longer matches the prompt's species, so the ground-truth parameters become misaligned with the visible animal; the paper's reported failures with lesser-known species make this testable.
- We infer that the pipeline's bottleneck will shift from image generation to pose distribution: the pose prior is built solely from dog images, which limits coverage of species-specific postures such as grooming; replacing dog-derived poses with a learned prior over SMAL poses from the generated images themselves could close the gap.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a synthetic-data-generation pipeline for 3D animal pose and shape estimation. The pipeline samples taxon, shape, and pose parameters from SMAL-related models (AWOL and BITE), renders depth and Canny control maps, and uses FLUX with ControlNet to synthesize photorealistic images conditioned on those maps and on text prompts. The authors introduce GenZoo, a one-million-image dataset with associated SMAL pose/shape labels, and train a ViTPose-based regressor solely on it. They report state-of-the-art performance on the Animal3D benchmark, introduce a separate synthetic test set (GenZoo-Felidae) for species generalization, and include ablations on control signals, captioning, and image-generation model. The central claim is that this pipeline combines visual realism, scalability, and controllable data production, rivaling traditional synthetic-data generators.
Significance. If the claims hold, this is a significant contribution to 3D animal pose and shape estimation: it offers a scalable alternative to graphics-engine pipelines, introduces a large public dataset, demonstrates strong benchmark results, and includes a thoughtful set of ablations and a perceptual study that questions the quality of existing pseudo-labeled ground truth. The paper also benefits from a clear pipeline description, a data-efficiency analysis, and a commitment to release the dataset and pipeline. The main weakness is that the core label-fidelity premise--that the generated images actually instantiate the sampled SMAL pose and shape--is not directly verified, and the ablation table appears to contradict the qualitative text about which control mode aligns better. The empirical SOTA claim therefore needs additional support before the results can be fully accepted.
major comments (4)
- [§3.6, §4.3, Table 2] The core premise of the pipeline is that the FLUX+ControlNet image faithfully realizes the SMAL parameters (β, θ) used to render the depth and Canny control maps, but this is never verified. The ablation evidence is internally inconsistent with the text: §4.3 states that Canny-only conditioning achieves the best alignment and depth-only the poorest, yet Table 2 on GenZoo-Felidae shows the opposite (depth-only '-Canny' gives S-V2V 57.7 and PA-V2V 39.1, better than Full at 59.3/50.2, while Canny-only '-Depth' gives 95.4/65.9). Because the SOTA transfer claim depends on the generated images carrying the intended 3D labels, the manuscript should provide a direct fidelity check (e.g., regress or optimize SMAL parameters from generated images and compare them with the conditioning parameters, or a human study comparing mesh overlays), and should reconcile the Table 2 numbers with the qualitative claim.
- [§4.1, Fig. 6] Animal3D's ground truth is itself produced by fitting SMAL to manual 2D labels, and the authors document physical implausibilities in it (Sec. 4.1, Fig. 6). Since GenZoo images are generated from SMAL renders, a model trained on GenZoo may be learning the statistics of SMAL fits rather than accurate 3D pose/shape transfer, and the reported S-MPJPE gain (374.9→160.1) could partly reflect this shared parameter space rather than true image-to-3D accuracy. The paper should evaluate on an independent benchmark with higher-fidelity ground truth, or provide per-sample analysis on Animal3D showing that improvements are not confined to cases where the pseudo-label is biased.
- [§4.1, Tables 1 and 2] All quantitative claims are based on single training runs without error bars or multiple seeds. Given the stochasticity of both image generation and network training, and the fact that the main SOTA claim rests on one S-MPJPE number, the paper should report mean and standard deviation over at least a few repeated runs (or an equivalent variance estimate) for the main comparisons and the ablations. This is especially important because Table 2 reports 100k-sample ablations while Table 1 reports the 1M-sample model, and the reader cannot assess whether the observed differences are significant.
- [§1, §3.6] The claim that the pipeline offers 'control comparable to traditional synthetic-data generators' (Sec. 1) is not quantified. Traditional renderers provide exact pose and shape by construction, while here the ControlNet conditions on depth and Canny proxies with unspecified strengths; the degree of pose/shape controllability should be measured (e.g., distribution of pose/shape errors under controlled variations) rather than asserted from qualitative samples.
minor comments (5)
- [§3.2] The text contains a typo: 'limitated shape-space expressivity' should be 'limited shape-space expressivity'.
- [§3.6] The exact ControlNet conditioning strengths for Canny and depth are not reported; 'reduced strength' is not reproducible without numerical values.
- [§3.5] The predefined lists for camera and scenery settings are not described in terms of size, sampling distribution, or examples; this hampers replication of the prompt-sampling procedure.
- [Table 2] The caption of Table 2 should state explicitly that these ablations use 100,000 training samples while Table 1 reports the 1M-sample model, to avoid apparent inconsistencies between the two tables.
- [Fig. 7] In the control-signal ablation figure, the 'Canny' and 'Canny+Depth' examples appear to show very similar images; a quantitative overlay with the conditioning render would make the alignment difference visible.
Circularity Check
No circular derivation: the GenZoo claim is empirical, and the authors' reuse of their own SMAL, AWOL, and BITE models is tool use, not circular evidence.
full rationale
The paper's derivation chain is not circular. The synthetic dataset is generated by sampling SMAL pose and shape parameters, rendering depth and Canny control maps, and prompting FLUX to produce images; the ground-truth labels are the sampled SMAL parameters by construction. The central empirical claim is that a regressor trained solely on these images transfers to Animal3D, a benchmark of real images with SMAL pseudo-labels produced independently of this pipeline. No equation in the paper defines a predicted quantity in terms of the fitted input, and no fitted parameter is renamed as a prediction. The pipeline reuses the authors' prior models (SMAL+, AWOL, BITE), but these are independently published tools with their own training data and assumptions; they are not invoked as uniqueness theorems or as self-contained justifications for the paper's conclusion. The shared SMAL parameterization between GenZoo labels and Animal3D pseudo-labels is a real limitation on what the S-MPJPE gain means, but it does not make the transfer result equivalent to the input by construction. The unverified premise that FLUX plus ControlNet faithfully realizes the sampled 3D parameters is an empirical validity concern, and the apparent inconsistency between Sec. 4.3's text and Table 2 (where depth-only '-Canny' gives better alignment metrics than Canny-only '-Depth') bears on that concern, but it is not circularity: the outcome is not contained in the premise by definition. Overall, the central claim rests on external benchmark performance and controlled ablations, not on a self-referential chain.
Assumptions & free parameters
free parameters (1)
- ControlNet conditioning strengths for Canny and depth =
Not reported
assumptions (5)
- domain assumption SMAL/SMAL+ is an adequate parameterization for the sampled quadrupedal mammals in Laurasiatheria.
- domain assumption AWOL reliably decodes CLIP embeddings into plausible SMAL shape parameters.
- domain assumption Dog poses extracted by BITE transfer to other quadrupedal species.
- domain assumption FLUX plus ControlNet yields images whose implied 3D pose and shape match the rendered control signals.
- domain assumption Animal3D pseudo-labels are sufficient for meaningful benchmark comparison.
Cite this review
Pith. "Pith review of Generative Zoo." pith.science (2026). https://pith.science/paper/VFQYAACO
@misc{pith2026241208101,
author = {Pith},
title = {Pith review of: Generative Zoo},
year = {2026},
howpublished = {\url{https://pith.science/paper/VFQYAACO}},
note = {Machine review of arXiv:2412.08101}
}
read the original abstract
The model-based estimation of 3D animal pose and shape from images enables computational modeling of animal behavior. Training models for this purpose requires large amounts of labeled image data with precise pose and shape annotations. However, capturing such data requires the use of multi-view or marker-based motion-capture systems, which are impractical to adapt to wild animals in situ and impossible to scale across a comprehensive set of animal species. Some have attempted to address the challenge of procuring training data by pseudo-labeling individual real-world images through manual 2D annotation, followed by 3D-parameter optimization to those labels. While this approach may produce silhouette-aligned samples, the obtained pose and shape parameters are often implausible due to the ill-posed nature of the monocular fitting problem. Sidestepping real-world ambiguity, others have designed complex synthetic-data-generation pipelines leveraging video-game engines and collections of artist-designed 3D assets. Such engines yield perfect ground-truth annotations but are often lacking in visual realism and require considerable manual effort to adapt to new species or environments. Motivated by these shortcomings, we propose an alternative approach to synthetic-data generation: rendering with a conditional image-generation model. We introduce a pipeline that samples a diverse set of poses and shapes for a variety of mammalian quadrupeds and generates realistic images with corresponding ground-truth pose and shape parameters. To demonstrate the scalability of our approach, we introduce GenZoo, a synthetic dataset containing one million images of distinct subjects. We train a 3D pose and shape regressor on GenZoo, which achieves state-of-the-art performance on a real-world animal pose and shape estimation benchmark, despite being trained solely on synthetic data. https://genzoo.is.tue.mpg.de
Figures
Figures from the paper (6 more)
Forward citations
Cited by 1 Pith paper
-
AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer
A family-aware Transformer with supervised contrastive learning and a diffusion-generated synthetic dataset achieves state-of-the-art 3D animal pose and shape estimation.
Reference graph
Works this paper leans on
-
[1]
David J. Anderson and Pietro Perona. Toward a science of computational ethology. Neuron, 84(1):18–31, 2014. 1
work page 2014
-
[2]
Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia, Mo- hammad Norouzi, and David J. Fleet. Synthetic data from diffusion models improves ImageNet classification. TMLR,
-
[3]
Marc Badger, Yufu Wang, Adarsh Modh, Ammon Perkes, Nikos Kolotouros, Bernd G. Pfrommer, Marc F. Schmidt, and Kostas Daniilidis. 3D bird reconstruction: A dataset, model, and shape recovery from a single view. In ECCV, page 1–17, Berlin, Heidelberg, 2020. Springer-Verlag. 2
work page 2020
-
[4]
Benjamin Biggs, Thomas Roddick, Andrew Fitzgibbon, and Roberto Cipolla. Creatures great and SMAL: Recovering the 1 Animal3D GenZoo-Felidae ↑ PCK@0.5 ↓ S-MPJPE ↓ PA-MPJPE ↓ S-V2V ↓ PA-V2V FLUX 97.1 166.9 118.4 59.3 50.2 Hunyuan-DiT 95.9 174.0 125.6 67.8 47.6 Stable Diffusion 3 97.5 178.3 127.8 85.2 61.1 Table 3. Image-Generation-Model Ablation Effects . We...
work page 2019
-
[5]
Who left the dogs out? 3D animal reconstruction with expectation maximization in the loop
Benjamin Biggs, Oliver Boyne, James Charles, Andrew Fitzgibbon, and Roberto Cipolla. Who left the dogs out? 3D animal reconstruction with expectation maximization in the loop. In ECCV, pages 195–211. Springer, 2020. 3
work page 2020
-
[6]
Black, Priyanka Patel, Joachim Tesch, and Jin- long Yang
Michael J. Black, Priyanka Patel, Joachim Tesch, and Jin- long Yang. BEDLAM: A synthetic dataset of bodies ex- hibiting detailed lifelike animated motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 8726–8737, 2023. 2, 3
work page 2023
-
[7]
Bola ˜nos, Dongsheng Xiao, Nancy L
Luis A. Bola ˜nos, Dongsheng Xiao, Nancy L. Ford, Jeff M. LeDue, Pankaj K. Gupta, Carlos Doebeli, Hao Hu, Helge Rhodin, and Timothy H. Murphy. A three-dimensional vir- tual mouse generates synthetic training data for behavioral analysis. Nature Methods, 18(4):378–381, 2021. 2, 3
work page 2021
-
[8]
Sofia Broom ´e, Marcelo Feighelstein, Anna Zamansky, Gabriel Carreira Lencioni, Pia Haubro Andersen, Francisca Pessanha, Marwa Mahmoud, Hedvig Kjellstr ¨om, and Al- bert Ali Salah. Going deeper than tracking: A survey of computer-vision based recognition of animal pain and emo- tions. IJCV, 131(2):572–590, 2023. 2
work page 2023
Show all 68 references
-
[9]
Cashman and Andrew W
Thomas J. Cashman and Andrew W. Fitzgibbon. What shape are dolphins? Building 3D morphable models from 2D im- ages. TPAMI, 35(1):232–244, 2013. 2
2013
-
[10]
Hanz Cuevas-Velasquez, Priyanka Patel, Haiwen Feng, and Michael J. Black. Toward human understanding with con- trollable synthesis, 2024. 3
2024
-
[11]
Mammal diversity database,
Mammal Diversity Database. Mammal diversity database,
-
[12]
Smith, Hannaneh Ha- jishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kemb- havi
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tri- pathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Chris Callison- Bur...
2024
-
[13]
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009. 6
2009
-
[14]
Scaling rec- tified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. Scaling rec- tified flow transformers for high-resolution image syn...
2024
-
[15]
ZooBuilder: 2D and 3D pose estimation for quadrupeds using synthetic data
Abassin Sourou Fangbemi, Yi Fei Lu, Maoyuan Xu, Xi- aowu Luo, Alexis Rolland, and Chedy Raissi. ZooBuilder: 2D and 3D pose estimation for quadrupeds using synthetic data. In Eurographics/ ACM SIGGRAPH Symposium on Computer Animation - Showcases . The Eurographics Asso- ciation...
2020
-
[16]
3D human recon- struction in the wild with synthetic data using generative models, 2024
Yongtao Ge, Wenjia Wang, Yongfan Chen, Yang Liu, Hao Chen, Xuan Wang, and Chunhua Shen. 3D human recon- struction in the wild with synthetic data using generative models, 2024. 3
2024
-
[17]
Learning with 3D rotations, a hitch- hiker’s guide to SO(3)
Andreas Ren ´e Geist, Jonas Frey, Mikel Zhobro, Anna Lev- ina, and Georg Martius. Learning with 3D rotations, a hitch- hiker’s guide to SO(3). In ICML, 2024. 1
2024
-
[18]
Humans in 4D: Reconstructing and tracking humans with transformers
Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4D: Reconstructing and tracking humans with transformers. In ICCV, pages 14783–14794, 2023. 6
2023
-
[19]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR,
-
[20]
Procedural humans for computer vision
Charlie Hewitt, Tadas Baltru ˇsaitis, Erroll Wood, Lohit Petikam, Louis Florentin, and Hanz Cuevas Velasquez. Procedural humans for computer vision. arXiv preprint arXiv:2301.01161, 2023. 2, 3
2023 arXiv
-
[21]
Cashman, Julien Valentin, Darren Cosker, et al
Charlie Hewitt, Fatemeh Saleh, Sadegh Aliakbarian, Lohit Petikam, Shideh Rezaeifar, Louis Florentin, Zafiirah Hose- nie, Thomas J. Cashman, Julien Valentin, Darren Cosker, et al. Look Ma, no markers: holistic performance capture without the hassle. arXiv preprint arXiv:2410.11...
-
[22]
Yuan-Ting Hu, Hong-Shuo Chen, Kexin Hui, Jia-Bin Huang, and Alexander G. Schwing. SAIL-VOS: Semantic amodal instance level video object segmentation-a synthetic dataset and baselines. In CVPR, pages 3105–3115, 2019. 2
2019
-
[23]
Human3.6M: Large scale datasets and predic- 2 tive methods for 3D human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large scale datasets and predic- 2 tive methods for 3D human sensing in natural environments. TPAMI, 36(7):1325–1339, 2014. 2
2014
-
[24]
Learning 3D deformation of animals from 2D images
Angjoo Kanazawa, Shahar Kovalsky, Ronen Basri, and David Jacobs. Learning 3D deformation of animals from 2D images. In Proceedings of the 37th Annual Conference of the European Association for Computer Graphics , page 365–374, Goslar, DEU, 2016. Eurographics Association. 2
2016
-
[25]
Efros, and Jitendra Malik
Angjoo Kanazawa, Shubham Tulsiani, Alexei A. Efros, and Jitendra Malik. Learning category-specific mesh reconstruc- tion from image collections. In ECCV, 2018. 2
2018
-
[26]
Black, and Silvia Zuffi
Peter Kulits, Michael J. Black, and Silvia Zuffi. Reconstruct- ing animals and the wild, 2024. 3
2024
-
[27]
Black Forest Labs. FLUX. https://github.com/ black-forest-labs/flux, 2022. 5, 1
2022
-
[28]
Black, Elin Hernlund, Hedvig Kjellstr ¨om, and Silvia Zuffi
Ci Li, Nima Ghorbani, Sofia Broom ´e, Maheen Rashid, Michael J. Black, Elin Hernlund, Hedvig Kjellstr ¨om, and Silvia Zuffi. hSMAL: Detailed horse shape and pose recon- struction for motion pattern recognition, 2021. 2, 3
2021
-
[29]
Black, Hao Li, and Javier Romero
Tianye Li, Timo Bolkart, Michael J. Black, Hao Li, and Javier Romero. Learning a model of facial shape and ex- pression from 4D scans. TOG, 36(6), 2017. 2
2017
-
[30]
Hunyuan-DiT: A powerful multi-resolution diffusion trans- former with fine-grained chinese understanding, 2024
Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, Dayou Chen, Jiajun He, Jiahao Li, Wenyue Li, Chen Zhang, Rongwei Quan, Jianxiang Lu, Jiabin Huang, Xiaoyan Yuan, Xiaoxiao Zheng, Yixuan Li, Ji...
2024
-
[31]
Lawrence Zitnick
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, pages 740–755, Cham, 2014. Springer International Publishing. 6
2014
-
[32]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi- person linear model. TOG, 34(6):1–16, 2015. 2, 3, 5
2015
-
[33]
Generating images with 3D annotations us- ing diffusion models
Wufei Ma, Qihao Liu, Jiahao Wang, Angtian Wang, Xiaod- ing Yuan, Yi Zhang, Zihao Xiao, Guofeng Zhang, Beijia Lu, Ruxiao Duan, Yongrui Qi, Adam Kortylewski, Yaoyao Liu, and Alan Yuille. Generating images with 3D annotations us- ing diffusion models. In ICLR, 2024. 3
2024
-
[34]
Troje, Ger- ard Pons-Moll, and Michael J
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Ger- ard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. In ICCV, pages 5442– 5451, 2019. 3
2019
-
[35]
Marshall, Tianqing Li, Joshua H
Jesse D. Marshall, Tianqing Li, Joshua H. Wu, and Timo- thy W. Dunn. Leaving Flatland: Advances in 3D behavioral measurement. Current Opinion in Neurobiology, 73:102522,
-
[36]
Pyrender
Matthew Matl. Pyrender. https://github.com/ mmatl/pyrender, 2022. 5
2022
-
[37]
Hager, and Alan L
Jiteng Mu, Weichao Qiu, Gregory D. Hager, and Alan L. Yuille. Learning from synthetic animals. In CVPR, pages 12386–12395, 2020. 3
2020
-
[38]
Huang, Joachim Tesch, David T
Priyanka Patel, Chun-Hao P. Huang, Joachim Tesch, David T. Hoffmann, Shashank Tripathi, and Michael J. Black. AGORA: Avatars in geography optimized for regres- sion analysis. In CVPR, pages 13468–13478, 2021. 3
2021
-
[39]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In CVPR, pages 10975– 10985, 2019. 3, 5
2019
-
[40]
Pereira, Joshua W
Talmo D. Pereira, Joshua W. Shaevitz, and Mala Murthy. Quantifying behavior to understand the brain. Nature Neu- roscience, 23(12):1537–1549, 2020. 2
2020
-
[41]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, pages 8748–8763. PMLR, 2021. 4, 5
2021
-
[42]
Adam Roberts, Hyung Won Chung, Gaurav Mishra, Anselm Levskaya, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Hait...
2023
-
[43]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. In CVPR, pages 10684– 10695, 2022. 3
2022
-
[44]
Nadine Rueegg, Silvia Zuffi, Konrad Schindler, and Michael J. Black. BARC: Learning to regress 3D dog shape from images by exploiting breed information. In CVPR, pages 3876–3884, 2022. 3
2022
-
[45]
Black, and Silvia Zuffi
Nadine R ¨uegg, Shashank Tripathi, Konrad Schindler, Michael J. Black, and Silvia Zuffi. BITE: Beyond priors for improved three-D dog pose estimation. In CVPR, pages 8867–8876, 2023. 3, 5
2023
-
[46]
Mitra, and David Novotny
Remy Sabathier, Niloy J. Mitra, and David Novotny. Animal avatars: Reconstructing animatable 3D animals from casual videos. In ECCV, pages 270–287. Springer Nature Switzer- land, 2025. 3
2025
-
[47]
SyDog-Video: A synthetic dog video dataset for temporal pose estimation
Moira Shooter, Charles Malleson, and Adrian Hilton. SyDog-Video: A synthetic dog video dataset for temporal pose estimation. IJCV, 132(6):1986–2002, 2024. 3
1986
-
[48]
Qwen2.5: A party of foundation models, 2024
Qwen Team. Qwen2.5: A party of foundation models, 2024. 5 3
2024
-
[49]
Black, Ivan Laptev, and Cordelia Schmid
G ¨ul Varol, Javier Romero, Xavier Martin, Naureen Mah- mood, Michael J. Black, Ivan Laptev, and Cordelia Schmid. Learning from synthetic humans. In CVPR, 2017. 3
2017
-
[50]
Black, Bodo Rosenhahn, and Gerard Pons-Moll
Timo von Marcard, Roberto Henschel, Michael J. Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering ac- curate 3D human pose in the wild using IMUs and a moving camera. In ECCV, 2018. 2
2018
-
[51]
Birds of a feather: Capturing avian shape models from images
Yufu Wang, Nikos Kolotouros, Kostas Daniilidis, and Marc Badger. Birds of a feather: Capturing avian shape models from images. In CVPR, pages 14739–14749, 2021. 3
2021
-
[52]
Diffusion-HPC: Synthetic data generation for human mesh recovery in challenging domains
Zhenzhen Weng, Laura Bravo-S ´anchez, and Serena Yeung- Levy. Diffusion-HPC: Synthetic data generation for human mesh recovery in challenging domains. In 3DV, pages 257–
-
[53]
MagicPony: Learning articu- lated 3D animals in the wild
Shangzhe Wu, Ruining Li, Tomas Jakab, Christian Rup- precht, and Andrea Vedaldi. MagicPony: Learning articu- lated 3D animals in the wild. In CVPR, pages 8792–8802,
-
[54]
DatasetDM: Synthesizing data with perception anno- tations using diffusion models
Weijia Wu, Yuzhong Zhao, Hao Chen, Yuchao Gu, Rui Zhao, Yefei He, Hong Zhou, Mike Zheng Shou, and Chunhua Shen. DatasetDM: Synthesizing data with perception anno- tations using diffusion models. NeurIPS, 36:54683–54695,
-
[55]
CASA: Category-agnostic skeletal ani- mal reconstruction
Yuefan Wu, Zeyuan Chen, Shaowei Liu, Zhongzheng Ren, and Shenlong Wang. CASA: Category-agnostic skeletal ani- mal reconstruction. In NeurIPS, pages 28559–28574. Curran Associates, Inc., 2022. 2
2022
-
[56]
Animal3D: A comprehensive dataset of 3D ani- mal pose and shape
Jiacong Xu, Yi Zhang, Jiawei Peng, Wufei Ma, Artur Jesslen, Pengliang Ji, Qixin Hu, Jiehua Zhang, Qihao Liu, Jiahao Wang, et al. Animal3D: A comprehensive dataset of 3D ani- mal pose and shape. In ICCV, pages 9099–9109, 2023. 2, 6, 1
2023
-
[57]
ViTPose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. In NeurIPS, pages 38571–38584. Curran Associates, Inc., 2022. 6
2022
-
[58]
Freeman, and Ce Liu
Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic, Forrester Cole, Huiwen Chang, Deva Ramanan, William T. Freeman, and Ce Liu. LASR: Learning articulated shape reconstruction from a monocular video. In CVPR, 2021. 2
2021
-
[59]
ViSER: Video-specific surface embeddings for articulated 3D shape reconstruction
Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vla- sic, Forrester Cole, Ce Liu, and Deva Ramanan. ViSER: Video-specific surface embeddings for articulated 3D shape reconstruction. In NeurIPS, 2021. 2
2021
-
[60]
BANMo: Build- ing animatable 3D neural models from many casual videos
Gengshan Yang, Minh V o, Natalia Neverova, Deva Ra- manan, Andrea Vedaldi, and Hanbyul Joo. BANMo: Build- ing animatable 3D neural models from many casual videos. In CVPR, pages 2853–2863, 2022. 2
2022
-
[61]
PPR: Physically plausible recon- struction from monocular videos
Gengshan Yang, Shuo Yang, John Z Zhang, Zachary Manch- ester, and Deva Ramanan. PPR: Physically plausible recon- struction from monocular videos. In ICCV, pages 3914– 3924, 2023. 2
2023
-
[62]
LASSIE: Learning articulated shapes from sparse image ensemble via 3D part discovery
Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Ru- binstein, Ming-Hsuan Yang, and Varun Jampani. LASSIE: Learning articulated shapes from sparse image ensemble via 3D part discovery. In NeurIPS, pages 15296–15308. Curran Associates, Inc., 2022. 2
2022
-
[63]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, pages 3836–3847, 2023. 3, 5
2023
-
[64]
Silvia Zuffi and Michael J. Black. AWOL: Analysis with- out synthesis using language. In ECCV, pages 1–19, Cham,
-
[65]
Jacobs, and Michael J
Silvia Zuffi, Angjoo Kanazawa, David W. Jacobs, and Michael J. Black. 3D menagerie: Modeling the 3D shape and pose of animals. In CVPR, 2017. 2, 3, 4
2017
-
[66]
in the wild
Silvia Zuffi, Angjoo Kanazawa, Tanya Berger-Wolf, and Michael J. Black. Three-D safari: Learning to estimate ze- bra pose, shape, and texture from images “in the wild”’. In ICCV, 2019. 3
2019
-
[67]
Silvia Zuffi, Ylva Mellbin, Ci Li, Markus Hoeschle, Hedvig Kjellstr¨om, Senya Polikovsky, Elin Hernlund, and Michael J. Black. V AREN: Very accurate and realistic equine network. In CVPR, pages 5374–5383, 2024. 2, 3 4
2024
-
[2025]
Springer Nature Switzerland. 3, 4
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.