Pith. sign in

REVIEW 1 major objections 30 references

VolHuMe: a High-Resolution Large Scale Dataset of Volumetric Human Meshes

T0 review · 1 major / 0 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read VolHuMe dataset uses close-range cameras with 64 RGB and 32 depth sensors to capture high-fidelity 4D human meshes from 104 subjects.

desk verdict VolHuMe is a new dataset of 104 subjects with close-range capture and rich annotations, but the paper gives no metrics to show it actually delivers higher fidelity than prior full-body datasets. read the letter →

arxiv 2606.23062 v1 pith:53U6A7S4 submitted 2026-06-22 cs.GR cs.CV

classification cs.GRcs.CV
keywords humanmeshdataset4Dreconstructionvolumetriccapturehigh-resolutionscanningSMPL-Xbodygarmentsegmentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces VolHuMe, a dataset of high-quality 4D human scans from 104 subjects captured in a volumetric studio. It relies on a close-range setup with 64 RGB and 32 depth cameras to preserve fine-grained body-part details. The dataset supplies extensive ground truth including SMPL-X fits, high-resolution meshes, multi-view images, rigged meshes, point clouds, garment segmentation, and detailed hand and face geometry. Benchmarks on 3D and 4D reconstruction tasks are presented to demonstrate the dataset quality and highlight shortcomings in existing testbeds.

What carries the argument

The close-range, high-resolution capture setup using 64 RGB and 32 depth cameras in a volumetric studio.

What would settle it

A controlled experiment showing no measurable gain in reconstruction accuracy or visual quality metrics when models are evaluated on VolHuMe versus prior datasets.

Watch

Extended reading notes

Core claim

VolHuMe is a dataset of high-quality 4D human scans captured with a state-of-the-art volumetric studio using 64 RGB and 32 depth cameras. It contains individual captures of 104 subjects and provides extensive ground truth, including SMPL-X, high-resolution meshes, multi-view RGB/depth images, rigged meshes, point clouds, garment segmentation, and detailed hand and facial geometry. Unlike prior datasets that primarily rely on full-body imagery, VolHuMe uses a close-range, high-resolution capture setup that preserves fine-grained body-part details, improving geometric fidelity and texture resolution.

Load-bearing premise

The state-of-the-art volumetric studio using 64 RGB and 32 depth cameras in a close-range setup produces ground-truth data of meaningfully higher geometric and texture fidelity than prior full-body capture methods.

Editorial extensions

If this is right

  • Enables more precise benchmarking of 3D and 4D human reconstruction algorithms.
  • Reveals limitations in current evaluation testbeds used for human reconstruction.
  • Supplies detailed annotations for hands, faces, and garments to support targeted method development.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The higher-resolution data may support training of models that better recover subtle surface details in uncontrolled environments.
  • Applications in animation and virtual reality could benefit from the included garment and rigging information.
  • The capture protocol could be adapted to other body-centric domains such as medical imaging or performance capture.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 0 minor

Summary. The manuscript introduces VolHuMe, a dataset of high-quality 4D human scans from 104 subjects captured in a volumetric studio using 64 RGB and 32 depth cameras. It supplies extensive ground truth including SMPL-X, high-resolution meshes, multi-view RGB/depth images, rigged meshes, point clouds, garment segmentation, and detailed hand/facial geometry. The paper claims that its close-range capture setup yields better geometric fidelity and texture resolution than prior full-body datasets, and it benchmarks state-of-the-art methods on 3D and 4D reconstruction tasks to demonstrate dataset quality and expose limitations of existing testbeds.

Significance. A large-scale dataset with rich annotations and claimed high fidelity could advance 3D/4D human reconstruction research by enabling better training and more challenging evaluation. The benchmarking effort is a positive step toward identifying method weaknesses. However, without demonstrated fidelity gains, the dataset's added value over existing resources remains potential rather than established.

major comments (1)
  1. [Abstract] Abstract: The claim that the close-range setup with 64 RGB + 32 depth cameras 'preserves fine-grained body-part details, improving geometric fidelity and texture resolution' is not accompanied by any quantitative support (e.g., reconstruction error metrics, resolution comparisons, or side-by-side results versus prior full-body capture datasets). This assumption is load-bearing for the paper's positioning as a higher-fidelity resource.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the detailed review and constructive comments. We address the single major comment below and will revise the manuscript accordingly to strengthen the presentation of the dataset's claimed advantages.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The claim that the close-range setup with 64 RGB + 32 depth cameras 'preserves fine-grained body-part details, improving geometric fidelity and texture resolution' is not accompanied by any quantitative support (e.g., reconstruction error metrics, resolution comparisons, or side-by-side results versus prior full-body capture datasets). This assumption is load-bearing for the paper's positioning as a higher-fidelity resource.

    Authors: We agree that the abstract claim would be strengthened by quantitative evidence or a more qualified statement. The positioning relies on the hardware configuration (close-range multi-view capture versus typical full-body setups), but the current manuscript does not include direct side-by-side metrics such as reconstruction error or resolution comparisons against prior datasets. In the revision we will either (a) add a concise quantitative comparison where feasible using available data or (b) revise the abstract wording to describe the setup's design intent without asserting unquantified improvements. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: dataset collection and benchmarking paper with no derivations or self-referential predictions

full rationale

The paper introduces a new volumetric human mesh dataset captured via a 64 RGB + 32 depth camera studio and benchmarks existing reconstruction methods on it. No equations, fitted parameters, predictions, or derivation chains appear in the provided text. Claims about improved geometric/texture fidelity are presented as direct consequences of the described capture setup rather than outputs of any self-referential logic, self-citation chain, or ansatz. This is a standard empirical contribution with no load-bearing steps that reduce to inputs by construction.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

No mathematical derivations, fitted parameters, or new theoretical entities are involved; the paper describes a data acquisition effort.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VolHuMe: a High-Resolution Large Scale Dataset of Volumetric Human Meshes." pith.science (2026). https://pith.science/paper/53U6A7S4

@misc{pith2026260623062,
  author       = {Pith},
  title        = {Pith review of: VolHuMe: a High-Resolution Large Scale Dataset of Volumetric Human Meshes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/53U6A7S4}},
  note         = {Machine review of arXiv:2606.23062}
}
read the original abstract

We introduce VolHuMe, a dataset of high-quality 4D human scans captured with a state-of-the-art volumetric studio using 64 RGB and 32 depth cameras. VolHuMe contains individual captures of 104 subjects and provides extensive ground truth, including SMPL-X, high-resolution meshes, multi-view RGB/depth images, rigged meshes, point clouds, garment segmentation, and detailed hand and facial geometry. Unlike prior datasets that primarily rely on full-body imagery, VolHuMe uses a close-range, high-resolution capture setup that preserves fine-grained body-part details, improving geometric fidelity and texture resolution. We benchmark VolHuMe on state-of-the-art methods across 3D and 4D human reconstruction tasks, showcasing the dataset's quality and exposing the limitations of current evaluation testbeds.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 3 canonical work pages

  1. [1]

    VolHuMe: a High-Resolution Large Scale Dataset of Volumetric Human Meshes

    INTRODUCTION 4D human reconstruction aims to generate realistic avatars of humans in motion, capturing geometry, appearance, clothing, and fine-grained articulation such as facial expressions and hand movements. These representations are central to many computer vision and graphics applications, including virtual environments, fashion, and embodied AI. Da...

  2. [2]

    In-the-wild datasets (e.g

    RELATED WORK 4D Human Datasets.Deep learning for human pose esti- mation (HPE), human mesh recovery (HMR), and 4D human modeling depends on datasets that trade off realism, capture conditions, and annotation quality. In-the-wild datasets (e.g. EMDB [3]) often exhibit noisy imagery and imperfect mesh annotations, while synthetic datasets (e.g., AGORA, BEDL...

  3. [3]

    When the participant stands in the area, the distance to the cameras Fig

    DATA ACQUISITION AND PROCESSING Sensing HardwareThe capture area forms a cylindrical volume with a 2.5-meter diameter and a 3-meter height, pro- viding ample space for subjects to move freely. When the participant stands in the area, the distance to the cameras Fig. 2: Details of the high-quality data. The rigged mesh (left), alongside with a sample frame...

  4. [4]

    View Synthesis and 3D Mesh Estimation We evaluate view synthesis and mesh estimation using all views from the first frame of each subject

    EXPERIMENTS 4.1. View Synthesis and 3D Mesh Estimation We evaluate view synthesis and mesh estimation using all views from the first frame of each subject. Our evaluation includes four categories of methods: photogrammetry-based methods (Reality Capture, Metashape, and Meshroom), NeRF-based methods (Nerfacto, V olNerfacto- big, and Depth-V olNerfacto), 3D...

  5. [5]

    Unlike existing benchmarks, which rely on full-body cap- ture configurations, V olHuMe adopts a sparse multi-view setup with close-range, partial observations

    CONCLUSIONS We presented V olHuMe, a large-scale dataset of high-resolution 4D human meshes acquired with a professional volumetric capture setup and enriched with comprehensive annotations. Unlike existing benchmarks, which rely on full-body cap- ture configurations, V olHuMe adopts a sparse multi-view setup with close-range, partial observations. The ex...

  6. [6]

    1https://civit.fi/

    ACKNOWLEDGEMENTS This work was carried out with the support of Centre for Im- mersive Visual Technologies (CIVIT) research infrastructure, Tampere University, Finland1. 1https://civit.fi/

  7. [7]

    Semi-automatic Human Parametric Model Registra- tion As discussed in Sec

    SUPPLEMENTARY MATERIAL 7.1. Semi-automatic Human Parametric Model Registra- tion As discussed in Sec. 2, most existing methods for 4D human registration rely on parametric mesh initialization using mod- els such as SMPL or SMPL-X. In our dataset, we introduce a novel, highly accurate method for the semi-automatic regis- tration of SMPL-X to align the high...

  8. [8]

    AGORA: Avatars in geography optimized for regression analysis,

    Priyanka Patel, Chun-Hao P. Huang, Joachim Tesch, David T. Hoffmann, Shashank Tripathi, and Michael J. Black, “AGORA: Avatars in geography optimized for regression analysis,” inCVPR, 2021. 1, 2, 3

Show all 30 references
  1. [9]

    Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion,

    Michael J Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang, “Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion,” inCVPR,

  2. [10]

    Emdb: The electromagnetic database of global 3d human pose and shape in the wild,

    Manuel Kaufmann, Jie Song, Chen Guo, Kaiyue Shen, Tianjian Jiang, Chengcheng Tang, Juan Jos´e Z´arate, and Otmar Hilliges, “Emdb: The electromagnetic database of global 3d human pose and shape in the wild,” in ICCV, 2023. 1, 2

  3. [11]

    Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,

    Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cris- tian Sminchisescu, “Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,”IEEE transactions on pattern analysis and machine intelligence, vol. 36, no. 7, pp. 1325–1339,

  4. [12]

    Panoptic studio: A massively multiview system for social motion capture,

    Hanbyul Joo, Hao Liu, Lei Tan, Lin Gui, Bart Nabbe, Iain Matthews, Takeo Kanade, Shohei Nobuhara, and Yaser Sheikh, “Panoptic studio: A massively multiview system for social motion capture,” inProceedings of the IEEE international conference on computer vision, 2015, pp. 3334–...

  5. [13]

    Human4d: A human- centric multimodal dataset for motions and immersive media,

    Anargyros Chatzitofis, Leonidas Saroglou, Prodromos Boutis, Petros Drakoulis, Nikolaos Zioulis, Shishir Sub- ramanyam, Bart Kevelham, Caecilia Charbonnier, Pablo Cesar, Dimitrios Zarpalas, et al., “Human4d: A human- centric multimodal dataset for motions and immersive media,”I...

  6. [14]

    Humbi: A large multiview dataset of human body expressions and benchmark challenge,

    Jae Shin Yoon, Zhixuan Yu, Jaesik Park, and Hyun Soo Park, “Humbi: A large multiview dataset of human body expressions and benchmark challenge,”IEEE Transac- tions on Pattern Analysis and Machine Intelligence, vol. 45, no. 1, pp. 623–640, 2021. 2, 3

  7. [15]

    Structured local radiance fields for human avatar modeling,

    Zerong Zheng, Han Huang, Tao Yu, Hongwen Zhang, Yandong Guo, and Yebin Liu, “Structured local radiance fields for human avatar modeling,” inCVPR, 2022. 2, 3

  8. [16]

    Avatarrex: Real-time expres- sive full-body avatars,

    Zerong Zheng, Xiaochen Zhao, Hongwen Zhang, Bon- ing Liu, and Yebin Liu, “Avatarrex: Real-time expres- sive full-body avatars,”ACM Transactions on Graphics (TOG), 2023. 2

  9. [17]

    Humanrf: High-fidelity neural radiance fields for humans in motion,

    Mustafa Is ¸ık, Martin R ¨unz, Markos Georgopoulos, Taras Khakhulin, Jonathan Starck, Lourdes Agapito, and Matthias Nießner, “Humanrf: High-fidelity neural radiance fields for humans in motion,”ACM Transac- tions on Graphics (TOG), 2023. 2, 3

  10. [18]

    Dna-rendering: A diverse neural actor repository for high-fidelity human- centric rendering,

    Wei Cheng, Ruixiang Chen, Siming Fan, Wanqi Yin, Keyu Chen, Zhongang Cai, Jingbo Wang, Yang Gao, Zhengming Yu, Zhengyu Lin, et al., “Dna-rendering: A diverse neural actor repository for high-fidelity human- centric rendering,” inICCV, 2023. 2, 3

  11. [19]

    Mvhumannet++: A large-scale dataset of multi-view daily dressing human captures with richer annotations for 3d human digitization,

    Chenghong Li, Hongjie Liao, Yihao Zhi, Xihe Yang, Zhengwentai Sun, Jiahao Chang, Shuguang Cui, and Xi- aoguang Han, “Mvhumannet++: A large-scale dataset of multi-view daily dressing human captures with richer annotations for 3d human digitization,”arXiv preprint arXiv:2505.018...

  12. [20]

    Expressive body capture: 3d hands, face, and body from a single image,

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black, “Expressive body capture: 3d hands, face, and body from a single image,” inCVPR,

  13. [21]

    Nerfs- tudio: A modular framework for neural radiance field development,

    Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, et al., “Nerfs- tudio: A modular framework for neural radiance field development,” inACM SIGGRAPH 2023 Conference Proceedings, 2023. 2

  14. [22]

    Animatable gaussians: Learning pose-dependent gaus- sian maps for high-fidelity human avatar modeling,

    Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu, “Animatable gaussians: Learning pose-dependent gaus- sian maps for high-fidelity human avatar modeling,” in CVPR, 2024. 2, 4

  15. [23]

    Gps-gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis,

    Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu, “Gps-gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis,” in CVPR, 2024. 2, 4

  16. [24]

    Radiant foam: Real-time differentiable ray tracing,

    Shrisudhan Govindarajan, Daniel Rebain, Kwang Moo Yi, and Andrea Tagliasacchi, “Radiant foam: Real-time differentiable ray tracing,”arXiv:2502.01157, 2025. 2, 3

  17. [25]

    Function4d: Real-time human volumetric capture from very sparse consumer rgbd sensors,

    Tao Yu, Zerong Zheng, Kaiwen Guo, Pengpeng Liu, Qionghai Dai, and Yebin Liu, “Function4d: Real-time human volumetric capture from very sparse consumer rgbd sensors,” inCVPR, 2021. 3

  18. [26]

    Deephuman: 3d human reconstruction from a single image,

    Zerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai, and Yebin Liu, “Deephuman: 3d human reconstruction from a single image,” inICCV, 2019. 3

  19. [27]

    Learning a model of facial shape and expression from 4d scans.,

    Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero, “Learning a model of facial shape and expression from 4d scans.,”ACM Transactions on graphics (TOG), vol. 36, no. 6, pp. 194–1, 2017. 3

  20. [28]

    Monocular expressive body regression through body-driven atten- tion,

    Vasileios Choutas, Georgios Pavlakos, Timo Bolkart, Dimitrios Tzionas, and Michael J Black, “Monocular expressive body regression through body-driven atten- tion,” inECCV, 2020. 1, 2

  21. [29]

    Estimating garment patterns from static scan data,

    Seungbae Bang, Maria Korosteleva, and Sung-Hee Lee, “Estimating garment patterns from static scan data,” Computer Graphics Forum, vol. 40, 05 2021. 1

  22. [30]

    CloSe: A 3D clothing segmentation dataset and model,

    Dimitrije Anti ´c, Garvita Tiwari, Batuhan Ozcomlekci, Riccardo Marin, and Gerard Pons-Moll, “CloSe: A 3D clothing segmentation dataset and model,” in3DV,

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.