REVIEW 4 major objections 4 minor 46 references
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper introduces JWB-DH-V1, a two-million-clip dataset and a three-track evaluation protocol meant to benchmark joint whole-body avatar and speech generation.
desk verdict Claims a large joint whole-body avatar benchmark but shows neither the dataset nor a joint evaluation; region-specific metrics are the one useful idea. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the curated dataset itself, with its region-specific annotations, paired with a three-track evaluation protocol: reference-free video metrics, co-speech fidelity metrics, and a Large-Audio-Language-Model win-rate for speech. The dataset is what enables the region-specific breakdown; the protocol is what turns raw clips and audio into comparable numerical scores.
What would settle it
Generate a clip in which the body and lip motion are frame-perfect but the audio track is shifted by a fixed delay; the benchmark's video and audio sub-scores would be unchanged while a human viewer would immediately see the synchronization is broken. If the protocol cannot flag this mismatch, it does not measure joint audio-video generation.
Extended reading notes
Core claim
The central claim is that the field lacks a benchmark tailored to joint whole-body avatar and speech generation, and that JWB-DH-V1 provides one. The dataset is described as containing 10,000 unique identities, each in roughly 200 scene configurations, for about 2 million video samples with annotations including segmentation, landmarks, bounding boxes for hands and legs, motion text, speech transcription, word boundaries, and ground truth audio; 20,000 samples are held out for evaluation. The protocol has three tracks: reference-free video metrics, co-speech fidelity metrics (FID, FVD, SSIM, PSNR, E-FID, CSIM), and speech audio evaluation using WER plus a Large-Audio-Language-Model win-rate. On this protocol, eight generative models are scored across face, hand, and whole-body regions, and the results show a consistent gap between face/hand quality and whole-body quality. The only joint audio-video model considered, Veo-3, is excluded because its outputs from a single frame were unstable.
Load-bearing premise
The benchmark assumes that a protocol which scores video and audio in separate tracks, and which excludes the only joint generation model, still measures joint audio-video generation quality.
Editorial extensions
If this is right
- Region-specific scoring shows consistent performance gaps between face/hand and whole-body regions, pointing to where generative models need work.
- The benchmark's 2 million annotated clips and 20,000 evaluation samples provide a common ground for comparing whole-body avatar generation systems.
- The video and speech protocols yield numerical scores that can be reported and compared without human ratings.
- Future versions that extend clips to 60 seconds would test long-horizon synchronization rather than short clips.
Reading between the lines
- The exclusion of the only joint model implies that current joint generation is too unstable to benchmark; a version that includes stable joint models would be needed to test the joint claim.
- Because the protocol scores video and audio separately, a model with perfect synchronization but mediocre per-track scores would not be recognized; adding a synchronization-aware metric or human study is a natural extension.
- The region annotations could let researchers ask whether models over-fit to faces at the expense of hands and whole-body, or whether whole-body errors stem from data imbalance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript introduces JWB-DH-V1, described as a large-scale benchmark for joint whole-body talking-avatar and speech generation, with a dataset of 10,000 identities, about 200 scene configurations per identity, and 2 million video clips, plus an evaluation protocol and an initial evaluation of several video-generation, talking-avatar, and TTS models. The paper reports separate video-quality metrics (SC, BC, MS, DD, AQ, IQ, FID, FVD, SSIM, PSNR, E-FID, CSIM) over whole-body, face, and hand regions (Table 2) and audio metrics WER and LALM-based win-rate (Table 1). Its stated conclusion is that current methods show consistent performance disparities between face/hand-centric and whole-body performance, indicating areas for future research. The dataset and evaluation tools are said to be publicly available at a GitHub repository.
Significance. Assessed as a benchmark paper, the contribution is currently a comparatively small empirical comparison rather than a validated resource. The authors deserve credit for making an associated GitHub repository available, for evaluating a broad set of recent video and TTS models, and for reporting results separately for whole-body, face, and hand regions, which is a useful format for diagnosing generation failures. If the dataset existed at the claimed scale and the protocol actually measured joint audio-video quality, JWB-DH-V1 would be a valuable addition. However, as described, the dataset curation is unverifiable and the evaluation protocol does not assess joint generation, so the main claims are not supported.
major comments (4)
- [Section 3.1] The central resource claim—'10,000 unique identities' and '2 million samples of video clips'—is presented without any description of data collection, annotation procedure, quality control, licensing, train/test split, or validation of the annotations shown in Figure 1. Because the paper's value as a benchmark depends on the dataset being curated and trustworthy, this omission makes the resource's existence and utility unassessable. A dataset card or a detailed curation section with annotation statistics is required.
- [Sections 3.2-3.4 and Abstract] The evaluation protocol is not joint. All twelve video metrics listed in Sections 3.2-3.3 and all audio metrics in Section 3.4 take only one modality as input; no metric compares a generated video with its paired generated audio, and no synchronization, lip-sync, or cross-modal consistency measure is defined. Section 4 then excludes Veo-3, the only joint audio-video model, 'due to instability in generating synchronized outputs from a single frame.' Consequently, the evaluation cannot substantiate the Abstract's claim of 'an evaluation protocol for assessing joint audio-video generation of whole-body animatable avatars.'
- [Table 2] Several entries are numerically implausible and suggest a preprocessing or scaling error. For instance, the whole-body row for Ha3/wo reports PSNR = 0.83 and SSIM = 0.016, with the hand region reporting PSNR = 0.49 and SSIM = 0.008, while other models in the same table report PSNR around 6-20 dB and SSIM 0.2-0.9. PSNR values below 1 dB are outside the normal range for images represented in [0,1] or [0,255], so the reported numbers cannot be trusted as meaningful quality scores. The conclusions in Section 4 rely on comparisons across these rows and must be recomputed.
- [Section 4] The protocol is not reproducible as written. The paper does not specify the number and selection of evaluation clips, video resolution and duration, generation prompts, random seeds, or the exact implementation and reference statistics for FID/FVD/E-FID. It also does not say how the region-specific masks (face, hand, whole body) are obtained or how many samples are used per model. These details are necessary for a benchmark paper.
minor comments (4)
- [Abstract] There are several typos and formatting issues: 'incidates' should be 'indicates', 'with10,000' needs a space, and 'Version I(JWB-DH-V1)' needs proper formatting.
- [Table 1 and Section 3.4] The baseline row gpt-4o-mini-tts has a dash for win-rate, but the text in Section 3.4 says candidates are compared against a strong baseline; define how ties and 'winner=0' are handled in the win-rate formula.
- [Section 3.1] The statement '20,000 samples are used for evaluation' is ambiguous; it should be clarified whether this is the test split, a subset, or the total evaluation set.
- [References] The reference list contains 'contentReference[oaicite:...]' artifacts (e.g., [1], [7], [25]), which should be removed before publication.
Circularity Check
No circularity found: JWB-DH-V1 is a dataset/benchmark paper with no derivation chain, so there is no claimed result that reduces to its inputs; the joint-protocol gap is a validity concern, not circularity.
full rationale
JWB-DH-V1 is a benchmark contribution rather than a derived result. Its load-bearing claims are existential and empirical: the dataset exists at the stated scale, and the protocol evaluates audio-video generation. Neither claim is obtained by deriving a conclusion from assumptions that already contain it. The evaluation sections specify standard, externally defined metrics: six reference-free video metrics (SC, BC, MS, DD, AQ, IQ), six frame-quality/temporal metrics (FID, FVD, SSIM, PSNR, E-FID, CSIM), and audio metrics (WER plus a LALM-based win-rate). The win-rate formula W(T_i) = (P(winner=index_i)+0.5*P(winner=0))/n is a definitional aggregation, not a fitted parameter presented as a prediction. No quantity is fitted on a subset of data and then relabeled as a prediction. The bibliography contains no citations authored by the present authors, so the self-citation load-bearing and uniqueness-imported-from-authors patterns cannot be substantiated. The strongest criticism of the paper—that the protocol is not actually joint because video-only and audio-only metrics are reported separately and the sole joint model Veo-3 is excluded 'due to instability in generating synchronized outputs from a single frame'—is a mismatch between the benchmark's stated purpose and its implemented evaluation. That is a correctness or validity concern, not circularity: it does not exhibit an equation that equates a claimed result to its inputs by construction. Under the hard rule that circularity must be shown by quoted reduction, no circular step is present, so the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Existing quality metrics (DINO, CLIP, RAFT, LAION aesthetic, MUSIQ) are valid proxies for perceptual quality and coherence of whole-body avatars.
- domain assumption Gemini 2.5 Pro, used as a LALM judge, reliably and fairly ranks TTS systems on prosody, pausing, and expressiveness.
- ad hoc to paper The dataset annotations (bounding boxes, landmarks, motion text, speech transcripts) are accurate and complete, and the dataset itself exists as claimed.
Cite this review
Pith. "Pith review of JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1." pith.science (2026). https://pith.science/paper/P67RDH4A
@misc{pith2026250720987,
author = {Pith},
title = {Pith review of: JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1},
year = {2026},
howpublished = {\url{https://pith.science/paper/P67RDH4A}},
note = {Machine review of arXiv:2507.20987}
}
read the original abstract
Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly generating whole-body motion and natural speech. Current approaches lack comprehensive evaluation frameworks that assess both visual and audio quality, and there are insufficient benchmarks for region-specific performance analysis. To address these gaps, we introduce the Joint Whole-Body Talking Avatar and Speech Generation Version I(JWB-DH-V1), comprising a large-scale multi-modal dataset with 10,000 unique identities across 2 million video samples, and an evaluation protocol for assessing joint audio-video generation of whole-body animatable avatars. Our evaluation of SOTA models reveals consistent performance disparities between face/hand-centric and whole-body performance, which incidates essential areas for future research. The dataset and evaluation tools are publicly available at https://github.com/deepreasonings/WholeBodyBenchmark.
Figures
Reference graph
Works this paper leans on
-
[1]
J. Betker. Better Speech Synthesis Through Scaling: Tor- toise, an expressive multi-voice tts system.arXiv preprint arXiv:2305.07243, May 2023. URLhttps://arxiv. org/abs/2305.07243. Introduces TorToiSe, combining autoregressive and diffusion models for high-quality multi- voice TTS; code released under Apache-2.0 :contentRefer- ence[oaicite:1]index=1. 3, 4
arXiv 2023
-
[2]
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y . Levi, Z. English, V . V oleti, A. Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127,
- [3]
-
[4]
J. Cui, H. Li, Y . Yao, H. Zhu, H. Shang, K. Cheng, H. Zhou, S. Zhu, and J. Wang. Hallo2: Long-duration and high- resolution audio-driven portrait image animation.arXiv preprint arXiv:2410.07718, 2024. 1, 3
arXiv 2024
-
[5]
J. Cui, H. Li, Y . Zhan, H. Shang, K. Cheng, Y . Ma, S. Mu, H. Zhou, J. Wang, and S. Zhu. Hallo3: Highly dynamic and realistic portrait image animation with diffusion transformer networks.arXiv preprint arXiv:2412.00733, 2024. 1, 3, 4
arXiv 2024
- [6]
-
[7]
Eleven Multilingual v2: A foundational multi- lingual text-to-speech model supporting 29 languages
ElevenLabs. Eleven Multilingual v2: A foundational multi- lingual text-to-speech model supporting 29 languages. Of- ficial ElevenLabs blog, Aug. 2023. Released out of beta August22,2023; supports emotionally rich AI voice in 29 languages while preserving speaker identity :contentRefer- ence[oaicite:1]index=1. 3, 4
work page 2023
-
[8]
J. Guan, Z. Zhang, H. Zhou, T. Hu, K. Wang, D. He, H. Feng, J. Liu, E. Ding, Z. Liu, et al. Stylesync: High-fidelity gen- eralized and personalized lip sync in style-based generator. InProceedings of the IEEE/CVF CVPR, pages 1505–1515,
Show all 46 references
-
[9]
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet. Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646, 2022. 1
2022
-
[10]
S. Hu, Y . Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y . Fang, Y . Huang, W. Zhao, X. Zhang, Z. Zhai, N. Ding, C. Jia, G. Zeng, D. Li, Z. Liu, and M. Sun. MiniCPM: Un- veiling the Potential of Small Language Models with Scal- able Training Strategies.arXiv preprint arXiv:2...
2024 arXiv
-
[11]
Huang, G
H. Huang, G. Ma, N. Duan, X. Chen, C. Wan, R. Ming, T. Wang, B. Wang, Z. Lu, A. Li, et al. Step-video-ti2v techni- cal report: A state-of-the-art text-driven image-to-video gen- eration model.arXiv preprint arXiv:2503.11251, 2025. 1, 4
2025 arXiv
-
[12]
Huynh-Thu and M
Q. Huynh-Thu and M. Ghanbari. The accuracy of psnr in predicting video quality for different video scenes and frame rates.Telecommunication systems, 49:35–48, 2012. 3
2012
-
[13]
X. Ji, X. Hu, Z. Xu, J. Zhu, C. Lin, Q. He, J. Zhang, D. Luo, Y . Chen, Q. Lin, et al. Sonic: Shifting focus to global audio perception in portrait animation.arXiv preprint arXiv:2411.16331, 2024. 1, 3
2024 arXiv
-
[14]
Jiang, C
J. Jiang, C. Liang, J. Yang, G. Lin, T. Zhong, and Y . Zheng. Loopy: Taming audio-driven portrait avatar with long-term motion dependency.arXiv preprint arXiv:2409.02634, 2024. 1
2024 arXiv
-
[15]
J. Ke, Q. Wang, Y . Wang, P. Milanfar, and F. Yang. Musiq: Multi-scale image quality transformer. InProceedings of the IEEE/CVF international conference on computer vision, pages 5148–5157, 2021. 3
2021
-
[16]
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al. Hunyuanvideo: A sys- tematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024. 1, 3, 4
2024 arXiv
-
[17]
Laion aesthetic predictor.https : / / github.com/LAION-AI/aesthetic-predictor,
LAION-AI. Laion aesthetic predictor.https : / / github.com/LAION-AI/aesthetic-predictor,
-
[18]
C. Li, C. Zhang, W. Xu, J. Xie, W. Feng, B. Peng, and W. Xing. Latentsync: Audio conditioned latent diffusion models for lip sync.arXiv preprint arXiv:2412.09262, 2024. 1
2024 arXiv
-
[19]
X. Li, W. Chu, Y . Wu, W. Yuan, F. Liu, Q. Zhang, F. Li, H. Feng, E. Ding, and J. Wang. Videogen: A reference- guided latent diffusion approach for high definition text-to- video generation.arXiv preprint arXiv:2309.00398, 2023. 1
2023 arXiv
-
[20]
Li, Z.-L
Z. Li, Z.-L. Zhu, L.-H. Han, Q. Hou, C.-L. Guo, and M.-M. Cheng. Amt: All-pairs multi-field transforms for efficient frame interpolation. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 9801–9810, 2023. 2
2023
-
[21]
K. Liu, W. Li, L. Chen, S. Wu, Y . Zheng, J. Ji, F. Zhou, R. Jiang, J. Luo, H. Fei, and T. Chua. Javis- DiT: Joint audio–video diffusion transformer with hierarchi- cal spatio-temporal prior synchronization.arXiv preprint arXiv:2503.23377, Mar. 2025. 1
2025
-
[22]
R. R. Manku, Y . Tang, X. Shi, M. Li, and A. Smola. Emergenttts-eval: Evaluating tts models on complex prosodic, expressiveness, and linguistic challenges using model-as-a-judge.arXiv preprint arXiv:2505.23009, May
-
[23]
R. Meng, X. Zhang, Y . Li, and C. Ma. Echomimicv2: To- wards striking, simplified, and semi-body human animation. arXiv preprint arXiv:2411.10061, 2024. 1, 3, 4
2024
-
[24]
D. Min, M. Song, E. Ko, and S. J. Hwang. Styletalker: One- shot style-based audio-driven talking head video generation. arXiv preprint arXiv:2208.10922, 2022. 1
2022 arXiv
-
[25]
GPT-4o-Mini- Audio-Preview: A compact, cost-efficient audio-capable 5 multimodal model
OpenAI & Microsoft Azure AI Foundry. GPT-4o-Mini- Audio-Preview: A compact, cost-efficient audio-capable 5 multimodal model. Azure AI Services / OpenAI Plat- form documentation, Dec. 2024. Public preview of audio- in/audio-out capabilities via Chat Completions API (model versi...
2024
-
[26]
X. Peng, Z. Zheng, and C. S.et al.Open-sora 2.0: Training a commercial-level video generation model in200k.arXiv preprint arXiv:2503.09642, 2025. 3, 4
2025 arXiv
-
[27]
Prajwal, R
K. Prajwal, R. Mukhopadhyay, V . P. Namboodiri, and C. Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. Inthe 28th ACM international conference on multimedia, pages 484–492, 2020. 1
2020
-
[28]
Z. Qin, R. Zheng, Y . Wang, T. Li, Z. Zhu, M. Yang, M. Yang, and L. Wang. Versatile multimodal controls for whole-body talking human animation.arXiv preprint arXiv:2503.08714,
-
[29]
D. Qiu, Z. Fei, R. Wang, J. Bai, C. Yu, M. Fan, G. Chen, and X. Wen. Skyreels-a1: Expressive portrait animation in video diffusion transformers.arXiv preprint arXiv:2502.10841,
-
[30]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learn- ing, pages 8748–8763. PmLR, 2021. 2
2021
-
[31]
Singer, A
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al. Make-a-video: Text-to-video generation without text-video data.arXiv preprint arXiv:2209.14792, 2022. 1
2022 arXiv
-
[32]
Bark: Text-prompted generative audio model
Suno AI. Bark: Text-prompted generative audio model. GitHub and Hugging Face repositories, Apr. 2023. Released under MIT license in April/2023; transformer-based model capable of generating realistic multilingual speech, music, sound effects, and nonverbal audio. 3, 4
2023
-
[33]
Teed and J
Z. Teed and J. Deng. Raft: Recurrent all-pairs field trans- forms for optical flow. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, pages 402–419. Springer. 3
2020
-
[34]
Google introduces stable gemini 2.5 flash and pro, previews gemini 2.5 flash-lite.The Economic Times (India)
The Economic Times. Google introduces stable gemini 2.5 flash and pro, previews gemini 2.5 flash-lite.The Economic Times (India). Announces general availability of Gemini2.5Pro alongside Gemini2.5Flash and preview of Flash-Lite. 4
-
[35]
S. Tu, Z. Xing, X. Han, Z.-Q. Cheng, Q. Dai, C. Luo, and Z. Wu. Stableanimator: High-quality identity-preserving hu- man image animation.arXiv preprint arXiv:2411.17697. 1
-
[36]
Unterthiner, S
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly. Towards accurate generative models of video: A new metric & challenges.arXiv preprint arXiv:1812.01717, 2018. 3
2018 arXiv
-
[37]
A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, J. Wang, J. Zhang, J. Zhou, J. Wang, J. Chen, K. Zhu, K. Zhao, K. Yan, L. Huang, M. Feng, N. Zhang, P. Li, P. Wu, R. Chu, R. Feng, S. Zhang, S. Sun, T. Fang, T. Wang, T. Gui, T. Weng, T. Shen...
2025 arXiv
-
[38]
Y . Wang, X. Chen, X. Ma, S. Zhou, Z. Huang, Y . Wang, C. Yang, Y . He, J. Yu, P. Yang, et al. Lavie: High- quality video generation with cascaded latent diffusion mod- els.IJCV, pages 1–20, 2024. 1
2024
-
[39]
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4): 600–612, 2004. 3
2004
-
[40]
J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y . Fan, K. Dang, B. Zhang, X. Wang, Y . Chu, and J. Lin. Qwen2.5-omni technical report.arXiv preprint arXiv:2503.20215, Mar. 2025. URLhttps://arxiv. org/abs/2503.20215. Describes Qwen2.5-Omni, an end-to-end streami...
2025 arXiv
-
[41]
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y . Yang, W. Hong, X. Zhang, G. Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 1
2024 arXiv
-
[42]
Zhang, J
S. Zhang, J. Wang, Y . Zhang, K. Zhao, H. Yuan, Z. Qin, X. Wang, D. Zhao, and J. Zhou. I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models. arXiv preprint arXiv:2311.04145, 2023. 1
2023 arXiv
-
[43]
L. Zhao, L. Feng, D. Ge, F. Yi, C. Zhang, X. Zhang, and X. Li. UniForm: A unified diffusion transformer for au- dio–video generation.arXiv preprint arXiv:2502.03897, Feb. 2025. 1
2025 arXiv
-
[44]
H. Zhou, Y . Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu. Pose-controllable talking face generation by implicitly mod- ularized audio-visual representation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4176–4186, 2021. 1 6
2021
-
[2022]
Accessed: 2025-04-21. 3
2025
-
[2025]
org / abs / 2505
URLhttps : / / arxiv . org / abs / 2505 . 23009. Introduces the EmergentTTS-Eval benchmark and a model-as-a-judge evaluation framework using a Large Audio Language Model (LALM). 4
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.