{"id":"05bf15aa-c8dc-42bd-885e-07cacefc32ae","arxiv_id":"2508.08891","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper previews a claimed 2M-clip multimodal benchmark for whole-body talking avatar video generation, with standard metrics and an initial evaluation of eight open-source models.","lead":"The paper previews WB-DH, a claimed dataset of two million video clips of talking whole-body avatars with multi-modal annotations, plus an evaluation using twelve standard metrics. It reviews eight open-source video and avatar models, but gives no dataset details or evidence that the data is publicly available.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset claim rests on an unverified external repository; the manuscript contains no internal evidence for the 2M-sample open-source benchmark assertion.","rationale":"I read the paper as a benchmark-dataset proposal; its central claim is the existence and availability of WB-DH with 2M annotated clips. The manuscript provides no way to verify that claim from within the paper. The reader's weakest assumption—that the dataset actually exists as advertised and the repository is live—is exactly the load-bearing condition. I also note an explicit internal caveat: the Conclusion says a 'formal Version 1' is still being developed, which strengthens the concern that the current artifact is preliminary. The evaluation tables do not rescue the paper because they depend on the dataset's existence. A single external check of the GitHub repository would settle the factual claim; until that check is performed, the manuscript as written does not support its central assertion. Therefore I agree with the reader's REJECT verdict and would not change it.","tokens_in":7474,"tokens_out":4291,"duration_ms":43740,"concrete_test":"At review time, visit https://github.com/deepreasonings/WholeBodyBenchmark and verify: (a) the repository resolves and contains dataset download links, annotation schemas, and evaluation code; (b) the dataset actually contains at least 2M video clips (or a documented sample manifest) with the claimed segmentation, landmark, bounding-box, motion-text, and transcription annotations; and (c) a 20K test split is defined. If any of these fail, the central claim is false. If the repository is live and complete, the core factual claim would be supported, though a data card and protocol details would still be needed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that WB-DH is an open-source, multimodal benchmark with approximately 2M video clips and detailed annotations, publicly accessible at a GitHub URL. The manuscript provides no data card, no collection or annotation protocol, no license information, no sample count verification, and no code or data in the submission. The only internal evidence is Figure 1, which shows annotations for a single keyframe, and Table 1, which reports evaluations but does not establish the dataset's existence or composition. The Conclusion further states that 'a formal Version 1 of the dataset' is still being developed, indicating the released artifact is explicitly preliminary. Therefore the entire contribution rests on an external repository that is never substantiated. If that repository is absent, empty, or contains fewer clips or different annotations than claimed, the paper's central claim fails. Evaluation details such as missing error bars and unspecified train/test splits are secondary; the dataset claim is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes WB-DH, claimed to be an open-source multi-modal benchmark containing approximately 2M video clips with annotations such as segmentation, landmarks, bounding boxes, motion text, and speech transcription. It defines two evaluation protocols (six reference-free video metrics and six reference-based metrics) and reports initial results for eight models across whole-body, face, and hand regions in Table 1. The paper is explicitly a preview: dataset construction is described in a single paragraph, and the conclusion states that a formal Version 1 of the dataset is still being developed.","tokens_in":7728,"tokens_out":6313,"duration_ms":61167,"significance":"If fully realized, a large-scale whole-body talking-avatar benchmark with multi-modal annotations would fill a genuine gap, because current benchmarks are head/upper-body focused. The combination of region-specific evaluation and co-speech metrics is sensible. However, the paper ships no verifiable dataset, no collection protocol, no code, and no statistical analysis. The manuscript's own statement that a formal Version 1 is being developed undermines the central claim as it stands. The significance cannot be assessed without the actual artifact and a reproducible evaluation.","major_comments":[{"comment":"The central claim that WB-DH is an open-source 2M-clip benchmark with detailed multi-modal annotations is not supported by the manuscript. Section 3.1 gives only the numbers (10,000 identities, 200 scene configurations, 2M clips, 20K test samples) and one illustrative keyframe in Figure 1. There is no collection protocol, annotation pipeline, inter-annotator agreement, quality control, license, or repository verification. The Conclusion further states that 'a formal Version 1 of the dataset' is being developed, implying the current artifact is preliminary. This is load-bearing: if the repository does not contain the advertised data, the paper's contribution fails. A data card and release status are required.","section":"Abstract and §3.1"},{"comment":"The evaluation is reported without any statistical support. No test-set size per model, no confidence intervals or error bars, and no significance testing are given. For example, the whole-body SC of 96.83% (Wan) vs 97.12% (OpenS) or FVD 750.51 vs 896.94 are treated as meaningful orderings without variance. Numerical claims of this precision are not interpretable. The authors should specify how the 20K samples are split across models and conditions and report repeated-run variance.","section":"§4, Table 1"},{"comment":"The metric definitions are too underspecified to be reproducible. 'Motion Smoothness (MS): Smoothness and physical plausibility of motion, computed using optical flow continuity' does not give the exact formula, frame window, or threshold. Similarly, SC/BC need the precise DINO/CLIP layers, pooling, and frame sampling. For FID/FVD/SSIM/PSNR/E-FID/CSIM, the paper does not state which real videos are used as reference, how generated videos are aligned, or which region crops are fed to the embedding networks. Code release or a precise protocol is needed for any of these numbers to be meaningful.","section":"§3.2-3.3"},{"comment":"The experimental setup for the eight models is not described: no prompts for text-driven models, no audio files or pose sequences for talking-avatar models, no resolution/duration, no preprocessing. The abbreviations ecv2/w and ecv2/wo are not defined in the text (only in the table). Table 1's 'GT' row also needs explanation: it is not clear whether these are ground-truth values and how a PSNR of ∞ or SSIM of 1 is used for comparison. Without this information, Table 1 cannot be reproduced or interpreted.","section":"§4"}],"minor_comments":[{"comment":"Typographical issues: 'over10,000', 'approximately200', '2M samples(video clips)', and '20K samplesare' are missing spaces. The sentence 'The annotation for each sample is demonstrated in Figure 1' should specify that Figure 1 shows one keyframe per column, not all samples.","section":"§3.1"},{"comment":"A caption should define GT, w/wo (with/without pose guidance), and make clear which direction is better for each metric. The table would benefit from bold-facing best non-GT values.","section":"Table 1"},{"comment":"Figure 2's qualitative comparison would be stronger with enlarged crops and a description of how frames were selected to avoid cherry-picking. The claim that Models 1-4 'show good performance' is not quantified in the figure.","section":"§3.2/Figure 2"},{"comment":"Reference [25] appears to be an arXiv preprint with incomplete bibliographic information; verify all URLs and preprint IDs.","section":"References"}],"recommendation":"reject","confidential_remarks":"This is a preview/position note rather than a completed benchmark paper. The central dataset claim is unverifiable from the manuscript, and the conclusion explicitly states that the formal dataset release is future work. I would encourage the authors to resubmit a full benchmark paper with a data card, anonymized repository verification, and a statistically grounded evaluation. I did not penalize the authors for the small scale of the initial evaluation, but the absence of any dataset documentation and statistical support is disqualifying."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a 5-page preview of a benchmark that, if real, would fill a genuine gap — but the paper itself provides no evidence the dataset exists. The 2M-clip, 10k-identity claim is a single sentence with no collection protocol, no data card, no annotation statistics, and no verification of the GitHub link. The conclusion even says a formal Version 1 is still in development, so the released artifact is explicitly preliminary.\n\nWhat's legitimately useful: the paper identifies a real gap — no standard benchmark for whole-body talking avatar video generation. Splitting evaluation into full body, face, and hands is a good idea, and the annotation schema in Figure 1 (segmentation, landmarks, boxes, motion text, transcription) is exactly the kind of multimodal signal you'd want. Running eight open-source models through standard metrics is a reasonable sanity check, and the table, while raw, gives a quick sense of relative performance.\n\nThe problems are proportional to the load-bearing central claim. First, there is no internal evidence for the dataset. No sample frames beyond one keyframe, no distribution over identities, scenes, or languages, no inter-annotator agreement, no mention of how clips were collected or filtered. The 20K test split is asserted. Second, Table 1 reports point estimates with no error bars, no test-set size, and no significance tests. Third, the metrics are all off-the-shelf; no new evaluation protocol is introduced beyond grouping by region. Fourth, the with/without pose-guidance runs for the talking-avatar models are not described well enough to interpret.\n\nNone of these flaws are disqualifying for a preview — but the paper's value hinges on the repository, and that is not substantiated. If the GitHub is live and the data is as described, this could become a useful resource. As written, it's an announcement, not a benchmark.\n\nWho would get value: practitioners who want a quick overview of how current models perform on whole-body tasks, but they should treat all numbers as provisional. I wouldn't cite it in its current form. For peer review, I'd suggest the authors submit a complete dataset paper with a data card, a sample viewer, and verification of the repository. Until then, I wouldn't spend a referee's time on it.","headline":"A dataset-preview paper whose entire utility rests on an unverified GitHub repo and a 2M-clip claim the manuscript does not back up; the gap it targets is real, but the paper is not yet a citable resource.","tokens_in":8145,"tokens_out":3208,"would_cite":false,"duration_ms":30186,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces WB-DH, an open two-million-clip benchmark for whole-body talking avatar generation, and reports initial results showing hands are the hardest region.","keywords":["whole-body talking avatar","video generation benchmark","multi-modal video dataset","diffusion video models","speech-gesture alignment","region-specific evaluation","reference-free metrics","digital human"],"falsifier":"Open the linked repository and verify the numbers: count the clips to check the total is 2,000,000, confirm more than 10,000 identities, and inspect that each clip carries body segmentation, landmarks, bounding boxes, motion text, and transcription, with the 20,000-clip test split present. Independently, rerun the twelve metrics on the released test split and see whether the model rankings match the paper's table.","tokens_in":7434,"feed_emoji":"🎬","tokens_out":10350,"duration_ms":100017,"temperature":0.7,"pith_summary":"This paper, framed as a preview, introduces WB-DH, an open benchmark meant to decide whether AI systems can generate whole-body talking avatars—a person who speaks while moving their body, hands, and face—from a single portrait. The authors' central claim is that existing talking-head datasets and metrics cannot measure this task because they stop at the shoulders, so WB-DH adds fine-grained multi-modal annotations (body segmentation, landmarks, body-part bounding boxes, motion text, and speech transcription) to a dataset totaling two million video clips across over ten thousand identities. To demonstrate the benchmark, the paper evaluates eight model configurations under a twelve-metric protocol that scores full body, face, and hands separately. The initial numbers indicate that video-generation models hold identity and motion coherence across the whole body, while talking-avatar models degrade sharply once hands and legs are in frame. If the dataset and tools are released as promised, the field gains a common yardstick for a task that currently has none.","feed_headline":"Two-million-clip dataset tests whole-body talking avatars","feed_subtitle":"The new benchmark scores face, hands, and full body separately, so failures are no longer hidden by head-only metrics.","key_machinery":"The load-bearing object is the WB-DH dataset plus its region-specific evaluation protocol. Each video clip is annotated with six aligned modalities—body segmentation, landmark positions, bounding boxes for hands, legs, and the whole body, motion-text descriptions, and speech transcription—so the same clip can guide or test speech-driven whole-body animation. The evaluation protocol then scores generated clips independently for the full body, the face, and the hands, combining six reference-free metrics (DINO-based subject consistency, CLIP-based background consistency, optical-flow motion smoothness and dynamic degree, aesthetic and imaging quality) with six co-speech metrics (FID, FVD, SSIM","core_discovery":"The central claim is that whole-body talking avatar generation is a distinct evaluation problem and that WB-DH is the benchmark for it. The dataset is described as containing over 10,000 unique identities, each appearing in roughly 200 scene configurations, for a total of 2 million video clips, with 20,000 clips reserved for testing. Each clip carries aligned annotations for body segmentation, landmarks, bounding boxes around hands, legs, and the whole body, a motion-text description of pose semantics, and a speech transcription. The evaluation framework pairs six reference-free video-generation metrics with six co-speech metrics and applies them independently to three spatial regions—full b","pith_inferences":["A testable extension not in the paper: crop hand patches and compute hand-specific FID or FVD, to check whether the large hand-region gap is an artifact of small image area or a real generative failure.","The paper's with- and without-pose-guidance comparison could be reused as an ablation: the gap between the two variants measures how much of talking-avatar quality comes from explicit body-pose conditioning.","If the claimed alignment between motion text and speech transcription holds, the dataset could be used to train gesture-speech coupling models directly, a use the paper only motivates.","The benchmark's two metric families often order models differently (pixel-level versus reference-free), so a reader should not assume they measure the same property."],"forward_implications":["WB-DH gives whole-body avatar systems a shared twelve-metric, three-region test, so different models can be compared on the same clips.","The initial results imply that speech-to-gesture quality and full-body motion cannot be judged by face-only metrics; hand-region scores separate the models far more sharply than face scores.","The multi-modal annotations let a system condition generation on motion text, pose boxes, and speech transcription at once, which the paper argues is necessary for fine-grained whole-body control.","Because six of the twelve metrics are reference-free, the benchmark can score generated videos even when no ground-truth video exists, the common case in open-domain generation.","The planned Version 1 extension to 60-second clips would let the benchmark test long-horizon audio-video coherence, not just short clips."],"supporting_citations":[{"why":"DINO self-supervised features, used to compute the Subject Consistency score.","marker":"[2]"},{"why":"CLIP features, used to compute the Background Consistency score.","marker":"[27]"},{"why":"Defines FID, E-FID, and CSIM for identity-sensitive portrait evaluation; the benchmark's co-speech metrics draw on it.","marker":"[25]"},{"why":"Defines Fréchet Video Distance for measuring temporal coherence of generated videos.","marker":"[32]"},{"why":"Hallo3 talking-avatar model, evaluated with and without body-pose guidance as one of the two avatar baselines.","marker":"[5]"},{"why":"EchoMimicV2 talking-avatar model, evaluated with and without body-pose guidance as the second avatar baseline.","marker":"[21]"},{"why":"Wan, one of the open video-generation models used in the initial evaluation.","marker":"[33]"},{"why":"Open-Sora 2.0, one of the open video-generation models used in the initial evaluation.","marker":"[23]"},{"why":"Step-Video-TI2V, one of the open video-generation baselines in the initial evaluation.","marker":"[10]"},{"why":"HunyuanVideo, one of the open video-generation baselines in the initial evaluation.","marker":"[15]"}],"fun_headline_variants":["2M-clip benchmark scores whole-body talking avatars","Whole-body avatar eval splits face, hands, body scores","New benchmark uncovers flaws head-only metrics miss","Open-source 2M-clip benchmark for whole-body avatars","Benchmark rates avatars by face, hands, and body"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The entire contribution stands on the dataset being exactly what is advertised: two million video clips, ten thousand plus identities, and six aligned annotations per clip, all publicly downloadable from the linked repository.","fun_headline_variants_meta":{"raw":{"variants":["2M-clip benchmark scores whole-body talking avatars","Whole-body avatar eval splits face, hands, body scores","New benchmark uncovers flaws head-only metrics miss","Open-source 2M-clip benchmark for whole-body avatars","Benchmark rates avatars by face, hands, and body"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1145,"prompt_tokens":626,"completion_tokens":519,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":370,"completion_tokens_details":{"reasoning_tokens":451}},"tokens_in":370,"tokens_out":519,"duration_ms":6403,"temperature":1.0,"reasoning_tokens":451,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:19:38.486124+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Open the linked repository and verify the numbers: count the clips to check the total is 2,000,000, confirm more than 10,000 identities, and inspect that each clip carries body segmentation, landmarks, bounding boxes, motion text, and transcription, with the 20,000-clip test split present. Independently, rerun the twelve metrics on the released test split and see whether the model rankings match the paper's table.","supporting_citations":[{"cited_title":"Versatile Multimodal Controls for Expressive Talking Human Animation","cited_arxiv_id":"2503.08714","evidence_quote":"Defines FID, E-FID, and CSIM for identity-sensitive portrait evaluation; the benchmark's co-speech metrics draw on it."}],"review_version":1}