{"id":"9406265d-4465-48ce-b17d-892f2bbaadfd","arxiv_id":"2608.10839","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In four large user studies, human motion capture beat all five gesture-generation systems on realism and alignment, and nearly all systems scored at chance on semantic and dyadic appropriateness.","lead":"This challenge paper evaluates five AI gesture-generation systems on a dataset of dyadic conversations, using over 23,000 human votes across four studies. It finds recorded human motion is far more realistic and better aligned to speech than any system, and that AI systems fail at semantically meaningful and interlocutor-responsive gestures.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The semantic mismatching study may measure visible articulation and emphasis timing rather than semantic expressivity, since T4 is not reported with a face-masking or articulation-removal control; this is the most load-bearing untested premise behind the 79% vs 8% semantic result.","rationale":"I focused on T4 because it is the novel methodology that carries the paper's strongest new claim. The realism and speech-alignment results (E1/E2) rest on established GENEA protocols and are internally consistent; they would survive even if the semantic study were invalid. The dyadic study (E3) inherits the 2023 protocol, and its main risk, system-level sampling noise, is explicitly discussed and partially bounded by the reported confidence intervals. T4, however, introduces a new text-mismatching task with no validation against a control baseline. The 79% identification rate is the only evidence that the grounded-gesture subset contains semantically expressive motion, and the 8% ceiling for systems is the only evidence that systems lack it. Both numbers are equally compatible with a low-level cue confound: the matched video contains visible articulation and an emphasized target word, while the mismatched sentence is randomly drawn and thus unlikely to share timing or stress; generated systems typically lack such articulatory detail. The reader's verdict already flags T4's interpretation as the weakest assumption; I agree in direction but not in the specific mechanism. The more dangerous confound is not iconic gesture vocabulary but non-semantic physical timing and articulation cues. A face-masked or articulation-free render of the same E4 stimuli is a single, cheap experiment that settles this. My recommendation is therefore to keep the paper conditional: accept if the authors run this control, or if they can show the existing renderings already mask articulation. If the control fails, the semantic part of the central claim should be withdrawn or substantially softened, while the T1/T2/T3 conclusions remain.","tokens_in":11016,"tokens_out":10477,"duration_ms":127797,"concrete_test":"Render the T4 stimuli with the face/mouth region covered (or with a headless avatar that has no articulator motion) and repeat E4 using the same voting interface and the same 257 segments. If the mocap matched-identification rate falls from 79% toward chance, the original score was inflated by visible articulation or emphasis timing rather than semantic gesture; if it stays near 79%, the semantic-mismatching result is robust to this confound.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"T4's central result is that test-takers can identify the matching transcript for mocap (79%) but not for systems (~8%). This interpretation requires that the only useful signal in the matched-mismatched comparison is the semantic content of the gestures. The paper does not report any control removing two non-semantic signals that are present in real motion capture and largely absent from generated body-only motion: visible mouth/jaw articulation in the muted video, and timing/emphasis of the intentionally marked target word. The mismatched sentence is drawn at random, so it will usually differ in length, stress, and target-word position, allowing the matched pair to be identified by low-level temporal matching even if the gestures carry no meaning. These cues are perfectly correlated with the matched condition for the mocap reference and are not generated by most submissions, so the observed 79% vs 8% gap is exactly what a non-semantic confound would produce. Because T4 is the paper's new methodology and the basis for a headline conclusion, this is the most load-bearing untested premise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the GENEA Challenge 2026, a large-scale crowdsourced evaluation of five speech-driven gesture-generation systems trained on the Seamless Interaction dataset. The authors describe four disentangled user studies: motion realism (E1), speech-motion alignment via audio mismatching (E2), dyadic alignment via interlocutor mismatching (E3), and a newly proposed semantic alignment task based on text mismatching (E4). Across more than 23,000 votes, the filtered motion-capture reference outperformed all submissions on every axis, with submissions near chance on speech, dyadic, and semantic alignment. The paper also introduces an updated appropriateness scaling centered at zero and releases the collected votes and outputs.","tokens_in":11203,"tokens_out":9155,"duration_ms":98511,"significance":"If the results hold, the paper makes a useful contribution as the first large-scale evaluation of gesture generation on the Seamless Interaction dataset and as a proposal for a semantic text-mismatching evaluation. Its strengths are the large vote counts, bootstrapped confidence intervals, pairwise disentangled design, JUICE justifications, and the public release of votes and outputs. However, the new semantic task is the most novel and headline-producing component, and its validity currently rests on an untested assumption that no non-semantic cues are available to test-takers; the semantic score also contains an apparent arithmetic inconsistency. The paper is therefore promising but needs additional control analyses before the central claims can be accepted.","major_comments":[{"comment":"The paper's headline semantic result—79% matched-text identification for motion capture versus no more than 8% for the submissions—is interpreted as semantic expressiveness, but the text-mismatching task as described does not isolate semantic content. In a single muted monadic video, two non-semantic cues are available if the renderings show the speaker's face, which the paper does not state: visible mouth/jaw articulation, and the timing/emphasis of the marked target word. Since the mismatched sentence is drawn at random, it will generally differ in length, prosody, and target-word position, so a test-taker can identify the matched sentence by low-level temporal or articulatory matching even if the gestures carry no meaning. These cues are perfectly correlated with the matched condition for the motion-capture reference and are not produced by body-only gesture generators, which is precisely the pattern reported. Please add a control condition, such as lower-face masking or distractor sentences matched on length and prosody, or otherwise demonstrate that the 79% versus 8% gap is not an artifact of these confounds.","section":"§3.1.1 (T4), Figs. 13–14"},{"comment":"Applying the stated formula Appr_semantic = (P − Pbar)/(P + Pbar + N + B) to the response shares reported for the motion-capture condition in Fig. 13 gives (0.79 − 0.07)/(0.79 + 0.07 + 0.09 + 0.05) = 0.72, whereas Fig. 14 reports 0.77. The displayed distribution and the displayed score are therefore inconsistent under the stated definition. Please reconcile the raw vote counts, the formula, and the reported score; if the discrepancy is due to rounding, state the exact counts.","section":"§4.4 (semantic appropriateness score), Figs. 13–14"},{"comment":"For challenge submissions, the matched and mismatched videos in the dyadic study contain different generated agent motions because the mismatched motion is newly generated from segment B's interlocutor inputs. The pairwise comparison can therefore be won on overall motion naturalness or generation quality rather than on responsiveness to the interlocutor, which would bias the appropriateness score even for a system with good dyadic behaviour. The paper does not report any check that separates generation-quality differences from dyadic responsiveness. Please add such an analysis, for example by correlating matched-versus-mismatched quality ratings or by including a system-generated condition with an unresponsive interlocutor as a sanity check, or soften the conclusion in Sec. 5 that 'systems were not yet able to generate motion that is responsive to the interlocutor.'","section":"§3.2.4 and §4.3"}],"minor_comments":[{"comment":"In Section 3.2.1, 'V oice Activity Detection' contains a stray space, and 'V AD' is similarly formatted in Section 3.2.3; please fix the spacing.","section":"§3.2.1"},{"comment":"The symbols \\bar{S} and \\bar{C} in the appropriateness score formulas are never explicitly defined; please define them the first time they appear.","section":"§4.2.1"},{"comment":"References [8] and [9] appear to be the same GENEA Challenge 2023 paper; please cite it once.","section":"References"},{"comment":"Since the T3 stimulus selection thresholds (30%–50% VAD overlap) are described only as 'determined empirically', a brief sensitivity analysis or a report of the number of candidate windows discarded at each threshold would strengthen reproducibility.","section":"§3.2.3"},{"comment":"The abstract and Table 1 report 'over 23,000 votes' and '869 test-takers'; the table sums to 23,210, so the statement is accurate, but the body text has missing spaces in 'over23,000' and '869test-takers'.","section":"Abstract and Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well within scope for a venue that publishes challenge evaluation papers. My main concern is the missing T4 control, which is load-bearing for the paper's most novel claim, and the apparent arithmetic inconsistency in the semantic score. I do not see grounds for rejection, but the paper needs additional control analyses and a reconciliation of the reported numbers before the central claims can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, the key thing to know: GENEA 2026 is the first large-scale attempt to evaluate semantic expressivity in gesture generation via text mismatching, and on that new axis the gap between motion capture and submissions is enormous (77% vs 8%). The evaluation is careful and mostly convincing for realism, speech alignment, and dyadic alignment. But the semantic result rests on an untested assumption, and I would not take the 79% vs 8% as strong evidence about semantics until the confound is addressed.\n\nWhat is genuinely new: a fourth GENEA-style evaluation task using the Grounded Gestures subset, text mismatching rather than audio mismatching, applied to the Seamless Interaction dataset. The scale (23k votes, 869 test-takers) and the bootstrapped confidence intervals give the headline numbers real weight. The updated appropriateness scaling is an explicit linear transform of the 2023 formula, so there is no circularity. The authors also openly discuss system-level and study-level randomness, which is more than most challenge reports do.\n\nThe soft spot is T4. Test-takers see a muted video of the speaker and choose which of two sentences the gestures express. The matched sentence is the one actually spoken, with a marked target word. Visible mouth/jaw articulation and the timing/emphasis of that target word will both correlate with the matched condition for the mocap videos, and both are essentially absent from generated body-only motion. The mismatched sentence is drawn at random, so it will usually differ in length, stress, and target-word position. That means a test-taker could identify the matched transcript from low-level temporal or articulatory cues without reading any semantic content in the gestures. The paper reports no control that removes the face or masks articulation. Given that this is the paper's new methodology and the basis for a headline conclusion, the 79% vs 8% gap is exactly what such a confound would produce. The dyadic study has a similar but smaller issue: the 30-50% VAD overlap thresholds are empirically chosen, and the segments are manually curated, so the ceiling comparison is shaped by the selection.\n\nThe T1–T3 results stand mostly on their own, though the promises that votes/outputs and system descriptions will be released later mean the reproducibility is currently prospective. If the promised data ship, this will be a useful benchmark paper for gesture-generation evaluators. I would send it to review, and I would cite it for the T1–T3 findings; I would be cautious about citing the semantic result without the missing control.\n\nRecommendation: accept the challenge report with revision, and require the authors to either report a face-masking/articulation-removal control for T4 or soften the semantic claim.","headline":"The new semantic mismatching evaluation is promising but likely confounded by visible articulation and timing cues; the T1–T3 results are solid.","tokens_in":11768,"tokens_out":2478,"would_cite":true,"duration_ms":26554,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human motion capture beats all five gesture-generation systems on every evaluated dimension.","keywords":["gesture generation","speech-driven animation","dyadic interaction","semantic gestures","mismatching evaluation","motion capture benchmark","user study evaluation","Seamless Interaction dataset"],"falsifier":"A control study in which the 'mismatched' audio is matched to the original on prosody, voice, and recording conditions (not just voice-activity structure) should reproduce the 62 percent ceiling for motion capture; if test-takers cannot tell matched from such a near-twin audio, the T2 score is inflated by non-semantic acoustics. Similarly for T4, replacing the mismatched sentence with one of matched length and similar keywords should drop the 79 percent mocap identification if the test measures meaning rather than superficial cue matching.","tokens_in":10821,"feed_emoji":"🕺","tokens_out":4446,"duration_ms":47383,"temperature":0.7,"pith_summary":"This paper reports the results of the fourth GENEA Challenge, a large-scale crowdsourced evaluation of five speech-driven gesture-generation systems trained on the Seamless Interaction dataset of dyadic conversations. Its central claim is that recorded human motion is substantially better than every submitted system on all four tested dimensions: motion realism, alignment with the input speech, responsiveness to the interlocutor, and semantic expressiveness of gestures. The mismatch-based evaluation methodology is designed so that an input-independent system should score zero, giving the results an interpretable floor; under that metric, all five systems score near zero on dyadic responsiveness and semantic expressiveness while motion capture reaches 65 and 77 percent respectively. The paper additionally proposes a new semantic mismatching methodology using the Grounded Gestures subset of the dataset, and argues that the dataset and procedure are suitable for benchmarking these harder capabilities in the future.","feed_headline":"Mocap beats every gesture AI on all four tests","feed_subtitle":"Five challenged systems trail human motion, scoring near chance on conversation and meaning.","key_machinery":"The central mechanism is a set of mismatching evaluations in which the attribute under test is isolated by holding everything else fixed. T2 pairs the same motion with matched versus mismatched audio chosen to have similar voice-activity structure; T3 pairs matched and mismatched interlocutor contexts while keeping both speakers' audio and the interlocutor's motion identical; T4, the paper's new contribution, shows a single muted video of a monadic gesture and asks test-takers to pick which of two sentences the gestures express. These designs convert the evaluation into a discrimination task with an interpretable chance level of zero for an input-independent system, and pairwise Likert or forced-choice voting is aggregated into Elo ratings and appropriateness scores in the range [-1, 1].","core_discovery":"The paper's discovery is that, on the Seamless Interaction dataset, the gap between human motion and state-of-the-art learned gesture generation remains wide on every axis, and is largest on exactly the capabilities that motivate dyadic data. In pairwise realism comparisons the motion-capture reference wins 68 to 95 percent of matches; in speech-mismatching tests it posts a 62 percent appropriateness score against 32 percent for the top submission, with the rest near the zero floor; in dyadic mismatching it scores 65 percent with all submissions at chance; and in the new semantic mismatching study test-takers identify the correct sentence from mocap video 79 percent of the time, while the best system scores 8 percent. The paper treats these numbers as evidence that current systems cannot yet produce interlocutor-responsive or semantically expressive motion, and that the newly introduced text-mismatching procedure gives a valid, interpretable measure of that shortfall.","pith_inferences":["The reported near-chance scores for systems also depend on how representative the five submissions are of the field; the stronger system that previously reached mocap-level alignment on the leaderboard was not among the submissions, so the challenge may underestimate the current state of the art.","The semantic mismatching method's ceiling of 79 percent on mocap, not 100, suggests the grounded-gesture clips still contain a sizable fraction of gestures that test-takers cannot map to a specific sentence; the method measures recognizable expressivity, not all communicative content.","A stronger control would withhold all transcript information and instead compare matched versus mismatched texts with identical keywords or identical prosody, which would tell whether the 79 percent reflects lexical iconicity or speaker-specific delivery."],"forward_implications":["If the results hold, Seamless Interaction with the mismatching methodology becomes a benchmark on which progress in dyadic and semantic gesture generation can be measured against interpretable zero-floor and mocap-ceiling anchors.","Any future system must clear the demonstrated gap between 8 percent and 77 percent on semantic appropriateness before it can credibly claim meaning-aware gesture generation.","The near-chance dyadic scores indicate that models trained on dyadic data are not yet using the interlocutor's audio or motion at generation time in a way test-takers perceive as responsive.","The speech-alignment results put the best submission at roughly half the mocap ceiling, implying that state-of-the-art co-speech rhythm still lags human timing."],"supporting_citations":[{"why":"Provides the Seamless Interaction dataset with the dyadic-conversations and Grounded Gestures subsets that all challenge systems train on and that the evaluation segments are drawn from.","marker":"[1]"},{"why":"Supplies the audio-mismatching methodology, the Elo and appropriateness scoring conventions, and the GENEA Leaderboard baseline results this challenge compares against.","marker":"[11]"},{"why":"Contributes the dyadic mismatching study design and the dyadic-appropriateness formulation reused for T3.","marker":"[8]"},{"why":"Introduces the JUICE follow-up voting procedure used to collect attribute-level justifications in E1-E3.","marker":"[4]"}],"fun_headline_variants":["Human motion trounces AI gesture systems on every test","Gesture AI falls far behind mocap in all four evaluations","Mocap wins every round against gesture AI in four studies","Gesture AI near chance on conversation and meaning in new study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The mismatched stimuli differ from the matched ones only in the attribute being measured, so that a preference for the matched video reflects dyadic responsiveness, speech alignment, or semantic expressiveness rather than some side difference such as voice, prosody, rendering, or motion quality.","fun_headline_variants_meta":{"raw":{"variants":["Human motion trounces AI gesture systems on every test","Gesture AI falls far behind mocap in all four evaluations","Mocap wins every round against gesture AI in four studies","Gesture AI near chance on conversation and meaning in new study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000697,"raw_usage":{"total_tokens":3209,"prompt_tokens":1061,"completion_tokens":2148,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":677,"completion_tokens_details":{"reasoning_tokens":2080}},"tokens_in":677,"tokens_out":2148,"duration_ms":16110,"temperature":1.0,"reasoning_tokens":2080,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:18:05.456708+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A control study in which the 'mismatched' audio is matched to the original on prosody, voice, and recording conditions (not just voice-activity structure) should reproduce the 62 percent ceiling for motion capture; if test-takers cannot tell matched from such a near-twin audio, the T2 score is inflated by non-semantic acoustics. Similarly for T4, replacing the mismatched sentence with one of matched length and similar keywords should drop the 79 percent mocap identification if the test measures meaning rather than superficial cue matching.","supporting_citations":[{"cited_title":"Towards reli- able human evaluations in gesture generation: Insights from a community-driven state-of-the-art benchmark","cited_arxiv_id":null,"evidence_quote":"Supplies the audio-mismatching methodology, the Elo and appropriateness scoring conventions, and the GENEA Leaderboard baseline results this challenge compares against."},{"cited_title":"The genea challenge 2023: A large-scale evaluation of gesture generation models in monadic and dyadic settings","cited_arxiv_id":null,"evidence_quote":"Contributes the dyadic mismatching study design and the dyadic-appropriateness formulation reused for T3."},{"cited_title":"Factorizing text-to-video generation by explicit image conditioning","cited_arxiv_id":null,"evidence_quote":"Introduces the JUICE follow-up voting procedure used to collect attribute-level justifications in E1-E3."}],"review_version":1}