{"id":"ee67d446-0d18-40f5-b8fd-8c53b6266d39","arxiv_id":"2412.16151","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In virtual crowds of twelve physics-based avatars, varying body shape did not help hide motion clones, but increasing motion variety reduced clone detection.","lead":"This study asked 22 people to spot cloned walking motions in twelve-avatar virtual crowds with different body shapes. Body shape variety did not change detection accuracy, while motion variety did, suggesting motion matters more than body shape for crowd realism.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on an arbitrary ordinal-to-interval conversion of the 5-point confidence scale; a robust ordinal reanalysis could flip the borderline MOTION effect (p=0.046), undermining the main conclusion.","rationale":"The reader's weakest assumption identifies the Likert-to-accuracy conversion, and I agree this is the single most load-bearing concern. The paper's key positive result (MOTION p=0.046) is borderline, and the entire statistical edifice is built on assigning equally spaced numbers to ordinal confidence labels. A reanalysis using ordinal models or even a simple sensitivity analysis could change the conclusion. The null BODY result is also expressed through this conversion, though the large p-value makes it less likely to flip. I considered whether the stimulus confound—the baseline always has 12 unique body shapes while the test side has 1/3/6/12—is more fundamental, since it means body diversity differences could influence responses independently of motion. However, that confound is a design limitation that affects interpretation rather than the statistical validity of the reported tests, and the paper already notes the masking trend for males. The measurement assumption is more directly load-bearing because it underpins every reported F and p value. The appropriate remedy is to require the authors to either provide raw data for reanalysis or justify the coding with a sensitivity analysis. This aligns with the reader's CONDITIONAL verdict, so no change is needed.","tokens_in":10096,"tokens_out":7061,"duration_ms":69481,"concrete_test":"Obtain the raw trial-level 5-point responses and re-run the analysis with an ordinal cumulative-link mixed model (participant as random intercept; fixed factors SEX, MOTION, BODY and interactions). Compare the MOTION main effect p-value and the BODY p-value to the reported ANOVA results. Also run a sensitivity analysis with alternative monotonic codings (e.g., treat undecided as missing, or use 1,2,4,5 with undecided at 3, or perform an aligned rank transform ANOVA). If the MOTION effect is no longer significant at α=0.05 in any reasonable ordinal model, the central claim's positive half is not robust; if BODY becomes significant, the null claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 ('Results') states that responses were converted to numeric accuracy by assigning 5/4 for correct definitely/probably, 3 for undecided, and 2/1 for incorrect probably/definitely. This treats a 5-point Likert-type confidence scale as an interval measurement where 'undecided' equals chance (3) and where the distance between 'probably correct' and 'undecided' is exactly one unit. No justification or sensitivity analysis is provided. This conversion is load-bearing because the central claim has two parts: (i) body shape has no effect (BODY F(3,60)=0.47, p=0.701) and (ii) motion variety has a significant effect (MOTION F(2.5,49.2)=3.05, p=0.046). The MOTION p-value is just below the 0.05 threshold, so if a more appropriate ordinal model (e.g., cumulative-link mixed model) or a different but defensible coding yields a slightly larger p-value, the paper's headline finding that 'motion variety had a greater impact on their perception' loses statistical support. The null BODY result is also affected by the coding: a different interval assignment could change the F statistic and effect size, although the very high p=0.701 suggests that the null may be robust. The paper does not report raw responses or a robustness check, so the reader cannot verify whether the reported ANOVA results are an artifact of the chosen numeric mapping. Because the statistical inference rests on an untested measurement assumption, and the significant effect is marginal, this is the most load-bearing weakness in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a perception experiment in which 22 participants viewed side-by-side video clips of small-scale virtual crowds (12 avatars) and were asked to identify which side contained all-unique motions versus a side containing motion clones. The stimuli varied the number of unique motions (1, 2, 3, 6) and the number of distinct body shapes (1, 3, 6, 12), for both male and female avatar crowds. The participants' directional confidence responses were converted to a 1-5 accuracy score, and a mixed ANOVA (SEX between-subjects; MOTION and BODY within-subjects) was run. The authors report a significant effect of MOTION (F(2.5,49.2)=3.05, p=0.046, partial η²=0.132) and no significant effect of BODY (F(3,60)=0.47, p=0.701). They conclude that body shape diversity does not influence motion clone detection and that motion variety has a greater impact on perceived crowd realism.","tokens_in":10375,"tokens_out":8599,"duration_ms":74699,"significance":"If the findings are robust, this work provides a useful extension of prior crowd-variety research by isolating body shape as an appearance cue in physics-based small-scale crowds. The experimental design has credible features: out-of-step clones to avoid trivial detection, randomized body-motion pairings, side-balanced presentation, and a mixed ANOVA with sphericity correction. The paper also ships a concrete, falsifiable claim (motion variety matters more than body shape for clone detection) that is directly relevant to crowd animation practitioners. However, the main positive finding rests on a marginal p-value (0.046) and on an untested ordinal-to-interval conversion; the null claim about body shape is stated more strongly than the design can support. These issues need to be addressed before the conclusions can be taken as established.","major_comments":[{"comment":"The conversion of the 5-point directional confidence response into a numeric accuracy score (5/4 for correct definitely/probably, 3 for undecided, 2/1 for incorrect probably/definitely) is an unexamined interval-scale assumption. In particular, assigning exactly 3 to 'Undecided' equates that response with chance performance, while the equal spacing between categories (e.g., 'probably correct' versus 'definitely correct') is not justified. This coding is load-bearing because the only significant effect (MOTION) has p=0.046, just below the 0.05 threshold, and a different, equally defensible coding (or an ordinal model such as a cumulative-link mixed model) could push this p-value above 0.05 and thereby undermine the headline conclusion that motion variety has a significant effect. Please provide a robustness analysis: for example, an ordinal mixed model on the raw ordered responses, a binary correct/incorrect logistic mixed model, or a sensitivity analysis over alternate numeric codings, and report whether the MOTION effect survives.","section":"§4, first paragraph of Results"},{"comment":"The statement that 'body shape diversity did not influence participants' ratings of motion clone detection' overstates the evidence. The statistical analysis only supports 'no statistically significant effect was found,' and with only 22 participants (11 per SEX group) the study has limited power to detect a small body-shape effect. The descriptive results even show non-overlapping standard errors for the male 1-body versus 12-body conditions at one motion level, which the authors themselves acknowledge in the Discussion as a possible masking effect. The conclusion should be tempered to 'no significant effect was found in this experiment,' ideally accompanied by a power analysis or an equivalence test, rather than the strong causal phrasing used in the abstract and title.","section":"Abstract and §5 (Discussion)"}],"minor_comments":[{"comment":"The sentence 'as the variety of motions increased, participants were more likely to detect differences between the virtual avatars' is inconsistent with the immediately following post-hoc finding that accuracy was highest with one motion and lowest with six motions. Increasing motion variety made the clone side look more varied, which reduced the ability to identify the baseline; please correct the direction of this sentence to avoid confusing readers.","section":"§4, Results, third paragraph"},{"comment":"Please report exact p values for all effects (e.g., p=0.046 rather than p<0.05) and specify whether the reported η² is partial eta-squared or generalized eta-squared, since this affects comparability with prior work.","section":"Table 1"},{"comment":"There are two spelling errors: 'Hyunh-Feldt' should be 'Huynh-Feldt', and 'Bonferonni' should be 'Bonferroni'.","section":"Throughout"},{"comment":"The claim that the results 'align with previous findings by Hoyet et al. [9], who found that motion variety had a greater impact on crowd perception than other visual elements' appears mismatched with the cited paper's title and content, which concerns the perceptual effect of shoulder motions. Please verify this citation or rephrase to refer to the appropriate prior work.","section":"§5, Discussion, second paragraph"},{"comment":"The figure uses between-subjects standard error bars for a within-subjects design. Consider plotting within-subject error bars (e.g., Cousineau-Morey intervals) to better represent the repeated-measures nature of the data.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The reader's stress-test concern about the ordinal-to-interval conversion is well-founded and is the primary barrier to acceptance. The paper's central claim about motion variety rests on a marginal p-value, and the paper does not currently provide enough information to verify that the ANOVA result is not an artifact of the coding. The requested robustness analysis is feasible within the manuscript's scope (the raw data or per-cell response distributions should be available), so this should be treated as a major revision rather than a rejection. I would also encourage the editor to ask the authors to soften the causal phrasing of the null body-shape result, as the design is underpowered for strong claims of 'no influence'."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what is genuinely useful here: the study tests an untested combination—body shape diversity versus motion clone detection—using physics-based avatars, which is a sensible and well-controlled way to vary shape while holding motion constant. The design is careful: side counterbalancing, out-of-step clones, shuffled pairings, and a practice phase. The null result for body shape is probably real: p=0.701 with a small effect size is not likely to flip under a different coding. The authors are also appropriately cautious, noting that body shape might still matter in interaction with other appearance cues.\n\nNow the soft spots. The big one is the conversion of the five-point confidence scale to a numeric accuracy score without justification. They assign 5/4 for correct, 3 for undecided, 2/1 for incorrect, and treat that as interval. This is a defensible coding, but it is not the only one, and the only significant effect in the paper—MOTION, p=0.046—sits right at the threshold. A cumulative-link mixed model or even a different but equally plausible coding could move that p-value over 0.05. Since the paper's secondary claim is that motion variety 'had a greater impact,' this matters. The null for BODY is robust, but the positive claim is not. Secondly, with 22 participants split between two sex groups, the per-group n is about 11; the lack of a significant SEX effect and the inconsistency in the female data (widely varying error bars) should be discussed more. Third, no data or code are provided, which makes the robustness question hard for a reader to resolve.\n\nOverall, I think the central half of the paper—the body-shape null—holds up. The motion-variety half is less secure and needs a reanalysis. This is a solid, modest empirical contribution for crowd perception researchers, and the topic deserves a serious referee, but I would not want to see it accepted without the measurement assumption addressed.","headline":"The body-shape null likely holds, but the motion-variety effect rests on an unjustified Likert-to-accuracy conversion and should be reanalyzed before being cited as evidence.","tokens_in":10911,"tokens_out":3308,"would_cite":false,"duration_ms":30229,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Body shape does not help hide cloned motions in crowds","keywords":["motion clone detection","body shape diversity","motion variety","virtual crowd perception","physics-based avatars","small-scale crowds","perceptual realism"],"falsifier":"Re-analyze the raw ordinal responses with a model that does not assume equal spacing, for example ordinal logistic regression with BODY and MOTION as predictors. If BODY becomes significant or the MOTION effect disappears, the paper's central null result is an artifact of the numeric conversion rather than a perceptual fact.","tokens_in":9878,"feed_emoji":"👥","tokens_out":5189,"duration_ms":44897,"temperature":0.7,"pith_summary":"The paper asks whether changing the body shapes of virtual avatars makes it harder for viewers to notice that several avatars in a small crowd are driven by the same recorded walking motion. In a side-by-side video experiment with twelve-avatar crowds, the authors find that body shape diversity has no significant overall effect on clone detection, while the number of distinct motions does. The conclusion is that, in small-scale crowds, motion variety is the stronger lever for perceived realism, and varying appearance alone will not hide repeated motions.","feed_headline":"Body shape does not help hide cloned motions in crowds","feed_subtitle":"Perception test finds motion variety, not body shape, drives perceived realism in 12-avatar crowds.","key_machinery":"The load-bearing apparatus is a clone-detection task: participants watch a baseline crowd in which every avatar has a unique motion and body shape side by side with a crowd containing cloned motions, and state on a five-point scale which side had all different motions. The stimuli use physics-based avatars, whose joint angles are driven by forces and torques from a proportional-derivative (PD) controller tracking captured walking motions, so identical motions can be applied consistently across different body shapes. Body shapes come from real actors' anthropometric measurements in three BMI categories, and the factors BODY (1, 3, 6, or 12 shapes) and MOTION (1, 2, 3, or 6 motions) are crossed inside twelve-avatar crowds, with repeated motions desynchronized so clones are not trivially obvious.","core_discovery":"The paper's central claim is that body shape diversity does not change how easily people detect motion clones in small-scale virtual crowds. The evidence is a mixed ANOVA on accuracy scores converted from five-point confidence responses, which yields a significant main effect of motion variety ($F(2.5,49.2)=3.05$, $p<0.05$, $\\eta^2=0.132$) and no significant effect of body shape ($F(3,60)=0.47$, $p=0.701$). The authors read this as supporting their hypothesis that increasing motion variety lowers clone detection, and as failing to support the hypothesis that more body shapes would mask cloned motions. They note that the male-avatar data show above-chance detection only when all twelve avatars share one motion, while the female-avatar responses were more variable.","pith_inferences":["The null result for body shape likely depends on the particular motion set: participants' comments point to arm and hand movement as the dominant cue, so a motion set with less distinctive arm swings might allow body-shape variety to become visible.","The study varies the number of body shapes but not their distinctiveness; more extreme or unusual proportions could still mask clones even though a broader set of average shapes does not.","Modeling the five-point confidence responses ordinally, rather than as a 5-4-3-2-1 numeric scale, might reveal a body-shape effect in decision confidence even where mean accuracy is flat.","The small-scale result may not transfer to larger crowds, where the sheer number of characters can make any single appearance cue less salient; testing at larger scales would show whether the conclusion is scale-dependent."],"forward_implications":["With two or more distinct motions distributed through a twelve-avatar crowd, clone detection often falls to chance, meaning a small number of motions can stand in for a fully varied crowd.","Rendering budgets for small-scale virtual crowds should prioritize motion variety over body-shape variety if the goal is perceptual realism.","Body shape diversification is not sufficient on its own to mask motion clones in animated small crowds, despite being a visible appearance cue.","The absence of a significant body-shape effect is qualified by high response variability for the female-avatar group, so sex-based differences in stimulus salience deserve separate testing."],"supporting_citations":[{"why":"Supplies the prior result that three evenly cloned motions are perceived as a fully varied crowd, the paradigm this paper extends to body shape.","marker":"[27]"},{"why":"Establishes that cloned motions are easier to detect than appearance clones, motivating the current clone-detection task.","marker":"[22]"},{"why":"Shows that even low motion variety can match unique crowds in large-scale settings, supporting the motion-variety result.","marker":"[1]"},{"why":"Earlier demonstration that motion cues dominate crowd perception, used to interpret why body shape did not matter.","marker":"[9]"},{"why":"Provides the captured walking motions and anthropometric measurements from which the twelve-avatar stimuli were built.","marker":"[10]"},{"why":"Reports that observers detect body-shape and motion inconsistencies in walking but not other motions, motivating the body-shape interaction hypothesis.","marker":"[29]"},{"why":"Supplies the feedback-control gain settings and body-shape and motion matching findings used to construct the physics-based avatars.","marker":"[36]"},{"why":"Shows that distinctive body shapes are noticeable in crowds, the appearance-variety manipulation this study adapts to animation.","marker":"[30]"}],"fun_headline_variants":["Motion variety shapes crowd realism more than body shape","Body shape won't hide cloned motions in virtual crowds","Clone detection: motion beats body shape in crowd perception","Crowd clones: Body shape fails to mask repeated motions","Small crowd study: body shape doesn't affect clone spotting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis treats the five-point confidence response as an interval accuracy score (5, 4, 3, 2, 1) and runs an ANOVA on those numbers, assuming 'undecided' equals chance and that the distances between response levels are equal.","fun_headline_variants_meta":{"raw":{"variants":["Motion variety shapes crowd realism more than body shape","Body shape won't hide cloned motions in virtual crowds","Clone detection: motion beats body shape in crowd perception","Crowd clones: Body shape fails to mask repeated motions","Small crowd study: body shape doesn't affect clone spotting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000344,"raw_usage":{"total_tokens":1855,"prompt_tokens":874,"completion_tokens":981,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":917}},"tokens_in":490,"tokens_out":981,"duration_ms":8329,"temperature":1.0,"reasoning_tokens":917,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:44:28.398888+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the raw ordinal responses with a model that does not assume equal spacing, for example ordinal logistic regression with BODY and MOTION as predictors. If BODY becomes significant or the MOTION effect disappears, the paper's central null result is an artifact of the numeric conversion rather than a perceptual fact.","supporting_citations":[{"cited_title":"Pražák and C","cited_arxiv_id":null,"evidence_quote":"Supplies the prior result that three evenly cloned motions are perceived as a fully varied crowd, the paradigm this paper extends to body shape."},{"cited_title":"McDonnell, M","cited_arxiv_id":null,"evidence_quote":"Establishes that cloned motions are easier to detect than appearance clones, motivating the current clone-detection task."},{"cited_title":"Adili, B","cited_arxiv_id":null,"evidence_quote":"Shows that even low motion variety can match unique crowds in large-scale settings, supporting the motion-variety result."},{"cited_title":"Hoyet, A.-H","cited_arxiv_id":null,"evidence_quote":"Earlier demonstration that motion cues dominate crowd perception, used to interpret why body shape did not matter."},{"cited_title":"Hoyet, K","cited_arxiv_id":null,"evidence_quote":"Provides the captured walking motions and anthropometric measurements from which the twelve-avatar stimuli were built."},{"cited_title":"Russell, S","cited_arxiv_id":null,"evidence_quote":"Reports that observers detect body-shape and motion inconsistencies in walking but not other motions, motivating the body-shape interaction hypothesis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the feedback-control gain settings and body-shape and motion matching findings used to construct the physics-based avatars."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that distinctive body shapes are noticeable in crowds, the appearance-variety manipulation this study adapts to animation."}],"review_version":1}