REVIEW 2 major objections 5 minor 40 references
Shape Shifters: Does Body Shape Change the Perception of Small-Scale Crowd Motions?
T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Body shape does not help hide cloned motions in crowds
desk verdict The body-shape null likely holds, but the motion-variety effect rests on an unjustified Likert-to-accuracy conversion and should be reanalyzed before being cited as evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing apparatus is a clone-detection task: participants watch a baseline crowd in which every avatar has a unique motion and body shape side by side with a crowd containing cloned motions, and state on a five-point scale which side had all different motions. The stimuli use physics-based avatars, whose joint angles are driven by forces and torques from a proportional-derivative (PD) controller tracking captured walking motions, so identical motions can be applied consistently across different body shapes. Body shapes come from real actors' anthropometric measurements in three BMI categories, and the factors BODY (1, 3, 6, or 12 shapes) and MOTION (1, 2, 3, or 6 motions) are crossed inside twelve-avatar crowds, with repeated motions desynchronized so clones are not trivially obvious.
What would settle it
Re-analyze the raw ordinal responses with a model that does not assume equal spacing, for example ordinal logistic regression with BODY and MOTION as predictors. If BODY becomes significant or the MOTION effect disappears, the paper's central null result is an artifact of the numeric conversion rather than a perceptual fact.
Extended reading notes
Core claim
The paper's central claim is that body shape diversity does not change how easily people detect motion clones in small-scale virtual crowds. The evidence is a mixed ANOVA on accuracy scores converted from five-point confidence responses, which yields a significant main effect of motion variety ($F(2.5,49.2)=3.05$, $p<0.05$, $\eta^2=0.132$) and no significant effect of body shape ($F(3,60)=0.47$, $p=0.701$). The authors read this as supporting their hypothesis that increasing motion variety lowers clone detection, and as failing to support the hypothesis that more body shapes would mask cloned motions. They note that the male-avatar data show above-chance detection only when all twelve avatars share one motion, while the female-avatar responses were more variable.
Load-bearing premise
The analysis treats the five-point confidence response as an interval accuracy score (5, 4, 3, 2, 1) and runs an ANOVA on those numbers, assuming 'undecided' equals chance and that the distances between response levels are equal.
Editorial extensions
If this is right
- With two or more distinct motions distributed through a twelve-avatar crowd, clone detection often falls to chance, meaning a small number of motions can stand in for a fully varied crowd.
- Rendering budgets for small-scale virtual crowds should prioritize motion variety over body-shape variety if the goal is perceptual realism.
- Body shape diversification is not sufficient on its own to mask motion clones in animated small crowds, despite being a visible appearance cue.
- The absence of a significant body-shape effect is qualified by high response variability for the female-avatar group, so sex-based differences in stimulus salience deserve separate testing.
Reading between the lines
- The null result for body shape likely depends on the particular motion set: participants' comments point to arm and hand movement as the dominant cue, so a motion set with less distinctive arm swings might allow body-shape variety to become visible.
- The study varies the number of body shapes but not their distinctiveness; more extreme or unusual proportions could still mask clones even though a broader set of average shapes does not.
- Modeling the five-point confidence responses ordinally, rather than as a 5-4-3-2-1 numeric scale, might reveal a body-shape effect in decision confidence even where mean accuracy is flat.
- The small-scale result may not transfer to larger crowds, where the sheer number of characters can make any single appearance cue less salient; testing at larger scales would show whether the conclusion is scale-dependent.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a perception experiment in which 22 participants viewed side-by-side video clips of small-scale virtual crowds (12 avatars) and were asked to identify which side contained all-unique motions versus a side containing motion clones. The stimuli varied the number of unique motions (1, 2, 3, 6) and the number of distinct body shapes (1, 3, 6, 12), for both male and female avatar crowds. The participants' directional confidence responses were converted to a 1-5 accuracy score, and a mixed ANOVA (SEX between-subjects; MOTION and BODY within-subjects) was run. The authors report a significant effect of MOTION (F(2.5,49.2)=3.05, p=0.046, partial η²=0.132) and no significant effect of BODY (F(3,60)=0.47, p=0.701). They conclude that body shape diversity does not influence motion clone detection and that motion variety has a greater impact on perceived crowd realism.
Significance. If the findings are robust, this work provides a useful extension of prior crowd-variety research by isolating body shape as an appearance cue in physics-based small-scale crowds. The experimental design has credible features: out-of-step clones to avoid trivial detection, randomized body-motion pairings, side-balanced presentation, and a mixed ANOVA with sphericity correction. The paper also ships a concrete, falsifiable claim (motion variety matters more than body shape for clone detection) that is directly relevant to crowd animation practitioners. However, the main positive finding rests on a marginal p-value (0.046) and on an untested ordinal-to-interval conversion; the null claim about body shape is stated more strongly than the design can support. These issues need to be addressed before the conclusions can be taken as established.
major comments (2)
- [§4, first paragraph of Results] The conversion of the 5-point directional confidence response into a numeric accuracy score (5/4 for correct definitely/probably, 3 for undecided, 2/1 for incorrect probably/definitely) is an unexamined interval-scale assumption. In particular, assigning exactly 3 to 'Undecided' equates that response with chance performance, while the equal spacing between categories (e.g., 'probably correct' versus 'definitely correct') is not justified. This coding is load-bearing because the only significant effect (MOTION) has p=0.046, just below the 0.05 threshold, and a different, equally defensible coding (or an ordinal model such as a cumulative-link mixed model) could push this p-value above 0.05 and thereby undermine the headline conclusion that motion variety has a significant effect. Please provide a robustness analysis: for example, an ordinal mixed model on the raw ordered responses, a binary correct/incorrect logistic mixed model, or a sensitivity analysis over alternate numeric codings, and report whether the MOTION effect survives.
- [Abstract and §5 (Discussion)] The statement that 'body shape diversity did not influence participants' ratings of motion clone detection' overstates the evidence. The statistical analysis only supports 'no statistically significant effect was found,' and with only 22 participants (11 per SEX group) the study has limited power to detect a small body-shape effect. The descriptive results even show non-overlapping standard errors for the male 1-body versus 12-body conditions at one motion level, which the authors themselves acknowledge in the Discussion as a possible masking effect. The conclusion should be tempered to 'no significant effect was found in this experiment,' ideally accompanied by a power analysis or an equivalence test, rather than the strong causal phrasing used in the abstract and title.
minor comments (5)
- [§4, Results, third paragraph] The sentence 'as the variety of motions increased, participants were more likely to detect differences between the virtual avatars' is inconsistent with the immediately following post-hoc finding that accuracy was highest with one motion and lowest with six motions. Increasing motion variety made the clone side look more varied, which reduced the ability to identify the baseline; please correct the direction of this sentence to avoid confusing readers.
- [Table 1] Please report exact p values for all effects (e.g., p=0.046 rather than p<0.05) and specify whether the reported η² is partial eta-squared or generalized eta-squared, since this affects comparability with prior work.
- [Throughout] There are two spelling errors: 'Hyunh-Feldt' should be 'Huynh-Feldt', and 'Bonferonni' should be 'Bonferroni'.
- [§5, Discussion, second paragraph] The claim that the results 'align with previous findings by Hoyet et al. [9], who found that motion variety had a greater impact on crowd perception than other visual elements' appears mismatched with the cited paper's title and content, which concerns the perceptual effect of shoulder motions. Please verify this citation or rephrase to refer to the appropriate prior work.
- [Figure 5] The figure uses between-subjects standard error bars for a within-subjects design. Consider plotting within-subject error bars (e.g., Cousineau-Morey intervals) to better represent the repeated-measures nature of the data.
Circularity Check
No significant circularity: the paper's central claims are empirical results from a user study, not derivations from fitted inputs or self-citations.
full rationale
This is a perception experiment, not a derivation or modeling paper. The hypotheses H1 and H2 are stated before data collection, and the reported ANOVA tests (Table 1) are computed from independent participant responses to randomized video stimuli. The central claim that body shape diversity did not influence motion clone detection and motion variety had a greater impact is supported by directly measured accuracy scores, not by any equation that reduces to its own inputs. The conversion of the 5-point Likert scale to numeric accuracy (Section 4) is a measurement assumption that could be challenged on ordinal-scale grounds, but it is not circular: the numeric coding is fixed in advance and is not fitted to produce the reported outcomes, nor is any 'prediction' derived from data already containing that outcome. The paper does cite prior work by one of the authors (e.g., Vyas et al. [36] for PD gain values, and Pražák and O'Sullivan [27] for related motion-variety findings), but these citations are used for implementation details and contextual motivation, not as the load-bearing evidence for the paper's own empirical conclusion. The main results stand on the participant data reported in the paper. Therefore, no circular step meeting the required standard—where a claim is equivalent by construction to its inputs, a fitted parameter is renamed as a prediction, or a self-citation chain is used to force a conclusion—was found.
Assumptions & free parameters
assumptions (4)
- domain assumption The PD-controlled ragdoll avatars faithfully reproduce the reference walking motions at a perceptual level sufficient for the experiment.
- domain assumption The selected BMI categories (low, average, high) and the four actors per category produce perceivably distinct body shapes at the tested levels.
- domain assumption Converting the 5-point Likert confidence response into a numerical accuracy score (5,4,3,2,1) yields an interval-scale measure suitable for ANOVA.
- domain assumption The task of identifying the side with all-unique motions is a valid proxy for perceived crowd realism.
Cite this review
Pith. "Pith review of Shape Shifters: Does Body Shape Change the Perception of Small-Scale Crowd Motions?." pith.science (2026). https://pith.science/paper/57PX554D
@misc{pith2026241216151,
author = {Pith},
title = {Pith review of: Shape Shifters: Does Body Shape Change the Perception of Small-Scale Crowd Motions?},
year = {2026},
howpublished = {\url{https://pith.science/paper/57PX554D}},
note = {Machine review of arXiv:2412.16151}
}
read the original abstract
The animation of realistic virtual avatars in crowd scenarios is an important element of immersive virtual environments. However, achieving this realism requires attention to multiple factors, such as their visual appearance and motion cues. We investigated how body shape diversity influences the perception of motion clones in virtual crowds. A physics-based model was used to simulate virtual avatars in a small-scale crowd of size twelve. Participants viewed side-by-side video clips of these virtual crowds: one featuring all unique motions (Baseline) and the other containing motion clones (i.e., the same motion used to animate two or more avatars in the crowd). We also varied the levels of body shape and motion diversity. Our findings revealed that body shape diversity did not influence participants' ratings of motion clone detection, and motion variety had a greater impact on their perception of the crowd. Further research is needed to investigate how other visual factors interact with motion in order to enhance the perception of virtual crowd realism.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
E. Alvarado, D. Rohmer, and M.-P. Cani. Generating upper-body mo- tion for real-time characters making their way through dynamic envi- ronments. Computer Graphics Forum, 41(8):169–181, 2022. 2, 3
work page 2022
-
[3]
D. Anguelov, P. Srinivasan, D. Koller, S. Thrun, J. Rodgers, and J. Davis. Scape: shape completion and animation of people. In ACM SIGGRAPH 2005 Papers, pp. 408–416. 2005. 2
work page 2005
-
[4]
T. Beardsworth and T. Buckner. The ability to recognize oneself from a video recording of one’s movements without seeing one’s body.Bul- letin of the Psychonomic Society, 18(1):19–22, 1981. 2
work page 1981
-
[5]
A. O. Bebko, A. Thaler, and N. F. Troje. bmlsup–a smpl unity player. In 2021 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW), pp. 573–574. IEEE, 2021. 3
work page 2021
-
[6]
B. C. Daniel, R. Marques, L. Hoyet, J. Pettré, and J. Blat. A perceptually-validated metric for crowd trajectory quality evaluation. Proceedings of the ACM on Computer Graphics and Interactive Tech- niques, 4(3):1–18, 2021. 3
work page 2021
-
[7]
P. Faloutsos, M. Van de Panne, and D. Terzopoulos. Composable con- trollers for physics-based character animation. In ACM SIGGRAPH 2001 Papers, pp. 251–260, 2001. 2
work page 2001
-
[8]
T. Geijtenbeek and N. Pronost. Interactive character animation using simulated physics: A state-of-the-art review. In Computer graphics forum, vol. 31, pp. 2492–2515. Wiley Online Library, 2012. 2
work page 2012
Show all 40 references
-
[9]
Hoyet, A.-H
L. Hoyet, A.-H. Olivier, R. Kulpa, and J. Pettré. Perceptual effect of shoulder motions on crowd animations. ACM Transactions on Graph- ics (TOG), 35(4):1–10, 2016. 1, 5
2016
-
[10]
Hoyet, K
L. Hoyet, K. Ryall, K. Zibrek, H. Park, J. Lee, J. Hodgins, and C. O’sullivan. Evaluating the distinctiveness and attractiveness of hu- man motions on realistic virtual bodies. ACM Transactions on Graph- ics (TOG), 32(6):1–11, 2013. 2, 3
2013
-
[11]
Huang and M
Y . Huang and M. Kallmann. Motion parameterization with inverse blending. In Motion in Games: Third International Conference, MIG 2010, Utrecht, The Netherlands, November 14-16, 2010. Proceedings 3, pp. 242–253. Springer, 2010. 1
2010
-
[12]
Johansson
G. Johansson. Visual perception of biological motion and a model for its analysis. Perception & psychophysics, 14:201–211, 1973. 2
1973
-
[13]
K. L. Johnson and L. G. Tassinary. Compatibility of basic social per- ceptions determines perceived attractiveness. Proceedings of the Na- tional Academy of Sciences, 104(12):5246–5251, 2007. 2
2007
-
[14]
Kenny, N
S. Kenny, N. Mahmood, C. Honda, M. J. Black, and N. F. Troje. Per- ceptual effects of inconsistency in human animations. ACM Transac- tions on Applied Perception (TAP), 16(1):1–18, 2019. 2
2019
-
[15]
L. T. Kozlowski and J. E. Cutting. Recognizing the sex of a walker from a dynamic point-light display. Perception & psychophysics , 21:575–580, 1977. 2
1977
-
[16]
Kulpa, A.-H
R. Kulpa, A.-H. Olivierxs, J. Ond ˇrej, and J. Pettré. Imperceptible relaxation of collision avoidance constraints in virtual crowds. InPro- ceedings of the 2011 SIGGRAPH Asia Conference , pp. 1–10, 2011. 3
2011
-
[17]
Levine and J
S. Levine and J. Popovi ´c. Physically plausible simulation for character animation. In Proceedings of the 11th ACM SIGGRAPH/Eurographics conference on Computer Animation, pp. 221–230, 2012. 2
2012
-
[18]
Loper, N
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, Oct. 2015. 2
2015
-
[19]
Magnenat-Thalmann and D
N. Magnenat-Thalmann and D. Thalmann. Handbook of virtual hu- mans. John Wiley & Sons, 2005. 1
2005
-
[20]
Mahmood, N
N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black. Amass: Archive of motion capture as surface shapes. In Pro- ceedings of the IEEE/CVF international conference on computer vi- sion, pp. 5442–5451, 2019. 3
2019
-
[21]
McDonnell, S
R. McDonnell, S. Jörg, J. K. Hodgins, F. Newell, and C. O’Sullivan. Virtual shapers & movers: form and motion affect sex perception. In Proceedings of the 4th Symposium on Applied Perception in Graphics and Visualization, pp. 7–10, 2007. 2
2007
-
[22]
McDonnell, M
R. McDonnell, M. Larkin, S. Dobbyn, S. Collins, and C. O’Sullivan. Clone attack! perception of crowd variety. In ACM SIGGRAPH 2008 papers, pp. 1–8. 2008. 1, 2
2008
-
[23]
McDonnell, M
R. McDonnell, M. Larkin, B. Hernández, I. Rudomin, and C. O’Sullivan. Eye-catching crowds: saliency based selective vari- ation. ACM Transactions on Graphics (TOG), 28(3):1–10, 2009. 2
2009
-
[24]
W. H. Organization et al. Physical status: The use of and interpre- tation of anthropometry, Report of a WHO Expert Committee. World Health Organization, 1995. 2
1995
-
[25]
Pavlakos, V
G. Pavlakos, V . Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp. 10975– 10985, 2019. 3
2019
-
[26]
X. B. Peng, P. Abbeel, S. Levine, and M. Van de Panne. Deepmimic: Example-guided deep reinforcement learning of physics-based char- acter skills. ACM Transactions On Graphics (TOG) , 37(4):1–14,
-
[27]
Pražák and C
M. Pražák and C. O’Sullivan. Perceiving human motion variety. In Proceedings of the ACM SIGGRAPH Symposium on Applied Percep- tion in Graphics and Visualization, pp. 87–92, 2011. 1, 2, 3, 5
2011
-
[28]
Romero, D
J. Romero, D. Tzionas, and M. J. Black. Embodied hands: Mod- eling and capturing hands and bodies together. arXiv preprint arXiv:2201.02610, 2022. 3
2022 arXiv
-
[29]
Russell, S
J. Russell, S. Saboor, M. Makkar, A. Barua, and A. Thaler. Detection of inconsistency between shape and motion in realistic female and male animation. STEM Fellowship Journal, 8(1):17–23, 2023. 2, 3, 5
2023
-
[30]
Y . Shi, J. Ondˇrej, H. Wang, and C. O’Sullivan. Shape up! perception based body shape variation for data-driven crowds. In 2017 ieee vir- tual humans and crowds for immersive environments (vhcie), pp. 1–7. IEEE, 2017. 1, 2
2017
-
[31]
D. B. d. Silva, C. A. Vidal, J. B. Cavalcante-Neto, I. N. Pessoa, and R. F. Nunes. Topple-free foot strategy applied to real-time motion capture data using kinect sensor. In Proceedings of the 36th Annual ACM Symposium on Applied Computing, pp. 98–106, 2021. 3
2021
-
[32]
G. C. Silva, A. T. da Silva, and M. da Silva Hounsell. Crowd genera- tion using morphological obesity criteria. In2019 18th Brazilian Sym- posium on Computer Games and Digital Entertainment (SBGames) , pp. 81–90. IEEE, 2019. 1
2019
-
[33]
Tecchia, C
F. Tecchia, C. Loscos, and Y . Chrysanthou. Image-based crowd ren- dering. IEEE computer graphics and applications, 22(2):36–43, 2002. 1
2002
-
[34]
N. F. Troje. Decomposing biological motion: A framework for analy- sis and synthesis of human gait patterns. Journal of vision, 2(5):2–2,
-
[35]
Ulicny, P
B. Ulicny, P. d. H. Ciechomski, and D. Thalmann. Crowdbrush: inter- active authoring of real-time crowd scenes. InProceedings of the 2004 ACM SIGGRAPH/Eurographics symposium on Computer animation, pp. 243–252, 2004. 1
2004
-
[36]
B. Vyas, L. Hoyet, and C. O’Sullivan. Exploring the perception of center of mass changes for vr avatars. In ICAT-EGVE 2023- International Conference on Artificial Reality and Telexistence & Eu- rographics Symposium on Virtual Environments, pp. 1–12, 2023. 2, 3, 5
2023
-
[37]
J. Won, D. Gopinath, and J. Hodgins. A scalable approach to control diverse behaviors for physically simulated characters. ACM Transac- tions on Graphics (TOG), 39(4):33–1, 2020. 2
2020
-
[38]
J. Won, D. Gopinath, and J. Hodgins. Control strategies for physically simulated characters performing two-player competitive sports. ACM Transactions on Graphics (TOG), 40(4):1–11, 2021. 2
2021
-
[39]
Won and J
J. Won and J. Lee. Learning body shape variation in physics-based characters. ACM Transactions on Graphics (TOG), 38(6):1–12, 2019. 2
2019
-
[40]
K. Yin, K. Loken, and M. Van de Panne. Simbicon: Simple biped lo- comotion control. ACM Transactions on Graphics (TOG), 26(3):105– es, 2007. 3
2007
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.