REVIEW 4 major objections 4 minor 31 references
Learning to gesticulate by observation using a deep generative approach
T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper claims that a GAN trained on human motion capture gives the Pepper robot a varied, natural repertoire of talking gestures.
desk verdict Realistic gesture data in, plausible Pepper motion out, but the naturalness claim rests only on videos—and the concatenation of independently sampled 4-pose units is exactly the spot where their earlier system turned jerky. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the 'unit of movement': a four-pose sequence of joint configurations, each pose being 14 float values covering head, arms, wrists, and finger opening. The GAN's discriminator learns the distribution of these 56-element samples, while the generator maps 100 uniform random values to new units; training runs for 2000 epochs with empirically tuned hyperparameters. This machinery does the work of replacing a fixed, hand-compiled gesture list with an open-ended sampler, so that each rendering of a talk can draw a fresh sequence of movements from the learned distribution. A separate direct-kinematic mapping transfers the captured human skeleton and approximate hand state into Pepper joint commands.
What would settle it
Record the robot's commanded joint angles during a generated utterance and compute the peak joint velocity or acceleration at each boundary between concatenated units; if those boundary spikes are much larger than the velocities within units, the random-chaining premise fails and the naturalness claim would need temporal smoothing to survive.
Extended reading notes
Core claim
On its own terms, the paper's central claim is that a GAN trained on 2018 units of movement—each unit a sequence of four consecutive poses, each pose holding 14 joint values for head, arms, wrists, and finger opening—can generate talking gestures for the Pepper robot that are appropriate and natural in the sense of being varied rather than repetitive. Human poses captured by a depth camera are translated joint-by-joint into Pepper's joint commands, with wrist yaw inferred from colored gloves and finger positions randomized at each frame; the GAN's generator, seeded with 100-dimensional uniform noise, produces new samples from this distribution. The robot's whole gesture sequence is built by concatenating these generated units one after another, with the number of units determined by the length of the speech audio. The paper reports that the resulting behavior is appropriate and that movement variability gives the robot naturalness.
Load-bearing premise
The load-bearing premise is that randomly chaining short, independently sampled four-pose units produces smooth, coherent natural motion; the paper provides no smoothing, no coherence metric, and no evaluation of the transitions between units.
Editorial extensions
If this is right
- The same trained model can accompany speech of any length, because the audio duration only decides how many generated units are chained together.
- Each rendering of a talk draws new sampled units, so repeated utterances need not repeat the same head and arm movements.
- Hand and finger motion is included even though the skeleton tracker cannot see it: wrist yaw is inferred from colored gloves and finger positions are randomized each frame.
- If the reported naturalness is real, the approach replaces manual choreography or gesture lookup with a generative model whose variability is the source of perceived naturalness.
Reading between the lines
- A testable extension not run in the paper: measure joint velocity and acceleration at the seams between concatenated units; if the seam spikes are no larger than within-unit motion, the chaining design is validated, and if they spike, temporal smoothing would be the needed fix.
- The paper leaves implicit that the data pool of five speakers bounds the vocabulary of learned gestures; recording unscripted talks or more speakers would directly test whether variety scales with the training distribution.
- A forced-choice user study comparing the GAN stream against the earlier random-concatenation baseline would make the naturalness claim falsifiable: the claim predicts listeners prefer the generated stream.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a system for generating talking gestures for a Pepper robot. Human upper-body motion is captured with a Kinect, mapped to the robot's joint space through direct kinematics, and used to train a GAN whose generator outputs short sequences of four poses, called units of movement (UMs). During robot execution, the number of UMs is chosen according to the intended speech duration and the UMs are reproduced one after another. The paper claims that the resulting robot behavior is appropriate and natural, and this claim is supported only by two linked videos, one showing motion capture and one showing the GAN-based behavior after training.
Significance. The idea of using a GAN trained on human motion capture to generate talking gestures for a social robot is timely and relevant, and the kinematic mapping procedure is described in sufficient detail to be a useful starting point for other researchers. The strength of the paper is its pipeline: motion capture, direct kinematic retargeting, GAN training on short movement units, and real-robot reproduction. However, the central claim of naturalness is not supported by any quantitative evaluation, baseline comparison, user study, or analysis of the concatenation process. If the naturalness claim were properly supported, the contribution would be a useful building block for spontaneous gesture generation, but in its current form the evidence is primarily anecdotal.
major comments (4)
- [Section 4] The full robot movement is formed by concatenating independently sampled four-pose UMs, as stated in Section 4: 'the execution of those units of movements, one after the other, defines the whole movement displayed by the robot.' Since each UM is generated from an independent uniform random seed, there is no mechanism described that aligns, filters, or blends the final pose of one UM with the initial pose of the next. No joint-position, velocity, or acceleration continuity is guaranteed across the boundaries. This is a load-bearing problem because Section 1 states that the previous random-concatenation approach 'resulted in unnatural jerky expression.' The manuscript does not provide any evaluation of inter-unit transitions, so the claim that the new GAN-based system avoids the same failure is untested. The authors should either describe an explicit transition-smoothing mechanism or provide a quantitative analysis of boundary continuity and its effect on perceived naturalness.
- [Section 5 and Results] The conclusion states that 'Results show that the obtained robot behavior is appropriate, and thanks to the movement variability the robot expresses itself with naturalness.' This is not supported by the presented evidence. The results section contains only two video links and no quantitative measures such as joint-velocity continuity, jerk, gesture-range statistics, or comparison with the previous random-concatenation method, nor any human evaluation. Since 'naturalness' is the central claim of the paper, the evaluation must include at least one objective or subjective measure that is reported in the manuscript; otherwise the conclusion is an assertion rather than a demonstrated result.
- [Section 3.1, Eqs. (4)-(6)] The mapping equations contain unspecified constants: the normalizing constant N in Eq. (4) and the gains K1 and K2 in Eqs. (5) and (6). These values are load-bearing because they directly determine the robot's joint ranges and therefore the character of the generated motion. The manuscript does not state how they were chosen, whether they are fixed across all participants, or how sensitive the resulting behavior is to their values. Without these details, the system is not fully reproducible and the reader cannot assess whether the gains were tuned to make the videos look favorable. At minimum, the authors should report the values and justify them; a brief sensitivity analysis would be even better.
- [Section 3.1, fingers] The sentence 'Regarding the fingers, as they cannot be tracked, their position is randomly set at each skeleton frame' introduces an unstated assumption that random per-frame finger positions do not degrade naturalness. Fingers are visible in the videos and contribute to perceived gesture naturalness, so random noise at every frame may cause visibly unnatural hand motion. The authors should provide evidence that this choice is acceptable, for example by comparing the generated hand motion with a fixed neutral hand posture or with the recorded human hand motion.
minor comments (4)
- [Conclusions] The phrase 'a GAN feeded with natural motion data' should be 'a GAN fed with natural motion data.' Also, 'The work presented here pretends to be the starting point' should be 'intends to be the starting point.'
- [Section 3.1, Eq. (3)] The derivation of the shoulder pitch angle is hard to follow: the statement 'z = 0 occurs with the arm extended at 90 degrees with respect to the torso' is ambiguous, and the notation '||A|| = LSEz (by definition)' does not clarify why the Z coordinate alone equals the vertical component. A diagram or a more explicit definition of the coordinate frame would improve readability.
- [Section 3.1, Eq. (5)] Equation (5) writes Hrobot_gamma = K1 * H_beta, but the text says the gain is applied to the human's yaw value. The variable names 'gamma' and 'beta' appear to be swapped or inconsistently used; please align the notation with the joint definitions in Figure 3.
- [Section 3.2] The dataset description reports 2018 UMs from five speakers over about nine minutes, but it does not specify the sampling rate, the number of UMs per speaker, or whether there was any train/validation split or data augmentation. Adding these details would help assess the risk of GAN overfitting and the diversity of the training data.
Circularity Check
No significant circularity: the system trains a GAN on captured human motion and maps generated poses to the robot by direct kinematics, so the central result is not equivalent to its inputs.
full rationale
The paper's derivation chain is self-contained with respect to its stated goal. Human talking gestures are captured with a Kinect, converted to Pepper joint angles through explicit kinematic equations (Eqs. 1-6), segmented into 2018 four-pose units of movement, and used to train a GAN whose generator outputs new 56-dimensional pose sequences from a 100-dimensional uniform latent input. The final robot motion is the concatenation of these generated units. No prediction is derived from a fitted parameter that was itself fit to that prediction, and no result is obtained by renaming an input as an output. The only self-citations are contextual: [25] describes a prior teleoperation setup, [24] describes an earlier random-concatenation approach, and [26] is a related GAN study, but the present paper does not rely on these citations as the load-bearing evidence for its naturalness claim. The authors acknowledge limitations, such as the lack of a fully objective database and the constraining effect of recording, and the absence of quantitative transition-smoothness evaluation is a legitimate correctness concern, not a circularity. Overall, the central claim is an empirical demonstration supported by videos and a trained generative model, with no circular definition, fitted input masquerading as prediction, or author-imported uniqueness theorem.
Assumptions & free parameters
free parameters (3)
- K1
- K2
- N
assumptions (7)
- standard math Trigonometric identities and dot products used in the joint angle calculations are valid.
- domain assumption The Kinect/OpenNI skeleton tracker provides accurate 3D joint positions for the 15 tracked joints.
- domain assumption Direct kinematics from human to Pepper's joints preserves enough gesture expressiveness for the GAN to learn natural motion.
- domain assumption The GAN, trained on 2018 four-pose units from 5 people over 2000 epochs, captures the distribution of natural talking gestures.
- domain assumption Concatenating independently generated units of movement produces smooth, natural long sequences.
- ad hoc to paper Randomly setting finger positions at each frame does not degrade naturalness.
- ad hoc to paper The head pitch correction formula, including an unspecified gain K2, is an appropriate heuristic.
Cite this review
Pith. "Pith review of Learning to gesticulate by observation using a deep generative approach." pith.science (2026). https://pith.science/paper/7PL4QHR7
@misc{pith2026190901768,
author = {Pith},
title = {Pith review of: Learning to gesticulate by observation using a deep generative approach},
year = {2026},
howpublished = {\url{https://pith.science/paper/7PL4QHR7}},
note = {Machine review of arXiv:1909.01768}
}
read the original abstract
The goal of the system presented in this paper is to develop a natural talking gesture generation behavior for a humanoid robot, by feeding a Generative Adversarial Network (GAN) with human talking gestures recorded by a Kinect. A direct kinematic approach is used to translate from human poses to robot joint positions. The provided videos show that the robot is able to use a wide variety of gestures, offering a non-dreary, natural expression level.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Journal of Intelligent & Robotic Systems 85(1), 27–45 (Jan 2017)
Alibeigi, M., Rabiee, S., Ahmadabadi, M.N.: Inverse kinematics based human mim- icking system using skeletal tracking technology. Journal of Intelligent & Robotic Systems 85(1), 27–45 (Jan 2017)
work page 2017
-
[2]
Creative Robot Dance with Variational Encoder
Augello, A., Cipolla, E., Infantino, I., Manfr` e, A., Pilato, G., Vella, F.: Creative robot dance with variational encoder. CoRR abs/1707.01489 (2017)
work page Pith review arXiv 2017
-
[3]
In: Social signal processing, chap
Beck, A., Yumak, Z., Magnenat-Thalmann, N.: Body movements generation for virtual characters and social robots. In: Social signal processing, chap. 20, pp. 273–286. Cambridge University Press (2017)
work page 2017
-
[4]
Intelligent Robotics and Autonomous Agents, MIT Press, Cambridge MA, USA (2004)
Breazeal, C.: Designing sociable robots. Intelligent Robotics and Autonomous Agents, MIT Press, Cambridge MA, USA (2004)
work page 2004
-
[5]
In: arXiv preprint arXiv:1812.08008 (2018)
Cao, Z., Hidalgo, G., Simon, T., Wei, S.E., Sheikh, Y.: OpenPose: realtime multi-person 2D pose estimation using Part Affinity Fields. In: arXiv preprint arXiv:1812.08008 (2018)
arXiv 2018
-
[6]
Cao, Z., Simon, T., Wei, S.E., Sheikh, Y.: Realtime multi-person 2D pose estima- tion using part affinity fields. In: CVPR (2017)
work page 2017
-
[7]
Expert Systems and Probabilistic Network Models
Enrique Castillo, J.M.G., Hadi, A.S.: Learning Bayesian Networks. Expert Systems and Probabilistic Network Models. Monographs in computer science. New York: Springer-Verlag (1997)
work page 1997
-
[8]
Everitt, B., Hand, D.: Finite mixture distributions. Chapman and Hall (1981)
work page 1981
Show all 31 references
-
[9]
In: International Confer- ence on Advanced Mechatronics, Intelligent Manufacture, and Industrial Automa- tion (ICAMIMIA)
Fadli, H., Machbub, C., Hidayat, E.: Human gesture imitation on NAO humanoid robot using Kinect based on inverse kinematics method. In: International Confer- ence on Advanced Mechatronics, Intelligent Manufacture, and Industrial Automa- tion (ICAMIMIA). IEEE (2015)
2015
-
[10]
ArXiv e-prints (Dec 2017)
Goodfellow, I.: NIPS Tutorial: Generative Adversarial Networks. ArXiv e-prints (Dec 2017)
2017
-
[11]
In: Advances in neural information processing systems
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. In: Advances in neural information processing systems. pp. 2672–2680 (2014)
2014
-
[12]
CoRR abs/1803.10892 (2018), http://arxiv.org/abs/1803.10892
Gupta, A., Johnson, J., Fei-Fei, L., Savarese, S., Alahi, A.: Social GAN: socially acceptable trajectories with Generative Adversarial Networks. CoRR abs/1803.10892 (2018), http://arxiv.org/abs/1803.10892
2018 arXiv
-
[13]
arXiv preprint arXiv:1412.6980 (2014) 10 U
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014) 10 U. Zabala et al
2014 arXiv
-
[14]
In: International Conference on Intelligent Robots and Systems (IROS)
Kwon, J., Park, F.C.: Using Hidden Markov Models to Generate Natural Humanoid Movement. In: International Conference on Intelligent Robots and Systems (IROS). IEEE/RSJ (2006)
2006
-
[15]
Nature 521(7553), 436–444 (2015)
LeCun, Y., Bengio, Y., Hinton, G.: Deep learning. Nature 521(7553), 436–444 (2015)
2015
-
[16]
MacCormick, J.: How does the Kinect work? http://pages.cs.wisc.edu/ ah- mad/kinect.pdf (accessed June 3, 2019)
2019
-
[17]
Biologically Inspired Cognitive Architectures 15, 1–9 (2016)
Manfr` e, A., Infantino, I., Vella, F., Gaglio, S.: An automatic system for humanoid dance creation. Biologically Inspired Cognitive Architectures 15, 1–9 (2016)
2016
-
[18]
University of Chicago press (1992)
McNeill, D.: Hand and mind: What gestures reveal about thought. University of Chicago press (1992)
1992
-
[19]
ACM Trans
Mehta, D., Sridhar, S., Sotnychenko, O., Rhodin, H., Shafiei, M., Seidel, H.P., Xu, W., Casas, D., Theobalt, C.: VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera. ACM Trans. Graph. 36(4), 44:1–44:14 (Jul 2017)
2017
-
[20]
In: International Conference on Robotics, Automation, Control and Embedded Systems (RACE)
Mukherjee, S., Paramkusam, D., Dwivedy, S.K.: Inverse kinematics of a NAO hu- manoid robot using Kinect to track and imitate human motion. In: International Conference on Robotics, Automation, Control and Embedded Systems (RACE). IEEE (2015)
2015
-
[21]
IEEE Transactions on Robotics 30, 771–778 (Jun 2014)
Okamoto, T., Shiratori, T., Kudoh, S., Nakaoka, S., Ikeuchi, K.: Toward a dancing robot with listening capability: Keypose-based integration of lower-, middle-, and upper-body motions for varying music tempos. IEEE Transactions on Robotics 30, 771–778 (Jun 2014). https://doi.o...
2014
-
[22]
Master’s thesis, Ecole Centrale de Nantes–Warsaw Uni- versity of Technology (2013)
Poubel, L.P.: Whole-body Online Human Motion Imitation by a Humanoid Robot Using Task Specification. Master’s thesis, Ecole Centrale de Nantes–Warsaw Uni- versity of Technology (2013)
2013
-
[23]
In: Proceedings of the IEEE
Rabiner, L.R.: A tutorial on Hidden Markov Models and selected applications in speech recognition. In: Proceedings of the IEEE. vol. 77, pp. 257–286 (1989)
1989
-
[24]
In: IEEE International Conference on Robotics and Automation (ICRA)
Rodriguez, I., Astigarraga, A., Ruiz, T., Lazkano, E.: Singing minstrel robots, a means for improving social behaviors. In: IEEE International Conference on Robotics and Automation (ICRA). pp. 2902–2907 (2016)
2016
-
[25]
In: International Conference on Humanoid Robots (Humanoids) (2014)
Rodriguez, I., Astigarraga, A., Jauregi, E., Ruiz, T., Lazkano, E.: Humanizing NAO robot teleoperation using ROS. In: International Conference on Humanoid Robots (Humanoids) (2014)
2014
-
[26]
Robotics and Autonomous Systems 114, 57 – 65 (2019)
Rodriguez, I., Mart´ ınez-Otzeta, J.M., Irigoien, I., Lazkano, E.: Spontaneous talking gestures using generative adversarial networks. Robotics and Autonomous Systems 114, 57 – 65 (2019)
2019
-
[27]
In: International Conference on Robotics and Automation (ICRA)
Schubert, T., Eggensperger, K., Gkogkidis, A., Hutter, F., Ball, T., Burgard, W.: Automatic bone parameter estimation for skeleton tracking in optical motion cap- ture. In: International Conference on Robotics and Automation (ICRA). IEEE (2016)
2016
-
[28]
Tanwani, A.K.: Generative Models for learning Robot Manipulation. Ph.D. thesis, ´Ecole Polytechnique F´ ed´ eral de Laussane (EPFL) (2018)
2018
-
[29]
PLOS ONE 13(7), 1–21 (Jul 2018)
Tits, M., Tilmanne, J., Dutoit, T.: Robust and automatic motion-capture data recovery using soft skeleton constraints and model averaging. PLOS ONE 13(7), 1–21 (Jul 2018)
2018
-
[30]
Chemometrics and intelligent laboratory systems 2(1-3), 37–52 (1987)
Wold, S., Esbensen, K., Geladi, P.: Principal component analysis. Chemometrics and intelligent laboratory systems 2(1-3), 37–52 (1987)
1987
-
[31]
Applied Sciences 8(10) (2018), https://www.mdpi.com/ 2076-3417/8/10/2005
Zhang, Z., Niu, Y., Yan, Z., Lin, S.: Real-time whole-body imitation by humanoid robots and task-oriented teleoperation using an analytical mapping method and quantitative evaluation. Applied Sciences 8(10) (2018), https://www.mdpi.com/ 2076-3417/8/10/2005
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.