REVIEW 3 major objections 7 minor 47 references
MAGE:A Multi-stage Avatar Generator with Sparse Observations
T0 review · 3 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read MAGE reconstructs full-body motion from three tracked joints by generating a 6-node skeleton first, then 11, then the final 22 joints, and reports that this progressive scheme beats one-stage mapping on accuracy and smoothness.
desk verdict Useful incremental architecture for sparse avatar generation, but the central coarse-to-fine claim is not isolated from model capacity; requires a parameter-matched ablation and a clearer description of the multi-scale skeletons. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The multi-scale human motion representation is the central object: three skeletons $S_1$ (6 composite nodes), $S_2$ (11 composite nodes), and $S_3$ (22 SMPL joints), obtained by iteratively merging adjacent joints in the SMPL kinematic tree. Generation runs in reverse, through a single diffusion model divided into three cascaded phases: it outputs $\hat{S}_1$, then $\hat{S}_2$ conditioned on $\hat{S}_1$, then $\hat{S}_3$ conditioned on $\hat{S}_2$. The objective is a weighted sum $\alpha L_1 + \beta L_2 + \gamma L_3$ of per-scale losses, and each stage concatenates the sparse observations $C$, the intermediate denoised features $F$, and the re-embedded stage output $F^{\mathrm{rec}}$. The coarse scales act as global constraints that narrow the inference space for the finer scales.
What would settle it
Run an ablation where the 6- and 11-node skeletons are replaced by a different fixed merging, such as anatomical limb groups or randomly chosen joint clusters; if the coarse-to-fine advantage on MPJRE, MPJVE, and Jitter disappears, the reported gains depend on the exact merging, which is never specified.
Extended reading notes
Core claim
The central discovery is a design strategy: full-body motion from sparse HMD observations should not be predicted as one direct 3-to-22 mapping, but progressively from a 6-composite-node skeleton that sets global motion, to an 11-composite-node refinement, to all 22 SMPL joints. Each stage's prediction is fed back as a condition for the next, so later stages see richer constraints and a narrower inference space. In the paper's experiments on AMASS D1, MAGE reports the best MPJRE (2.40 deg), MPJPE (3.21 cm), MPJVE (16.71 cm/s), and Jitter (6.27) among compared methods; on D2 it reports the best MPJRE, MPJVE, and Jitter. The paper interprets these results as evidence that coarse-to-fine generation reduces cumulative kinematic error, especially in the lower body, and balances static accuracy against dynamic continuity.
Load-bearing premise
The load-bearing premise is that the hand-designed merging of SMPL joints into the 6- and 11-node composite skeletons keeps enough kinematic structure that coarse-stage predictions constrain rather than mislead finer stages, and the paper never states the merging rule.
Editorial extensions
If this is right
- A headset and two handheld controllers are enough to drive a full avatar in real time: MAGE reports about 0.36 ms per frame with 4-step DDIM sampling, far beyond AR/VR frame-rate requirements.
- Lower-body prediction, a weak spot of one-stage methods, improves because the coarse 6-node stage fixes global motion before leg details are added; on D1 MAGE reports lower-body position error of 5.93 cm versus 6.01 cm for the next-best compared method.
- Generated motions are smoother as well as more accurate: on D1 Jitter drops to 6.27 from 6.55 for the closest compared method, and on D2 MAGE reports 5.81 versus 7.13 for SAGE.
- The multi-scale supervision acts as a regularizer, so the model can generalize across the different training and test motion distributions in D2 without overfitting the direct 3-to-22 mapping.
- The staged design gives a flexible place to insert extra constraints or priors in future work, since each intermediate representation is itself a supervised prediction that can be checked or guided.
Reading between the lines
- Editorial inference: The coarse-to-fine recipe is not tied to the specific 3-joint input; it could be applied to other sparse tracker layouts, and the paper's own S0 ablation suggests the benefit is not monotonic in the number of scales.
- Editorial inference: Because the paper never specifies how SMPL joints are merged into the 6- and 11-node skeletons, a learned or data-driven hierarchy is a natural next step that could make the method less sensitive to the hand-chosen grouping.
- Editorial inference: Directly predicting the clean target at each stage, as MAGE does, creates multiple error-correction points, so a failure at a coarse stage need not be fatal if a finer stage can revise it; this is a testable design principle beyond avatar generation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses full-body avatar generation from sparse HMD observations (head and two wrists). The proposed method, MAGE, is a conditional diffusion model that replaces the standard direct 3-to-22-joint mapping with three progressive stages: it first predicts a coarse 6-node composite skeleton, then an 11-node intermediate skeleton, then the full 22-joint SMPL pose; each stage's output is embedded and concatenated with the sparse observations to condition the next stage, and the final training loss is a weighted sum of per-stage regression losses. The authors evaluate on two AMASS splits (D1, D2) and report state-of-the-art or competitive results on MPJRE, MPJPE, MPJVE, and Jitter, together with region-specific errors (Hand/Upper/Lower/Root PE), ablations over the scale set, the multi-stage architecture, and the feature fusion scheme, and a real-time inference claim (0.36 ms/frame with 4-step DDIM). The central claim is that coarse-to-fine factorization reduces the inference space and cumulative kinematic errors, leading to better accuracy and temporal consistency.
Significance. If the central claim holds, the contribution is a useful and well-motivated empirical idea: factorizing an ill-posed 3-to-22 joint mapping through coarse composite skeletons is natural given the SMPL kinematic tree, and the paper supplies consistent ranking improvements on two benchmarks, a cross-distribution generalization test (D2), and three focused ablations. Credit is due for the clear formulation of the stage losses, the attempt to isolate the contribution of the multi-scale design, and the real-time performance report. The main caveat is that the headline mechanism is supported by an ablation (Table 3) that does not control for model capacity, and the coarse skeleton construction (Section 3.2) is not specified at a level that permits reproduction. The method is empirical and supervised by AMASS ground truth, so there is no circularity in the stage losses. The practical significance for AR/VR applications is real, provided the gains survive capacity-matched controls.
major comments (3)
- [Section 4.1, Table 3 (scale-set ablation)] The scale-set ablation in Table 3 is confounded with model capacity. Section 4.1 fixes 12 denoiser blocks per stage (a=b=c=12), so the S1+S2+S3 model contains 36 blocks, whereas the S3-only row contains 12 blocks and the S1+S3 and S2+S3 rows contain 24; the improvement from S3-only to S1+S2+S3 is therefore in the same direction as the increase in parameters and compute. The S0+S1+S2+S3 row (48 blocks) shows the benefit is not strictly monotone in capacity, but the decisive comparison remains unperformed: no capacity-matched control (a 36-block S3-only model, or a multi-stage variant with shared weights giving 12 effective blocks) is reported, even though Section 4.4 (Table 4) demonstrates that the authors match total layer counts when they want a fair architectural comparison. As it stands, the experiment attributed to the coarse-to-fine mechanism could equally be explained by a larger denoiser stack.
- [Section 3.2, Figure 2, Eqs. (6)-(8)] The construction of the composite skeletons is under-specified to the point of non-reproducibility. The text states that adjacent SMPL joints are 'iteratively merge[d]' into the 6-node skeleton S1 and the 11-node skeleton S2, but it never specifies which joints form which composite nodes, nor how a composite node's 6D rotation is derived from the rotations of its member joints. This choice matters: averaging 6D rotation vectors is not a geometrically well-defined rotation average, while taking a member joint's rotation (e.g., the parent) would make the coarse losses L1 and L2 carry no information about the orientations of the merged children. Since L1 and L2 are regression targets that shape everything the later stages can constrain, the explicit 22-to-11-to-6 mapping and the composite rotation (and position) definition must be given.
- [Abstract, Tables 1 and 2] The quantitative claims in the abstract are not consistently supported by the tables. The stated improvements of 5% MPJRE, 10% MPJVE, and 11% Jitter match D1 only for selected baselines (5.1% MPJRE relative to SAGE; 10.1% MPJVE relative to AGRoL), and no baseline yields the 11% Jitter figure (4.3% relative to SAGE and 13.6% relative to AGRoL on D1; 18.5% relative to SAGE on D2). In addition, several reported advantages are small relative to what a single run can resolve: on D1, MPJPE is 3.21 for MAGE versus 3.28 for SAGE; on D2, MPJRE is 4.26 for MAGE versus 4.30 for both AvatarJLM and AGRoL. No error bars, no number of training seeds, and no significance tests are reported. The claim that MAGE 'significantly outperforms state-of-the-art methods' should be backed by variance information or a significance test on the main comparison tables.
minor comments (7)
- [Eqs. (6)-(8)] It is not stated whether the previous stages' outputs S1-hat and S2-hat are detached from the computation graph inside L2 and L3; the gradient-flow decision changes training dynamics and should be specified for reproducibility.
- [Table 1, Conclusion, Section 2] Table 1's entry 'Avatorposer' should read 'AvatarPoser'; the Conclusion contains 'obsevations' for 'observations'; and 'V AEHMD' and 'VAE-HMD' are used inconsistently in Section 2.
- [Sections 3.1 and 4.1] The condition notation is inconsistent: Section 3.1 writes C^{1:N} in R^{N by 18 by M}, while Section 4.1 writes R^{N by 18 by 3}; since M=3, the two agree, but the identification should be stated once.
- [Figure 2] Figure 2 labels the three scales only by node count; annotating the composite-node groupings (or giving them in the caption) would resolve part of the reproducibility gap noted in Major Comment 2.
- [Section 4.2, Tables 1-2] The provenance of baseline numbers should be documented: it should be stated whether all baseline results are taken from the original papers, from re-reported values in later papers, or from the authors' own reimplementation (only the D2 AvatarPoser result is currently attributed).
- [Section 4.1] The timing measurement behind the 0.36 ms/frame (2778 FPS) claim should state the batch size, sequence length, and whether it includes the overlapping-generation overhead.
- [References] A few references are incomplete: the Kingma and Welling entry lacks a venue and year, and the Denoising Diffusion Implicit Models entry appears as a title-only citation.
Circularity Check
No significant circularity: the multi-stage targets are external AMASS ground truth, not the model's own outputs.
full rationale
MAGE's derivation chain is self-contained and externally anchored. The conditioning features are computed from the observed head/wrist signals (Eqs. 1-2), and the supervised targets are the 22 SMPL joint rotations from AMASS. The coarse 6- and 11-node targets are deterministic merges of the same ground-truth skeleton (Sec. 3.2, Fig. 2), and losses L1-L3 (Eqs. 6-8) compare stage outputs with those external targets; no fitted parameter is renamed as a prediction, and no output is used to define the quantity it is supposed to predict. The citations to co-authors' work are contextual and not load-bearing, and RepIn [Du et al. 2023] is an architectural block, not an imported uniqueness result. The scale-set ablation (Table 3) is not a circular reduction: full MAGE uses more denoiser blocks than S3-only, so the coarse-to-fine gain is not fully isolated, but that is an experimental-control concern rather than equation-level circularity. Similarly, the under-specified joint-merging rule in Sec. 3.2 is a reproducibility limitation, not a circular step. Nothing in the paper reduces, by construction or by self-citation, to its own inputs.
Assumptions & free parameters
free parameters (6)
- Stage loss weights alpha, beta, gamma =
Not reported
- Denoiser blocks per stage =
12
- Sequence length N =
120
- DDIM inference steps =
4
- Historical frames for overlapping generation =
12
- Latent dimension =
512
assumptions (4)
- standard math Diffusion processes as defined by Ho et al. (2020), including the forward noise schedule and the reverse denoising objective, with direct X0 prediction as in Ramesh et al. (2022)
- domain assumption SMPL model: the first 22 joints and their local rotations fully determine full-body pose, and facial/hand joints can be discarded
- ad hoc to paper The coarse 6-node and 11-node representations preserve enough motion information to constrain fine-grained prediction
- domain assumption AMASS ground-truth data and the D1/D2 splits are valid for measuring sparse-input full-body tracking quality
invented entities (1)
-
Multi-scale composite skeletons S1 (6 nodes) and S2 (11 nodes)
Cite this review
Pith. "Pith review of MAGE:A Multi-stage Avatar Generator with Sparse Observations." pith.science (2026). https://pith.science/paper/2EUPQZ4I
@misc{pith2026250506411,
author = {Pith},
title = {Pith review of: MAGE:A Multi-stage Avatar Generator with Sparse Observations},
year = {2026},
howpublished = {\url{https://pith.science/paper/2EUPQZ4I}},
note = {Machine review of arXiv:2505.06411}
}
read the original abstract
Inferring full-body poses from Head Mounted Devices, which capture only 3-joint observations from the head and wrists, is a challenging task with wide AR/VR applications. Previous attempts focus on learning one-stage motion mapping and thus suffer from an over-large inference space for unobserved body joint motions. This often leads to unsatisfactory lower-body predictions and poor temporal consistency, resulting in unrealistic or incoherent motion sequences. To address this, we propose a powerful Multi-stage Avatar GEnerator named MAGE that factorizes this one-stage direct motion mapping learning with a progressive prediction strategy. Specifically, given initial 3-joint motions, MAGE gradually inferring multi-scale body part poses at different abstract granularity levels, starting from a 6-part body representation and gradually refining to 22 joints. With decreasing abstract levels step by step, MAGE introduces more motion context priors from former prediction stages and thus improves realistic motion completion with richer constraint conditions and less ambiguity. Extensive experiments on large-scale datasets verify that MAGE significantly outperforms state-of-the-art methods with better accuracy and continuity.
Figures
Reference graph
Works this paper leans on
-
[1]
Advanced Computing Center for the Arts and Design . ACCAD MoCap dataset
-
[2]
Ijaz Akhter and Michael J. Black. Pose-conditioned joint angle limits for 3D human pose reconstruction. In IEEE Conf. on Computer Vision and Pattern Recognition ( CVPR ) 2015 , June 2015
work page 2015
-
[3]
Sadegh Aliakbarian, Pashmina Cameron, Federica Bogo, Andrew Fitzgibbon, and Thomas J. Cashman. FLAG : Flow-Based 3D Avatar Generation From Sparse Observations . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , pages 13253--13262, 2022
work page 2022
-
[4]
Rhythm is a Dancer : Music-Driven Motion Synthesis With Global Structure
Andreas Aristidou, Anastasios Yiannakidis, Kfir Aberman, Daniel Cohen-Or , Ariel Shamir, and Yiorgos Chrysanthou. Rhythm is a Dancer : Music-Driven Motion Synthesis With Global Structure . IEEE Transactions on Visualization and Computer Graphics , 29(8):3519--3534, August 2023
work page 2023
- [5]
-
[6]
BoDiffusion : Diffusing Sparse Observations for Full-Body Human Motion Synthesis
Angela Castillo, Maria Escobar, Guillaume Jeanneret, Albert Pumarola, Pablo Arbel \'a ez, Ali Thabet, and Artsiom Sanakoyeu. BoDiffusion : Diffusing Sparse Observations for Full-Body Human Motion Synthesis . In Proceedings of the IEEE / CVF International Conference on Computer Vision , pages 4221--4231, 2023
work page 2023
-
[7]
Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion Models Beat GANs on Image Synthesis . In Advances in Neural Information Processing Systems , volume 34, pages 8780--8794. Curran Associates, Inc., 2021
2021
-
[8]
Andrea Dittadi, Sebastian Dziadzio, Darren Cosker, Ben Lundell, Thomas J. Cashman, and Jamie Shotton. Full- Body Motion From a Single Head-Mounted Device : Generating SMPL Poses From Partial Observations . In Proceedings of the IEEE / CVF International Conference on Computer Vision , pages 11687--11697, 2021
work page 2021
Show all 47 references
-
[9]
Avatars Grow Legs : Generating Smooth Human Motion From Sparse Tracking Inputs With Diffusion Model
Yuming Du, Robin Kips, Albert Pumarola, Sebastian Starke, Ali Thabet, and Artsiom Sanakoyeu. Avatars Grow Legs : Generating Smooth Human Motion From Sparse Tracking Inputs With Diffusion Model . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recogni...
2023
-
[10]
Stratified Avatar Generation from Sparse Observations
Han Feng, Wenchao Ma, Quankai Gao, Xianwei Zheng, Nan Xue, and Huijuan Xu. Stratified Avatar Generation from Sparse Observations . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , pages 153--163, 2024
2024
-
[11]
Recurrent Network Models for Human Dynamics
Katerina Fragkiadaki, Sergey Levine, Panna Felsen, and Jitendra Malik. Recurrent Network Models for Human Dynamics . In Proceedings of the IEEE International Conference on Computer Vision , pages 4346--4354, 2015
2015
-
[12]
GUESS : GradUally Enriching SyntheSis for Text-Driven Human Motion Generation
Xuehao Gao, Yang Yang, Zhenyu Xie, Shaoyi Du, Zhongqian Sun, and Yang Wu. GUESS : GradUally Enriching SyntheSis for Text-Driven Human Motion Generation . IEEE Transactions on Visualization and Computer Graphics , 30(12):7518--7530, December 2024
2024
-
[13]
Saeed Ghorbani, Kimia Mahdaviani, Anne Thaler, Konrad Kording, Douglas James Cook, Gunnar Blohm, and Nikolaus F. Troje. MoVi : A large multipurpose motion and video dataset. arXiv preprint arXiv: 2003.01888 , 2020
2003 arXiv
-
[14]
TM2T : Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts
Chuan Guo, Xinxin Zuo, Sen Wang, and Li Cheng. TM2T : Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts . In Shai Avidan, Gabriel Brostow, Moustapha Ciss \'e , Giovanni Maria Farinella, and Tal Hassner, editors, Computer Vision -- EC...
2022
-
[15]
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Diffusion Probabilistic Models . In Advances in Neural Information Processing Systems , volume 33, pages 6840--6851. Curran Associates, Inc., 2020
2020
-
[16]
Black, Otmar Hilliges, and Gerard Pons-Moll
Yinghao Huang, Manuel Kaufmann, Emre Aksan, Michael J. Black, Otmar Hilliges, and Gerard Pons-Moll . Deep inertial poser: Learning to reconstruct human pose from sparse inertial measurements in real time. ACM Trans. Graph. , 37(6):185:1--185:15, December 2018
2018
-
[17]
Zamir, Silvio Savarese, and Ashutosh Saxena
Ashesh Jain, Amir R. Zamir, Silvio Savarese, and Ashutosh Saxena. Structural- RNN : Deep Learning on Spatio-Temporal Graphs . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 5308--5317, 2016
2016
-
[18]
AvatarPoser : Articulated Full-Body Pose Tracking from Sparse Motion Sensing
Jiaxi Jiang, Paul Streli, Huajian Qiu, Andreas Fender, Larissa Laich, Patrick Snape, and Christian Holz. AvatarPoser : Articulated Full-Body Pose Tracking from Sparse Motion Sensing . In Shai Avidan, Gabriel Brostow, Moustapha Ciss \'e , Giovanni Maria Farinella, and Tal Hassn...
2022
-
[19]
Winkler, and C
Yifeng Jiang, Yuting Ye, Deepak Gopinath, Jungdam Won, Alexander W. Winkler, and C. Karen Liu. Transformer Inertial Poser : Real-time Human Motion Reconstruction from Sparse IMUs with Simultaneous Terrain Generation . In SIGGRAPH Asia 2022 Conference Papers , pages 1--9, Daegu...
2022
-
[20]
Auto- Encoding Variational Bayes
Diederik P Kingma and Max Welling. Auto- Encoding Variational Bayes
-
[21]
Ross, and Angjoo Kanazawa
Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa. AI Choreographer : Music Conditioned 3D Dance Generation With AIST ++. In Proceedings of the IEEE / CVF International Conference on Computer Vision , pages 13401--13412, 2021
2021
-
[22]
DanceFormer : Music Conditioned 3D Dance Generation with Parametric Motion Transformer
Buyu Li, Yongchi Zhao, Shi Zhelun, and Lu Sheng. DanceFormer : Music Conditioned 3D Dance Generation with Parametric Motion Transformer . Proceedings of the AAAI Conference on Artificial Intelligence , 36(2):1272--1279, June 2022
2022
-
[23]
Loper, Naureen Mahmood, and Michael J
Matthew M. Loper, Naureen Mahmood, and Michael J. Black. MoSh : Motion and shape capture from sparse markers. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia) , 33(6):220:1--220:13, November 2014
2014
-
[24]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll , and Michael J. Black. SMPL : A skinned multi-person linear model. ACM Trans. Graph. , 34(6), October 2015
2015
-
[25]
Eyes JAPAN Co. Ltd. Eyes japan MoCap dataset
-
[26]
Troje, Gerard Pons-Moll , and Michael Black
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll , and Michael Black. AMASS : Archive of Motion Capture As Surface Shapes . In 2019 IEEE / CVF International Conference on Computer Vision ( ICCV ) , pages 5441--5450, Seoul, Korea (South), October 2019. IEEE
2019
-
[27]
The KIT whole-body human motion database
Christian Mandery, \"O mer Terlemez, Martin Do, Nikolaus Vahrenkamp, and Tamim Asfour. The KIT whole-body human motion database. In International Conference on Advanced Robotics ( ICAR ) , pages 329--336, 2015
2015
-
[28]
u ller, T. R \
M. M \"u ller, T. R \"o der, M. Clausen, B. Eberhardt, B. Kr \"u ger, and A. Weber. Documentation mocap database HDM05 . Technical Report CG-2007-2, Universit \"a t Bonn, June 2007
2007
-
[29]
Improved Denoising Diffusion Probabilistic Models
Alexander Quinn Nichol and Prafulla Dhariwal. Improved Denoising Diffusion Probabilistic Models . In Proceedings of the 38th International Conference on Machine Learning , pages 8162--8171. PMLR, July 2021
2021
-
[30]
GLIDE : Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models , December 2021
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE : Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models , December 2021
2021
-
[31]
Hierarchical Text-Conditional Image Generation with CLIP Latents , April 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical Text-Conditional Image Generation with CLIP Latents , April 2022
2022
-
[32]
Variational Inference with Normalizing Flows
Danilo Rezende and Shakir Mohamed. Variational Inference with Normalizing Flows . In Proceedings of the 32nd International Conference on Machine Learning , pages 1530--1538. PMLR, June 2015
2015
-
[33]
Final IK , 2018
RootMotion. Final IK , 2018
2018
-
[34]
Sigal, A
L. Sigal, A. Balan, and M. J. Black. HumanEva : Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International Journal of Computer Vision , 87(1):4--27, March 2010
2010
-
[35]
Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Jascha Sohl-Dickstein , Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep Unsupervised Learning using Nonequilibrium Thermodynamics . In Proceedings of the 32nd International Conference on Machine Learning , pages 2256--2265. PMLR, June 2015
2015
-
[36]
DENOISING DIFFUSION IMPLICIT MODELS
Jiaming Song, Chenlin Meng, and Stefano Ermon. DENOISING DIFFUSION IMPLICIT MODELS . 2021
2021
-
[37]
Local motion phases for learning multi-contact character movements
Sebastian Starke, Yiwei Zhao, Taku Komura, and Kazi Zaman. Local motion phases for learning multi-contact character movements. ACM Trans. Graph. , 39(4):54:54:1--54:54:13, August 2020
2020
-
[38]
Nikolaus F. Troje. Decomposing biological motion: A framework for analysis and synthesis of human gait patterns. Journal of Vision , 2(5):2--2, September 2002
2002
-
[39]
Total Capture : 3D human pose estimation fusing video and inertial sensors
Matt Trumble, Andrew Gilbert, Charles Malleson, Adrian Hilton, and John Collomosse. Total Capture : 3D human pose estimation fusing video and inertial sensors. In 2017 British Machine Vision Conference ( BMVC ) , 2017
2017
-
[40]
SFU motion capture database
Simon Fraser University and National University of Singapore . SFU motion capture database
-
[41]
von Marcard , B
T. von Marcard , B. Rosenhahn, M. J. Black, and G. Pons-Moll . Sparse Inertial Poser : Automatic 3D Human Pose Estimation from Sparse IMUs . Computer Graphics Forum , 36(2):349--360, 2017
2017
-
[42]
LoBSTr : Real-time Lower-body Pose Prediction from Sparse Upper-body Tracking Signals
Dongseok Yang, Doyeon Kim, and Sung-Hee Lee. LoBSTr : Real-time Lower-body Pose Prediction from Sparse Upper-body Tracking Signals . Computer Graphics Forum , 40(2):265--275, May 2021
2021
-
[43]
TransPose : Real-time 3D human translation and pose estimation with six inertial sensors
Xinyu Yi, Yuxiao Zhou, and Feng Xu. TransPose : Real-time 3D human translation and pose estimation with six inertial sensors. ACM Trans. Graph. , 40(4), July 2021
2021
-
[44]
Physical Inertial Poser ( PIP ): Physics-Aware Real-Time Human Motion Tracking From Sparse Inertial Sensors
Xinyu Yi, Yuxiao Zhou, Marc Habermann, Soshi Shimada, Vladislav Golyanik, Christian Theobalt, and Feng Xu. Physical Inertial Poser ( PIP ): Physics-Aware Real-Time Human Motion Tracking From Sparse Inertial Sensors . In Proceedings of the IEEE / CVF Conference on Computer Visi...
2022
-
[45]
Realistic Full-Body Tracking from Sparse Observations via Joint-Level Modeling
Xiaozheng Zheng, Zhuo Su, Chao Wen, Zhou Xue, and Xiaojie Jin. Realistic Full-Body Tracking from Sparse Observations via Joint-Level Modeling . In Proceedings of the IEEE / CVF International Conference on Computer Vision , pages 14678--14688, 2023
2023
-
[46]
On the Continuity of Rotation Representations in Neural Networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the Continuity of Rotation Representations in Neural Networks . In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , pages 5745--5753, 2019
2019
-
[47]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence '...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.