Pith. sign in

REVIEW 2 cited by

HiT-DVAE: Human Motion Generation via Hierarchical Transformer Dynamical VAE

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.01565 v1 pith:E5S4L2EH submitted 2022-04-04 cs.CV

classification cs.CV
keywords generationhumanhit-dvaelatentmethodsmotionspaceattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Studies on the automatic processing of 3D human pose data have flourished in the recent past. In this paper, we are interested in the generation of plausible and diverse future human poses following an observed 3D pose sequence. Current methods address this problem by injecting random variables from a single latent space into a deterministic motion prediction framework, which precludes the inherent multi-modality in human motion generation. In addition, previous works rarely explore the use of attention to select which frames are to be used to inform the generation process up to our knowledge. To overcome these limitations, we propose Hierarchical Transformer Dynamical Variational Autoencoder, HiT-DVAE, which implements auto-regressive generation with transformer-like attention mechanisms. HiT-DVAE simultaneously learns the evolution of data and latent space distribution with time correlated probabilistic dependencies, thus enabling the generative model to learn a more complex and time-varying latent space as well as diverse and realistic human motions. Furthermore, the auto-regressive generation brings more flexibility on observation and prediction, i.e. one can have any length of observation and predict arbitrary large sequences of poses with a single pre-trained model. We evaluate the proposed method on HumanEva-I and Human3.6M with various evaluation methods, and outperform the state-of-the-art methods on most of the metrics.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EquiFusion: Kinematics-Agnostic Human Motion Prediction via Equivariant Latent Diffusion

    cs.CV 2026-07 accept novelty 7.5 of 10

    A permutation-equivariant latent diffusion model treats skeleton connectivity as input, enabling the first kinematics-agnostic stochastic human motion predictor that generalizes zero-shot to unseen and partial skeletons.

  2. Robust Monitoring of Arc Welding Processes: A Generalizable Framework with DVAE and Particle Filter

    eess.IV 2026-07 conditional novelty 6.0 of 10

    A DVAE-PF framework learns a low-dimensional latent state from weld pool images and fuses it with process dynamics to monitor weld penetration, tested on GTAW and GMAW.

Pith tools