A one-stage Transformer with placeholder tokens generates continuous 2D pose sequences from a single image and text, avoiding autoregressive error accumulation.
Long-term Human Motion Prediction with Scene Context
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Human movement is goal-directed and influenced by the spatial layout of the objects in the scene. To plan future human motion, it is crucial to perceive the environment -- imagine how hard it is to navigate a new room with lights off. Existing works on predicting human motion do not pay attention to the scene context and thus struggle in long-term prediction. In this work, we propose a novel three-stage framework that exploits scene context to tackle this task. Given a single scene image and 2D pose histories, our method first samples multiple human motion goals, then plans 3D human paths towards each goal, and finally predicts 3D human pose sequences following each path. For stable training and rigorous evaluation, we contribute a diverse synthetic dataset with clean annotations. In both synthetic and real datasets, our method shows consistent quantitative and qualitative improvements over existing methods.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Towards Consistent Long-Term Pose Generation
A one-stage Transformer with placeholder tokens generates continuous 2D pose sequences from a single image and text, avoiding autoregressive error accumulation.