Pith. sign in

Long-term Human Motion Prediction with Scene Context

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Human movement is goal-directed and influenced by the spatial layout of the objects in the scene. To plan future human motion, it is crucial to perceive the environment -- imagine how hard it is to navigate a new room with lights off. Existing works on predicting human motion do not pay attention to the scene context and thus struggle in long-term prediction. In this work, we propose a novel three-stage framework that exploits scene context to tackle this task. Given a single scene image and 2D pose histories, our method first samples multiple human motion goals, then plans 3D human paths towards each goal, and finally predicts 3D human pose sequences following each path. For stable training and rigorous evaluation, we contribute a diverse synthetic dataset with clean annotations. In both synthetic and real datasets, our method shows consistent quantitative and qualitative improvements over existing methods.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Towards Consistent Long-Term Pose Generation

cs.CV · 2025-07-24 · conditional · novelty 5.0

A one-stage Transformer with placeholder tokens generates continuous 2D pose sequences from a single image and text, avoiding autoregressive error accumulation.

citing papers explorer

Showing 1 of 1 citing paper.

  • Towards Consistent Long-Term Pose Generation cs.CV · 2025-07-24 · conditional · none · ref 5 · internal anchor

    A one-stage Transformer with placeholder tokens generates continuous 2D pose sequences from a single image and text, avoiding autoregressive error accumulation.