Pith. sign in

ROAM: Robust and Object-Aware Motion Generation Using Neural Pose Descriptors

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Existing automatic approaches for 3D virtual character motion synthesis supporting scene interactions do not generalise well to new objects outside training distributions, even when trained on extensive motion capture datasets with diverse objects and annotated interactions. This paper addresses this limitation and shows that robustness and generalisation to novel scene objects in 3D object-aware character synthesis can be achieved by training a motion model with as few as one reference object. We leverage an implicit feature representation trained on object-only datasets, which encodes an SE(3)-equivariant descriptor field around the object. Given an unseen object and a reference pose-object pair, we optimise for the object-aware pose that is closest in the feature space to the reference pose. Finally, we use l-NSM, i.e., our motion generation model that is trained to seamlessly transition from locomotion to object interaction with the proposed bidirectional pose blending scheme. Through comprehensive numerical comparisons to state-of-the-art methods and in a user study, we demonstrate substantial improvements in 3D virtual character motion and interaction quality and robustness to scenarios with unseen objects. Our project page is available at https://vcai.mpi-inf.mpg.de/projects/ROAM/.

fields

cs.CV 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

GIRAF: Towards Generalizable Human Interactions with Articulated Objects

cs.CV · 2026-07-08 · conditional · novelty 6.0

A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat adapted baselines on contact and pose metrics.

citing papers explorer

Showing 1 of 1 citing paper.

  • GIRAF: Towards Generalizable Human Interactions with Articulated Objects cs.CV · 2026-07-08 · conditional · none · ref 82 · internal anchor

    A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat adapted baselines on contact and pose metrics.