Vimorag: Video-based retrieval-augmented 3d motion generation for motion language models

Haidong Xu, Guangwei Xu, Zhedong Zheng, Xiatian Zhu, Wei Ji, Xiangtai Li, Ruijie Guo, Meishan Zhang, Hao Fei, et al · 2025 · arXiv 2508.12081

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

read on arXiv browse 2 citing papers

citation-role summary

background 2

citation-polarity summary

background 2

representative citing papers

Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization

cs.CV · 2026-05-08 · unverdicted · novelty 7.0

Retrieval from motion datasets combined with LLM task parsing and reward-guided noise initialization enables training-free diffusion optimization to satisfy severe spatiotemporal constraints in human motion generation.

SentiAvatar: Towards Expressive and Interactive Digital Humans

cs.CV · 2026-04-03 · unverdicted · novelty 7.0

SentiAvatar generates expressive interactive 3D avatars in real time by combining a 37-hour mocap dialogue dataset with a pre-trained motion foundation model and an audio-aware plan-then-infill architecture that separates semantic planning from prosody-driven frame interpolation.

citing papers explorer

Showing 2 of 2 citing papers.

Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization cs.CV · 2026-05-08 · unverdicted · none · ref 46
Retrieval from motion datasets combined with LLM task parsing and reward-guided noise initialization enables training-free diffusion optimization to satisfy severe spatiotemporal constraints in human motion generation.
SentiAvatar: Towards Expressive and Interactive Digital Humans cs.CV · 2026-04-03 · unverdicted · none · ref 62
SentiAvatar generates expressive interactive 3D avatars in real time by combining a 37-hour mocap dialogue dataset with a pre-trained motion foundation model and an audio-aware plan-then-infill architecture that separates semantic planning from prosody-driven frame interpolation.

Vimorag: Video-based retrieval-augmented 3d motion generation for motion language models

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer