An end-to-end audio-conditioned model generates 3D human motion from spoken instructions, with performance close to text-conditioned baselines on synthetic speech datasets.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Semantics-Aware Human Motion Generation from Audio Instructions
An end-to-end audio-conditioned model generates 3D human motion from spoken instructions, with performance close to text-conditioned baselines on synthetic speech datasets.