DyPES-VLA trains a single cross-embodiment policy using future-video prediction to share dynamics and a per-embodiment Mixture-of-Experts head for control, reporting state-of-the-art simulation and real-world success rates.
Unified diffusion vla: Vision-language-action model via joint discrete denosing diffusion process
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
DyPES-VLA trains a single cross-embodiment policy using future-video prediction to share dynamics and a per-embodiment Mixture-of-Experts head for control, reporting state-of-the-art simulation and real-world success rates.