A vision-language-action robot policy that embeds text and image features in hyperbolic space with a soft expert-routing module reports higher LIBERO success than Dita and other baselines.
Specifically, we utilized four datasets: Spatial, Object, Goal, and LONG
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models
A vision-language-action robot policy that embeds text and image features in hyperbolic space with a soft expert-routing module reports higher LIBERO success than Dita and other baselines.