GF-VLA builds temporally ordered interaction graphs from a human demo video and converts them, via a vision-language-action model, into dual-arm assembly policies that generalize to novel arrangements.
HYPERmotion: Learning hybrid behavior planning for autonomous loco-manipulation,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Graph-Fused Vision-Language-Action for Policy Reasoning in Multi-Arm Robotic Manipulation
GF-VLA builds temporally ordered interaction graphs from a human demo video and converts them, via a vision-language-action model, into dual-arm assembly policies that generalize to novel arrangements.