MV-Actor proposes a multi-view framework using semantic interaction, semantic-spatial token interaction, and guided depth repair to reach 87.8% success on the PerAct2 bimanual benchmark and outperform RGB/RGB-D baselines in real-world tests.
Peafowl: Perception-enhanced multi-view vision-language-action for bimanual manip- ulation
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 2years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
GuidedVLA improves VLA generalization by supervising individual attention heads with manually defined auxiliary signals for three task-relevant factors.
citing papers explorer
-
MV-Actor: Aligning Multi-View Semantics and Spatial Awareness for Bimanual Manipulation
MV-Actor proposes a multi-view framework using semantic interaction, semantic-spatial token interaction, and guided depth repair to reach 87.8% success on the PerAct2 bimanual benchmark and outperform RGB/RGB-D baselines in real-world tests.
-
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
GuidedVLA improves VLA generalization by supervising individual attention heads with manually defined auxiliary signals for three task-relevant factors.