Attention weights in a cross-modal transformer trained on furniture assembly cluster into distinct patterns for different manipulation primitives, suggesting the possibility of unsupervised skill segmentation, but the paper does not implement or evaluate such segmentation.
Language models are few-shot learners,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Modality Selection and Skill Segmentation via Cross-Modality Attention
Attention weights in a cross-modal transformer trained on furniture assembly cluster into distinct patterns for different manipulation primitives, suggesting the possibility of unsupervised skill segmentation, but the paper does not implement or evaluate such segmentation.