Octopi-1.5 combines a Qwen2-VL 7B language model with a tactile encoder and retrieval-augmented generation to describe, identify, and sort objects from touch alone.
A touch, vision, and language dataset for multimodal alignment
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Demonstrating the Octopi-1.5 Visual-Tactile-Language Model
Octopi-1.5 combines a Qwen2-VL 7B language model with a tactile encoder and retrieval-augmented generation to describe, identify, and sort objects from touch alone.