A frozen MLLM with LoRA-on-language and a discrete token waypoint interface navigates four unseen real-world environments after just 8.7 hours of training data.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.RO 2years
2026 2representative citing papers
TAP-VLA improves VLA performance in contact-rich manipulation by visually annotating tactile shear fields onto input images, reaching 78% success versus under 50% for vision-only and other tactile methods.
citing papers explorer
-
GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
A frozen MLLM with LoRA-on-language and a discrete token waypoint interface navigates four unseen real-world environments after just 8.7 hours of training data.
-
TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models
TAP-VLA improves VLA performance in contact-rich manipulation by visually annotating tactile shear fields onto input images, reaching 78% success versus under 50% for vision-only and other tactile methods.