A benchmark of 15 multimodal LLMs on grid path planning reports modest success on 8x8 grids and near-failure on 20x20 grids, but its visual-vs-text comparison is confounded by prompt differences.
V oice con- trol interface for surgical robot assistants
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning
A benchmark of 15 multimodal LLMs on grid path planning reports modest success on 8x8 grids and near-failure on 20x20 grids, but its visual-vs-text comparison is confounded by prompt differences.