Probing layer-wise embeddings with three prompt-variant families reveals a consistent grounding-reasoning-decoding structure in LLaVA-1.5, LLaVA-Next, and Qwen2-VL, with base LLM architecture shifting layer allocation.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
Probing layer-wise embeddings with three prompt-variant families reveals a consistent grounding-reasoning-decoding structure in LLaVA-1.5, LLaVA-Next, and Qwen2-VL, with base LLM architecture shifting layer allocation.