CoRoI injects a chain of language-guided image regions into LLM hidden layers and reports improved MLLM benchmark scores at 7B-34B scale.
https://www-cdn.anthropic.com/ de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Visual Instruction Tuning with Chain of Region-of-Interest
CoRoI injects a chain of language-guided image regions into LLM hidden layers and reports improved MLLM benchmark scores at 7B-34B scale.