A semantic-similarity weighting module for sub-images improves high-resolution vision-language model performance, pending clarification of whether evaluation benchmarks overlap with training data.
Monkey: Image resolution and text label are important things for large multi-modal models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models
A semantic-similarity weighting module for sub-images improves high-resolution vision-language model performance, pending clarification of whether evaluation benchmarks overlap with training data.