CARVE contrasts attention maps from a general prompt and a specific question to mask out visual noise, then re-asks the question on the cropped and enlarged image, improving VQA accuracy by up to 75% on some benchmarks.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
CARVE contrasts attention maps from a general prompt and a specific question to mask out visual noise, then re-asks the question on the cropped and enlarged image, improving VQA accuracy by up to 75% on some benchmarks.