Llama 3.2 Vision reaches 83.3% on e-SNLI-VE after fine-tuning, but high explanation scores persist with black images, showing VE accuracy and BERTScore are weak evidence of visual grounding.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
Llama 3.2 Vision reaches 83.3% on e-SNLI-VE after fine-tuning, but high explanation scores persist with black images, showing VE accuracy and BERTScore are weak evidence of visual grounding.