Both visual interfaces, textualized diagrams and image-native VLM input, beat text-only difficulty prediction on point estimates, but their difference is not statistically reliable, and image-native gains depend on broadening LoRA adaptation.
Qwen-VL and PaliGemma use their packaged image preprocessing; InternVL uses a448× 448image transform and the model’s image-context tokens
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling
Both visual interfaces, textualized diagrams and image-native VLM input, beat text-only difficulty prediction on point estimates, but their difference is not statistically reliable, and image-native gains depend on broadening LoRA adaptation.