A new expert-annotated benchmark of students' handwritten math responses shows current vision language models perform poorly, especially on correctness and error-detection questions.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
A new expert-annotated benchmark of students' handwritten math responses shows current vision language models perform poorly, especially on correctness and error-detection questions.