A full-manuscript CS figure benchmark (6,308 expert-rated images) and staged SFQ-Agent judge beat direct VLM scoring, reaching 0.418 MAE and 93.4% within-1 on eval1200.
Journal of Natural Language Processing , volume =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
A full-manuscript CS figure benchmark (6,308 expert-rated images) and staged SFQ-Agent judge beat direct VLM scoring, reaching 0.418 MAE and 93.4% within-1 on eval1200.