SciCUEval provides a multi-domain, multi-modality benchmark for LLM scientific context understanding, and evaluation results show reasoning-augmented models outperform specialized scientific models, but the dataset is not yet released and some numbers are inconsistent.
Given the follow- ing four materials: mp-xxxxx, mp-xxxxx, mp-xxxxx, mp-xxxxx
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
SciCUEval provides a multi-domain, multi-modality benchmark for LLM scientific context understanding, and evaluation results show reasoning-augmented models outperform specialized scientific models, but the dataset is not yet released and some numbers are inconsistent.