DeepSeek-R1 achieves 75.9% zero-shot and 81.3% few-shot accuracy on filtered SciBench physics problems, far above general-purpose chat models, and correct answers tend to use symbolic derivation.
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Item difficulty plays a crucial role in adaptive testing. However, few works have focused on generating questions of varying difficulty levels, especially for multiple-choice (MC) cloze tests. We propose training pre-trained language models (PLMs) as surrogate models to enable item response theory (IRT) assessment, avoiding the need for human test subjects. We also propose two strategies to control the difficulty levels of both the gaps and the distractors using ranking rules to reduce invalid distractors. Experimentation on a benchmark dataset demonstrates that our proposed framework and methods can effectively control and evaluate the difficulty levels of MC cloze tests.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
DeepSeek-R1 achieves 75.9% zero-shot and 81.3% few-shot accuracy on filtered SciBench physics problems, far above general-purpose chat models, and correct answers tend to use symbolic derivation.