Pith. sign in

REVIEW 1 cited by

Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01456 v1 pith:XGS2MSII submitted 2024-03-03 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords difficultyitemlevelsmodelsassessmentclozecontroldistractors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Item difficulty plays a crucial role in adaptive testing. However, few works have focused on generating questions of varying difficulty levels, especially for multiple-choice (MC) cloze tests. We propose training pre-trained language models (PLMs) as surrogate models to enable item response theory (IRT) assessment, avoiding the need for human test subjects. We also propose two strategies to control the difficulty levels of both the gaps and the distractors using ranking rules to reduce invalid distractors. Experimentation on a benchmark dataset demonstrates that our proposed framework and methods can effectively control and evaluate the difficulty levels of MC cloze tests.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs

    cs.CL 2025-07 conditional novelty 5.0 of 10

    DeepSeek-R1 achieves 75.9% zero-shot and 81.3% few-shot accuracy on filtered SciBench physics problems, far above general-purpose chat models, and correct answers tend to use symbolic derivation.

Pith tools