Text embeddings recover 57-63% of the reliable variance in exam-item difficulty, and apparent differences in predictability across IRT parameters are mostly artifacts of calibration noise rather than text signal.
Predicting the difficulty of multiple choice questions in a high-stakes medical exam , booktitle =
2 Pith papers cite this work, alongside 45 external citations. Polarity classification is still indexing.
2
Pith papers citing it
45
external citations · OpenAlex
fields
cs.CL 2years
2026 2representative citing papers
Fine-tuned transformers with multi-task learning recover substantial wording-derived signal for item difficulty at small sample sizes typical in applied testing.
citing papers explorer
-
From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings
Text embeddings recover 57-63% of the reliable variance in exam-item difficulty, and apparent differences in predictability across IRT parameters are mostly artifacts of calibration noise rather than text signal.
-
Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning
Fine-tuned transformers with multi-task learning recover substantial wording-derived signal for item difficulty at small sample sizes typical in applied testing.