Item response theory applied to 17 LLMs on SciEntsBank and Beetle reveals that models with similar overall scores differ sharply in robustness to difficult responses, with errors clustering on partial-credit labels.
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 3years
2026 3representative citing papers
Text embeddings recover 57-63% of the reliable variance in exam-item difficulty, and apparent differences in predictability across IRT parameters are mostly artifacts of calibration noise rather than text signal.
A difficulty-aware conversational knowledge tracing framework that combines LLMs with Item Response Theory to produce interpretable student performance predictions in tutor dialogues.
citing papers explorer
-
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
Item response theory applied to 17 LLMs on SciEntsBank and Beetle reveals that models with similar overall scores differ sharply in robustness to difficult responses, with errors clustering on partial-credit labels.
-
From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings
Text embeddings recover 57-63% of the reliable variance in exam-item difficulty, and apparent differences in predictability across IRT parameters are mostly artifacts of calibration noise rather than text signal.
-
Interpretable Difficulty-Aware Knowledge Tracing in Tutor-Student Dialogues
A difficulty-aware conversational knowledge tracing framework that combines LLMs with Item Response Theory to produce interpretable student performance predictions in tutor dialogues.