REVIEW 2 major objections 6 minor 1 cited by
Item response theory for AI evaluation is only trustworthy with at least 100 models and a balanced capability distribution, simulations on six LLM benchmarks show.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 23:53 UTC pith:3EQ7MLOC
load-bearing objection Useful, field-relevant simulation study with a credible sample-size finding and an overclaimed, confounded skewness result; worth refereeing after revision. the 2 major comments →
Can We Trust Item Response Theory for AI Evaluation?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central finding is that data-regime mismatch, not the IRT model itself, is what breaks IRT-based AI evaluation. At N=30 evaluated models, difficulty and discrimination recovery fell below 0.50–0.60 across all four estimators, with item-level error much higher than at N=100 or above. Ranking recovery, measured by Kendall's tau, stayed above 0.85 for near-symmetric capability distributions (absolute skew under 0.5) but dropped below 0.60 for the most skewed conditions (absolute skew over 2.0), with estimator choice playing a secondary role. Aggregated full-benchmark scores were comparatively robust, and IRT-based selection of a 10% short benchmark consistently beat random item selection in
What carries the argument
The argument rests on a simulation study calibrated to real AI benchmark data. True item parameters (difficulty, discrimination, guessing) and true model-ability distributions are obtained by fitting variational inference IRT models to binary response matrices from six OpenLLM benchmarks; these fitted values serve as the ground truth. Simulated response matrices are then generated under 1PL, 2PL, and 3PL models at N=30, 100, 180, 400, and 1000, and recovered parameters from four estimators—MML-EM, MCMC, variational inference, and a neural pseudo-Siamese estimator—are compared against that ground truth on feasibility, parameter recovery, aggregate score error, and ranking recovery.
Load-bearing premise
The load-bearing assumption is that the variational-inference fitted values from the OpenLLM response matrices are the true item parameter and ability values; if those fitted values are biased in the small-sample or skewed regimes, all recovery metrics measure error against a biased target.
What would settle it
Recompute the same six benchmarks' item parameters with a full-sample MCMC or with a set of anchor items whose difficulties are known from independent human or expert labels, then rerun the N=30 recovery analysis against that truth; if recovery looks much better than the paper's thresholds, the reported N≥100 and skewness rules are artifacts of the VI ground truth.
If this is right
- Item-level claims (difficulty, discrimination, item quality) from fewer than 100 evaluated models should not be trusted; the simulations show recovery collapses at N=30 and improves clearly at N≥100.
- Ranking claims require checking the shape of the model-capability distribution: absolute skew above 0.5 visibly degrades Kendall recovery, and above 2.0 it can fall below 0.6 for every estimator tested.
- Classical estimators are not a safe default: MML-EM failed in about 69% of runs and could not handle benchmarks with more than about 5,000 items, while MCMC exceeded 72-hour limits in many large conditions.
- Scalable estimators are not uniformly safe either: variational inference was fast but gave unreliable difficulty recovery on several benchmarks at small N, and the neural estimator gave no uncertainty quantification.
- Despite these problems, IRT-based short-form benchmark compression outperformed random item subsets across conditions, suggesting that item selection for cheaper evaluation is a comparatively robust use.
Where Pith is reading between the lines
- The ground-truth parameters are themselves VI estimates. If VI is biased in exactly the small-N or skewed regimes studied, the reported thresholds would be distorted; an independent anchor-item or alternative-estimator check on the same benchmarks would settle this.
- A natural extension is to turn the paper's advice into a routine diagnostic: benchmark reports using IRT could publish N, absolute skewness of estimated ability, estimator, and convergence status alongside item and ranking results.
- Because the paper only varied sample size and distribution shape, real evaluations that combine small N with long benchmarks and 3PL models are the most dangerous corner; those combinations should be treated as untested until further simulations.
- Short-form compression's insensitivity to skewness and sample size, noted but unexplained in the paper, suggests a testable hypothesis: item selection needs only coarse item ordering, which degrades more slowly than exact parameter recovery.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper asks whether item response theory (IRT) can be trusted when applied to AI benchmark data, where the number of evaluated models (N) is often small (tens to low hundreds) and the number of items (J) is large (hundreds to tens of thousands), and where model capability distributions may be skewed, clustered, or multimodal. The authors first present a scoping review of 19 studies that apply IRT to AI evaluation, documenting heterogeneous estimation choices and data regimes. They then construct a simulation study: using six OpenLLM benchmark response matrices, they fit 1PL/2PL/3PL models with variational inference (VI) to obtain 'true' item parameters and ability distributions; sample N ∈ {30,100,180,400,1000} abilities from these distributions; generate binary response matrices (50 replications per condition, 90 conditions); and apply four estimators—MML-EM, MCMC, VI, and a neural pseudo-Siamese estimator (PSN). They evaluate computational feasibility (failure rates, runtime), ranking recovery (Kendall's tau), aggregate score error, item parameter recovery, and short-form ranking recovery. The main results are that MML-EM and MCMC are frequently infeasible on large benchmarks or samples; VI has a 10.71% failure rate concentrated in 3PL; PSN never fails; item recovery is poor at N=30 and improves markedly at N≥100; and ranking recovery degrades with the absolute skewness of the capability distribution. The paper concludes with recommendations on estimator choice, minimum
Significance. If the central claims hold, this is a valuable and timely contribution: it is the first systematic simulation study of IRT estimation under AI benchmarking regimes, and it offers concrete, actionable thresholds (N≥100 for item-level analysis; caution when |skewness|>2 for rankings). The computational feasibility findings (MML-EM 69.45% failure rate, MCMC timeouts, VI 3PL failures, PSN 0% failure) are crisp and useful, and the study design follows the ADEMP framework with 50 replications and reported variability. The item-recovery result at N=30 vs N≥100 appears robust. However, two load-bearing issues weaken the paper as it stands: the ground truth is generated by VI, one of the estimators being evaluated, without validation; and the skewness–ranking conclusion is confounded with benchmark and IRT-model factors. These issues must be addressed before the thresholds and recommendations can be fully credited.
major comments (2)
- [Sec. 4.1] The simulation's 'true' item and ability parameters are the outputs of a variational-inference fit (py-irt) to the full OpenLLM response matrices, and VI is itself one of the four estimators evaluated in Section 4.2. The paper does not validate these full-data VI estimates against any independent criterion, alternative estimator, or known-parameter simulation. If VI is biased in exactly the regimes studied (small N, skewed/multimodal ability, long item pools), all recovery metrics in Sections 5.2-5.4 are measured relative to a potentially biased target, and the recommended thresholds (N>=100, |gamma|<2) could be distorted. I recommend adding a validation study: e.g., generate data from known parameters with the same VI implementation and check full-data recovery; or compare VI full-data estimates to MCMC/MML on a tractable subset; or perform sensitivity analyses with ground truth generat
- [Sec. 5.2, Fig. 3] The claim that skewness is 'the dominant factor' for ranking recovery rests on 18 condition-level points (6 benchmarks x 3 IRT models) pooled over sample sizes; skewness is an observed property of the VI-fit distributions, not a manipulated factor. The points differ simultaneously in J (644-12,508), item parameter distributions, saturation filters, and IRT model. No regression with controls, matched design, or within-benchmark manipulation of skewness is provided, and pooling across N may hide sample-size effects. The paper motivates multimodality but analyzes only absolute skewness. The abstract's 'non-normally distributed model sets' conclusion is thus under-identified. I recommend reporting N-separated results, adding ability-distribution transformations within a benchmark, or fitting a regression with J and benchmark as controls.
minor comments (6)
- [Abstract and Sec. 7] '18,000 simulation conditions' is incorrect; there are 90 conditions with 50 replications each, and 4 estimators applied to each, yielding 18,000 estimation runs.
- [Sec. 5.3] The sentence 'PSN, in general, showed more reliable parameter recovery than MML-EM and PSN' appears to contain a typo; it should likely read 'than MML-EM and VI'.
- [Fig. 15] The caption says 'Item parameter recovery under 1PL' but describes guessing parameter c error, which only appears in 3PL; the caption should say 'under 3PL'.
- [Sec. 5.4] 'Supposing' should be 'suggesting' in the sentence about coarse discrimination estimates.
- [Fig. 2] The runtime comparison is based on mixed hardware (CPU for MML-EM, MCMC, VI; GPU for PSN). This is stated in the text but should also be made explicit in the figure caption.
- [Sec. 4] No code or data release is mentioned. For a large simulation study, providing code would substantially improve reproducibility.
Circularity Check
Simulation ground truth is produced by VI, one of the estimators under test, making key recovery comparisons partially self-referential.
specific steps
-
fitted input called prediction
[Sec. 4.1 (Data generation) and Sec. 4.2 (Estimation)]
"For each benchmark, we fit the 1PL, 2PL and 3PL models to the empirical response matrix using variational inference ... The resulting item parameter estimates were treated as the true item parameters in the simulation. ... In each replication, N true capability values were sampled from the empirical distribution of the estimated abilities ... We evaluated four IRT estimatiors ... MML-EM, MCMC, VI, and PSN."
The 'true' item parameters and the empirical ability distribution used to generate every simulated response matrix are VI estimates, and VI is one of the four estimators whose recovery is measured. Item-difficulty recovery, item-mean error, and ranking recovery therefore measure how well the estimators (including VI itself) reproduce VI's full-data solution rather than an externally validated ground truth. If VI is biased in the small-N or non-normal regimes the paper studies, those biases are inherited by the simulation target, so the headline thresholds (N>=100; skewness degrades rankings) are not statistically independent of the estimator under test. The coupling is partial because simulated responses are regenerated and recovery is not algebraically forced to 1, but the paper never val
full rationale
The paper is a large empirical simulation study, not a derivation from its own equations, and most results (runtime, failure rates, short-form vs random baseline, MCMC vs MML behavior) are self-contained measurements. The main circularity concern is the single load-bearing design choice in Sec. 4.1: the ground-truth parameters are obtained by fitting VI to OpenLLM matrices and are then used as the truth against which VI and the other estimators are scored. This is a genuine partial self-reference, since the evaluation target is an output of one of the methods under test, and the abstract's central claims about item-level and ranking unreliability inherit that coupling. However, the recovery values are computed from newly simulated response matrices, so they are not tautologically equal to the input; the paper's N=30 vs N>=100 item-recovery finding and its computational-feasibility findings would still be meaningful even with an independent ground truth. I find no load-bearing self-citation, no imported uniqueness theorem, no ansatz smuggled in by citation, and no renaming of a known result. The skewness-versus-ranking analysis is confounded with benchmark and IRT-model factors, but that is a validity concern rather than a circularity. Score 4 reflects one substantial partial self-reference without the central claim reducing to its input by construction.
Axiom & Free-Parameter Ledger
free parameters (5)
- True item parameters (a_j, b_j, c_j) per benchmark and IRT model =
Not reported numerically; VI estimates from OpenLLM matrices
- Empirical ability distributions per benchmark and IRT model =
Not reported numerically; histograms in App. B
- Data preprocessing cutoffs =
bottom 0.1% models removed; items with SD<1% or mean>95% excluded
- MCMC prior hyperparameters =
θ∼N(0,1) for 2PL/3PL; σ_b∼N+(0,3); log a_j priors per Table 1
- PSN architecture and training hyperparameters =
Not reported
axioms (4)
- domain assumption Unidimensional 1PL/2PL/3PL models are the correct generative models for AI benchmark response matrices.
- standard math Responses to different items are conditionally independent given ability θ.
- ad hoc to paper VI estimates from the full OpenLLM matrices are close enough to the unknown true parameters to serve as simulation ground truth.
- domain assumption The six OpenLLM benchmarks and their preprocessing are representative of conditions under which IRT is used for AI evaluation.
read the original abstract
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or non-normally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.
Figures
Forward citations
Cited by 1 Pith paper
-
BayesAME: Bayesian Active Model Evaluation
A sequential Bayesian method automatically grows a coreset until performance estimate and uncertainty stabilize, outperforming adapted baselines and showing active selection beats random when reference signals are rich.
Reference graph
Works this paper leans on
-
[1]
J. H. Albert. Bayesian estimation of normal ogive item response curves using gibbs sampling. Journal of Educational Statistics, 17(3):251–269, 1992
1992
-
[2]
J. H. Albert and S. Chib. Bayesian analysis of binary and polychotomous response data.Journal of the American Statistical Association, 88(422):669–679, 1993
1993
-
[3]
M. A. Barton and F. M. Lord. An upper asymptote for the three-parameter logistic item-response model.ETS Research Report Series, 1981(1):i–8, 1981
1981
-
[4]
R. D. Bock and M. Aitkin. Marginal maximum likelihood estimation of item parameters: Application of an EM algorithm.Psychometrika, 46(4):443–459, 1981
1981
-
[5]
P.-C. Bürkner. Bayesian item response modeling in r with brms and stan.Journal of statistical software, 100:1–54, 2021
2021
-
[6]
Byrd and S
M. Byrd and S. Srivastava. Predicting difficulty and discrimination of natural language questions. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 119–130, Stroudsburg, PA, USA, 2022. Association for Computational Linguistics
2022
-
[7]
P.-C. Bürkner. brms: An r package for bayesian multilevel models using stan.Journal of Statistical Software, 80(1):1–28, 2017
2017
-
[8]
R. P. Chalmers. mirt: A multidimensional item response theory package for the r environment. Journal of Statistical Software, 48(6):1–29, 2012
2012
-
[9]
J. Chen, C. Wang, G. Zhang, P. Ye, L. Bai, W. Hu, Y . Qu, and S. Hu. Learning compact representations of LLM abilities via item response theory.arXiv [cs.AI], Oct. 2025
2025
-
[10]
Clark, I
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
2018
-
[11]
Cobbe, V
K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman. Training verifiers to solve math word problems, 2021
2021
-
[12]
R. J. De Ayala.The theory and practice of item response theory. Guilford Publications, 2013
2013
-
[13]
DeMars.Item Response Theory
C. DeMars.Item Response Theory. Understanding Statistics: Measurement. Oxford University Press, New York, NY , 2010
2010
-
[14]
˙I. E. Deveci and D. Ataman. The ouroboros of benchmarking: Reasoning evaluation in an era of saturation. InNeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025
2025
-
[15]
Ethayarajh and D
K. Ethayarajh and D. Jurafsky. Utility is in the eye of the user: A critique of NLP leaderboards. In B. Webber, T. Cohn, Y . He, and Y . Liu, editors,Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4846–4853, Online, Nov
2020
-
[16]
Frick, A
S. Frick, A. Krivosija, and A. Munteanu. Scalable learning of item response theory models. In S. Dasgupta, S. Mandt, and Y . Li, editors,Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, volume 238 ofProceedings of Machine Learning Research, pages 1234–1242. PMLR, 02–04 May 2024
2024
-
[17]
A. Gill, A. Ravichander, and A. Marasovic. What has been lost with synthetic evaluation? In C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . Peng, editors,Findings of the Association for Computational Linguistics: EMNLP 2025, pages 9902–9945, Suzhou, China, Nov. 2025. Association for Computational Linguistics
2025
-
[18]
Gulliksen.Theory of Mental Tests
H. Gulliksen.Theory of Mental Tests. Wiley Publications in Psychology. John Wiley & Sons, Hoboken, NJ, 1950
1950
-
[19]
R. K. Hambleton and R. W. Jones. Comparison of classical test theory and item response theory and their applications to test development.Educational Measurement: Issues and Practice, 12(3):38–47, 1993
1993
-
[20]
R. K. Hambleton, H. Swaminathan, and H. J. Rogers.Fundamentals of Item Response Theory, volume 2 ofMeasurement Methods for the Social Sciences. Sage Publications, Thousand Oaks, CA, 1991
1991
-
[21]
Heineman, V
D. Heineman, V . Hofmann, I. Magnusson, Y . Gu, N. A. Smith, H. Hajishirzi, K. Lo, and J. Dodge. Signal and noise: A framework for reducing uncertainty in language model evaluation. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026
2026
-
[22]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021
2021
-
[23]
Hofmann, D
V . Hofmann, D. Heineman, I. Magnusson, K. Lo, J. Dodge, M. Sap, P. W. Koh, C. Wang, H. Hajishirzi, and N. A. Smith. Fluid language model benchmarking. InSecond Conference on Language Modeling, 2025
2025
-
[24]
Jiang, X
H. Jiang, X. Yi, Z. Wei, Z. Xiao, S. Wang, and X. Xie. Raising the bar: Investigating the values of large language models via generative evolving testing. InForty-second International Conference on Machine Learning, June 2025
2025
-
[25]
Jiang, S
H. Jiang, S. Zhang, X. Yi, X. Xie, and Z. Xiao. Position: Science of ai evaluation requires item-level benchmark data, 2026
2026
-
[26]
Jiang, C
S. Jiang, C. Wang, and D. J. Weiss. Sample size requirements for estimation of item parameters in the multidimensional graded response model.Frontiers in Psychology, V olume 7 - 2016, 2016
2016
-
[27]
D. N. Joanes and C. A. Gill. Comparing measures of sample skewness and kurtosis.Journal of the Royal Statistical Society. Series D (The Statistician), 47(1):183–189, 1998
1998
-
[28]
Kipnis, K
A. Kipnis, K. V oudouris, L. M. Schulze Buschoff, and E. Schulz. metabench - a sparse bench- mark of reasoning and knowledge in large language models. InThe Thirteenth International Conference on Learning Representations, Oct. 2024
2024
-
[29]
Kirisci, T.-c
L. Kirisci, T.-c. Hsu, and L. Yu. Robustness of item parameter estimation programs to assump- tions of unidimensionality and normality.Applied Psychological Measurement, 25(2):146–162, 2001
2001
-
[30]
J. P. Lalor and P. Rodriguez. py-irt: A scalable item response theory library for python. INFORMS J. on Computing, 35(1):5–13, Jan. 2023
2023
-
[31]
J. P. Lalor, H. Wu, T. Munkhdalai, and H. Yu. Understanding deep learning performance through an examination of test set difficulty: A psychometric case study.Proc. Conf. Empir. Methods Nat. Lang. Process., 2018:4711–4716, Oct. 2018. 11
2018
-
[32]
J. P. Lalor, H. Wu, and H. Yu. Building an evaluation scale using item response theory. In J. Su, K. Duh, and X. Carreras, editors,Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 648–657, Austin, Texas, Nov. 2016. Association for Computational Linguistics
2016
-
[33]
J. P. Lalor, H. Wu, and H. Yu. Learning latent parameters without human response patterns: Item response theory with artificial crowds.Proc. Conf. Empir. Methods Nat. Lang. Process., 2019:4240–4250, Nov. 2019
2019
-
[34]
Liang, R
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y . Zhang, D. Narayanan, Y . Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Re, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. WANG, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha...
2023
-
[35]
L. Liao, Q. Zhang, R. Wu, and G. Fang. Toward a unified framework for data-efficient evaluation of large language models.arXiv [cs.AI], Oct. 2025
2025
-
[36]
S. Lin, J. Hilton, and O. Evans. TruthfulQA: Measuring how models mimic human falsehoods. In S. Muresan, P. Nakov, and A. Villavicencio, editors,Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics
2022
-
[37]
F. M. Lord.Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum Associates, Hillsdale, NJ, 1980
1980
-
[38]
F. M. Lord and M. R. Novick.Statistical Theories of Mental Test Scores. Addison-Wesley, Reading, MA, 1968
1968
-
[39]
F. M. Lord, M. R. Novick, and A. Birnbaum. Some latent train models and their use in inferring an examinee’s ability. 1966
1966
-
[40]
Martínez-Plumed, R
F. Martínez-Plumed, R. B. C. Prudêncio, A. Martínez-Usó, and J. Hernández-Orallo. Item response theory in AI: Analysing machine learning classifiers at the instance level.Artif. Intell., 271:18–42, June 2019
2019
-
[41]
T. P. Morris, I. R. White, and M. J. Crowther. Using simulation studies to evaluate statistical methods.Statistics in Medicine, 38(11):2074–2102, 2019
2074
-
[42]
Myrzakhan, S
A. Myrzakhan, S. M. Bsharat, and Z. Shen. Open-llm-leaderboard: From multi-choice to open-style questions for llms evaluation, benchmark, and arena, 2024
2024
-
[43]
S. Ott, A. Barbosa-Silva, K. Blagec, J. Brauner, and M. Samwald. Mapping global dynamics of benchmark creation and saturation in artificial intelligence.Nature Communications, 13(1):6793, 2022
2022
-
[44]
F. M. Polo, L. Weber, L. Choshen, Y . Sun, G. Xu, and M. Yurochkin. tinyBenchmarks: evaluating LLMs with fewer examples. InInternational Conference on Machine Learning, pages 34303–34326. PMLR, July 2024
2024
-
[45]
Rasch.Probabilistic Models for Some Intelligence and Attainment Tests
G. Rasch.Probabilistic Models for Some Intelligence and Attainment Tests. Nielsen & Lydiche, Copenhagen, Denmark, 1960
1960
-
[46]
Robertson
Z. Robertson. Identity-link IRT for label-free LLM evaluation: Preserving additivity in TVD-MI scores.arXiv [cs.LG], Oct. 2025
2025
-
[47]
Rodriguez, J
P. Rodriguez, J. Barrow, A. Hoyle, J. P. Lalor, R. Jia, and J. Boyd-Graber. Evaluation examples are not equally informative: How should that change NLP leaderboards? In C. Zong, F. Xia, W. Li, and R. Navigli, editors,Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natur...
2021
-
[48]
Sakaguchi, R
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y . Choi. Winogrande: an adversarial winograd schema challenge at scale.Commun. ACM, 64(9):99–106, Aug. 2021
2021
-
[49]
Samejima
F. Samejima. Estimation of latent ability using a response pattern of graded scores.Psychome- trika, 34(4, Pt. 2):1–97, 1969. Psychometrika Monograph Supplement No. 17
1969
-
[50]
Schilling-Wilhelmi, N
M. Schilling-Wilhelmi, N. Alampara, and K. M. Jablonka. Lifting the benchmark iceberg with item-response theory. InAI for Accelerated Materials Design - ICLR 2025, 2025
2025
-
[51]
Schroeders and T
U. Schroeders and T. Gnambs. Sample-size planning in item-response theory: A tutorial. Advances in Methods and Practices in Psychological Science, 8(1), 2025
2025
-
[52]
Sen and A
S. Sen and A. S. Cohen. The impact of sample size and various other factors on estimation of dichotomous mixture IRT models.Educ. Psychol. Meas., 83(3):520–555, June 2023
2023
-
[53]
B. S. Siepe, F. Bartoš, T. P. Morris, A.-L. Boulesteix, D. W. Heck, and S. Pawel. Simulation studies for methodological research in psychology: A standardized template for planning, preregistration, and reporting.Psychological Methods, 2024
2024
-
[54]
Siska, K
C. Siska, K. Marazopoulou, M. Ailem, and J. Bono. Examining the robustness of LLM evalua- tion to the distributional assumptions of benchmarks. In L.-W. Ku, A. Martins, and V . Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Association for Computational Linguis- tics (Volume 1: Long Papers), pages 10406–10421, Bangkok, Thailand, Aug. 2024...
2024
-
[55]
C. A. Stone. Recovery of marginal maximum likelihood estimates in the two-parameter logistic response model: An evaluation of MULTILOG.Applied Psychological Measurement, 16(1):1– 16, 1992
1992
-
[56]
S. T. Truong, Y . Tu, P. Liang, B. Li, and S. Koyejo. Reliable and efficient amortized model-based evaluation. InForty-second International Conference on Machine Learning, June 2025
2025
-
[57]
Vania, P
C. Vania, P. M. Htut, W. Huang, D. Mungra, R. Y . Pang, J. Phang, H. Liu, K. Cho, and S. R. Bowman. Comparing test sets with item response theory. In C. Zong, F. Xia, W. Li, and R. Navigli, editors,Proceedings of the 59th Annual Meeting of the Association for Computa- tional Linguistics and the 11th International Joint Conference on Natural Language Proce...
2021
-
[58]
C. M. Woods and D. Thissen. Item response theory with estimation of the latent population distribution using spline-based densities.Psychometrika, 71(2):281–301, 2006
2006
-
[59]
M. Wu, R. L. Davis, B. W. Domingue, C. Piech, and N. Goodman. Variational item response theory: Fast, accurate, and expressive, 2020
2020
-
[60]
Z. Xu, J. Liu, Y . Wang, and Y . Gu. Latency-response theory model: Evaluating large language models via response accuracy and chain-of-thought length.arXiv [stat.ME], Dec. 2025
2025
-
[61]
Zellers, A
R. Zellers, A. Holtzman, Y . Bisk, A. Farhadi, and Y . Choi. HellaSwag: Can a machine really finish your sentence? In A. Korhonen, D. Traum, and L. Màrquez, editors,Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy, July 2019. Association for Computational Linguistics
2019
-
[62]
H. Zhou, H. Huang, Z. Zhao, L. Han, H. Wang, K. Chen, M. Yang, W. Bao, J. Dong, B. Xu, C. Zhu, H. Cao, and T. Zhao. Lost in benchmarks? rethinking large language model benchmark- ing with item response theory.Proceedings of the AAAI Conference on Artificial Intelligence, 40(41):35085–35093, Mar. 2026
2026
-
[63]
we were unable to find the license for the dataset we used
Y . Zhuang, Q. Liu, Y . Ning, W. Huang, R. Lv, Z. Huang, G. Zhao, Z. Zhang, Q. Mao, S. Wang, and E. Chen. Efficiently measuring the cognitive ability of LLMs: An adaptive testing perspec- tive. https://openreview.net/forum?id=s6X3s3rBPW, Oct. 2023. Accessed: 2026-4-16. 13 A Estimator introduction MML-EM.Marginal maximum likelihood (MML) is the conventiona...
2023
-
[65]
Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects
Institutional review board (IRB) approvals or equivalent for research with human subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals (or an equivalent approval/review based on the requirements of your country or ...
-
[2020]
Association for Computational Linguistics. 10
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.