An IRT-based adaptive testing framework, ATLAS, estimates LLM ability with 30-89 items per benchmark, matching whole-bank ability estimates and re-ranking 23-31% of models relative to accuracy.
hub
Title resolution pending
1 Pith paper cite this work, alongside 1,208 external citations. Polarity classification is still indexing.
1
Pith paper citing it
1,208
external citations · OpenAlex
hub tools
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
An IRT-based adaptive testing framework, ATLAS, estimates LLM ability with 30-89 items per benchmark, matching whole-bank ability estimates and re-ranking 23-31% of models relative to accuracy.