Pith. sign in

REVIEW 4 cited by

Position: AI Evaluation Should Learn from How We Test Humans

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.10512 v4 pith:VOPSUG2X submitted 2023-06-18 cs.CL

classification cs.CL
keywords evaluationtestparadigmpositionpsychometricshumanitemsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As AI systems continue to evolve, their rigorous evaluation becomes crucial for their development and deployment. Researchers have constructed various large-scale benchmarks to determine their capabilities, typically against a gold-standard test set and report metrics averaged across all items. However, this static evaluation paradigm increasingly shows its limitations, including high evaluation costs, data contamination, and the impact of low-quality or erroneous items on evaluation reliability and efficiency. In this Position, drawing from human psychometrics, we discuss a paradigm shift from static evaluation methods to adaptive testing. This involves estimating the characteristics or value of each test item in the benchmark, and tailoring each model's evaluation instead of relying on a fixed test set. This paradigm provides robust ability estimation, uncovering the latent traits underlying a model's observed scores. This position paper analyze the current possibilities, prospects, and reasons for adopting psychometrics in AI evaluation. We argue that psychometrics, a theory originating in the 20th century for human assessment, could be a powerful solution to the challenges in today's AI evaluations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fluid Language Model Benchmarking

    cs.CL 2025-09 conditional novelty 8.0 of 10

    Fluid Benchmarking, combining IRT-based ability estimation with Fisher-information-based adaptive item selection, improves LM evaluation across efficiency, validity, variance, and saturation in pretraining settings.

  2. Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

    cs.CL 2025-10 conditional novelty 6.0 of 10

    An IRT-based adaptive testing framework, ATLAS, estimates LLM ability with 30-89 items per benchmark, matching whole-bank ability estimates and re-ranking 23-31% of models relative to accuracy.

  3. AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An agent-driven framework adaptively selects a small subset of benchmark questions for MLLMs, preserving over 90% ranking accuracy with roughly 4-5% of the data.

  4. Psychometric-Based Evaluation for Theorem Proving with Large Language Models

    cs.AI 2025-02 conditional novelty 5.0 of 10

    The authors annotate miniF2F theorems with LLM-computed difficulty and discrimination scores, then use adaptive testing to rank 10 theorem-proving LLMs using only about 23% of the theorems.

Pith tools