Pith. sign in

REVIEW 2 cited by

Evaluating General-Purpose AI with Psychometrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16379 v2 pith:X5AAW5UG submitted 2023-10-25 cs.AI cs.CY

classification cs.AIcs.CY
keywords evaluationtasksgeneral-purposeperformancepsychometricsspecificsystemsbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their capabilities. Current evaluation methodology, mostly based on benchmarks of specific tasks, falls short of adequately assessing these versatile AI systems, as present techniques lack a scientific foundation for predicting their performance on unforeseen tasks and explaining their varying performance on specific task items or user inputs. Moreover, existing benchmarks of specific tasks raise growing concerns about their reliability and validity. To tackle these challenges, we suggest transitioning from task-oriented evaluation to construct-oriented evaluation. Psychometrics, the science of psychological measurement, provides a rigorous methodology for identifying and measuring the latent constructs that underlie performance across multiple tasks. We discuss its merits, warn against potential pitfalls, and propose a framework to put it into practice. Finally, we explore future opportunities of integrating psychometrics with the evaluation of general-purpose AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

    cs.CY 2025-02 accept novelty 5.0 of 10

    The paper proposes that GenAI evaluation be standardized with a four-level social science measurement framework that separates concept definition from measurement and emphasizes validity testing.

  2. Evaluating Generative AI Systems is a Social Science Measurement Challenge

    cs.CY 2024-11 conditional novelty 4.0 of 10

    The paper argues that generative AI evaluation should adopt a social science measurement framework with an explicit 'systematized concept' between high-level ideas and concrete instruments.

Pith tools