Pith. sign in

REVIEW 9 cited by

Evaluating Generative AI Systems is a Social Science Measurement Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.10939 v1 pith:BNBUFXES submitted 2024-11-17 cs.CY

classification cs.CY
keywords measurementevaluatingsystemsframeworkgenaimeasurementssocialtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult. We argue that these measurement tasks are highly reminiscent of measurement tasks found throughout the social sciences. With this in mind, we present a framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, impacts, opportunities, and risks of GenAI systems. The framework distinguishes between four levels: the background concept, the systematized concept, the measurement instrument(s), and the instance-level measurements themselves. This four-level approach differs from the way measurement is typically done in ML, where researchers and practitioners appear to jump straight from background concepts to measurement instruments, with little to no explicit systematization in between. As well as surfacing assumptions, thereby making it easier to understand exactly what the resulting measurements do and do not mean, this framework has two important implications for evaluating evaluations: First, it can enable stakeholders from different worlds to participate in conceptual debates, broadening the expertise involved in evaluating GenAI systems. Second, it brings rigor to operational debates by offering a set of lenses for interrogating the validity of measurement instruments and their resulting measurements.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

    cs.CY 2025-09 conditional novelty 7.0 of 10

    A new benchmark finds low to moderate human agency support in 20 LLM assistants across six dimensions.

  2. Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

    cs.CY 2026-05 conditional novelty 6.0 of 10

    Healthcare LLM benchmarks overlook implicit assumptions about user behavior that split into task assumptions testable from conversation data and outcome assumptions requiring behavioral studies, shown by reanalyzing a...

  3. Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Proposes a three-level taxonomy of Cultural Awareness, Cultural Sensitivity, and Cultural Competence for AI evaluation, grounded in intercultural communication scholarship to improve validity in multicultural contexts.

  4. From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    A practical evaluation protocol for AI pentesting agents that uses validated vulnerability discovery, LLM semantic matching, and bipartite scoring to assess performance in realistic, complex targets.

  5. From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

    cs.AI 2026-05 conditional novelty 6.0 of 10

    A compact OMR model trained only on normalized synthetic **kern data reaches 18.46% OMR-NED on Verovio scores and 63.97% on real Polish scans, outperforming larger baselines.

  6. From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

    cs.AI 2026-05 conditional novelty 6.0 of 10

    An evaluation protocol for AI pentesting agents that scores validated vulnerability discovery using LLM-based semantic matching and bipartite resolution.

  7. Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation

    cs.CL 2026-04 unverdicted novelty 4.0 of 10

    Closure of the Perspective API exposes structural dependence on a single proprietary toxicity scorer, leaving non-updatable benchmarks and irreproducible results while risking continued reliance on closed LLMs.

  8. Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop

    cs.HC 2025-08 conditional novelty 4.0 of 10

    A synthesis of the CHI 2025 workshop maps research and design opportunities for understanding, protecting, and augmenting human cognition with generative AI.

  9. Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

    cs.CL 2025-08 conditional novelty 4.0 of 10

    An argument, grounded in social-science measurement theory, that LLM-as-judge adoption has outpaced validity and reliability testing, with an analysis of four underlying assumptions.

Pith tools