Pith. sign in

REVIEW 18 cited by

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.00561 v2 pith:JQDGEE2V submitted 2025-02-01 cs.CY

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

classification cs.CY
keywords systemsevaluatinggenaimeasurementsocialpositionchallengeconceptual
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from learning from and drawing on the social sciences when developing and using measurement instruments for evaluating GenAI systems. Specifically, our position is that evaluating GenAI systems is a social science measurement challenge. We present a four-level framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, behaviors, and impacts of GenAI systems. This framework has two important implications: First, it can broaden the expertise involved in evaluating GenAI systems by enabling stakeholders with different perspectives to participate in conceptual debates. Second, it brings rigor to both conceptual and operational debates by offering a set of lenses for interrogating validity.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. On the Convergent Validity of Offline Evaluation Designs for Recommender Systems

    cs.IR 2026-07 conditional novelty 6.0

    Sparse offline recommender rankings correlate only weakly—and sometimes negatively—with dense ground-truth rankings, and no evaluation design is uniformly best.

  2. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  3. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0

    On the same 720 replies, scoring exposure versus manifestation shifts the auditor-judge gap by ~0.2 AUROC and can reverse their ranking, so single detection AUROCs are under-specified.

  4. Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

    cs.SE 2026-06 accept novelty 6.0

    Coding benchmarks conflate the model with the harness, environment, and verifier into a single end-to-end score, which is misaligned with agentic software engineering.

  5. Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios

    cs.HC 2026-05 unverdicted novelty 6.0

    A repeatable worksheet and human-reviewed expansion process turns expert-elicited AI use cases into 107 grounded scenarios to support consistent human-centered evaluations.

  6. "I Just Don't Want My Work Being Fed Into The AI Blender": Queer Artists on Refusing and Resisting Generative AI

    cs.HC 2026-04 unverdicted novelty 6.0

    Queer artists largely refuse and resist generative AI, seeing it as anti-relational and disruptive to the community-oriented, identity-forming nature of their art practices, with only limited acceptance for surreal im...

  7. From Ground Truth to Measurement: A Statistical Framework for Human Labeling

    stat.ME 2026-04 unverdicted novelty 6.0

    A statistical framework decomposes human annotation outcomes into four interpretable variation sources and extends classical measurement-error models to handle both shared and individualized notions of truth.

  8. Grounded Chess Reasoning in Language Models via Master Distillation

    cs.AI 2026-03 unverdicted novelty 6.0

    Master Distillation turns opaque chess-engine search into natural-language CoT, lifting a 4B model (C1) from near-zero to 48.1% accuracy that beats its teacher and most larger LMs.

  9. RLHF May Not Reflect Genuine Preferences

    cs.HC 2026-01 unverdicted novelty 6.0

    RLHF preference measurement is a social science validity problem because annotators routinely produce non-attitudes, constructed responses, and artifacts rather than stable values.

  10. RLHF May Not Reflect Genuine Preferences

    cs.HC 2026-01 conditional novelty 6.0

    RLHF annotations frequently lack stable underlying preferences; consistency diagnostics on PRISM and PluriHarms show that removing inconsistent annotators flips majority harm classifications for 18.6% of prompts.

  11. Responsible Evaluation of AI for Mental Health

    cs.CY 2026-01 unverdicted novelty 6.0

    Proposes an interdisciplinary framework and taxonomy for responsible evaluation of AI mental health tools based on analysis of 135 publications identifying gaps in metrics, expert involvement, safety, and equity.

  12. Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

    cs.AI 2026-08 accept novelty 5.0

    The paper sets out a research agenda for NLP to measure how prolonged language-model use changes human behavior over long time horizons, replacing single-session safety evaluations with longitudinal tracking.

  13. Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer

    cs.CL 2026-06 conditional novelty 5.0

    An LLM annotation pipeline for Schwartz human values in Russian social media, calibrated against expert labels and transferred to a smaller encoder, changes measured value prevalence but not necessarily overall trend ...

  14. Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents

    cs.HC 2025-09 unverdicted novelty 5.0

    Industry markets AI agents for orchestration, creation, and insight, but a usability study with 31 participants reveals users face challenges from capability misalignment and lack of meta-cognition in tools like Opera...

  15. Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

    cs.SE 2026-06 unverdicted novelty 4.0

    Coding benchmarks misalign with agentic software engineering because they conflate model and harness, grade against single references, and provide no component-level iteration signals.

  16. Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer

    cs.CL 2026-06 unverdicted novelty 4.0

    Calibrated LLM annotations of Schwartz human values in non-English social media text transfer to encoder models via soft-label training while retaining theoretical alignment and uncertainty signals.

  17. AI as a Tool for Simulation-Based Experiments in Literary Studies

    cs.CL 2026-06 unverdicted novelty 4.0

    Proposes AI-driven simulations for literary-historical experiments and reports preliminary text-generation results claiming the first limited in-distribution outputs matching human novels.

  18. Making AI Evaluation Deployment Relevant Through Context Specification

    cs.AI 2026-03 unverdicted novelty 4.0

    Context specification is a process that turns diffuse stakeholder perspectives into explicit definitions of properties, behaviors, and outcomes to guide context-aware AI evaluations.