REVIEW 18 cited by
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
read the original abstract
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from learning from and drawing on the social sciences when developing and using measurement instruments for evaluating GenAI systems. Specifically, our position is that evaluating GenAI systems is a social science measurement challenge. We present a four-level framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, behaviors, and impacts of GenAI systems. This framework has two important implications: First, it can broaden the expertise involved in evaluating GenAI systems by enabling stakeholders with different perspectives to participate in conceptual debates. Second, it brings rigor to both conceptual and operational debates by offering a set of lenses for interrogating validity.
Forward citations
Cited by 18 Pith papers
-
On the Convergent Validity of Offline Evaluation Designs for Recommender Systems
Sparse offline recommender rankings correlate only weakly—and sometimes negatively—with dense ground-truth rankings, and no evaluation design is uniformly best.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
On the same 720 replies, scoring exposure versus manifestation shifts the auditor-judge gap by ~0.2 AUROC and can reverse their ranking, so single detection AUROCs are under-specified.
-
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Coding benchmarks conflate the model with the harness, environment, and verifier into a single end-to-end score, which is misaligned with agentic software engineering.
-
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
A repeatable worksheet and human-reviewed expansion process turns expert-elicited AI use cases into 107 grounded scenarios to support consistent human-centered evaluations.
-
"I Just Don't Want My Work Being Fed Into The AI Blender": Queer Artists on Refusing and Resisting Generative AI
Queer artists largely refuse and resist generative AI, seeing it as anti-relational and disruptive to the community-oriented, identity-forming nature of their art practices, with only limited acceptance for surreal im...
-
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
A statistical framework decomposes human annotation outcomes into four interpretable variation sources and extends classical measurement-error models to handle both shared and individualized notions of truth.
-
Grounded Chess Reasoning in Language Models via Master Distillation
Master Distillation turns opaque chess-engine search into natural-language CoT, lifting a 4B model (C1) from near-zero to 48.1% accuracy that beats its teacher and most larger LMs.
-
RLHF May Not Reflect Genuine Preferences
RLHF preference measurement is a social science validity problem because annotators routinely produce non-attitudes, constructed responses, and artifacts rather than stable values.
-
RLHF May Not Reflect Genuine Preferences
RLHF annotations frequently lack stable underlying preferences; consistency diagnostics on PRISM and PluriHarms show that removing inconsistent annotators flips majority harm classifications for 18.6% of prompts.
-
Responsible Evaluation of AI for Mental Health
Proposes an interdisciplinary framework and taxonomy for responsible evaluation of AI mental health tools based on analysis of 135 publications identifying gaps in metrics, expert involvement, safety, and equity.
-
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
The paper sets out a research agenda for NLP to measure how prolonged language-model use changes human behavior over long time horizons, replacing single-session safety evaluations with longitudinal tracking.
-
Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer
An LLM annotation pipeline for Schwartz human values in Russian social media, calibrated against expert labels and transferred to a smaller encoder, changes measured value prevalence but not necessarily overall trend ...
-
Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents
Industry markets AI agents for orchestration, creation, and insight, but a usability study with 31 participants reveals users face challenges from capability misalignment and lack of meta-cognition in tools like Opera...
-
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Coding benchmarks misalign with agentic software engineering because they conflate model and harness, grade against single references, and provide no component-level iteration signals.
-
Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer
Calibrated LLM annotations of Schwartz human values in non-English social media text transfer to encoder models via soft-label training while retaining theoretical alignment and uncertainty signals.
-
AI as a Tool for Simulation-Based Experiments in Literary Studies
Proposes AI-driven simulations for literary-historical experiments and reports preliminary text-generation results claiming the first limited in-distribution outputs matching human novels.
-
Making AI Evaluation Deployment Relevant Through Context Specification
Context specification is a process that turns diffuse stakeholder perspectives into explicit definitions of properties, behaviors, and outcomes to guide context-aware AI evaluations.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.