Pith. sign in

REVIEW 8 cited by

Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03100 v4 pith:OMZ77NWP submitted 2023-06-01 cs.HC cs.AI

classification cs.HCcs.AI
keywords evaluationmodelmethodssocio-technicalchallengescommunitydiversehomogenization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent development of generative large language models (LLMs) poses new challenges for model evaluation that the research community and industry have been grappling with. While the versatile capabilities of these models ignite much excitement, they also inevitably make a leap toward homogenization: powering a wide range of applications with a single, often referred to as ``general-purpose'', model. In this position paper, we argue that model evaluation practices must take on a critical task to cope with the challenges and responsibilities brought by this homogenization: providing valid assessments for whether and how much human needs in diverse downstream use cases can be satisfied by the given model (\textit{socio-technical gap}). By drawing on lessons about improving research realism from the social sciences, human-computer interaction (HCI), and the interdisciplinary field of explainable AI (XAI), we urge the community to develop evaluation methods based on real-world contexts and human requirements, and embrace diverse evaluation methods with an acknowledgment of trade-offs between realisms and pragmatic costs to conduct the evaluation. By mapping HCI and current NLG evaluation methods, we identify opportunities for evaluation methods for LLMs to narrow the socio-technical gap and pose open questions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild

    cs.SE 2026-01 conditional novelty 7.0 of 10

    Qualitative study of 19 practitioners reveals ten LLM product evaluation practices and introduces the results-actionability gap as a key barrier to turning findings into improvements.

  2. TS-Arena -- A Live Forecast Pre-Registration Platform

    cs.LG 2025-12 conditional novelty 7.0 of 10

    TS-Arena is a live pre-registration platform that evaluates time series forecasts on future data streams to eliminate information leakage.

  3. Position: Evaluation Scores Are Perishable Knowledge Claims

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Evaluation scores should be treated as perishable knowledge claims with explicit formality, scope, and validity windows; weakest-link (min) aggregation is the proposed conservative endpoint for combining them.

  4. The Foreign Policy AI Evaluation Gap

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.

  5. Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations

    cs.AI 2026-04 conditional novelty 6.0 of 10

    Improvements in LLM Theory of Mind on static benchmarks do not reliably improve performance in dynamic, first-person human-AI interactions across goal-oriented and experience-oriented tasks.

  6. Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners

    cs.HC 2025-09 unverdicted novelty 6.0 of 10

    A human-centered design workshop with journalism practitioners yields an evaluation cookbook and design requirements for contextualized, value-aligned generative AI benchmarks.

  7. Computational Hermeneutics: Evaluating generative AI as a cultural technology

    cs.AI 2026-03 unverdicted novelty 5.0 of 10

    Generative AI should be evaluated through computational hermeneutics using iterative, human-inclusive benchmarks that measure cultural context rather than isolated model outputs.

  8. Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms

    cs.CL 2026-04 unverdicted novelty 4.0 of 10

    AI explanations of slang improve non-native speakers' writing competence more than definitions or rewrites according to native raters, but users overestimate their skill and a performance gap with natives remains.

Pith tools