Pith. sign in

REVIEW 8 cited by

Toward an Evaluation Science for Generative AI Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.05336 v3 pith:VN4LIJQZ submitted 2025-03-07 cs.AI cs.LG

classification cs.AIcs.LG
keywords evaluationgenerativesystemsmustsafetysciencechallengesengineering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks face validity challenges, and ad hoc case-by-case audits rarely scale. In this piece, we advocate for maturing an evaluation science for generative AI systems. While generative AI creates unique challenges for system safety engineering and measurement science, the field can draw valuable insights from the development of safety evaluation practices in other fields, including transportation, aerospace, and pharmaceutical engineering. In particular, we present three key lessons: Evaluation metrics must be applicable to real-world performance, metrics must be iteratively refined, and evaluation institutions and norms must be established. Applying these insights, we outline a concrete path toward a more rigorous approach for evaluating generative AI systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

    cs.CY 2025-09 conditional novelty 7.0 of 10

    A new benchmark finds low to moderate human agency support in 20 LLM assistants across six dimensions.

  2. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  3. From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    A practical evaluation protocol for AI pentesting agents that uses validated vulnerability discovery, LLM semantic matching, and bipartite scoring to assess performance in realistic, complex targets.

  4. Correlated Errors in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Large language models from different providers and architectures often make the same errors, and more accurate models are especially likely to share mistakes.

  5. Adultification Bias in LLMs and Text-to-Image Models

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Large language and text-to-image models show measurable adultification bias, portraying Black girls as more mature, culpable, and sexualized than White girls in several tested models.

  6. Participatory AI: A Scandinavian Approach to Human-Centered AI

    cs.HC 2025-09 conditional novelty 5.0 of 10

    Participatory AI applies five Scandinavian Participatory Design principles to four AI design challenges, illustrated through five diverse case studies.

  7. A Conceptual Framework for AI Capability Evaluations

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.

  8. NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The authors propose a competition with new scoring metrics to find benchmarks that give clean early-training signals for small language models, and show MMLU-var outperforms MMLU as a baseline.

Pith tools