Pith. sign in

REVIEW 7 cited by

Audit Cards: Contextualizing AI Evaluations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.13839 v2 pith:MU43ZSKO submitted 2025-04-18 cs.CY

Audit Cards: Contextualizing AI Evaluations

classification cs.CY
keywords reportingauditcardsevaluationscontextframeworksaccessanalysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

AI governance frameworks increasingly rely on audits, yet the results of their underlying evaluations require interpretation and context to be meaningfully informative. Even technically rigorous evaluations can offer little useful insight if reported selectively or obscurely. Current literature focuses primarily on technical best practices, but evaluations are an inherently sociotechnical process, and there is little guidance on reporting procedures and context. Through literature review, stakeholder interviews, and analysis of governance frameworks, we propose "audit cards" to make this context explicit. We identify six key types of contextual features to report and justify in audit cards: auditor identity, evaluation scope, methodology, resource access, process integrity, and review mechanisms. Through analysis of existing evaluation reports, we find significant variation in reporting practices, with most reports omitting crucial contextual information such as auditors' backgrounds, conflicts of interest, and the level and type of access to models. We also find that most existing regulations and frameworks lack guidance on rigorous reporting. In response to these shortcomings, we argue that audit cards can provide a structured format for reporting key claims alongside their justifications, enhancing transparency, facilitating proper interpretation, and establishing trust in reporting.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

    cs.AI 2026-06 unverdicted novelty 7.0

    Introduces the first community-governed unified JSON schema and crowdsourced repository for AI evaluation results, with converters and a database spanning 22,235 models and 2,273 benchmarks.

  2. Defeater Cards: Characterizing and Managing Safety Assurance Case Defeaters

    cs.SE 2026-06 unverdicted novelty 6.0

    Defeater Cards introduce a new 5W1H-structured documentation method for systematically characterizing defeaters in safety assurance cases, supported by an open repository and demonstrated in cross-domain case studies.

  3. Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

    cs.AI 2026-06 unverdicted novelty 6.0

    EvalCards is a composable reporting schema and monitoring tool for AI evaluations, derived from 52 papers and 10 interviews, and applied to 5,816 models and 101,843 results to surface reporting gaps.

  4. The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems

    cs.CY 2026-02 accept novelty 6.0

    The 2025 AI Agent Index catalogs technical and safety details for 30 deployed AI agents and finds low developer transparency on safety, evaluations, and societal impacts.

  5. Computational Hermeneutics: Evaluating generative AI as a cultural technology

    cs.AI 2026-03 unverdicted novelty 5.0

    Generative AI should be evaluated through computational hermeneutics using iterative, human-inclusive benchmarks that measure cultural context rather than isolated model outputs.

  6. Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents

    cs.HC 2025-09 unverdicted novelty 5.0

    Industry markets AI agents for orchestration, creation, and insight, but a usability study with 31 participants reveals users face challenges from capability misalignment and lack of meta-cognition in tools like Opera...

  7. Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends

    cs.RO 2026-06 unverdicted novelty 3.0

    Automation in embodied benchmark construction shifts costs from acquisition toward validation, auditability, version control, and long-term governance instead of simply lowering total cost.