Pith. sign in

REVIEW 3 major objections 5 minor 5 cited by

Audit Cards: Contextualizing AI Evaluations

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read AI evaluation results are uninterpretable without context, so the paper proposes a standardized 'audit card' template—three principles and six features—to make that context explicit.

desk verdict A genuinely useful audit-card template and descriptive gap survey, weakened by an untested claim that the cards will improve interpretation and trust. read the letter →

arxiv 2504.13839 v2 pith:MU43ZSKO submitted 2025-04-18 cs.CY

classification cs.CY
keywords auditcardsAIevaluationreportingauditstransparencygovernanceframeworksconflictofinterestcontextsociotechnical
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that AI audit results can be rigorous yet uninformative or misleading when the surrounding context is not reported. It proposes "audit cards": a standardized template that requires disclosing three cross-cutting principles (justification, assumptions, limitations) and six contextual features (who evaluated, what was evaluated, how, resource access, process integrity, and review mechanisms). The authors analyze 24 existing evaluation reports and 21 governance frameworks, finding that most reports omit auditor backgrounds, conflicts of interest, and levels of model access, while most regulations give little guidance on reporting. They argue that this structured format would enhance transparency, help readers interpret results correctly, and build trust in AI governance.

What carries the argument

The central object is the audit card template itself: a checklist organized around three principles (justification, assumptions, limitations) that cut across six features (auditor identity, evaluation scope, methodology, resource access, process integrity, and review mechanisms). The template converts diffuse norms about transparency into concrete disclosure questions, such as how auditors were selected, what conflicts of interest existed, what model access was granted, and when the evaluation becomes obsolete. These structured fields are what allow readers to assess whether an evaluation is credible and what it actually covers.

What would settle it

Run a controlled study in which readers receive the same evaluation results with and without an audit card, then measure whether the card changes their interpretation accuracy, their confidence in conclusions, or their trust in the evaluator; if no measurable improvement appears, the paper's core benefit claim collapses.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that AI evaluations are sociotechnical processes whose findings are shaped by non-technical details, and that current reporting practice largely fails to document those details. It synthesizes prior literature into a disclosure template consisting of three principles and six features, then shows through manual scoring that real evaluation reports from developers and third parties are highly inconsistent: scope and procedure are usually covered, but integrity, resource access, and auditor identity are rarely disclosed, and no examined report states obsolescence criteria. The paper further finds that existing governance frameworks rarely require or recommend such disclosures. It concludes that audit cards can close this reporting gap and thereby enable meaningful interpretation of evaluation results.

Load-bearing premise

The paper assumes, without empirical evidence, that adding these disclosure fields to evaluation reports will actually improve how readers interpret and trust the results, rather than being ignored or merely adding reporting burden.

Editorial extensions

If this is right

  • Evaluation reports from different organizations would become directly comparable because the same contextual fields are filled in each time.
  • Regulators and users could judge auditor independence and conflicts of interest at a glance, reducing the risk of selective or cosmetic reporting.
  • Requiring obsolescence criteria would prevent outdated evaluations from being cited as current evidence about a model.
  • Governance frameworks, which currently specify that audits should happen but not how they should be reported, could adopt the audit card as a concrete compliance template.
  • If adopted broadly, audit cards would make the evaluation process itself auditable, shifting accountability from technical execution alone to reporting integrity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's central benefit claim could be tested directly: present readers with identical evaluation metrics either with or without audit-card context and measure whether their interpretation accuracy, confidence calibration, and trust judgments change.
  • A registry of audit cards, one per model evaluation, could create a public ledger that makes opinion shopping visible and enables longitudinal tracking of whether disclosure quality improves over time.
  • The template could be extended to cover downstream consumers' needs, since the paper notes that policymakers currently read reports at a high level; prioritization of fields per audience is suggested but not developed.
  • If disclosure burden proves high, a plausible resolution is tiered audit cards: minimal required fields for all evaluations and fuller disclosure for higher-stakes or pre-deployment audits.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes 'audit cards,' a structured reporting template for documenting the context of AI evaluations, organized around three principles (justification, assumptions, limitations) and six features (auditor identity, evaluation scope, methodology, resource access, process integrity, review mechanisms). The framework is derived from a literature review of 28 papers, a manual scoring of 24 evaluation reports, a binary scoring of 21 governance frameworks, and ten expert stakeholder interviews. The authors report that most existing evaluation reports omit contextual details such as auditor background, conflicts of interest, and model access, and that most governance frameworks lack specific reporting requirements. They argue that audit cards would enhance transparency, facilitate proper interpretation, and establish trust in AI evaluation reporting.

Significance. The paper addresses a real and understudied problem: AI evaluation results are often uninterpretable without contextual information, and current reporting practices are inconsistent. The contribution is a concrete, actionable template with a checklist, backed by a multi-method empirical survey of reports and governance frameworks and by stakeholder interviews. Strengths include the systematic annotation effort across 24 reports and 21 frameworks, the detailed appendix with the full audit-card checklist and interview protocol, and the transparent discussion of limitations (e.g., limited sample size, trade-offs with audit burden). The proposal is falsifiable in the sense that its effects on interpretation and trust could be tested with reader studies, though the present manuscript does not conduct such tests. The paper is a useful step toward standardizing evaluation reporting, but its central claim about improving trust and interpretation currently rests on assertion rather than evidence.

major comments (3)
  1. [Section 7 and Abstract] The central claim that audit cards 'enable meaningful interpretation of evaluation results' and establish trust is a causal, behavioral claim, but the manuscript provides no outcome measure, comparison group, pilot deployment, or reader study. Section 7 asserts this benefit without evidence. Moreover, Section 6 reports that policymakers often read evaluation reports only at a high level (P6, P9), which undercuts the assumption that detailed contextual fields will reach key decision-makers. The authors should either reframe the contribution as a documentation standard whose effects on interpretation and trust remain open questions, or provide empirical evidence from a user study or pilot deployment.
  2. [Section 4 / Appendix B] The empirical claims about reporting gaps (e.g., 'contextual details are much less consistently reported,' average 0.94; only 4 of 24 reports disclose integrity-related information) rest on manual 0-2 or 0-1 scoring, but no inter-rater reliability statistics are reported. Appendix B indicates two annotators discussed discrepancies, but the absence of a kappa-like measure means the quantitative framing of the descriptive findings is not fully supported. Report inter-rater reliability, or present the survey results as qualitative observations rather than quantitative scores.
  3. [Section 3 / Table 1] The audit-card framework is presented as a 'comprehensive and exhaustive overview' (Table 1 caption), but the underlying literature selection is not systematic: Appendix B states that relevance was initially assessed by titles and yielded only eight papers, prompting expansion to adjacent fields. No formal search protocol, inclusion/exclusion criteria, or screening process is reported. One feature (obsolescence criteria) cites Joaquin et al. 2025, which includes a co-author of this paper, and the scoring is at risk of confirmation bias. The authors should either document a systematic review methodology or temper the completeness claim.
minor comments (5)
  1. [Section 3] The sentence 'This includes, what access auditors have to a system... and which resources available are available' contains a duplicated phrase; it should read 'which resources are available.'
  2. [Section 3] The phrase 'It's aim is to be a comprehensive and exhaustive overview' uses the contraction 'It's' where the possessive 'Its' is intended.
  3. [Section 4] The text states 'twelce do so comprehensively' — 'twelce' should be 'twelve.'
  4. [Table 2] The lower section of Table 2 appears to list individual reports with average scores, but the table is not clearly separated from the summary statistics; consider adding a visual divider or a separate table for per-report scores.
  5. [Appendix C] The checklist is very long and would benefit from a short 'how to use' prefatory paragraph explaining which questions are mandatory versus optional, and how to prioritize when resources are limited.

Circularity Check

1 steps flagged · score 2.0 of 10

Minor self-citation in one checklist item; no central circularity.

  1. other [Section 3, 'What is evaluated' feature bullet and footnote 5; Appendix C checklist]
    "as well as reporting on the evaluation obsolescence criteria (Joaquin et al. 2025)."

    The template's 'What is evaluated' feature includes obsolescence criteria, and the sole citation provided for this design choice is Joaquin et al. 2025, a paper co-authored by this paper's author Staufer. Thus this component of the proposed audit card is justified by the authors' own prior work rather than by independent external evidence. This is a minor, non-load-bearing self-citation: the central framework, the empirical scoring of reports and frameworks, and the transparency argument do not depend on this single checklist item, and the other principles and features are grounded in a 28-paper literature review plus stakeholder interviews.

full rationale

The paper's derivation chain is: (1) a literature review of 28 papers and stakeholder interviews produce a template of three principles and six contextual features; (2) the template is applied to score 24 evaluation reports and 21 governance frameworks; (3) the authors argue that structuring reports this way 'enable[s] meaningful interpretation of evaluation results' (Section 7). There is no fitted parameter and no equation; the template is defined before the scoring, so the survey findings are measurements rather than predictions forced by construction. The central claim that audit cards enhance transparency, interpretation, and trust is a normative argument, not a derived result, so it cannot reduce to its inputs. The only self-citation of note is the obsolescence criterion (Joaquin et al. 2025, co-authored by Staufer), which is a minor checklist item and not load-bearing; other self-citations (Casper et al. 2024/2025, Reuel et al. 2024b, Maslej et al. 2024) are ordinary literature sources without uniqueness claims. Not engaging with Ananny and Crawford's warning that transparency can be insufficient is an evidence or argumentation gap, not circularity. Overall circularity is negligible.

Assumptions & free parameters 1 free parameters · 4 assumptions · 1 invented entities

No numerical free parameters are fit in this qualitative paper. The analysis rests on hand-chosen scoring rubrics and a small, non-random sample of reports, frameworks, and interviews, which are listed as axioms and assumptions above. The proposed audit card itself is an invented artifact with no external validation.

free parameters (1)
  • Report scoring thresholds (0,1,2) and framework scoring thresholds (0,1)
    Hand-chosen by the authors to balance granularity and reliability; not fitted to data. They determine the quantitative findings in Tables 2 and 3.
assumptions (4)
  • domain assumption AI evaluations are inherently sociotechnical and cannot be meaningfully interpreted without contextual reporting.
    Stated in Section 1 and 3; the paper takes this as the premise for audit cards. It is not empirically established by the paper.
  • ad hoc to paper The three principles (justification, assumptions, limitations) and six features (who, what, how, access, integrity, review) identified from 28 selected papers are the complete set of relevant contextual features.
    Defined in Section 3 and Table 1. The selection of 28 papers and the mapping of 24 initial aspects to these categories is the authors' synthesis, not a proven taxonomy.
  • domain assumption Manual annotation with a 0-2 or 0-1 scale reliably captures reporting thoroughness and framework guidance.
    Described in Section 4 and 5 and Appendix B. No inter-rater reliability or validation against other measures is reported.
  • domain assumption Eight analyzed stakeholder interviews are representative of the AI evaluation ecosystem.
    Section 6 and Appendix E. Ten experts were interviewed, two were excluded, and the sample is a convenience sample.
invented entities (1)
  • Audit cards
    purpose: A structured reporting template and checklist for documenting the context of AI audits, including auditor identity, scope, methodology, access, integrity, and review.
    The paper proposes this artifact but provides no external validation that using it improves transparency, interpretation, or trust.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Audit Cards: Contextualizing AI Evaluations." pith.science (2026). https://pith.science/paper/MU43ZSKO

@misc{pith2026250413839,
  author       = {Pith},
  title        = {Pith review of: Audit Cards: Contextualizing AI Evaluations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MU43ZSKO}},
  note         = {Machine review of arXiv:2504.13839}
}
read the original abstract

AI governance frameworks increasingly rely on audits, yet the results of their underlying evaluations require interpretation and context to be meaningfully informative. Even technically rigorous evaluations can offer little useful insight if reported selectively or obscurely. Current literature focuses primarily on technical best practices, but evaluations are an inherently sociotechnical process, and there is little guidance on reporting procedures and context. Through literature review, stakeholder interviews, and analysis of governance frameworks, we propose "audit cards" to make this context explicit. We identify six key types of contextual features to report and justify in audit cards: auditor identity, evaluation scope, methodology, resource access, process integrity, and review mechanisms. Through analysis of existing evaluation reports, we find significant variation in reporting practices, with most reports omitting crucial contextual information such as auditors' backgrounds, conflicts of interest, and the level and type of access to models. We also find that most existing regulations and frameworks lack guidance on rigorous reporting. In response to these shortcomings, we argue that audit cards can provide a structured format for reporting key claims alongside their justifications, enhancing transparency, facilitating proper interpretation, and establishing trust in reporting.

Figures

Figures reproduced from arXiv: 2504.13839 by the authors.

Figure 1
Figure 1. Our audit card template requires reporting on three [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Silent Updates: Measuring and Closing the Post-Deployment Disclosure Gap

    cs.CY 2026-08 conditional novelty 6.0 of 10

    No AI provider in the studied sample exposes a publicly verifiable link between the model it serves and the model it evaluated, so published safety results cannot be reliably tied to deployed systems.

  2. Deprecating Benchmarks: Criteria and Framework

    cs.CY 2025-07 conditional novelty 6.0 of 10

    A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.

  3. Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.

  4. BiMi Sheets: Infosheets for bias mitigation methods

    cs.LG 2025-05 conditional novelty 6.0 of 10

    BiMi Sheets provide a uniform documentation format for bias mitigation methods, with six standardized sections and 24 pre-filled example sheets.

  5. A Conceptual Framework for AI Capability Evaluations

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.

Reference graph

Works this paper leans on

87 extracted references · 76 canonical work pages · cited by 5 Pith papers

  1. [1]

    Towards Publicly Accountable Frontier LLMs (An- derljung et al. 2023)

  2. [2]

    Science of Evals (Apollo Research 2024)

  3. [3]

    Declare and Justify: Explicit assumptions in AI evalua- tions are necessary for effective regulation (Barnett and Thiergart 2024a)

  4. [4]

    What AI evaluations for preventing catastrophic risks can and cannot do (Barnett and Thiergart 2024b)

  5. [5]

    AI auditing: The Broken Bus on the Road to AI Account- ability (Birhane et al. 2024)

  6. [6]

    Structured Access for Third-Party Research (Bucknall and Trager 2023)

  7. [7]

    Rethink reporting of evaluation results in AI (Burnell et al. 2023)

  8. [8]

    Black-Box Access is Insufficient for Rigorous AI Audits (Casper et al. 2024)

Show all 87 references
  1. [9]

    A Survey on Evaluation of Large Language Models (Chang et al. 2023)

  2. [10]

    Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem (Costanza-Chock, Raji, and Buolamwini 2022)

  3. [11]

    Hard choices in artificial intelligence (Dobbe, Krendl Gilbert, and Mintz 2021)

  4. [12]

    Dimensions of Generative AI Evaluation Design (Dow et al. 2024)

  5. [13]

    Can We Trust AI Benchmarks? (Eriksson et al. 2025)

  6. [14]

    The TRIPOD-LLM Reporting Guideline for Studies Us- ing Large Language Models (Gallifant et al. 2025)

  7. [15]

    Datasheets for Datasets (Gebru et al. 2021)

  8. [16]

    Devising ML Metrics (Hendrycks and Woodside 2024)

  9. [17]

    Responsible Reporting for Frontier AI Development (Kolt et al. 2024)

  10. [18]

    Holistic Evaluation of Language Models (Liang et al. 2023)

  11. [19]

    Task Development Guide (METR b)

  12. [20]

    Auditing Large Language Models: A Three-Layered Ap- proach (M¨okander et al. 2024)

  13. [21]

    Reasons to Doubt the Impact of AI Risk Evaluations (Mukobi 2024)

  14. [22]

    Towards AI Accountability Infrastructure (Ojewale et al. 2024)

  15. [23]

    Closing the AI Accountability Gap: Defining an End-to- End Framework for Internal Algorithmic Auditing (Raji et al. 2020)

  16. [24]

    BetterBench (Reuel et al. 2024b)

  17. [25]

    Fairness and Abstraction in Sociotechnical Systems (Selbst et al. 2019)

  18. [26]

    Model evaluation for extreme risks (Shevlane et al. 2023)

  19. [27]

    Toward an evaluation science for generative AI systems (Weidinger et al. 2025)

  20. [28]

    2024) Mapping of initial aspects Our initial literature review yielded 24 aspects listed below that were mapped (→ ) to the principles and features:

    Evaluatology: The Science and Engineering of Evalua- tion (Zhan et al. 2024) Mapping of initial aspects Our initial literature review yielded 24 aspects listed below that were mapped (→ ) to the principles and features:

  21. [33]

    Level of access auditors had→ Access and Resources

  22. [34]

    How did they work with the model developers → In- tegrity

  23. [35]

    Were they trained by the developers to evaluate the model → Who are the auditors

  24. [36]

    How much compute given to auditors, which manufac- turers/chip model→ Access and Resources

  25. [37]

    Expertise / background of people involved → Who are the auditors

  26. [38]

    What was their initial goal for the evaluation → What is evaluated

  27. [39]

    What were their documentation requirements → Review and communication

  28. [40]

    What was the contract structure / How were they paid/rewarded→ Integrity

  29. [41]

    What human input was used / who were they → How is it evaluated

  30. [42]

    Peer review process / QA testing→ Review and commu- nication

  31. [43]

    How does capability or concept translate to benchmark task→ How is it evaluated

  32. [44]

    Involvement of domain experts during design→ Who are the auditors

  33. [45]

    Feedback channel (openness + how to contact) → Re- view and communication

  34. [46]

    How should scores be interpreted or used → How is it evaluated

  35. [47]

    Assumptions of normative properties are documented→ Assumptions

  36. [48]

    Non-normative assumptions justified→ Assumptions

  37. [49]

    Evals as a process should consider eval’s limitations (not that evals are limited)→ Limitations

  38. [50]

    Interdisciplinary teams / including multiple perspectives → Who are the auditors

  39. [51]

    What is the technical goal of the model being evaluated → What is evaluated

  40. [52]

    What are the principles behind the development of the model→ What is evaluated

  41. [53]

    value hierarchies, diverse populations)→ Assump- tions

    Consider context in which the AI system will be used (e.g. value hierarchies, diverse populations)→ Assump- tions

  42. [54]

    Justification of evaluation methods used / using qualita- tive evaluations as well→ Justification

  43. [55]

    Timely, continuous evals→ What is evaluated

  44. [56]

    opinion shopping,

    Interpretable, easy to read, accessible to a wide audience → Review and communication Appendix C Audit cards Details and justifications of features This section provides an overview of audit card features, in- cluding a concise summary of what to report and why this contextual...

  45. [57]

    Major jurisdiction: broadly selected by geopolitical and economic centrality, especially in the AI industry (Stan- ford HAI staff 2024; Bianzino et al. 2023)

  46. [58]

    METR was the exception because their evaluation reports have been highly influential in the field, and often involve co-authorship from researchers in scaling labs and AISIs

    Authority of the institution: within each jurisdiction, we chose government and non-government institutions. METR was the exception because their evaluation reports have been highly influential in the field, and often involve co-authorship from researchers in scaling labs and AISIs

  47. [59]

    harmonized standards

    Direct influence on model evaluations. For model devel- opers, these tended to be their Responsible Scaling Poli- cies or Frontier Safety Frameworks (see METR a). For government bodies, we selected documents that tracked with more scale and influence (i.e. more national/federa...

  48. [61]

    The organization and role has been reported up to the level of detail interviewees were comfortable with

    Describe the work you do in relation to evaluations, try to be specific about your responsibilities, identities of peo- ple/orgs you relate to, your role in the evaluations field Participant Organization Role P1 Third-party evaluator Evaluation designer, developer & writer P2 ...

  49. [62]

    developers (internal, ex- ternal), research, governance, or also the general public

    Who is the evaluation report you [build/analyze/use] for? Who is the target audience? E.g. developers (internal, ex- ternal), research, governance, or also the general public

  50. [63]

    What part of the evaluation process which, if not done well/reported well, would make you doubt the quality of the eval?

  51. [64]

    How do imagine evaluation best practices best becom- ing reality? Do you see codification—whether in law or Industry norm—as a good/valid way to quality control?

  52. [65]

    In your view, what are the biggest challenge preventing evaluations from being better? More useful for improv- ing safety, or whatever their key goal is? Questions to ask specific categories of people [12 mins] Evaluation developer

  53. [66]

    Do you consider the limitations and assumptions of your evaluations process? Where does that thinking get cap- tured (if it does)?

  54. [67]

    What is the background/expertise of auditors—does it depend on the type of evaluation or something else, and who makes that decision?

  55. [68]

    Can you give us examples of evaluations where the evalu- ations process differed and whether it affected the quality of eval?

  56. [69]

    Do you follow any internal best practices?

  57. [70]

    What is the current training process both internally (with evaluation developers, human baseliners) and with your engagement partner? Evaluation report writer

  58. [71]

    type of info/detail level to pub- lish and why)? (b) What’s the right balance of transparency in reporting evaluations, and what risks surround the achieving of that?

    What is the most important information to share in the evaluation report? How do you decide that? (a) Can you walk through two situations in which you de- cided differently (i.e. type of info/detail level to pub- lish and why)? (b) What’s the right balance of transparency in r...

  59. [72]

    From our skim of evaluation reports, we found these tend to be underspecified/not specified—why do you think that might be? e.g. assumptions, auditors’ background, integrity and resources of process, peer review, commit- ments to take action based on evaluation results, and st...

  60. [73]

    specific downstream application context on which the benchmark is contingent

    Do you consider the limitations and assumptions of your evaluations? Where does that thinking get captured (if it does)? (a) e.g. specific downstream application context on which the benchmark is contingent

  61. [74]

    How do you select the right auditors (external / internal)? Walk through a recent eval’s selection?

  62. [75]

    How much autonomy does your organization get in de- ciding access/resources (e.g. time/compute) for evalua- tions [of a lab’s models]? What does this depend on? (a) Do you include different stakeholders in the evalua- tions process—who are they, and how are they en- gaged?

  63. [76]

    What is the current training process both internally (with evaluation developers, human baseliners) and with your engagement partner?

  64. [77]

    what makes an auditor ’qualified’ (back- ground/training), conflict of interest, funding/contract structures

    What is the right balance of flexibility and specificity in regulation? E.g. what makes an auditor ’qualified’ (back- ground/training), conflict of interest, funding/contract structures

  65. [78]

    If you wanted to, how easy would it be for you to be mis- leading when communicating evaluation results? Why? Evaluation operations support (decisions getting access to resources that evaluation developers need)

  66. [79]

    What resources do you usually require for an eval, and what influences that? Do you document resource require- ments, and where do you share this?

  67. [80]

    Recent Cyber CBRN Agent evaluations re- port used blue/red anonymized models at UK AISI

    What are the conditions of you being able to access them? E.g. Recent Cyber CBRN Agent evaluations re- port used blue/red anonymized models at UK AISI

  68. [81]

    Could you describe the different engagement processes, and why some have been easier than others? Evaluation used for recommendations - policy re- searcher/think tank, funders, lobbying, advocacy

  69. [82]

    If you have had to use an evaluation report (could be your org or another org.) to make a recommendation, what fea- tures of the report have helped you to do so?

  70. [83]

    Do you have examples of/from evaluation reports you found good and/or bad, and why?

  71. [84]

    internal different labs OpenAI, external: Apollo, AISI) AI safety communicators

    How much do you trust evaluation results depending on the funding source (e.g. internal different labs OpenAI, external: Apollo, AISI) AI safety communicators

  72. [85]

    How do you understand evaluations (as distinct from audits) and what type of evaluation report- ing/communication are you aware of?

  73. [86]

    What insights are people most often seeking from evalu- ations, and how are evaluations (not) meeting that need?

  74. [87]

    Re target audience question: Why do you think [previ- ous answer] is the relevant audience? Other relevant au- dience segments, how does the messaging change? How can evaluations most strongly communicate their goal?

  75. [2019]

    In Proceedings of the Conference on Fairness, Accountability, and Trans- parency, FAT* ’19, 220–229

    Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Trans- parency, FAT* ’19, 220–229. New York, NY , USA: Associ- ation for Computing Machinery. ISBN 978-1-4503-6125-5. M¨okander, J. 2023. Auditing of AI: Legal, ethical and tech-...

  76. [2020]

    AI auditing

    Closing the AI Accountability Gap: Defining an End- to-End Framework for Internal Algorithmic Auditing. In Proceedings of the 2020 Conference on Fairness, Account- ability, and Transparency, 33–44. Barcelona Spain: ACM. ISBN 978-1-4503-6936-7. Raji, I. D.; Xu, P.; Honigsberg, ...

  77. [2023]

    arXiv:2307.03109

    A Survey on Evaluation of Large Language Models. arXiv:2307.03109. Costanza-Chock, S.; Raji, I. D.; and Buolamwini, J. 2022. Who Audits the Auditors? Recommendations from a Field Scan of the Algorithmic Auditing Ecosystem. In Proceed- ings of the 2022 ACM Conference on Fairnes...

  78. [2025]

    Nature Medicine, 31(1): 60–69

    The TRIPOD-LLM Reporting Guideline for Stud- ies Using Large Language Models. Nature Medicine, 31(1): 60–69. Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J. W.; Wallach, H.; Iii, H. D.; and Crawford, K. 2021. Datasheets for Datasets. arXiv:1803.09010. Gemini Team. 2024....

  79. [2026]

    In- tegrity

    Governance of AI models and systems appears to be done by a keystone AI Act with effect from 2026 (National Assembly of the Republic of Korea 2024). – US: Given the lack of a US federal-wide AI regulation, we examined the Trump Administration’s AI Action Plan issued in July 20...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.