{"id":"7573032a-1bfc-49fb-90c8-a2a68ab964b0","arxiv_id":"2504.13839","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A proposed 'audit card' reporting standard would require AI evaluators to disclose auditor identity, scope, methods, access, integrity, and review, filling gaps found in existing reports and frameworks.","lead":"This paper proposes 'audit cards', a structured checklist for reporting the context behind AI evaluations, including who ran them, what access they had, and how they were reviewed. It surveys 24 evaluation reports and 21 governance frameworks, finding that most omit key context such as auditor conflicts of interest and model access.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central trust/interpretation claim is asserted, not demonstrated; no outcome data shows audit cards change reader comprehension or trust.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the paper asserts, without empirical test, that adding audit card fields will improve interpretation and trust. This is indeed the central claim's most insecure link. The descriptive contributions (literature synthesis, report/framework surveys, stakeholder interviews) are valuable and plausibly correct, but they only establish that contextual information is often missing. They do not establish that supplying it in a structured format changes reader comprehension, trust calibration, or decision-making. The paper's own interview finding that policymakers read reports at a high level (Section 6) further weakens the assumed mechanism. A randomized experiment directly testing comprehension and trust with and without audit cards would settle whether the central claim lands. Given the proposal is reasonable and well-grounded but unvalidated, the reader's CONDITIONAL verdict is appropriate; no change is needed. We do not see an internal inconsistency or a fatal flaw requiring rejection.","tokens_in":26841,"tokens_out":3369,"duration_ms":33725,"concrete_test":"Run a preregistered randomized experiment with participants from the target audiences (policymakers, ML practitioners, journalists). Give all participants the same evaluation results for a frontier model; randomly assign half to receive an audit card with the six features and three principles and half to receive the original report without it. Measure (a) accuracy in identifying stated limitations and conflicts of interest, (b) calibration of confidence in the reported results, and (c) trust ratings. If the audit card condition does not significantly improve comprehension or calibration over the control, the central claim is unsupported. Optionally, have two independent raters re-score Table 2 and report inter-rater reliability (e.g., Cohen's kappa) to assess the robustness of the descriptive evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that audit cards enhance transparency, facilitate proper interpretation, and establish trust in AI evaluation reporting. This is a causal, empirical claim about reader behavior, yet the paper provides no outcome measure, no comparison group, and no pilot deployment. Section 7 states that audit cards 'enable meaningful interpretation of evaluation results' without evidence. The cited literature includes Ananny and Crawford (2018), which warns that transparency alone is insufficient and can even obscure, but the paper does not engage with the possibility that adding structured fields could be ignored, cause information overload, or be strategically gamed. Section 6 reports that policymakers read evaluation reports at a high level (P6, P9), which undercuts the assumption that detailed contextual fields will reach key decision-makers. The descriptive surveys in Tables 2 and 3 establish a reporting gap, but a gap does not imply that filling it produces the claimed benefits. The central claim therefore rests on an untested mechanism: that readers will attend to, correctly interpret, and trust the disclosed context. If audit cards merely add reporting burden without changing interpretation or trust, the core contribution reduces to a documentation checklist rather than a governance intervention.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes 'audit cards,' a structured reporting template for documenting the context of AI evaluations, organized around three principles (justification, assumptions, limitations) and six features (auditor identity, evaluation scope, methodology, resource access, process integrity, review mechanisms). The framework is derived from a literature review of 28 papers, a manual scoring of 24 evaluation reports, a binary scoring of 21 governance frameworks, and ten expert stakeholder interviews. The authors report that most existing evaluation reports omit contextual details such as auditor background, conflicts of interest, and model access, and that most governance frameworks lack specific reporting requirements. They argue that audit cards would enhance transparency, facilitate proper interpretation, and establish trust in AI evaluation reporting.","tokens_in":27222,"tokens_out":2109,"duration_ms":21057,"significance":"The paper addresses a real and understudied problem: AI evaluation results are often uninterpretable without contextual information, and current reporting practices are inconsistent. The contribution is a concrete, actionable template with a checklist, backed by a multi-method empirical survey of reports and governance frameworks and by stakeholder interviews. Strengths include the systematic annotation effort across 24 reports and 21 frameworks, the detailed appendix with the full audit-card checklist and interview protocol, and the transparent discussion of limitations (e.g., limited sample size, trade-offs with audit burden). The proposal is falsifiable in the sense that its effects on interpretation and trust could be tested with reader studies, though the present manuscript does not conduct such tests. The paper is a useful step toward standardizing evaluation reporting, but its central claim about improving trust and interpretation currently rests on assertion rather than evidence.","major_comments":[{"comment":"The central claim that audit cards 'enable meaningful interpretation of evaluation results' and establish trust is a causal, behavioral claim, but the manuscript provides no outcome measure, comparison group, pilot deployment, or reader study. Section 7 asserts this benefit without evidence. Moreover, Section 6 reports that policymakers often read evaluation reports only at a high level (P6, P9), which undercuts the assumption that detailed contextual fields will reach key decision-makers. The authors should either reframe the contribution as a documentation standard whose effects on interpretation and trust remain open questions, or provide empirical evidence from a user study or pilot deployment.","section":"Section 7 and Abstract"},{"comment":"The empirical claims about reporting gaps (e.g., 'contextual details are much less consistently reported,' average 0.94; only 4 of 24 reports disclose integrity-related information) rest on manual 0-2 or 0-1 scoring, but no inter-rater reliability statistics are reported. Appendix B indicates two annotators discussed discrepancies, but the absence of a kappa-like measure means the quantitative framing of the descriptive findings is not fully supported. Report inter-rater reliability, or present the survey results as qualitative observations rather than quantitative scores.","section":"Section 4 / Appendix B"},{"comment":"The audit-card framework is presented as a 'comprehensive and exhaustive overview' (Table 1 caption), but the underlying literature selection is not systematic: Appendix B states that relevance was initially assessed by titles and yielded only eight papers, prompting expansion to adjacent fields. No formal search protocol, inclusion/exclusion criteria, or screening process is reported. One feature (obsolescence criteria) cites Joaquin et al. 2025, which includes a co-author of this paper, and the scoring is at risk of confirmation bias. The authors should either document a systematic review methodology or temper the completeness claim.","section":"Section 3 / Table 1"}],"minor_comments":[{"comment":"The sentence 'This includes, what access auditors have to a system... and which resources available are available' contains a duplicated phrase; it should read 'which resources are available.'","section":"Section 3"},{"comment":"The phrase 'It's aim is to be a comprehensive and exhaustive overview' uses the contraction 'It's' where the possessive 'Its' is intended.","section":"Section 3"},{"comment":"The text states 'twelce do so comprehensively' — 'twelce' should be 'twelve.'","section":"Section 4"},{"comment":"The lower section of Table 2 appears to list individual reports with average scores, but the table is not clearly separated from the summary statistics; consider adding a visual divider or a separate table for per-report scores.","section":"Table 2"},{"comment":"The checklist is very long and would benefit from a short 'how to use' prefatory paragraph explaining which questions are mandatory versus optional, and how to prioritize when resources are limited.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable fit for a sociotechnical AI governance venue. The main concern is that the authors overstate the policy-relevant effect of their proposed artifact; the descriptive gap analysis is solid, but the causal claim needs to be downgraded or tested. The selection of literature and the self-citation in Table 1 may raise questions about independence, though I do not see evidence of deliberate bias. The revision should focus on recalibrating the claims and adding reliability metrics for the annotation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful thing in this paper is the template and the descriptive survey. The six-feature, three-principle audit card is a sensible synthesis of scattered recommendations, and the scoring of 24 evaluation reports and 21 governance frameworks gives a concrete picture of how little context is currently disclosed. That part is real work, and the paper is honest about its own limits: small non-random samples, simple scoring, no inter-rater reliability check, and a clear statement that reporting burden is a trade-off.\n\nThe soft spot is the central claim. The abstract and Section 7 say audit cards \"enable meaningful interpretation\" and \"establish trust,\" but there is no evidence that readers actually attend to, understand, or trust these fields. The paper itself reports in Section 6 that policymakers read evaluation reports at a high level (P6, P9), which cuts against the idea that detailed contextual fields will reach key decision-makers. The citation to Ananny and Crawford (2018) in the introduction warns that transparency can be ineffective or misleading, but the paper does not grapple with that possibility when making its own positive claims. So the contribution is best read as a governance proposal plus a descriptive gap analysis, not a demonstrated intervention. That is still worth publishing if framed that way.\n\nThe citation pattern is mostly fine. There is some self-citation (Reuel, Casper, and co-authors appear several times), but those are genuinely relevant prior works, and the obsolescence criteria feature cites a paper co-authored by a team member. That is not fatal, but it is the kind of thing that should be disclosed at review.\n\nWho is this for? People working on AI evaluation reporting standards, AISI-type evaluation guidelines, or transparency policy. They will get a structured checklist and a useful snapshot of current omissions. A reader expecting proof that audit cards change interpretation or trust will be disappointed.\n\nI would send this to peer review. The empirical claims are modest enough if described as descriptive, and the template is a plausible starting point for standardization. The authors should be pushed to either soften the causal language or do a small user study.","headline":"A genuinely useful audit-card template and descriptive gap survey, weakened by an untested claim that the cards will improve interpretation and trust.","tokens_in":27548,"tokens_out":1703,"would_cite":true,"duration_ms":15967,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI evaluation results are uninterpretable without context, so the paper proposes a standardized 'audit card' template—three principles and six features—to make that context explicit.","keywords":["audit cards","AI evaluation reporting","AI audits","transparency","governance frameworks","conflict of interest","evaluation context","sociotechnical evaluation"],"falsifier":"Run a controlled study in which readers receive the same evaluation results with and without an audit card, then measure whether the card changes their interpretation accuracy, their confidence in conclusions, or their trust in the evaluator; if no measurable improvement appears, the paper's core benefit claim collapses.","tokens_in":26661,"feed_emoji":"📋","tokens_out":3189,"duration_ms":32098,"temperature":0.7,"pith_summary":"The paper claims that AI audit results can be rigorous yet uninformative or misleading when the surrounding context is not reported. It proposes \"audit cards\": a standardized template that requires disclosing three cross-cutting principles (justification, assumptions, limitations) and six contextual features (who evaluated, what was evaluated, how, resource access, process integrity, and review mechanisms). The authors analyze 24 existing evaluation reports and 21 governance frameworks, finding that most reports omit auditor backgrounds, conflicts of interest, and levels of model access, while most regulations give little guidance on reporting. They argue that this structured format would enhance transparency, help readers interpret results correctly, and build trust in AI governance.","feed_headline":"Most AI audit reports omit who evaluated and why","feed_subtitle":"A proposed 'audit card' template would require disclosure of access, conflicts, and review—making evaluations interpretable.","key_machinery":"The central object is the audit card template itself: a checklist organized around three principles (justification, assumptions, limitations) that cut across six features (auditor identity, evaluation scope, methodology, resource access, process integrity, and review mechanisms). The template converts diffuse norms about transparency into concrete disclosure questions, such as how auditors were selected, what conflicts of interest existed, what model access was granted, and when the evaluation becomes obsolete. These structured fields are what allow readers to assess whether an evaluation is credible and what it actually covers.","core_discovery":"On its own terms, the paper establishes that AI evaluations are sociotechnical processes whose findings are shaped by non-technical details, and that current reporting practice largely fails to document those details. It synthesizes prior literature into a disclosure template consisting of three principles and six features, then shows through manual scoring that real evaluation reports from developers and third parties are highly inconsistent: scope and procedure are usually covered, but integrity, resource access, and auditor identity are rarely disclosed, and no examined report states obsolescence criteria. The paper further finds that existing governance frameworks rarely require or recommend such disclosures. It concludes that audit cards can close this reporting gap and thereby enable meaningful interpretation of evaluation results.","pith_inferences":["The paper's central benefit claim could be tested directly: present readers with identical evaluation metrics either with or without audit-card context and measure whether their interpretation accuracy, confidence calibration, and trust judgments change.","A registry of audit cards, one per model evaluation, could create a public ledger that makes opinion shopping visible and enables longitudinal tracking of whether disclosure quality improves over time.","The template could be extended to cover downstream consumers' needs, since the paper notes that policymakers currently read reports at a high level; prioritization of fields per audience is suggested but not developed.","If disclosure burden proves high, a plausible resolution is tiered audit cards: minimal required fields for all evaluations and fuller disclosure for higher-stakes or pre-deployment audits."],"forward_implications":["Evaluation reports from different organizations would become directly comparable because the same contextual fields are filled in each time.","Regulators and users could judge auditor independence and conflicts of interest at a glance, reducing the risk of selective or cosmetic reporting.","Requiring obsolescence criteria would prevent outdated evaluations from being cited as current evidence about a model.","Governance frameworks, which currently specify that audits should happen but not how they should be reported, could adopt the audit card as a concrete compliance template.","If adopted broadly, audit cards would make the evaluation process itself auditable, shifting accountability from technical execution alone to reporting integrity."],"supporting_citations":[{"why":"Defines model cards, the prior reporting artifact that audit cards extend and differentiate from.","marker":"Mitchell et al. 2019"},{"why":"Supplies the case for third-party audit ecosystems, the governance context that makes audit reporting necessary.","marker":"Raji et al. 2022"},{"why":"Establishes that evaluations are a social science measurement challenge, the premise for focusing on context rather than only technical rigor.","marker":"Wallach et al. 2025"},{"why":"Provides the technical best practices for benchmark design that audit cards complement with reporting guidance.","marker":"Reuel et al. 2024b"},{"why":"Grounds the recommendation that evaluation scope, threat models, and capability-to-goal translations be reported.","marker":"Shevlane et al. 2023"},{"why":"Supplies the \"who audits the auditors\" concern that motivates the integrity and auditor-identity features.","marker":"Costanza-Chock, Raji, and Buolamwini 2022"},{"why":"Motivates the need for structured context by documenting the limits of transparency as a governance ideal.","marker":"Ananny and Crawford 2018"}],"fun_headline_variants":["AI audits rarely say who ran them","Proposed 'audit cards' would expose AI evaluation gaps","Why most AI audit reports miss key context","A template to make AI evaluations interpretable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes, without empirical evidence, that adding these disclosure fields to evaluation reports will actually improve how readers interpret and trust the results, rather than being ignored or merely adding reporting burden.","fun_headline_variants_meta":{"raw":{"variants":["AI audits rarely say who ran them","Proposed 'audit cards' would expose AI evaluation gaps","Why most AI audit reports miss key context","A template to make AI evaluations interpretable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000109,"raw_usage":{"total_tokens":1013,"prompt_tokens":869,"completion_tokens":144,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":86}},"tokens_in":485,"tokens_out":144,"duration_ms":2258,"temperature":1.0,"reasoning_tokens":86,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:58:09.920160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled study in which readers receive the same evaluation results with and without an audit card, then measure whether the card changes their interpretation accuracy, their confidence in conclusions, or their trust in the evaluator; if no measurable improvement appears, the paper's core benefit claim collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the case for third-party audit ecosystems, the governance context that makes audit reporting necessary."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the recommendation that evaluation scope, threat models, and capability-to-goal translations be reported."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the \"who audits the auditors\" concern that motivates the integrity and auditor-identity features."}],"review_version":1}