REVIEW 3 major objections 4 minor 12 references
This paper argues that giving an audit-writing LLM access to a model's internal evidence makes reports cite more but not be more valid; its shuffled-evidence control shows a report can be mostly right about the policy passage while citing i
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:21 UTC pith:OSEZYRNB
load-bearing objection A genuinely useful controlled study of evidence interfaces for LLM audit reports, but the human-scoring layer needs a condition-blind check before the headline numbers carry weight. the 3 major comments →
White Box Evidence Packages for Policy Audit Reports
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is that evidence uptake is not evidence validity. When the same policy passage, rubric, and auditor model are held fixed and only the evidence interface changes, shifting from black-box surface evidence to combined white-box evidence leaves average correctness nearly unchanged (4.60 vs 4.68) but cuts passage grounding from 4.52 to 3.25, lowers usefulness from 4.68 to 4.00, and raises evidence misuse from 1.00 to 2.50; misuse is higher in all 60 paired cases. The shuffled relevance control strengthens the point: 39 of 60 reports score at least 4 on correctness while also scoring at least 4 on misuse, meaning a report can be mostly right about th
What carries the argument
The central mechanism is the controlled evidence-interface comparison: the same policy passage, audit rubric, and auditor model are held fixed, and only the evidence package changes across ten conditions. The load-bearing instrument is the shuffled relevance control, which keeps the white-box package's format identical but draws the internal evidence from a different case; this isolates whether reports check relevance or merely cite plausible-looking evidence. The paper also uses a behavior-locked activation-patching diagnostic on forced-choice microtasks to test whether cited internal features are causally localized, with random-region thresholds and rendering controls.
Load-bearing premise
The load-bearing premise is that the five human reviewers' scores are unbiased ground truth for comparing the evidence conditions; the protocol does not state that reviewers were blind to which condition a report came from, and the small 20-case agreement check was low for correctness.
What would settle it
Run the same 60-case review with five fresh reviewers scoring reports with all condition-identifying cues removed, and check whether combined white-box evidence still weakens grounding in 53 of 60 cases and raises misuse in all 60; if the pattern mostly disappears, the central claim is not supported.
If this is right
- If white-box internal evidence is added without surface passage anchors, audit reports keep their correctness but lose grounding and increase evidence misuse; internal access should not be reported as transparency on its own.
- Citation volume should not be used as a quality metric: both the combined white-box condition and the shuffled relevance control produce heavy citation, but human review separates useful support from invalid support.
- Any audit workflow that evaluates LLM-generated evidence should include a relevance-broken shuffled control, because plausible-sounding reports can still misuse evidence from another case.
- Hybrid packages (passage surface evidence plus internal evidence) are the most useful interface tested on average, but the test does not isolate anchoring from package size; a capped hybrid condition is needed to decide whether the gain is causal.
Where Pith is reading between the lines
- Beyond the paper: if correctness and evidence misuse can diverge this cleanly, audit procedures that rely on an LLM's own cautionary notes or confidence scores are likely to miss exactly the dangerous failures; testing human auditors' ability to catch shuffled-evidence errors is a natural next step.
- Beyond the paper: the reported correlation between package size and misuse suggests a testable prediction: a volume-matched hybrid with few internal items should show a smaller usefulness gain, while a capped hybrid should separate anchoring from quantity.
- Beyond the paper: the narrow, prompt-sensitive causal localization found by patching suggests that readable interpretability labels may be reused by downstream models without corresponding mechanism-level support; this predicts that label-masked or numeric-only evidence would reduce overtrust.
- Beyond the paper: for governance, this reframes transparency from 'granting access' to 'designing checkable evidence'; one could require per-citation provenance and relevance metadata as standard report fields.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether LLM-generated policy audit reports can be distinguished from evidence-supported reports when the evidence interface, rather than the passage, rubric, or auditor model, is varied. It constructs 60 AGORA policy cases and uses a Gemma 2 2B target model and a Qwen 2.5 7B auditor to produce 600 structured reports under ten evidence conditions. Four primary conditions receive full human validation review on correctness, passage grounding, diagnostic usefulness, and evidence misuse: black-box surface evidence, combined white-box evidence, a hybrid package, and a shuffled white-box relevance control. The central claims are that internal evidence changes citation behavior but does not by itself improve validity; that the combined white-box condition weakens grounding and increases misuse relative to the surface baseline; that hybrid packages are promising but confounded by package size; and that the shuffled control can produce substantively plausible reports that cite irrelevant internal evidence. A residual-stream patching diagnostic is reported separately to characterize the narrow, prompt-sensitive causal localization.
Significance. If the main results hold, the paper makes a useful contribution: it operationalizes evidence packages as an object of audit design rather than assuming mechanistic interpretability outputs are self-validating. The strengths are real: gold briefs are written before reports or evidence packages are seen; the shuffled control is a genuine negative control for relevance; and the artifact repository is positioned for reproducibility. The paper also carefully disclaims causal claims about hybrid packages and about mechanistic localization. The study is therefore valuable as a controlled framework and a cautionary result. The limitation that currently gates the significance is the reliability of the human validation scores, since the headline quantitative conclusions are carried by those scores.
major comments (3)
- [§4 and Appendix E] The validation-scoring stage is not demonstrated to be condition-blind. The protocol explicitly blinds gold-brief writing, but the diagnostic review stage does not state that reviewers were masked to the evidence condition when scoring the 240 reports. Since reports cite tool-generated evidence entries, the condition is visually identifiable from the report itself. If expectations about white-box misuse leaked into scoring, the headline differences in Table 3—grounding 3.25 vs. 4.52, usefulness 4.00 vs. 4.68, misuse 2.50 vs. 1.00—and the all-60-case misuse increase in Table 4 could be inflated. This is load-bearing for the central claim. Please add a masked, fully crossed rescoring of at least a subset (e.g., redacted or normalized citations) and report whether the paired differences persist; also state how the five reviewers were assigned to the 240 reports.
- [Table 8 and RQ2] Inter-rater agreement is low on the axes that carry the central claims: quadratic weighted kappa is 0.25 for correctness, 0.50 for grounding, and 0.46 for usefulness, with only misuse high (0.88). The agreement diagnostics are computed only on the original 20-case subset with two reviewers, and full-set rater-level agreement is not reported. The paper acknowledges this in the limitations section, but the quantitative claims in RQ2 are paired mean differences on these same axes. Low agreement does not disprove the direction, but it means the magnitudes could shift materially under a different scoring draw. Please report full-set agreement, per-rater means, and rater-level paired analyses, or a bootstrap treating rater as a random effect, so the paired claims can be evaluated.
- [§5 RQ3, Table 3] The misuse comparison between the combined/hybrid interfaces and the surface baseline is confounded with evidence quantity: the hybrid package contains 22.97 evidence items on average versus 3.27 for surface evidence, and the combined package contains 19.70. The paper notes the Spearman correlation of 0.455 between package size and misuse and disclaims causality, but the statement that misuse is higher in all 60 cases is presented as an interface effect. The shuffled control helps separate volume from irrelevance because its misuse score is 5.00 despite citation volume similar to the combined condition. Still, a volume-controlled condition or an explicit regression controlling for item count would materially strengthen the quantitative misuse claims. As written, the all-60 misuse increase should be read as an interface-plus-volume effect, not as pure interface effect.
minor comments (4)
- [Table 4] The B/T/W formatting appears to omit counts or rows. For example, the combined-white-box Correctness row shows only two counts (39/13) under a column specified as better/tied/worse, and the hybrid section has no Grounding row even though the text reports grounding for hybrid. Please check that each B/T/W triple is complete and all four axes are listed.
- [Appendix C / Table 6] The residual patching diagnostic uses only 6 validation pairs per family in the main layer plot. The caveats about prompt-shell and answer-order sensitivity are appropriately stated, but the small n should be repeated in the main text near Figure 4, not only in the appendix, to prevent over-reading of layer 7/8 localization.
- [Section 3, Table 1] The naming of the shuffled condition as 'white box relevance control' is clear, but the reader must infer that the shuffled package is hidden from the auditor. Please state explicitly in Table 1 or its caption that the auditor is not informed that the evidence belongs to another case.
- [Abstract / Intro] The abstract says 'Five human reviewers evaluate the primary interfaces,' but the validation stage includes gold writing, verification, diagnostic scoring, and agreement checks across five reviewers. Clarify how many reviewers actually scored each of the 240 reports, since the number of scorers per report is not stated in the main text.
Circularity Check
No significant circularity; the central results are externally grounded and the paper explicitly disclaims the one confounded advantage.
full rationale
No load-bearing step reduces to its own inputs. Evidence packages are generated deterministically from fixed interpretability tools applied to Gemma 2 2B, and the auditor is a fixed Qwen 2.5 7B model; no parameter is fitted to the validation outcomes. The validation layer is external: gold briefs are written before seeing reports or evidence packages (Section 4, Appendix E), and the four axes are human judgments against those briefs. The shuffled relevance control is a genuine negative control: it holds package format fixed while breaking case relevance, and the central warning is the conjunction of gold-brief correctness >=4 with misuse >=4, where correctness is not defined by the shuffle. The paper also distinguishes structural citation validity (V_i,c in Eq. 5) from substantive relevance, and notes that perfect citation validity can coexist with high misuse, so no definitional collapse occurs. The residual stream patching diagnostic is behavior-locked with its own gate and random-region threshold, and the paper explicitly narrows the claim to a rendering-sensitive late answer-readout signal. The hybrid usefulness advantage is explicitly labeled as confounded with package size (22.97 vs 3.27 items), so it is not presented as a causal prediction. The acknowledged limitations—reviewer blinding not stated for the scoring stage and agreement measured only on 20 cases—are measurement-validity concerns, not circularity; they do not make any reported quantity equal to an input by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- Causal gate threshold for residual patching =
family random-region 95th percentile + 0.05 (deontic direct canonical: 0.0935)
- Steering scale alpha =
0.35 in Eq. (2)
- Surface evidence cue lexicon and top-5 sentence selection =
20 audit cue terms; five highest-scoring sentences
- Pre-patching behavior gate =
rendered accuracy >=0.70 and >=20 stable pairs out of 30
axioms (7)
- domain assumption AGORA policy passages and blind gold briefs are a valid operationalization of passage-anchored policy audits.
- domain assumption Qwen 2.5 7B structured-output auditor behavior is informative about LLM-assisted audit report writing.
- domain assumption Gemma 2 2B residual-stream evidence from Gemma Scope SAEs, logit lens, and steering directions is meaningful 'white-box evidence' for an audit report.
- domain assumption Five reviewers' gold briefs and scores are valid ground truth for correctness, grounding, usefulness, and misuse.
- domain assumption Residual activation patching recovery R = (m_patched - m_corrupted)/(m_clean - m_corrupted) measures causal localization of a behavior.
- domain assumption Report normalization/cleanup changes only formatting and citation consistency, not substantive findings.
- standard math Exact sign tests, Wilcoxon signed-rank approximations, and bootstrap intervals are appropriate descriptive statistics for ordinal 1-5 scores.
read the original abstract
As AI governance moves from benchmark scores toward auditable oversight, a central question is how reviewers can tell whether an LLM-generated audit report is actually supported by evidence. This paper studies that question in passage-anchored policy audits, where a report must interpret a given policy passage and cite evidence for its claims. We introduce a controlled evaluation framework that holds the passage, rubric, and auditor model fixed while changing only the evidence interface supplied to the auditor. Across 60 AGORA policy cases, we generate 600 structured reports under ten evidence conditions, including passage-based evidence, internal model evidence, a hybrid package, and a shuffled control that preserves evidence format while breaking case relevance. Five human reviewers evaluate the primary interfaces for correctness, passage grounding, diagnostic usefulness, and evidence misuse. The results show that internal evidence changes how reports cite and reason about evidence, but more internal citations do not by themselves make a report more valid. A white-box diagnostic explains the failure mode: causal localization is narrow, while reports readily reuse broader readable labels and token directions. The hybrid interface is the most useful on average, while the shuffled control exposes a key governance risk: reports can sound substantively plausible while citing irrelevant internal evidence. This study reframes internal model access as an evidence design problem for audit workflows, rather than as a guarantee of transparency.
Figures
Reference graph
Works this paper leans on
-
[5]
URL https://arxiv.org/abs/2401.14446
doi: 10.1145/3630106.3659037. URL https://arxiv.org/abs/2401.14446. Cen, S. H. and Alur, R. From transparency to accountability and back: A discussion of access and evidence in AI auditing.arXiv preprint arXiv:2410.04772,
-
[6]
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L
URL https://arxiv.org/abs/2410.04772. Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L. Sparse autoencoders find highly interpretable features in language models. InInternational Conference on Learning Representations,
-
[11]
URL https:// arxiv.org/abs/2402.17861. 11 White Box Evidence Packages for Policy Audit Reports Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., Anderljung, M., Garfinkel, B., Heim, L., Trask, A., Mukobi, G., Schaeffer, R., Baker, M., Hooker, S., So- laiman, I., Luccioni, A. S., Rajkumar, N...
-
[12]
Sheshadri, A., Ewart, A., Fronsdal, K., Gupta, I., Bowman, S
URL https:// arxiv.org/abs/2310.13548. Sheshadri, A., Ewart, A., Fronsdal, K., Gupta, I., Bowman, S. R., Price, S., Marks, S., and Wang, R. Auditbench: Evaluating alignment auditing techniques on models with hidden behaviors.arXiv preprint arXiv:2602.22755,
-
[14]
URL https://arxiv. org/abs/2601.18405. Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N., McDougall, C., MacDi- armid, M., Freeman, C. D., Sumers, T., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T. Scal- ing monosemanticity: Extrac...
-
[16]
Yoran, O., Wolfson, T., Ram, O., and Berant, J
URL https://arxiv.org/abs/2404.04500. Yoran, O., Wolfson, T., Ram, O., and Berant, J. Making retrieval-augmented language models robust to irrelevant context.arXiv preprint arXiv:2310.01558,
-
[17]
URL https://arxiv.org/abs/2310.01558. Zhang, F. and Nanda, N. Towards best practices of activation patching in language models. arXiv preprint,
-
[22]
For each pair, we replace the corrupted prompt residual vector at a candidate layer and token role with the corresponding clean prompt vector
The main layer plot uses the six validation pairs per family under the direct shell and canonical answer order. For each pair, we replace the corrupted prompt residual vector at a candidate layer and token role with the corresponding clean prompt vector. Gemma 2 2B contributes 26 layers, and the diagnostic tests changed, actor, predicate or modal, thresho...
1968
-
[2013]
Activation oracles: Train- ing and evaluating LLMs as general-purpose activation explainers
Karvonen, A., Chua, J., Dumas, C., Fraser-Taliente, K., Kan- tamneni, S., Minder, J., Ong, E., Sen Sharma, A., Wen, D., Evans, O., and Marks, S. Activation oracles: Train- ing and evaluating LLMs as general-purpose activation explainers. arXiv preprint arXiv:2512.15674, 2025a. Karvonen, A., Rager, C., Marks, S., et al. SAEBench: A comprehensive benchmark ...
-
[2022]
URL https:// arxiv.org/abs/2207.01482. Belrose, N., Furman, H., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J. Elicit- ing latent predictions from transformers with the tuned lens.arXiv preprint arXiv:2303.08112,
-
[2025]
URL https://arxiv. org/abs/2505.11577. Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobb- hahn, M., Sharkey, L., Krishna, S., V on Hagen, M., Alberti, S., Chan, A., Sun, Q., Gerovitch, M., Bau, D., Tegmark, M., Krueger, D., and Hadfield-Menell, D. Black-box access is insufficient for rigorou...
-
[2026]
URL https://arxiv. org/abs/2601.16398. Harack, B., Trager, R. F., Reuel, A., Manheim, D., Brundage, M., Aarne, O., Scher, A., Pan, Y ., Xiao, J., Loke, K., Adan, S. N., Bas, G., Caputo, N. A., Morse, J. C., Ahuja, J., Duan, I., Egan, J., Bucknall, B., Rosen, B., Araujo, R., Boulanin, V ., Lall, R., Barez, F., Alvira, S., Katzke, C., Atamli, A., and Awad, ...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.