Pith. sign in

REVIEW 3 major objections 4 minor 12 references

This paper argues that giving an audit-writing LLM access to a model's internal evidence makes reports cite more but not be more valid; its shuffled-evidence control shows a report can be mostly right about the policy passage while citing i

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 07:21 UTC pith:OSEZYRNB

load-bearing objection A genuinely useful controlled study of evidence interfaces for LLM audit reports, but the human-scoring layer needs a condition-blind check before the headline numbers carry weight. the 3 major comments →

arxiv 2607.21462 v1 pith:OSEZYRNB submitted 2026-07-23 cs.CY

White Box Evidence Packages for Policy Audit Reports

classification cs.CY
keywords AI governancepolicy audit reportswhite-box evidenceevidence interfaceevidence misuseshuffled relevance controlLLM auditpassage grounding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether giving an LLM that writes policy audit reports access to a model's internal evidence makes those reports more trustworthy. It tries to show that the answer is no in general: combined white-box evidence increases citation volume but weakens grounding, lowers usefulness, and increases evidence misuse relative to a passage-only surface baseline. The key demonstration is a shuffled relevance control in which internal evidence is drawn from a different case while keeping the same format; 39 of 60 reports remain mostly correct about the passage while heavily misusing the irrelevant evidence. The paper also argues that hybrid packages (surface passage anchors plus internal evidence) are a promising interface but that their advantage is confounded with package size, and that a stricter causal diagnostic localizes only a narrow, prompt-sensitive readout signal. The contribution is a reproducible evaluation framework for treating internal access as an evidence-design problem, not a transparency guarantee.

Core claim

On the paper's own terms, the central discovery is that evidence uptake is not evidence validity. When the same policy passage, rubric, and auditor model are held fixed and only the evidence interface changes, shifting from black-box surface evidence to combined white-box evidence leaves average correctness nearly unchanged (4.60 vs 4.68) but cuts passage grounding from 4.52 to 3.25, lowers usefulness from 4.68 to 4.00, and raises evidence misuse from 1.00 to 2.50; misuse is higher in all 60 paired cases. The shuffled relevance control strengthens the point: 39 of 60 reports score at least 4 on correctness while also scoring at least 4 on misuse, meaning a report can be mostly right about th

What carries the argument

The central mechanism is the controlled evidence-interface comparison: the same policy passage, audit rubric, and auditor model are held fixed, and only the evidence package changes across ten conditions. The load-bearing instrument is the shuffled relevance control, which keeps the white-box package's format identical but draws the internal evidence from a different case; this isolates whether reports check relevance or merely cite plausible-looking evidence. The paper also uses a behavior-locked activation-patching diagnostic on forced-choice microtasks to test whether cited internal features are causally localized, with random-region thresholds and rendering controls.

Load-bearing premise

The load-bearing premise is that the five human reviewers' scores are unbiased ground truth for comparing the evidence conditions; the protocol does not state that reviewers were blind to which condition a report came from, and the small 20-case agreement check was low for correctness.

What would settle it

Run the same 60-case review with five fresh reviewers scoring reports with all condition-identifying cues removed, and check whether combined white-box evidence still weakens grounding in 53 of 60 cases and raises misuse in all 60; if the pattern mostly disappears, the central claim is not supported.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If white-box internal evidence is added without surface passage anchors, audit reports keep their correctness but lose grounding and increase evidence misuse; internal access should not be reported as transparency on its own.
  • Citation volume should not be used as a quality metric: both the combined white-box condition and the shuffled relevance control produce heavy citation, but human review separates useful support from invalid support.
  • Any audit workflow that evaluates LLM-generated evidence should include a relevance-broken shuffled control, because plausible-sounding reports can still misuse evidence from another case.
  • Hybrid packages (passage surface evidence plus internal evidence) are the most useful interface tested on average, but the test does not isolate anchoring from package size; a capped hybrid condition is needed to decide whether the gain is causal.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if correctness and evidence misuse can diverge this cleanly, audit procedures that rely on an LLM's own cautionary notes or confidence scores are likely to miss exactly the dangerous failures; testing human auditors' ability to catch shuffled-evidence errors is a natural next step.
  • Beyond the paper: the reported correlation between package size and misuse suggests a testable prediction: a volume-matched hybrid with few internal items should show a smaller usefulness gain, while a capped hybrid should separate anchoring from quantity.
  • Beyond the paper: the narrow, prompt-sensitive causal localization found by patching suggests that readable interpretability labels may be reused by downstream models without corresponding mechanism-level support; this predicts that label-masked or numeric-only evidence would reduce overtrust.
  • Beyond the paper: for governance, this reframes transparency from 'granting access' to 'designing checkable evidence'; one could require per-citation provenance and relevance metadata as standard report fields.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies whether LLM-generated policy audit reports can be distinguished from evidence-supported reports when the evidence interface, rather than the passage, rubric, or auditor model, is varied. It constructs 60 AGORA policy cases and uses a Gemma 2 2B target model and a Qwen 2.5 7B auditor to produce 600 structured reports under ten evidence conditions. Four primary conditions receive full human validation review on correctness, passage grounding, diagnostic usefulness, and evidence misuse: black-box surface evidence, combined white-box evidence, a hybrid package, and a shuffled white-box relevance control. The central claims are that internal evidence changes citation behavior but does not by itself improve validity; that the combined white-box condition weakens grounding and increases misuse relative to the surface baseline; that hybrid packages are promising but confounded by package size; and that the shuffled control can produce substantively plausible reports that cite irrelevant internal evidence. A residual-stream patching diagnostic is reported separately to characterize the narrow, prompt-sensitive causal localization.

Significance. If the main results hold, the paper makes a useful contribution: it operationalizes evidence packages as an object of audit design rather than assuming mechanistic interpretability outputs are self-validating. The strengths are real: gold briefs are written before reports or evidence packages are seen; the shuffled control is a genuine negative control for relevance; and the artifact repository is positioned for reproducibility. The paper also carefully disclaims causal claims about hybrid packages and about mechanistic localization. The study is therefore valuable as a controlled framework and a cautionary result. The limitation that currently gates the significance is the reliability of the human validation scores, since the headline quantitative conclusions are carried by those scores.

major comments (3)
  1. [§4 and Appendix E] The validation-scoring stage is not demonstrated to be condition-blind. The protocol explicitly blinds gold-brief writing, but the diagnostic review stage does not state that reviewers were masked to the evidence condition when scoring the 240 reports. Since reports cite tool-generated evidence entries, the condition is visually identifiable from the report itself. If expectations about white-box misuse leaked into scoring, the headline differences in Table 3—grounding 3.25 vs. 4.52, usefulness 4.00 vs. 4.68, misuse 2.50 vs. 1.00—and the all-60-case misuse increase in Table 4 could be inflated. This is load-bearing for the central claim. Please add a masked, fully crossed rescoring of at least a subset (e.g., redacted or normalized citations) and report whether the paired differences persist; also state how the five reviewers were assigned to the 240 reports.
  2. [Table 8 and RQ2] Inter-rater agreement is low on the axes that carry the central claims: quadratic weighted kappa is 0.25 for correctness, 0.50 for grounding, and 0.46 for usefulness, with only misuse high (0.88). The agreement diagnostics are computed only on the original 20-case subset with two reviewers, and full-set rater-level agreement is not reported. The paper acknowledges this in the limitations section, but the quantitative claims in RQ2 are paired mean differences on these same axes. Low agreement does not disprove the direction, but it means the magnitudes could shift materially under a different scoring draw. Please report full-set agreement, per-rater means, and rater-level paired analyses, or a bootstrap treating rater as a random effect, so the paired claims can be evaluated.
  3. [§5 RQ3, Table 3] The misuse comparison between the combined/hybrid interfaces and the surface baseline is confounded with evidence quantity: the hybrid package contains 22.97 evidence items on average versus 3.27 for surface evidence, and the combined package contains 19.70. The paper notes the Spearman correlation of 0.455 between package size and misuse and disclaims causality, but the statement that misuse is higher in all 60 cases is presented as an interface effect. The shuffled control helps separate volume from irrelevance because its misuse score is 5.00 despite citation volume similar to the combined condition. Still, a volume-controlled condition or an explicit regression controlling for item count would materially strengthen the quantitative misuse claims. As written, the all-60 misuse increase should be read as an interface-plus-volume effect, not as pure interface effect.
minor comments (4)
  1. [Table 4] The B/T/W formatting appears to omit counts or rows. For example, the combined-white-box Correctness row shows only two counts (39/13) under a column specified as better/tied/worse, and the hybrid section has no Grounding row even though the text reports grounding for hybrid. Please check that each B/T/W triple is complete and all four axes are listed.
  2. [Appendix C / Table 6] The residual patching diagnostic uses only 6 validation pairs per family in the main layer plot. The caveats about prompt-shell and answer-order sensitivity are appropriately stated, but the small n should be repeated in the main text near Figure 4, not only in the appendix, to prevent over-reading of layer 7/8 localization.
  3. [Section 3, Table 1] The naming of the shuffled condition as 'white box relevance control' is clear, but the reader must infer that the shuffled package is hidden from the auditor. Please state explicitly in Table 1 or its caption that the auditor is not informed that the evidence belongs to another case.
  4. [Abstract / Intro] The abstract says 'Five human reviewers evaluate the primary interfaces,' but the validation stage includes gold writing, verification, diagnostic scoring, and agreement checks across five reviewers. Clarify how many reviewers actually scored each of the 240 reports, since the number of scorers per report is not stated in the main text.

Circularity Check

0 steps flagged

No significant circularity; the central results are externally grounded and the paper explicitly disclaims the one confounded advantage.

full rationale

No load-bearing step reduces to its own inputs. Evidence packages are generated deterministically from fixed interpretability tools applied to Gemma 2 2B, and the auditor is a fixed Qwen 2.5 7B model; no parameter is fitted to the validation outcomes. The validation layer is external: gold briefs are written before seeing reports or evidence packages (Section 4, Appendix E), and the four axes are human judgments against those briefs. The shuffled relevance control is a genuine negative control: it holds package format fixed while breaking case relevance, and the central warning is the conjunction of gold-brief correctness >=4 with misuse >=4, where correctness is not defined by the shuffle. The paper also distinguishes structural citation validity (V_i,c in Eq. 5) from substantive relevance, and notes that perfect citation validity can coexist with high misuse, so no definitional collapse occurs. The residual stream patching diagnostic is behavior-locked with its own gate and random-region threshold, and the paper explicitly narrows the claim to a rendering-sensitive late answer-readout signal. The hybrid usefulness advantage is explicitly labeled as confounded with package size (22.97 vs 3.27 items), so it is not presented as a causal prediction. The acknowledged limitations—reviewer blinding not stated for the scoring stage and agreement measured only on 20 cases—are measurement-validity concerns, not circularity; they do not make any reported quantity equal to an input by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 7 axioms · 0 invented entities

The central claims rest on the human validation layer and on the interpretation of interpretability-tool outputs as evidence. The key free parameters are hand-set thresholds in the residual patching diagnostic and steering evidence. No new physical or mathematical entities are introduced.

free parameters (4)
  • Causal gate threshold for residual patching = family random-region 95th percentile + 0.05 (deontic direct canonical: 0.0935)
    Hand-chosen criterion for whether a layer/token region 'passes'; determines the headline narrow-localization claim in Section 5 and Table 6.
  • Steering scale alpha = 0.35 in Eq. (2)
    Fixed multiplier for steering evidence; no sensitivity analysis or prior derivation is provided.
  • Surface evidence cue lexicon and top-5 sentence selection = 20 audit cue terms; five highest-scoring sentences
    Defines the black-box surface baseline; hand-authored and no robustness check is reported.
  • Pre-patching behavior gate = rendered accuracy >=0.70 and >=20 stable pairs out of 30
    Threshold for including microtask families in the residual patching diagnostic; affects which families are analyzed.
axioms (7)
  • domain assumption AGORA policy passages and blind gold briefs are a valid operationalization of passage-anchored policy audits.
    Section 4 and Appendix E; all conclusions are about this task family, and the paper limits its claims accordingly.
  • domain assumption Qwen 2.5 7B structured-output auditor behavior is informative about LLM-assisted audit report writing.
    Section 4; only one auditor model is used, acknowledged as a limitation in Section 7.
  • domain assumption Gemma 2 2B residual-stream evidence from Gemma Scope SAEs, logit lens, and steering directions is meaningful 'white-box evidence' for an audit report.
    Section 4 and Appendix B; SAE label quality, logit lens noise, and steering direction quality are not calibrated (acknowledged in Section 7).
  • domain assumption Five reviewers' gold briefs and scores are valid ground truth for correctness, grounding, usefulness, and misuse.
    Appendix E; agreement is low on correctness (QWK 0.25) on the 20-case subset, so this is a load-bearing assumption.
  • domain assumption Residual activation patching recovery R = (m_patched - m_corrupted)/(m_clean - m_corrupted) measures causal localization of a behavior.
    Appendix C; standard activation-patching assumption from Vig et al. 2020 and Zhang & Nanda 2024, but interpreting passing regions as a narrow readout is an additional interpretive step.
  • domain assumption Report normalization/cleanup changes only formatting and citation consistency, not substantive findings.
    Section 4 and Appendix K; reviewers score cleaned reports, so if cleanup altered findings the human scores would be affected.
  • standard math Exact sign tests, Wilcoxon signed-rank approximations, and bootstrap intervals are appropriate descriptive statistics for ordinal 1-5 scores.
    Section 4 and Appendix E; used for paired comparisons; no parametric claims are made.

pith-pipeline@v1.3.0-alltime-deepseek · 16723 in / 16647 out tokens · 155403 ms · 2026-08-01T07:21:42.046229+00:00 · methodology

0 comments
read the original abstract

As AI governance moves from benchmark scores toward auditable oversight, a central question is how reviewers can tell whether an LLM-generated audit report is actually supported by evidence. This paper studies that question in passage-anchored policy audits, where a report must interpret a given policy passage and cite evidence for its claims. We introduce a controlled evaluation framework that holds the passage, rubric, and auditor model fixed while changing only the evidence interface supplied to the auditor. Across 60 AGORA policy cases, we generate 600 structured reports under ten evidence conditions, including passage-based evidence, internal model evidence, a hybrid package, and a shuffled control that preserves evidence format while breaking case relevance. Five human reviewers evaluate the primary interfaces for correctness, passage grounding, diagnostic usefulness, and evidence misuse. The results show that internal evidence changes how reports cite and reason about evidence, but more internal citations do not by themselves make a report more valid. A white-box diagnostic explains the failure mode: causal localization is narrow, while reports readily reuse broader readable labels and token directions. The hybrid interface is the most useful on average, while the shuffled control exposes a key governance risk: reports can sound substantively plausible while citing irrelevant internal evidence. This study reframes internal model access as an evidence design problem for audit workflows, rather than as a guarantee of transparency.

Figures

Figures reproduced from arXiv: 2607.21462 by Seunghyun Yoo.

Figure 1
Figure 1. Figure 1: Evidence package audit design from policy passage to target model evidence, evidence package, auditor report, and human review. Internal states are summarized through SAE, logit lens, steering, and activation explanation surrogate records. The unit of comparison is the evidence interface, not a retrieval system, and the shuffled control preserves interface form while breaking case relevance. The displayed … view at source ↗
Figure 2
Figure 2. Figure 2: Validation review scores for black box surface evidence, combined white box evidence, hybrid surface and white box evidence, and the shuffled white box relevance control over the 60 case evaluation set. Scores use a 1 to 5 scale; higher is better except evidence misuse. 4 5 6 7 8 9 10 evidence citations per report 3.2 3.4 3.6 3.8 4.0 4.2 4.4 4.6 4.8 mean human score Human correctness 4 5 6 7 8 9 10 evidenc… view at source ↗
Figure 3
Figure 3. Figure 3: Citation volume plotted against validation correctness and evidence misuse. The x axis reports mean evidence citations per report over the 60 case run; the y axes report validation review scores over the same cases. 7 [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Residual stream activation patching diagnostic for deontic force under the direct prompt shell and canonical answer order. Bars show mean patch recovery by layer at the final prompt token; error bars show 95% bootstrap intervals, and the dotted line marks the empirical causal gate. 8 [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Structural metrics across all ten evidence conditions. Single tool conditions change citation behavior and formatting burden, but these measures are not validation scores. Logit lens evidence and raw AutoInterp sparse autoencoder evidence show high citation volume, while the shuffled white box relevance control shows that a combined white box shaped package can be cited heavily even when relevance is broke… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

12 extracted references · 8 linked inside Pith

  1. [5]

    URL https://arxiv.org/abs/2401.14446

    doi: 10.1145/3630106.3659037. URL https://arxiv.org/abs/2401.14446. Cen, S. H. and Alur, R. From transparency to accountability and back: A discussion of access and evidence in AI auditing.arXiv preprint arXiv:2410.04772,

  2. [6]

    Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L

    URL https://arxiv.org/abs/2410.04772. Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L. Sparse autoencoders find highly interpretable features in language models. InInternational Conference on Learning Representations,

  3. [11]

    URL https:// arxiv.org/abs/2402.17861. 11 White Box Evidence Packages for Policy Audit Reports Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., Anderljung, M., Garfinkel, B., Heim, L., Trask, A., Mukobi, G., Schaeffer, R., Baker, M., Hooker, S., So- laiman, I., Luccioni, A. S., Rajkumar, N...

  4. [12]

    Sheshadri, A., Ewart, A., Fronsdal, K., Gupta, I., Bowman, S

    URL https:// arxiv.org/abs/2310.13548. Sheshadri, A., Ewart, A., Fronsdal, K., Gupta, I., Bowman, S. R., Price, S., Marks, S., and Wang, R. Auditbench: Evaluating alignment auditing techniques on models with hidden behaviors.arXiv preprint arXiv:2602.22755,

  5. [14]

    org/abs/2601.18405

    URL https://arxiv. org/abs/2601.18405. Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N., McDougall, C., MacDi- armid, M., Freeman, C. D., Sumers, T., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T. Scal- ing monosemanticity: Extrac...

  6. [16]

    Yoran, O., Wolfson, T., Ram, O., and Berant, J

    URL https://arxiv.org/abs/2404.04500. Yoran, O., Wolfson, T., Ram, O., and Berant, J. Making retrieval-augmented language models robust to irrelevant context.arXiv preprint arXiv:2310.01558,

  7. [17]

    Zhang, F

    URL https://arxiv.org/abs/2310.01558. Zhang, F. and Nanda, N. Towards best practices of activation patching in language models. arXiv preprint,

  8. [22]

    For each pair, we replace the corrupted prompt residual vector at a candidate layer and token role with the corresponding clean prompt vector

    The main layer plot uses the six validation pairs per family under the direct shell and canonical answer order. For each pair, we replace the corrupted prompt residual vector at a candidate layer and token role with the corresponding clean prompt vector. Gemma 2 2B contributes 26 layers, and the diagnostic tests changed, actor, predicate or modal, thresho...

  9. [2013]

    Activation oracles: Train- ing and evaluating LLMs as general-purpose activation explainers

    Karvonen, A., Chua, J., Dumas, C., Fraser-Taliente, K., Kan- tamneni, S., Minder, J., Ong, E., Sen Sharma, A., Wen, D., Evans, O., and Marks, S. Activation oracles: Train- ing and evaluating LLMs as general-purpose activation explainers. arXiv preprint arXiv:2512.15674, 2025a. Karvonen, A., Rager, C., Marks, S., et al. SAEBench: A comprehensive benchmark ...

  10. [2022]

    Belrose, N., Furman, H., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J

    URL https:// arxiv.org/abs/2207.01482. Belrose, N., Furman, H., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J. Elicit- ing latent predictions from transformers with the tuned lens.arXiv preprint arXiv:2303.08112,

  11. [2025]

    org/abs/2505.11577

    URL https://arxiv. org/abs/2505.11577. Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobb- hahn, M., Sharkey, L., Krishna, S., V on Hagen, M., Alberti, S., Chan, A., Sun, Q., Gerovitch, M., Bau, D., Tegmark, M., Krueger, D., and Hadfield-Menell, D. Black-box access is insufficient for rigorou...

  12. [2026]

    org/abs/2601.16398

    URL https://arxiv. org/abs/2601.16398. Harack, B., Trager, R. F., Reuel, A., Manheim, D., Brundage, M., Aarne, O., Scher, A., Pan, Y ., Xiao, J., Loke, K., Adan, S. N., Bas, G., Caputo, N. A., Morse, J. C., Ahuja, J., Duan, I., Egan, J., Bucknall, B., Rosen, B., Araujo, R., Boulanin, V ., Lall, R., Barez, F., Alvira, S., Katzke, C., Atamli, A., and Awad, ...