{"id":"03260b19-3248-413b-ac84-d902e9340ca0","arxiv_id":"2607.21462","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In a 60-case controlled audit study, adding white-box model evidence to an LLM auditor increased citation volume but weakened passage grounding and raised evidence misuse, while a shuffled control showed reports can sound plausible while citing irrelevant internal evidence.","lead":"This paper tests whether giving an LLM auditor extra 'white-box' internal model evidence makes its policy audit reports more grounded in the policy text. It finds that more internal citations do not mean more valid reports, and that even shuffled, irrelevant internal evidence can yield plausible-sounding but mis-cited audits.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation scoring may not be condition-blind and inter-rater agreement is low on the axes carrying the central claim; the reported interface gaps need a masked, fully crossed rescoring to be load-bearing.","rationale":"The paper is methodologically stronger than most interpretability-tool evaluations: it uses a fixed passage, rubric, and auditor; blind gold briefs; 60 paired cases; and a shuffled relevance control that isolates format from relevance. The claim that internal evidence changes citation behavior is well supported by structural data and does not depend on human scoring. The claim that more citations do not imply validity is also conceptually well motivated and supported by the shuffled-control design. However, the quantitative magnitudes that carry the paper's central message—grounding loss, usefulness loss, and misuse increase—are human-validation scores. The reader correctly identified the weakest point: scoring is not stated to be condition-blind, and the only reported agreement is low to moderate on the axes that matter most, measured on a 20-case subset. This is not an accusation of bias; it is a measurement-validity concern. If the effects are real, a fully crossed condition-masked rescoring should preserve them. If not, the central empirical claims about grounding and misuse are not yet secure. The hybrid advantage is explicitly acknowledged as confounded with package size, so it is not the load-bearing issue. The shuffled-control misuse finding benefits from high QWK (0.88), but the 39/60 threshold also depends on correctness, which has QWK 0.25. Thus the same missing evidence—full-set, condition-blind, rater-modeled scores—undermines both main quantitative results. The verdict should remain CONDITIONAL: the framework and negative control are valuable, but the headline effect sizes require the proposed measurement check before the claims should be accepted as stable.","tokens_in":17217,"tokens_out":3810,"duration_ms":43808,"concrete_test":"Re-score all 240 reports under a condition-masked protocol: render every cited evidence entry as a neutral identifier (e.g., E1, E2, ...) with tool names and package labels removed, randomize report order, and have all five reviewers score every report against the gold briefs. Then fit an ordinal mixed-effects model with rater random intercepts and condition fixed effects, and compute full-set quadratic weighted kappa. The central claim should be considered supported only if (a) the combined-white-box grounding deficit and the misuse gap remain with confidence intervals excluding zero, and (b) the masked full-set QWK for grounding and misuse does not fall below the Table 8 values. If the gaps collapse under masking, the headline result is a scoring-expectation artifact rather than an evidence-interface effect.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central result rests on the 240 human validation scores: combined white-box evidence weakens grounding (3.25 vs. 4.52) and increases misuse (2.50 vs. 1.00), and the shuffled control yields 39/60 reports with correctness >=4 and misuse >=4. The protocol (Section 4 and Appendix E) blinds gold-brief writing to reports and evidence, but it does not state that the five reviewers were blind to evidence condition when scoring the reports. Reports reveal their condition through cited tool evidence, so reviewer expectations could systematically inflate the misuse and grounding differences. The reported agreement is limited to the original 20-case subset with two reviewers (Table 8): QWK is only 0.25 for correctness, 0.50 for grounding, 0.46 for usefulness, and high only for misuse (0.88). Full-set rater-level agreement and the assignment of reviewers to the 240 reports are not reported. Since the paper's own limitations section acknowledges that agreement was not measured on the full set, this is an explicitly open validity threat. If condition expectations leaked into scoring, the paired differences—especially the all-60-case misuse increase—could be inflated, and the shuffled-control count could shift materially. The conclusion that citation uptake is not evidence validity is plausible, but its quantitative support depends on measurement stability that has not yet been demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether LLM-generated policy audit reports can be distinguished from evidence-supported reports when the evidence interface, rather than the passage, rubric, or auditor model, is varied. It constructs 60 AGORA policy cases and uses a Gemma 2 2B target model and a Qwen 2.5 7B auditor to produce 600 structured reports under ten evidence conditions. Four primary conditions receive full human validation review on correctness, passage grounding, diagnostic usefulness, and evidence misuse: black-box surface evidence, combined white-box evidence, a hybrid package, and a shuffled white-box relevance control. The central claims are that internal evidence changes citation behavior but does not by itself improve validity; that the combined white-box condition weakens grounding and increases misuse relative to the surface baseline; that hybrid packages are promising but confounded by package size; and that the shuffled control can produce substantively plausible reports that cite irrelevant internal evidence. A residual-stream patching diagnostic is reported separately to characterize the narrow, prompt-sensitive causal localization.","tokens_in":17445,"tokens_out":5517,"duration_ms":60099,"significance":"If the main results hold, the paper makes a useful contribution: it operationalizes evidence packages as an object of audit design rather than assuming mechanistic interpretability outputs are self-validating. The strengths are real: gold briefs are written before reports or evidence packages are seen; the shuffled control is a genuine negative control for relevance; and the artifact repository is positioned for reproducibility. The paper also carefully disclaims causal claims about hybrid packages and about mechanistic localization. The study is therefore valuable as a controlled framework and a cautionary result. The limitation that currently gates the significance is the reliability of the human validation scores, since the headline quantitative conclusions are carried by those scores.","major_comments":[{"comment":"The validation-scoring stage is not demonstrated to be condition-blind. The protocol explicitly blinds gold-brief writing, but the diagnostic review stage does not state that reviewers were masked to the evidence condition when scoring the 240 reports. Since reports cite tool-generated evidence entries, the condition is visually identifiable from the report itself. If expectations about white-box misuse leaked into scoring, the headline differences in Table 3—grounding 3.25 vs. 4.52, usefulness 4.00 vs. 4.68, misuse 2.50 vs. 1.00—and the all-60-case misuse increase in Table 4 could be inflated. This is load-bearing for the central claim. Please add a masked, fully crossed rescoring of at least a subset (e.g., redacted or normalized citations) and report whether the paired differences persist; also state how the five reviewers were assigned to the 240 reports.","section":"§4 and Appendix E"},{"comment":"Inter-rater agreement is low on the axes that carry the central claims: quadratic weighted kappa is 0.25 for correctness, 0.50 for grounding, and 0.46 for usefulness, with only misuse high (0.88). The agreement diagnostics are computed only on the original 20-case subset with two reviewers, and full-set rater-level agreement is not reported. The paper acknowledges this in the limitations section, but the quantitative claims in RQ2 are paired mean differences on these same axes. Low agreement does not disprove the direction, but it means the magnitudes could shift materially under a different scoring draw. Please report full-set agreement, per-rater means, and rater-level paired analyses, or a bootstrap treating rater as a random effect, so the paired claims can be evaluated.","section":"Table 8 and RQ2"},{"comment":"The misuse comparison between the combined/hybrid interfaces and the surface baseline is confounded with evidence quantity: the hybrid package contains 22.97 evidence items on average versus 3.27 for surface evidence, and the combined package contains 19.70. The paper notes the Spearman correlation of 0.455 between package size and misuse and disclaims causality, but the statement that misuse is higher in all 60 cases is presented as an interface effect. The shuffled control helps separate volume from irrelevance because its misuse score is 5.00 despite citation volume similar to the combined condition. Still, a volume-controlled condition or an explicit regression controlling for item count would materially strengthen the quantitative misuse claims. As written, the all-60 misuse increase should be read as an interface-plus-volume effect, not as pure interface effect.","section":"§5 RQ3, Table 3"}],"minor_comments":[{"comment":"The B/T/W formatting appears to omit counts or rows. For example, the combined-white-box Correctness row shows only two counts (39/13) under a column specified as better/tied/worse, and the hybrid section has no Grounding row even though the text reports grounding for hybrid. Please check that each B/T/W triple is complete and all four axes are listed.","section":"Table 4"},{"comment":"The residual patching diagnostic uses only 6 validation pairs per family in the main layer plot. The caveats about prompt-shell and answer-order sensitivity are appropriately stated, but the small n should be repeated in the main text near Figure 4, not only in the appendix, to prevent over-reading of layer 7/8 localization.","section":"Appendix C / Table 6"},{"comment":"The naming of the shuffled condition as 'white box relevance control' is clear, but the reader must infer that the shuffled package is hidden from the auditor. Please state explicitly in Table 1 or its caption that the auditor is not informed that the evidence belongs to another case.","section":"Section 3, Table 1"},{"comment":"The abstract says 'Five human reviewers evaluate the primary interfaces,' but the validation stage includes gold writing, verification, diagnostic scoring, and agreement checks across five reviewers. Clarify how many reviewers actually scored each of the 240 reports, since the number of scorers per report is not stated in the main text.","section":"Abstract / Intro"}],"recommendation":"major_revision","confidential_remarks":"The core design is sound and the negative control is genuinely informative. The main gating issue is the reliability and blindness of the human validation scores. I am not asking for a fundamentally new study: a masked rescoring on a representative subset, full-set agreement statistics, and a volume-controlled misuse analysis would likely resolve the concern. If the effects persist under masked scoring, the paper could be a strong contribution; as it stands, the quantitative central claims rest on measurement properties that are not yet established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: the paper's core contribution is real — a controlled comparison of evidence interfaces for LLM-generated policy audit reports, with a shuffled relevance control that genuinely tests whether reports can sound plausible while citing irrelevant internal evidence. That control alone is worth the price of admission. The paper also does the right thing by treating the hybrid advantage as a design signal, not a causal proof, and by framing the residual-patching result as a narrow, prompt-sensitive diagnostic rather than broad mechanistic support.\n\nWhat's soft: the empirical backbone is the 240 human validation scores, and the protocol never states that the five reviewers were blind to evidence condition while scoring. Reports reveal their condition through the tool citations, so expectations could inflate the misuse/grounding gaps. The agreement diagnostics on the 20-case subset are weak where it matters: QWK 0.25 for correctness, 0.50 for grounding. Full-set rater agreement isn't reported. So the all-60-cases-worse misuse finding is not as secure as it looks. The paper itself flags that agreement was only measured on the subset, which is honest but doesn't fix it.\n\nThe hybrid-usefulness gain is also confounded with package size (22.97 items vs 3.27), as the paper admits. And the patching diagnostic uses a hand-set gate threshold, but the authors don't lean on it heavily, so I'd call that minor.\n\nOverall: this is a careful, clearly-written study that doesn't overclaim. The shuffled-control result is important for AI governance, and the framing of white-box access as an evidence-design problem is useful. But the quantitative claims depend on human scoring that hasn't been shown to be condition-blind or reliably agreed upon. That's fixable — a masked rescoring and full-set rater analysis would do it — and I'd want that before trusting the precise numbers.\n\nThe paper deserves a serious referee. It's the kind of work people working on audit tooling and interpretability evaluation will want to engage with, and the negative control is a reusable idea. My recommendation: send it to peer review, but ask for the scoring-blinding fix and a size-controlled hybrid before publication.","headline":"A genuinely useful controlled study of evidence interfaces for LLM audit reports, but the human-scoring layer needs a condition-blind check before the headline numbers carry weight.","tokens_in":18003,"tokens_out":4015,"would_cite":true,"duration_ms":41645,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that giving an audit-writing LLM access to a model's internal evidence makes reports cite more but not be more valid; its shuffled-evidence control shows a report can be mostly right about the policy passage while citing i","keywords":["AI governance","policy audit reports","white-box evidence","evidence interface","evidence misuse","shuffled relevance control","LLM audit","passage grounding"],"falsifier":"Run the same 60-case review with five fresh reviewers scoring reports with all condition-identifying cues removed, and check whether combined white-box evidence still weakens grounding in 53 of 60 cases and raises misuse in all 60; if the pattern mostly disappears, the central claim is not supported.","tokens_in":16999,"feed_emoji":"🧾","tokens_out":8425,"duration_ms":76752,"temperature":0.7,"pith_summary":"The paper asks whether giving an LLM that writes policy audit reports access to a model's internal evidence makes those reports more trustworthy. It tries to show that the answer is no in general: combined white-box evidence increases citation volume but weakens grounding, lowers usefulness, and increases evidence misuse relative to a passage-only surface baseline. The key demonstration is a shuffled relevance control in which internal evidence is drawn from a different case while keeping the same format; 39 of 60 reports remain mostly correct about the passage while heavily misusing the irrelevant evidence. The paper also argues that hybrid packages (surface passage anchors plus internal evidence) are a promising interface but that their advantage is confounded with package size, and that a stricter causal diagnostic localizes only a narrow, prompt-sensitive readout signal. The contribution is a reproducible evaluation framework for treating internal access as an evidence-design problem, not a transparency guarantee.","feed_headline":"More white-box citations don't make AI audit reports more valid","feed_subtitle":"A shuffled-evidence test found 39 of 60 reports mostly correct while citing internal evidence from another case.","key_machinery":"The central mechanism is the controlled evidence-interface comparison: the same policy passage, audit rubric, and auditor model are held fixed, and only the evidence package changes across ten conditions. The load-bearing instrument is the shuffled relevance control, which keeps the white-box package's format identical but draws the internal evidence from a different case; this isolates whether reports check relevance or merely cite plausible-looking evidence. The paper also uses a behavior-locked activation-patching diagnostic on forced-choice microtasks to test whether cited internal features are causally localized, with random-region thresholds and rendering controls.","core_discovery":"On the paper's own terms, the central discovery is that evidence uptake is not evidence validity. When the same policy passage, rubric, and auditor model are held fixed and only the evidence interface changes, shifting from black-box surface evidence to combined white-box evidence leaves average correctness nearly unchanged (4.60 vs 4.68) but cuts passage grounding from 4.52 to 3.25, lowers usefulness from 4.68 to 4.00, and raises evidence misuse from 1.00 to 2.50; misuse is higher in all 60 paired cases. The shuffled relevance control strengthens the point: 39 of 60 reports score at least 4 on correctness while also scoring at least 4 on misuse, meaning a report can be mostly right about th","pith_inferences":["Beyond the paper: if correctness and evidence misuse can diverge this cleanly, audit procedures that rely on an LLM's own cautionary notes or confidence scores are likely to miss exactly the dangerous failures; testing human auditors' ability to catch shuffled-evidence errors is a natural next step.","Beyond the paper: the reported correlation between package size and misuse suggests a testable prediction: a volume-matched hybrid with few internal items should show a smaller usefulness gain, while a capped hybrid should separate anchoring from quantity.","Beyond the paper: the narrow, prompt-sensitive causal localization found by patching suggests that readable interpretability labels may be reused by downstream models without corresponding mechanism-level support; this predicts that label-masked or numeric-only evidence would reduce overtrust.","Beyond the paper: for governance, this reframes transparency from 'granting access' to 'designing checkable evidence'; one could require per-citation provenance and relevance metadata as standard report fields."],"forward_implications":["If white-box internal evidence is added without surface passage anchors, audit reports keep their correctness but lose grounding and increase evidence misuse; internal access should not be reported as transparency on its own.","Citation volume should not be used as a quality metric: both the combined white-box condition and the shuffled relevance control produce heavy citation, but human review separates useful support from invalid support.","Any audit workflow that evaluates LLM-generated evidence should include a relevance-broken shuffled control, because plausible-sounding reports can still misuse evidence from another case.","Hybrid packages (passage surface evidence plus internal evidence) are the most useful interface tested on average, but the test does not isolate anchoring from package size; a capped hybrid condition is needed to decide whether the gain is causal."],"fun_headline_variants":["White-box access doesn't raise audit report validity","AI audits: more citations, not more validity","Shuffled evidence exposes audit report misuse","Internal evidence changes audits without improving them"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the five human reviewers' scores are unbiased ground truth for comparing the evidence conditions; the protocol does not state that reviewers were blind to which condition a report came from, and the small 20-case agreement check was low for correctness.","fun_headline_variants_meta":{"raw":{"variants":["White-box access doesn't raise audit report validity","AI audits: more citations, not more validity","Shuffled evidence exposes audit report misuse","Internal evidence changes audits without improving them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1288,"prompt_tokens":777,"completion_tokens":511,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":455}},"tokens_in":521,"tokens_out":511,"duration_ms":5988,"temperature":1.0,"reasoning_tokens":455,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:21:42.046229+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 60-case review with five fresh reviewers scoring reports with all condition-identifying cues removed, and check whether combined white-box evidence still weakens grounding in 53 of 60 cases and raises misuse in all 60; if the pattern mostly disappears, the central claim is not supported.","supporting_citations":[],"review_version":1}