{"id":"1cbd8c90-f123-4c02-93b2-b8c61db5f069","arxiv_id":"2607.24539","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A task-conditional faithfulness audit compares self-reported, ablation-derived, and engineering-registered evidence use in multimodal grid LLMs and corrects failures via evidence-gated re-generation.","lead":"A new audit checks whether multimodal LLMs diagnosing power grids actually use topology, measurements, and text the way engineers require—not just whether answers look right. It can flag shortcut reliance and force corrected, re-tested answers without tanking accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The \"independent re-audit\" is not independent of the correction: evidence-gated regeneration mechanically steers the same ablation-measured pathway that ϕ_Abl,Ref then scores, so the +0.164 grounding gain is close to guaranteed by construction.","rationale":"The reader flagged the author-chosen reference g as the weakest assumption; I agree that is a real risk but locate the sharper instantiation one step downstream: because correction and scoring share the same registered target E_j/g, even a perfectly valid g would not rescue the correction claim from its built-in circularity, and the claimed independence of the re-audit is procedural only. Conversely, if g is mis-specified, the correction actively trains models toward the wrong target — so the two concerns compound. The detection-side contribution (validity gating, shams, responsiveness, signed effects, explicit inconclusive handling) is unusually careful for this venue and survives the critique, which is why I do not move to REJECT. The sham-control experiment is cheap (the infrastructure already exists: 89 paired cases, re-ablation pipeline, bootstrap), directly falsifies the mechanism I am worried about, and would convert the correction result from suggestive to credible. Verdict stays CONDITIONAL, with the sham-correction arm and release of the w_j tables/thresholds as the gating conditions.","tokens_in":6478,"tokens_out":1899,"duration_ms":73615,"concrete_test":"Add a sham-correction arm to Subcase A: regenerate the same 89 failed stressed cases with a format- and length-matched prompt that omits the evidence gating (no citation constraint to E_j, no masking of L_j), then independently re-ablate identically. Compare paired ∆ϕ_Abl,Ref and ∆P between evidence-gated and sham arms with the same family bootstrap. If the sham arm recovers a substantial fraction of the +0.164 grounding gain (or of +0.442 performance), the correction claim reduces to regeneration-on-selection and §II.D needs a control condition before the 'correct failures' claim stands. A secondary check: re-audit corrected responses on counterfactual siblings whose registered g differs; if corrected models keep citing the originally gated evidence, the fix is compliance, not task-conditional grounding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim says the framework can \"correct faithfulness failures... via evidence-gated regeneration plus independent re-ablation.\" But §II.D's correction levers — constraining citations to the registered admissible evidence E_j, masking prohibited fields L_j, provenance-constrained generation — act directly on the input/output surface that the behavioral reliance vector b (Eq. 5) measures. If you forbid the model from using shortcut text X and require it to cite topology/measurements, then re-ablating T or S will of course change the answer more, raising δ_{T}, δ_{S} and hence cosine(b, g). The re-audit is \"independent\" only procedurally (new ablation calls), not in the sense that matters: the correction was optimized against the same g-target that scores it. The claimed gain is measured on the 89 cases selected for failing gate (8), which includes p < τ_P, so the +0.442 performance jump is partly selection-on-failures plus a prompt that re-injects the correct evidence — a strong baseline effect that is never controlled against a sham regeneration. There is no arm showing a generic \"try again, be careful\" prompt wouldn't produce similar gains. This does not invalidate the detection results (the SR–Abl–Ref mismatch finding stands on its own), but the \"correct\" half of the detect–diagnose–correct triad is currently evidence of prompt compliance with the audit's own target, not of improved task-conditional evidence use.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"This letter proposes a task-conditional faithfulness audit for multimodal LLMs applied to grid diagnosis. Before testing, the authors register a \"task contract\" (Eq. 2) that fixes admissible evidence, prohibited outcome-bearing fields, thresholds, and an engineering reference importance vector g (Eq. 3) built from modality availability and preregistered task weights. The audit then compares three quantities per task–scenario pair: self-reported reliance s, behavioral reliance b derived from matched single-modality ablations with validity screening (Eqs. 4–5), and g, via three cosine similarities (Eq. 7). Signed per-modality utility effects (Eq. 6), responsiveness filtering, sham edits, and positive controls guard the measurement. Cases failing a composite gate (Eq. 8) are diagnosed into four categories and corrected by evidence-gated regeneration, then re-ablated and accepted only if the lower CI on Δϕ_Abl,Ref is positive with noninferior performance. A proof of concept on IEEE 39/118-bus scenarios with three quantized open LLMs (2606 registered calls) finds a small but significant stressed SR–Abl gap (0.024 [0.011, 0.036]), model-varying shortcut/conflict effects, gate balanced accuracy .869–.943, and paired correction gains of +0.442 performance and +0.164 grounding over 89 eligible failures.","tokens_in":6871,"tokens_out":4067,"duration_ms":147726,"significance":"If the results hold, this is a useful contribution to LLM evaluation in safety-relevant domains: the preregistered contract separating engineering requirements from model outputs, the explicit intervention-validity screening (I★) with sham edits and positive controls, the responsiveness rule that declares low-signal samples inconclusive rather than imputing uniform reliance, signed modality effects that distinguish useful from harmful influence, and the joint reporting of coverage alongside alignment are all genuinely careful measurement design and are more rigorous than typical \"explanation faithfulness\" evaluations. The framework is also falsifiable (gates and acceptance criteria are stated in advance) and portable by construction (adapter split in §II.E). The honest reporting of G12 adapter attrition (Table III) and of non-uniformly-harmful conflict effects strengthens credibility. The main qualification is that the \"correct\" third of the detect–diagnose–correct claim is presently weaker than the detection half, for reasons detailed below; the detection/diagnosis contribution stands on its own.","major_comments":[{"comment":"The correction result is evaluated without a control regeneration arm, and the re-audit is not independent of the scoring target. The correction levers (constrain citations to E_j, mask L_j, provenance-constrained generation) act directly on the input/output surface that b (Eq. 5) measures; re-ablating a response that was forced to cite T/S and blocked from X will mechanically raise δ_T, δ_S and hence ϕ_Abl,Ref. The 'independent re-audit' is independent only procedurally (fresh ablation calls), since acceptance is scored against the same g the correction steers toward. Separately, the +0.442 performance gain is computed on the 89 cases selected for failing gate (8), which includes p < τ_P, so part of it is selection-on-failures plus re-injection of the correct evidence. A sham-regeneration arm (generic 'reconsider carefully' prompt with no evidence constraints) is needed to show the gain","section":"§II.D, Fig. 2(c), §III.A"},{"comment":"The validity of g as external ground truth is assumed, not established. w_j(z_n) are hand-specified under structural/generative/conditional rules, and both the detection metric ϕ_Abl,Ref (Eq. 7) and the pass gate (Eq. 8) inherit this dependence. §II.B states that disagreement among engineering rules is 'retained as a reference interval,' but no interval appears anywhere in the reported results — all ϕ values are point estimates against a single g. The paper should (i) report how the weights were set and by whom, with some inter-rater or expert-agreement evidence, and (ii) include a sensitivity analysis showing that the qualitative conclusions (sign of the SR–Abl gap, per-modality signed effects, correction acceptance) are stable under plausible perturbations of w_j(z_n). Without this, high ϕ_Abl,Ref demonstrates agreement with author-chosen weights, which is not identical to task-appropr","section":"§II.B, Eq. (3)"},{"comment":"The pooled correction statistics mask substantial heterogeneity and rest on small per-cell samples. Table I reports per-cell pair counts from n=6 (Q4/39) to n=19 (G12/39), with ΔP^c ranging from +.147 to +.833; the headline +0.442 [0.324, 0.543] pools 75 pairs across three models, two systems, and five tasks, and the bootstrap is family-stratified, capturing scenario but not model/task variation. The distribution of the 89 failures across the five tasks is never given, so the reader cannot tell whether correction works uniformly or is driven by one task (e.g., event classification, where masking X is most likely to flip a copied label). Please report per-task (or at least per-model) correction CIs and the failure distribution, and state explicitly which claims are supported at cell level versus only in the pool.","section":"§III, Table I"}],"minor_comments":[{"comment":"The calibrated threshold values are never reported. τ_r, τ_P, τ_B, ε_P, τ_j, and the calibration-set procedure that produces them (§II.C–D) are load-bearing for gate (8) and for reproducibility; please tabulate them per task.","section":"§II.C–D"},{"comment":"Decoding details are missing: the number of repetitions L for stochastic decoding (§II.C), temperature/sampling settings, and the exact quantized checkpoint versions (only 'Qwen3 4B Instruct (Q4)' etc. with the Ollama library as citation [7]) are not pinned. An execution manifest is described in §II.E, but the letter does not state whether manifests, prompts, or code are released.","section":"§III"},{"comment":"Detection performance (BA .869–.943, Sp .779–.938) is measured against author-constructed aligned/shortcut/conflict regimes, so the labels are ground truth by construction. This is acceptable for a proof of concept, but the letter should state explicitly that gate performance in uncontrolled settings is unknown and discuss expected failure modes.","section":"§III.B, Table III"},{"comment":"Eq. (5): b is undefined when the denominator Σ I_j,n,ℓ δ_j,n,ℓ = 0; the τ_r inconclusive rule covers this only partially (it uses the same sum). Please state the tie/zero handling explicitly, and clarify whether ϕ_Abl,Ref is ever computed on samples with I★ = 0 but partially valid interventions.","section":"§II.C, Eq. (5)"},{"comment":"Notation: 'sim' in Eq. (7) is identified as cosine only in the following sentence; since s, b, g all live on the simplex, cosine lies in [0,1] and is insensitive to magnitude — worth one sentence justifying cosine over, e.g., L1 or KL, given that near-uniform and near-one-hot vectors can have high cosine yet very different reliance.","section":"§II.D, Eq. (7)"},{"comment":"Fig. 2(b) legend labels ('X shortcut', 'X conflict') do not match Table II column headers (Δ_X^Sh, Δ_X^Cf); please harmonize. In Table I, the ΔP^c / Δϕ_Abl,Ref^c columns should be cross-referenced to the pooled Fig. 2(c) numbers, and the 'A/E' subscript convention repeated in the caption is easy to misread.","section":"Fig. 2, Tables I–II"},{"comment":"Related-work positioning is thin: the audit's b and δ are close in spirit to sufficiency/comprehensiveness and erasure-based faithfulness metrics, and to multimodal ablation audits beyond [6]. One or two sentences situating the SR–Abl–Ref triangle relative to these would help readers see what is new (the preregistered external reference and the validity gating, as I understand it).","section":"§I"},{"comment":"Grammar/typography: Abstract 'a general framework in order to conduct task-conditional faithfulness audit' and 'validate the framework ability' need articles/possessives; author block has stray spaces in membership designations; §III header 'CASESTUDY' missing space.","section":"Abstract, title block, §III"}],"recommendation":"major_revision","confidential_remarks":"The measurement methodology is unusually careful for this venue and the detection/diagnosis results are, in my reading, sound. My reservation is confined to the correction claim, which is currently guaranteed-in-part by construction (correction steers the same surface the re-audit scores) and lacks a sham-regeneration control; the fix (one additional arm plus a g-sensitivity analysis) is within the scope of the existing pipeline. There is no code/data availability statement; given the reproducibility emphasis (execution manifests, frozen contracts), the editor may wish to require release of prompts and manifests as a condition of publication. Venue fit seems good if the correction arm is tightened."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful part of this letter is not another claim that LLMs can be unfaithful—that is already on the table from Turpin and Chen. What is new is a packaged, validity-aware audit for multimodal grid diagnosis: a preregistered task contract, matched full/ablated calls with shams and positive controls, an SR–Abl–Ref triangle, signed modality effects, and explicit gates on intervention validity and responsiveness so a model cannot look faithful by failing to answer.\n\nThey execute it cleanly enough for a proof of concept. Three local quantized models on IEEE 39/118, five task families, bootstrap intervals, and honest reporting that G12 loses coverage on structured output. The stressed gap between claimed and behavioral alignment is small but real, and shortcut text showing positive utility with zero registered importance is a concrete diagnostic finding. The measurement design is more careful than most XAI-for-grids letters.\n\nThe soft spot is the correct half of detect–diagnose–correct. Engineering importance g is author-registered weights, not an external oracle. Correction then constrains citations to admissible evidence E_j and masks L_j—the same surface the ablation vector b measures—so raising cosine(b, g) on re-ablation is close to what you would expect by construction. Procedural independence of the re-ablation calls does not fix that. The large performance jump on the 89 gate failures also lacks a sham “try again” arm, so some of it is selection plus re-injecting the right evidence. Detection and signed-effect diagnosis still stand; the correction claim should be dialed back to “prompt compliance with the registered contract under a noninferiority gate.”\n\nMinor limits for a letter: free thresholds and w_j are not numerically disclosed, systems are small, no code/data. Circularity is moderate, not vacuous, because performance noninferiority and validity are separate.\n\nThis is for people building multimodal diagnosis tools who need something stricter than answer accuracy. It deserves a serious referee, ideally with a request for weight tables, threshold calibration, a sham-regeneration control, and released adapters. I would engage the detection framework; I would not treat the correction deltas as settled evidence of deeper grounding.","headline":"Solid detection pipeline for task-conditional faithfulness in grid LLMs; the correction gains are real but partly built into the target they score against.","tokens_in":7856,"tokens_out":550,"would_cite":false,"duration_ms":23267,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A task-conditional audit can detect and fix when multimodal grid LLMs give right answers from the wrong evidence.","keywords":["Explainable AI","large language models","model faithfulness","multimodal fusion","power systems","grid diagnosis","modality ablation"],"falsifier":"On held-out stressed grid scenarios, if evidence-gated correction fails to produce a positive lower confidence bound on the change in behavioral-to-reference alignment while keeping performance noninferior, or if ablation-derived reliance systematically fails to track the preregistered modality weights on tasks with known single-source answers, the central claim does not hold.","tokens_in":7564,"feed_emoji":"⚡","tokens_out":867,"duration_ms":22435,"temperature":0.7,"pith_summary":"Correct answers from multimodal large language models used for power-grid diagnosis do not prove the models used the evidence each task actually requires. This letter proposes a general audit that first registers, for each diagnostic task, which modalities matter in engineering terms, then compares that reference against both the model’s self-reported reliance and its real behavioral dependence under controlled single-modality ablations. When those three disagree, an evidence-gated correction regenerates the answer under explicit evidence constraints and independently re-ablates it, accepting the fix only if grounding improves and performance does not fall. Case studies on IEEE 39- and 118-bus systems with three differently scaled models show the framework can surface shortcut and conflict use of incident text, diagnose the failure type, and raise both performance and behavioral grounding on failed stressed cases. A sympathetic reader cares because grid tasks such as topology identification, N–1 screening, and congestion impact estimation demand different evidence pathways; trusting accuracy alone can hide unsafe reliance on narrative shortcuts.","feed_headline":"Grid LLMs can be right for the wrong reasons—and fixed","feed_subtitle":"A task-conditional audit catches shortcut evidence use and re-grounds answers without hurting accuracy.","key_machinery":"The SR–Abl–Ref measurement triangle (cosine similarities among self-report s, ablation-derived behavioral reliance b, and preregistered engineering importance g), gated by validity, responsiveness, and performance, together with diagnosis-specific evidence-gated correction that is accepted only after independent re-ablation shows improved grounding and noninferior performance.","core_discovery":"The authors show that a preregistered task-conditional faithfulness audit—jointly comparing self-reported reliance, intervention-derived behavioral reliance, and engineering reference importance—can detect, diagnose, and, via evidence-gated regeneration plus independent re-ablation, correct task-conditional faithfulness failures in multimodal grid LLMs without material performance loss.","pith_inferences":["The same preregistered contract-plus-ablation pattern could audit multimodal assistants in other safety-critical domains where narrative context can dominate sensor or topology evidence.","Persistent model-specific attrition on structured outputs (as seen with the largest model here) implies that faithfulness scores must be conditioned on parse validity or they will favor models that simply refuse hard cases.","If engineering reference weights disagree across operators, the framework’s retained reference intervals could become a formal way to encode expert disagreement rather than forcing a single ground truth."],"forward_implications":["Grid LLM deployments can report validity, responsiveness, and signed modality effects alongside accuracy rather than treating correct labels as sufficient.","Shortcut or conflicting incident text can be diagnosed as distinct failure modes and routed to masking, citation constraints, or physics tools before answers are accepted.","The same frozen task contract and gates can transfer across grid systems and model backends by swapping only lightweight adapters.","Failed cases that cannot be corrected under the gate become explicit human-review or tool-verification candidates rather than silent errors."],"fun_headline_variants":["Task-conditional audit finds and fixes unfaithful grid LLM answers","Grid LLMs can answer right on wrong evidence—audit corrects it","Faithfulness audit detects shortcut evidence in multimodal grid LLMs","Evidence-gated re-audit corrects grounding without accuracy loss","Preregistered audit aligns self-report, behavior, and engineering importance"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The hand-specified engineering importance weights registered before testing are treated as the true measure of which evidence each grid task should use.","fun_headline_variants_meta":{"raw":{"variants":["Task-conditional audit finds and fixes unfaithful grid LLM answers","Grid LLMs can answer right on wrong evidence—audit corrects it","Faithfulness audit detects shortcut evidence in multimodal grid LLMs","Evidence-gated re-audit corrects grounding without accuracy loss","Preregistered audit aligns self-report, behavior, and engineering importance"]},"model":"grok-4.5","effort":"low","cost_usd":0.00272,"raw_usage":{"total_tokens":953,"prompt_tokens":695,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":27204000,"prompt_tokens_details":{"text_tokens":695,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":186,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":695,"tokens_out":72,"duration_ms":4302,"temperature":1.0,"reasoning_tokens":186,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T12:10:39.112934+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On held-out stressed grid scenarios, if evidence-gated correction fails to produce a positive lower confidence bound on the change in behavioral-to-reference alignment while keeping performance noninferior, or if ablation-derived reliance systematically fails to track the preregistered modality weights on tasks with known single-source answers, the central claim does not hold.","supporting_citations":[],"review_version":1}