{"id":"80b3627d-4160-40ea-9903-e766b8fd9ee6","arxiv_id":"2607.20827","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LLM agents respond to source-authority cues yet show measurable sensitivity to unauthorized context under controlled tests: 5.4% action discordance under competition and a 2.4% retained-invalid error pattern.","lead":"This paper introduces a target-specific authorization audit that tests whether LLM agents let untrusted context influence which tool or argument value they choose, even when the final action is correct. Across 450 controlled next-action tasks, trusted-versus-untrusted source changes alter generated actions in 5.4% of competing cases, and removing valid evidence exposes residual sensitivity to unauthorized context.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2.4% retained-invalid degradation rate rests on author labels that external annotators match only 78.9% (κ=0.626); the degradation prompt endpoint was never rerun on consensus labels, leaving the headline residual-sensitivity estimate unvalidated at its most label-dependent point.","rationale":"The reader's weakest assumption and my load-bearing concern coincide: the 2.4% retained-invalid endpoint is directly tied to author-provided valid/invalid labels, and the paper itself states that consensus-label validation was applied only to coalition metrics, not to the degradation prompts. This is the single most load-bearing issue because the 2.4% figure is a headline quantitative claim in the abstract and conclusion, used to support the central conclusion that models retain residual sensitivity to unauthorized context. The matched source-intervention results (Table 4, Table 5) provide independent, comparatively label-light evidence for the core 'models respond to source cues' claim, so the paper's central argument does not collapse entirely if the 2.4% turns out to be fragile. However, the degradation endpoint is the most direct behavioral demonstration of 'retained unauthorized influence,' and its label-dependence is real, acknowledged, and untested. The paper deserves credit for explicitly flagging this limitation and for running many other controls (same-proposition source-only control, role-matched boundaries, consensus-label coalition checks), which is why I do not recommend REJECT or even a harsher verdict. CONDITIONAL is the appropriate status, with the condition being exactly this consensus-label degradation rerun. My verdict is therefore UNCHANGED relative to the reader's CONDITIONAL. If the rerun were to show a large drop in the pattern rate, I would move to CONDITIONAL with a tighter scope or possibly UNVERDICTED for that specific endpoint; but that is a future adjustment, not a present one.","tokens_in":31210,"tokens_out":3447,"duration_ms":37739,"concrete_test":"Rerun the valid-evidence degradation experiment on the 300 externally relabeled targets (1,200 target–model rows) using the adjudicated consensus labels. For each consensus-Valid factor, build full, mixed (remove that valid factor), and clean (remove that valid factor plus all consensus-Invalid factors for the target) prompts; decode with the four degradation models. Compute the retained-invalid pattern rate on this subset and compare it with the same subset's author-label rate. Report a cluster-bootstrap 95% CI for the difference (or ratio). If the consensus-label pattern rate falls within the author-label CI, the endpoint is robust; if it drops materially (e.g., below 1% or to zero), the 2.4% headline is label-dependent and the residual-sensitivity claim must be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central quantitative claim includes a 2.4% [2.1,3.0] rate of full-correct/mixed-error/clean-correct retention of unauthorized competition. This endpoint is defined entirely by which factors the author team labeled Valid and Invalid (Section: Controlled Valid-Evidence Degradation). The paper's own external relabeling (Table 7) shows only 78.9% agreement between author labels and adjudicated consensus (κ=0.626 for invalid vs non-invalid), meaning 21.1% of target–factor pairs change labels. Table 18 reruns coalition metrics on consensus labels, but the degradation prompts are explicitly not rerun: 'On the relabeled subset, consensus labels rerun only coalition metrics, not the author-label-derived degradation prompts' (Dataset Construction and Annotation). If the label disagreements are systematically related to the degraded contexts—e.g., if the author team over-labels stale or neighboring records as Invalid precisely in cases where removing them makes the clean prompt return to the reference—the 2.4% rate could be inflated, possibly to near zero. The paper acknowledges this as a general limitation (Appendix A), but the headline estimate itself lacks the direct robustness check that the coalition analysis received. This is the weakest link in the chain from 'models respond to source cues' to 'unauthorized evidence still influences actions'.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a target-specific authorization audit for LLM next-action selection. It labels each context factor separately for each tool/argument target as Valid, Invalid, or Neutral under an application-level authorization rule, then runs three tests: (1) matched source-authority interventions that hold task, proposition, position, and policy fixed while changing only the trusted/untrusted source marker; (2) valid-evidence degradation comparing full, mixed (remove one valid factor), and clean (remove that valid factor plus all Invalid factors) contexts; (3) an exact coalition analysis over all 256 subsets using Harsanyi/Shapley interactions. Across 450 controlled tasks and multiple open-weight models, the paper reports that generated actions differ between trusted and untrusted variants in 5.4% of competing cases vs 1.7% of supporting cases, and that a full-correct/mixed-error/clean-correct retained-invalid pattern occurs in 2.4% [2.1, 3.0] of degradation comparisons. The coalition analysis shows positive fractional excess under partial evidence but near-null full-context behavior, and the paper explicitly identifies strict role-matching as a boundary condition. The overall conclusion is that models respond to textual source-authority cues but do not fully isolate their actions from unauthorized evidence.","tokens_in":31454,"tokens_out":5851,"duration_ms":68883,"significance":"If the result holds, the audit is a valuable complement to correctness-based agent benchmarks: it separates outcome correctness from provenance dependence and extends prompt-injection-style evaluation to benign residues such as stale memory, neighboring records, and previous-action state. The paper's strengths include exact enumeration of all factor subsets, cluster-bootstrapped confidence intervals, and an unusually extensive set of controls: random-label matching, content-type matching, first-order-salience matching, value-bearing and role-matched controls, factorization coarsening, placeholders, factor-order shuffles, prompt-wording variants, and external relabeling on a stratified subset with adjudication. The authors are also appropriately cautious about the distinction between controlled stress-set rates and deployment prevalence, and about the limits of interaction diagnostics. The main quantitative claim, especially the 2.4% retained-invalid rate, is clearly identified and methodologically positioned.","major_comments":[{"comment":"The 2.4% retained-invalid degradation endpoint is constructed entirely from author-assigned Invalid labels, but the paper's own external relabeling (Table 7) shows only 78.9% agreement between author labels and adjudicated consensus (kappa=0.626 for invalid vs non-invalid). The text explicitly states: 'On the relabeled subset, consensus labels rerun only coalition metrics, not the author-label-derived degradation prompts.' Since the clean context is defined by deleting all factors the author team labeled Invalid, a 21.1% label disagreement rate could systematically alter which factors are removed and thus which rows exhibit the full-correct/mixed-error/clean-correct pattern. The paper acknowledges this as a general limitation in Appendix A, but the headline estimate itself lacks the direct robustness check that the coalition analysis received. Please rerun the full/mixed/clean degradatio","section":"Dataset Construction and Annotation / Controlled Valid-Evidence Degradation"}],"minor_comments":[{"comment":"The column headers 'Base', 'S err.', 'C err.', and 'S disc.' are not all defined. Please define 'Base' (base error rate), 'S' and 'C' (supporting/competing relations), and 'disc.' (paired target discordance) directly in the caption.","section":"Table 5 caption"},{"comment":"The summed target-token log-odds score may be sensitive to target string length; the robustness checks in Table 32 address this, but consider adding one sentence in the main text noting that the score is summed and that length robustness is reported in Appendix E.","section":"Equation (1)"},{"comment":"The subfigure labels (a)-(d) are not referenced in the text. Please add explicit references (e.g., 'Figure 1a shows the factor decomposition') so the reader can navigate the schematic.","section":"Figure 1"},{"comment":"Tables 3 and 16 report overlapping agreement statistics. Consider consolidating them or adding a cross-reference to avoid apparent duplication.","section":"Tables 3 and 16"},{"comment":"The API-served proxy result is appropriately framed as a feasibility check, but the text could state more explicitly that the closed-weight evidence is not part of the main claim, since only GPT-4o-mini separates high from low.","section":"Table 55 / API proxy"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically strong and the central finding is credible; the sole load-bearing gap is the missing consensus-label rerun of the degradation endpoint. If the authors can provide that rerun (or a convincingly argued substitute), I would expect this to become a minor-revision or accept. I found no evidence of circular reasoning or misrepresentation of the limitations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper introduces a target-specific authorization audit for LLM agent action selection, labeling each context factor separately for every tool and argument. The main design is strong: matched source-only interventions, controlled valid-evidence degradation, and an exact coalition interaction diagnostic. The source-response result—trusted vs. untrusted markers shift target scores and change actions more for competing than supporting propositions (5.4% vs. 1.7%)—is cleanly identified. The 450-task suite, with adapted Tau2 and BFCL tasks, is a useful contribution, and the external relabeling protocol is unusually honest for this literature.\n\nThe soft spot is exactly where your stress-test points: the 2.4% retained-invalid pattern. That endpoint depends on author-assigned Invalid labels, and external annotators agree with those labels on only 78.9% of pairs, kappa=0.626. The degradation prompts were not rerun with consensus labels; only the coalition metrics were. So the specific 2.4% [2.1, 3.0] number is not robustly tied to the label consensus. That is a genuine weakness and the right reason to be conditional. But it is not load-bearing for the paper's central conclusion: even if the 2.4% were inflated toward zero, the mixed–clean discordance of 24.6% and the matched source-sensitivity results already demonstrate residual unauthorized influence. The paper is also careful to call these controlled stress-set rates, not deployment prevalence. The 2.4% is one piece of evidence, not the whole argument.\n\nA minor caveat: the textual 'TRUSTED_CURRENT' vs 'UNTRUSTED_EXTERNAL' marker is a controlled proxy for source authority, not an operational provenance channel such as retrieval or tool transport. The paper acknowledges this. The role-matched controls that limit the coalition diagnostic's authorization-specific interpretation are a strength, not a deflection.\n\nWho is this for? Researchers working on agent evaluation, prompt-injection defenses, or provenance tracking. It gives a new evaluation axis worth testing, and the released code, data, seeds, and relabeling records make it reproducible.\n\nRecommendation: send to peer review. The label-dependence question is a well-scoped revision point—re-estimate the 2.4% rate on the externally labeled subset, or at least report sensitivity to consensus labels. This deserves referee time, not a desk rejection.","headline":"A careful, genuinely useful audit of provenance sensitivity in LLM agents; the headline 2.4% degradation number is softer than it looks, but the core claim holds and the paper deserves a serious referee.","tokens_in":776,"tokens_out":2079,"would_cite":true,"duration_ms":37827,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM agents respond to source-authority cues but still let unauthorized context influence their tool and argument choices.","keywords":["target-specific authorization audit","provenance sensitivity","LLM agents","action selection","source authority","Harsanyi interactions","valid-evidence degradation","tool and argument targets"],"falsifier":"Rerun the valid-evidence degradation experiment replacing the author labels with the adjudicated consensus labels from the external annotators. If the full-correct/mixed-error/clean-correct pattern rate drops to near zero (or to the clean-context error rate), the retained-invalid claim is an artifact of label assignment rather than model behavior.","tokens_in":31012,"feed_emoji":"🔍","tokens_out":3828,"duration_ms":38167,"temperature":0.7,"pith_summary":"LLM agents pick tools and argument values from prompts that mix user requests, trusted lookups, memory, and untrusted notes. This paper argues that correctness on the final action is not enough: support for a target action can rest on evidence that was not authorized to decide it. To expose this, the authors build a target-specific audit that labels each context factor as valid, invalid, or neutral for each tool and argument separately, then hold the task fixed while changing only the source authority of a proposition. Across 450 controlled tasks and several open-weight models, trusted and untrusted variants produced different actions in 5.4% of competing cases versus 1.7% of supporting cases, and a retained-invalid error pattern appeared in 2.4% [2.1,3.0] of degradation comparisons. The conclusion is that models respond to source-authority cues without fully isolating their decisions from unauthorized evidence.","feed_headline":"Unauthorized context still steers LLM agents' actions","feed_subtitle":"A target-specific audit finds residual leaks: 5.4% action changes under competition and a 2.4% retained-invalid error pattern.","key_machinery":"The audit's central object is the target-specific authorization label: each context factor is judged for whether its source is permitted to determine a given tool or argument target (authority) and whether its proposition supports, competes, or does neither (relation). Valid factors are authorized and supportive; Invalid factors are unauthorized and supply a concrete alternative; the rest are Neutral. This labeling enables three test instruments: matched source interventions that flip only the trusted/untrusted marker of a proposition; controlled valid-evidence degradation that removes a valid factor with and without the invalid set; and exact Harsanyi/Shapley interactions over all subsets o","core_discovery":"The central claim is that LLM agent action selection remains sensitive to unauthorized context even when models respond to source-authority cues. Using a matched intervention that changes only the source marker of a proposition, trusted competitors reduced target support more than their untrusted variants (a pooled gap of +1.15 log-odds); yet untrusted competitors still caused large score drops, and generated actions flipped between trusted and untrusted variants in 5.4% of competing cases versus 1.7% of supporting cases. When valid evidence was removed while unauthorized competitors stayed, the full-correct/mixed-error/clean-correct pattern occurred in 2.4% [2.1,3.0] of comparisons, associa","pith_inferences":["If the 2.4% retained-invalid rate reproduces under independently relabeled degradation prompts, it would provide a direct behavioral handle for guardrail testing; the current paper leaves that rerun undone.","The same audit design could be applied to retrieved context in RAG pipelines, where the authorization chain is the retrieval policy rather than a fixed task policy.","The near-null full-context transform suggests that unauthorized influence is partially masked by valid evidence; this implies that robustness tests should stress partial-evidence states, not just full prompts.","A testable extension: measuring whether provenance-tracking models or retrieval filters that append source metadata actually reduce the 5.4% competition discordance, which would connect the audit to deployment-time defenses."],"forward_implications":["Outcome benchmarks that only check final action correctness will miss actions that happen to be correct but are grounded in unauthorized context.","An explicit authority policy improves pooled correctness against untrusted competitors (from 14.9% to 9.5% error) but does not act as an invariant provenance filter; models remain source-sensitive.","Tool targets and argument slots need separate authorization audits because the same factor can be neutral for the tool but invalid for a slot.","Agents should be evaluated along a provenance-sensitivity axis, not just a correctness axis; the audit provides a way to score that axis.","Tests that remove one factor at a time can understate unauthorized influence because invalid factors often exert their effect only in higher-order combinations with valid evidence."],"fun_headline_variants":["LLM agents act on unauthorized context despite authority cues","5.4% of LLM agent actions flip when only source authority changes","Provenance audit finds LLM agents leak untrusted evidence","Source-authority cues don't stop untrusted context from influencing LLMs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the authors' target-specific authorization labels — which factors are 'invalid' for each tool or argument — are correct enough; if those labels systematically misassign invalid status, the 2.4% degradation pattern could be inflated.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents act on unauthorized context despite authority cues","5.4% of LLM agent actions flip when only source authority changes","Provenance audit finds LLM agents leak untrusted evidence","Source-authority cues don't stop untrusted context from influencing LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000639,"raw_usage":{"total_tokens":2781,"prompt_tokens":748,"completion_tokens":2033,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":1959}},"tokens_in":492,"tokens_out":2033,"duration_ms":13980,"temperature":1.0,"reasoning_tokens":1959,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:14:18.168591+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the valid-evidence degradation experiment replacing the author labels with the adjudicated consensus labels from the external annotators. If the full-correct/mixed-error/clean-correct pattern rate drops to near zero (or to the clean-context error rate), the retained-invalid claim is an artifact of label assignment rather than model behavior.","supporting_citations":[],"review_version":1}