{"id":"f1ae25a3-4513-4ade-8186-091392934287","arxiv_id":"2607.23999","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Terminal policy labels are insufficient: two containment policies with identical zero-harm endpoints still differ in 73.5% of trajectories and in authorized-work completion.","lead":"This paper introduces ContainmentBench, a trace-based benchmark for measuring what happens to LLM agents after a prompt injection, instead of only whether the final attack succeeded. It shows that two defenses with identical zero-harm end results can differ sharply in how much authorized work gets blocked, so security evaluations should report trajectories and utility separately.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: endpoint-insufficiency claim is independently carried by utility evidence, so the taint-log completeness caveat does not move the verdict.","rationale":"The reader's weakest assumption identifies a genuine measurement boundary, but it is not the load-bearing point. The central claim has two supports: endpoint-equal pairs differ in authorized utility, and they differ in concrete logged actions. The first is independent of taint-label completeness; the second is visible in matched calendar traces where v1 blocks and v2 commits. The paper already states that Eq. (2) is an instrumentation definition, not a semantic proof of influence, and that logged spread is not semantic influence. Finding 3's denominator sensitivity further shows the authors do not overclaim propagation rankings. The proposed semantic audit would provide a worthwhile external calibration of the propagation axis, but its failure would not overturn the main conclusion. The oracle-ledger assumption is conditional and scoped, and the single-model/synthetic limitations are explicitly stated. The manuscript's internal consistency is strong: the frozen 17,640-rollout protocol, positive controls, parser diagnostics, paired and cluster bootstraps, and outcome-conditioned matching support the endpoint-insufficiency claim as stated. No internally inconsistent step in Eqs. (1)–(12) invalidates the matched-pair logic. Therefore the reader's ACCEPT stands unchanged; the only refinement is to keep logged-spread comparisons labeled as operational rather than semantic, which the paper already does.","tokens_in":23327,"tokens_out":11629,"duration_ms":117289,"concrete_test":"Run a NeuroTaint-style semantic/causal auditor over the 600 active-tainted matched pairs and compare its reconstructed influence traces with the logged taint-spread/trajectory projection. If the 73.5% divergence and the 0.6925 utility delta persist under the semantic audit, the logging-completeness caveat is empirically bounded and the headline claim is confirmed. If the divergence collapses, the logged-spread axis should be re-labeled as purely operational, but the utility-based endpoint-insufficiency finding would still stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After checking the outcome-conditioned argument, I find no load-bearing defect. The reader's weakest assumption—that logged taint labels may undercount implicit semantic influence—is real and explicitly acknowledged in §3.5 and §4.4, but it attaches to the diagnostic propagation axis, not to the central sufficiency claim. The utility leg is computed from authorized structured-action completion and the structured oracle B(τ), not from taint-label coverage: v1 completes 0.1642 vs v2's 0.8567 (delta 0.6925, cluster CI [0.6160, 0.7637]). Since the same 441/600 pairs drive both trajectory and utility divergence, and Cases 2–3 show v1 blocks the authorized invite while v2 commits it, the equal-endpoint concealment is grounded in observable action differences. The paper explicitly scopes logged spread as operational rather than semantic, and Table 14 shows the authors do not overclaim propagation rankings. The oracle-ledger assumption is also explicitly conditional. I therefore see no reason to move the reader's ACCEPT.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"ContainmentBench proposes a trace-based evaluation protocol for post-exposure containment in tool-using LLM agents, decomposing containment into four observables: benchmark-defined endpoint policy compliance, instrumented logged propagation, recovery instrumentation, and authorized structured-action completion. In a pre-specified, hashed 17,640-rollout study with Qwen2.5-7B-Instruct, the paper reports that all 600 matched active-tainted pairs under a taint-only firewall (v1) and an intent-ledger firewall (v2) share the same zero committed-harm endpoint, yet 441/600 (73.5%) differ on the logged-trajectory or utility axes; authorized tainted-action completion is 0.1642 for v1 versus 0.8567 for v2 (scenario-cluster paired delta 0.6925, 95% CI [0.6160, 0.7637]), with a strong tool-boundary baseline at 0.9233. A third finding is that aggregate logged-spread rankings are stage- and denominator-sensitive. The paper concludes that a terminal policy label is not a sufficient statistic for operational post-exposure containment and recommends reporting endpoint, stage-stratified trajectory, and utility evidence separately. The full-scale evidence is explicitly scoped as synthetic, single-model, and, for the policy case study, conditional on a correct structured authorization ledger.","tokens_in":23541,"tokens_out":19566,"duration_ms":200129,"significance":"The central claim is well supported and, if adopted, would change evaluation practice: endpoint-equivalent policies can differ sharply on authorized utility and on logged trajectory, and propagation comparisons are only interpretable when stage composition and denominators are fixed or stratified. The paper's strengths should be credited explicitly: the evaluation is pre-registered in the frozen, hashed configuration; the no-defense condition supplies a nonzero positive control in the security stage (0.3354) and in the active-tainted stage; parser-repair diagnostics with invalid-tool non-execution close a measurement-validity gap; inference uses paired and scenario-cluster bootstraps; the authors explicitly decline to make comparative claims from non-discriminative endpoints (secret leakage, memory reinfection, rollback); and the artifact includes integrity checks and smoke reproduction. The outcome-conditioned design is notably direct: endpoint equality holds for every matched pair in Table 10, so the divergence numbers are not an artifact of subset selection. The claim is falsifiable, and its main evidentiary weight is carried by the utility axis, which does not depend on taint-l","major_comments":[],"minor_comments":[{"comment":"The headline 73.5% trajectory divergence for the v1/v2 contrast coincides exactly with the utility divergence (intersection = union = 441/600). As presented, the trajectory axis adds no independent information in this particular contrast; the independent trajectory evidence comes from the tool-boundary/v2 rows (16.23% all-stage; 5.10% security-stage), where utility is tied or differs far less. The coincidence is disclosed in the text, but I suggest making the evidentiary division of labor explicit, and optionally reporting divergence among utility-outcome-matched pairs, so that the four-axis decomposition is not read as crediting the trajectory axis with independent detection in the headline number.","section":"§7.1, Table 10"},{"comment":"The trajectory-divergence and logged-spread statistics are defined over logged taint-propagating edges, and §3.5 concedes that paraphrased or implicit influence can be under-counted. Since the headline 73.5% divergence is partly computed from taint-derived projection fields, state explicitly that the measured divergence is a lower bound under incomplete logging, and that the utility leg (0.1642 vs 0.8567) is the logging-independent carrier of the endpoint-insufficiency claim. This would preempt the natural objection that the spread/trajectory rankings reflect logging coverage rather than containment differences.","section":"§3.5, §4.4"},{"comment":"All full-scale operating points come from Qwen2.5-7B-Instruct; the Mistral diagnostic cannot carry a policy comparison because parse failure remains 0.63–0.66. The scope statements in §3.7 and §8.4 cover this, but the conclusion re-states the three findings without reminding the reader that the concrete values (0.1642/0.8567/0.9233) are single-model, synthetic-scenario measurements, whereas the measurement claim is protocol-level. One sentence in the conclusion echoing the abstract's scope would harden that boundary.","section":"§6.1, Table 22, §8.4"},{"comment":"The seed-variance table reports 'Std.' and 'Range' columns with dashes for the security, memory/recovery, and clean rows. Specify whether those dashes mean 'not computed', 'all seeds identical', or 'omitted as non-applicable', since this table is the main evidence that the headline results are not dominated by a single rollout seed.","section":"Table 26"},{"comment":"Figure 2's bars have no y-axis label or gridlines. Adding a labeled y axis (completion proportion) would make the figure legible without forcing the reader to cross-reference Table 11.","section":"Fig. 2"},{"comment":"The two adaptive artifacts (no-defense-optimized transfer diagnostic and target-preserving qualification) serve different protocol roles and support different claims. Consider adding a one-row-per-diagnostic table that maps each artifact to its optimization target, evaluation condition, and the claim it is permitted to support, since Tables 17–19 are easy to conflate.","section":"§6.3, Tables 17–19"}],"recommendation":"minor_revision","confidential_remarks":"The paper is unusually careful for this literature: frozen and hashed configuration, explicit positive controls, parser diagnostics, cluster-bootstrap inference, and a visible effort to avoid overclaiming non-discriminative endpoints. My recommendation of minor_revision rather than accept reflects only the local clarifications listed above; I regard the central measurement claim as sound. The single-model, synthetic full-scale evidence is a real limitation but is scoped appropriately by the authors. No concern about novelty or fit: the paper positions itself precisely relative to Fides, AgentDojo, and NeuroTaint."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. This is a solid benchmark paper that earns a careful referee. The new thing is not taint tracking or intent-aware authorization—those exist—but the measurement contract: a stage-structured trace protocol that separates endpoint compliance, logged propagation, recovery, and authorized utility, with explicit denominators and validity prerequisites. The empirical centerpiece is the outcome-conditioned comparison: 600 matched pairs with identical zero committed-harm endpoint, yet 73.5% diverge on trajectory or utility, and authorized completion jumps from 0.164 to 0.857 between taint-only and intent-ledger policies. That is a real, non-tautological demonstration that endpoint labels hide operational differences.\n\nThe paper is unusually careful about its own validity. The 17,640-rollout study is frozen and hashed, with no-defense positive controls, parser repair diagnostics, invalid-tool non-execution, and both paired and scenario-cluster bootstraps. They also decline to promote non-discriminative endpoints—secret leakage, memory reinfection, rollback—into comparative claims. The artifact and reproduction scripts appear substantive.\n\nSoft spots are real but mostly scoped. The full-scale evidence is one model and synthetic workflows. The intent ledger is an oracle assumption; the paper says plainly the result is conditional on a correct ledger. Taint-log completeness is acknowledged as an instrumentation definition, not a semantic proof, and it mainly affects the propagation diagnostics, not the utility-based sufficiency argument. The Mistral diagnostic and AgentDojo adapter are honestly reported as mechanical portability checks, not cross-model or external validation. None of this undermines the central measurement claim.\n\nI slightly disagree with the reader's emphasis on taint-log completeness: it's a caveat, not a threat to the endpoint-insufficiency finding, because the utility leg and the case-study traces carry that claim independently. The bigger limitation is single-model, but that is an extension, not a flaw in what they claim.\n\nWho benefits: anyone evaluating LLM-agent defenses or designing agent-security benchmarks. A serious referee should engage; I'd expect accept after moderate revision, mostly adding external-model evidence and perhaps more discussion of oracle-ledger sensitivity. Not a paradigm shift, but a durable contribution to how we measure containment.","headline":"A solid, carefully scoped benchmark paper that makes a real measurement point—endpoint-only scores hide trajectory and utility differences—supported by a frozen, well-controlled study; the caveats are single-model and synthetic workflows, both admitted.","tokens_in":24018,"tokens_out":2690,"would_cite":true,"duration_ms":28055,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A terminal security label is not enough to judge post-injection containment; equal zero-harm endpoints can hide both severe over-blocking of authorized work and stage-sensitive differences in how far untrusted influence spreads.","keywords":["prompt injection","LLM agents","containment","trace-based evaluation","taint tracking","authorization ledger","endpoint insufficiency","runtime enforcement"],"falsifier":"Run the frozen 17,640-rollout protocol with an independent semantic/causal auditor applied to the same traces; if the 73.5% outcome-conditioned divergence rate largely disappears when implicit influence is counted, the trajectory-insufficiency claim would collapse. Alternatively, perturb the structured ledger with omissions or over-authorizations; if the utility repair vanishes or endpoint violations appear, the policy case study would be overturned.","tokens_in":23220,"feed_emoji":"🔒","tokens_out":4512,"duration_ms":34347,"temperature":0.7,"pith_summary":"Tool-using LLM agents can be redirected by untrusted content they read, and existing benchmarks judge the damage by the final outcome: did the injection cause an unauthorized commit? This paper argues that the final endpoint is not enough. It introduces a trace-based benchmark that separately records endpoint policy compliance, logged influence spread, recovery behavior, and whether authorized work over external content still gets done. In a frozen 17,640-rollout study, all 600 matched policy pairs had the same zero-committed-harm outcome, yet 73.5% diverged in logged trajectory or utility; a taint-only firewall that looked secure completed only 16% of authorized tainted workflows, while an intent-ledger policy reached 86%. The paper concludes that containment is a trajectory property, and evaluations should report endpoint, stage-stratified path, and utility evidence separately.","feed_headline":"73.5% of equal-outcome rollouts diverge in trajectory or utility","feed_subtitle":"Final security labels can call a policy safe while it blocks most authorized work over untrusted content.","key_machinery":"The central object is the logged trace graph Gτ with nodes for sources, messages, tools, memory, authorization decisions, and commits, and the taint closure Tτ defined as reachability over taint-propagating edges only. On top of this, the benchmark separates four measurement axes: endpoint policy compliance, instrumented logged propagation (normalized blast radius), recovery instrumentation, and authorized structured-action completion. The intent-ledger policy adds a representation Is = {(a,t,φ)} of authorized actions, targets, and argument predicates, with an exact-match authorization predicate that blocks tainted authority expansion while allowing tainted content inside an authorized actio","core_discovery":"The paper's central claim is that a terminal policy label is not a sufficient statistic for operational post-exposure containment. Using a sandboxed benchmark that logs every message, tool proposal, memory write, authorization decision, and commit as a directed trace graph, it shows that two runtime policies can produce identical zero-violation endpoints while differing sharply in how far untrusted influence traveled and how much authorized work was blocked. In the matched active-tainted comparison, taint-only enforcement and intent-aware authorization both ended with zero committed harm, but 73.5% of the 600 pairs differed on the shared trajectory projection or utility, with authorized comp","pith_inferences":["If the trace-completeness limitation were fixed, for example by pairing online labels with an offline semantic/causal audit, the same benchmark could estimate how much implicit, paraphrased influence currently escapes the taint labels.","The endpoint-insufficiency principle likely extends beyond prompt injection to any agent safety evaluation where a blocker can mask failures; the four-axis separation could serve as a template for content-safety and tool-use audits.","The intent-ledger result suggests a testable design target: a runtime policy that combines tool-boundary checks with field-level provenance might reach both the 0.92 utility of the boundary baseline and lower spread, a hypothesis the current data does not resolve.","Because the full study is single-model and synthetic, the natural next experiment is to run the same frozen protocol on additional models and on an external dynamic task environment once positive controls hold."],"forward_implications":["A defense that reports zero committed violations can still be harmful in deployment if it blocks most authorized work that legitimately uses external content.","Containment evaluations should report at least endpoint, stage-stratified trajectory, and utility evidence as separate numbers, not one aggregate score.","Aggregate logged-spread rankings are only meaningful with fixed evidence-stage composition and denominator; the same policies can trade places under different normalizations.","A trusted structured ledger can repair over-blocking without lowering the observed security endpoint, but may still sit below a simpler tool-boundary policy on utility.","Recovery-related metrics (secret leakage, memory reinfection, rollback superiority) should only be used comparatively once a no-defense positive control shows they can discriminate."],"fun_headline_variants":["Zero harm endpoints hide 73.5% trace divergence","Same safe outcome, 73.5% different traces","Terminal labels miss 73.5% containment differences","When 'safe' blocks 84% of authorized work","Endpoint compliance ≠ true containment: 73.5% differ"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The trajectory and spread results assume the logged taint labels capture every semantically relevant path of untrusted influence; if paraphrased or implicit influence bypasses the instrumentation, logged spread and divergence measure logging coverage rather than containment.","fun_headline_variants_meta":{"raw":{"variants":["Zero harm endpoints hide 73.5% trace divergence","Same safe outcome, 73.5% different traces","Terminal labels miss 73.5% containment differences","When 'safe' blocks 84% of authorized work","Endpoint compliance ≠ true containment: 73.5% differ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1301,"prompt_tokens":797,"completion_tokens":504,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":422}},"tokens_in":541,"tokens_out":504,"duration_ms":5128,"temperature":1.0,"reasoning_tokens":422,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:17:06.444049+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the frozen 17,640-rollout protocol with an independent semantic/causal auditor applied to the same traces; if the 73.5% outcome-conditioned divergence rate largely disappears when implicit influence is counted, the trajectory-insufficiency claim would collapse. Alternatively, perturb the structured ledger with omissions or over-authorizations; if the utility repair vanishes or endpoint violations appear, the policy case study would be overturned.","supporting_citations":[],"review_version":1}