{"id":"e2ddd7f3-d1d6-41ba-9f25-4b2cc601ae30","arxiv_id":"2607.07405","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Deterministic read-only pre-execution gates raise τ²-bench airline success from 29.6% to 42.0% on gpt-4o-mini by blocking silent policy-violating tool writes, with the lift replicated on disjoint seeds.","lead":"LLM agents can break domain rules while tools silently accept the bad writes, leaving a wrong system state with no error. Simple deterministic pre-execution checks recover a large share of those silent failures on an airline agent benchmark and raise success rates.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the reader's already-flagged same-task gate authorship; the bounded claim holds under the disclosed evidence.","rationale":"The Reader correctly identifies the paper as a careful, bounded reliability study whose measured recovery of task success (not the enforcement idea itself) is the contribution, and correctly isolates same-task gate authorship without held-out freeze as the weakest assumption. After re-reading the full manuscript, no stronger internal inconsistency, statistical artifact, or unacknowledged confound appears that would move the verdict. The firing-share decomposition, per-gate precision audit (Table 6), seed-level replication, and negative controls already bound the mechanism tightly enough for a workshop evaluation track. The concrete held-out-task freeze test is the natural next verification step the paper itself recommends; until it is run (or code ships so others can run it), CONDITIONAL remains the right posture. No change to the Reader's verdict is warranted.","tokens_in":12042,"tokens_out":556,"duration_ms":5678,"concrete_test":"Freeze the four-gate suite exactly as published, randomly hold out 15 of the 50 airline tasks, re-author or re-select nothing, and re-run the full n=5 (or n=15) comparison only on the held-out tasks; if the gated-vs-vanilla Δ remains positive and statistically non-zero on the held-out set (and cancellation_eligibility precision stays high), the task-overfitting concern is substantially reduced.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No additional load-bearing concern is identified that would overturn or further weaken the paper's central claim beyond what the Reader already flags. The strongest claim is carefully scoped to policy-permissive tools with state-decidable rules and final-state evaluation; the +12.4pp lift (replicated +12.3pp on disjoint seeds), firing-stratum concentration (+19.2pp on 26/50 tasks), 100% precision of the load-bearing cancellation_eligibility gate over 161 fires, and negative controls (retail, BFCL) jointly support that deterministic pre-execution gates recover a measurable fraction of silent wrong-state failures while blocking the forbidden write. The same-task-set authorship of gates without a held-out-task freeze (§7(8)) remains the primary structural risk, but it is already disclosed, the dominant gate's precision audit supplies partial counter-evidence against pure task-fitting, and the claim does not assert domain generality or automatic gate correctness. Other disclosed limits (unreplicated frontier arm, missing prompting/reflection baselines, no recovery guarantee) are correctly treated as non-central.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper identifies silent policy-violating writes on policy-permissive tools as a distinct trust failure in tool-using LLM agents: a well-formed mutating call executes, the final state is wrong, and no tool error appears in the trace. On τ²-bench airline with gpt-4o-mini, 78% of observed failures are of this class. The intervention is a lightweight suite of deterministic, read-only pre-execution gates that inspect the proposed call and current database state before allowing a write. A four-gate suite raises full-benchmark pass1 from 29.6% to 42.0% (+12.4pp; paired task-level bootstrap P=0.0012), and the lift reproduces on a disjoint 15-seed set (+12.3pp; P=0.0008). Lift concentrates on the 26/50 firing tasks (+19.2pp); non-firing movement does not exclude zero. A per-gate audit shows cancellation_eligibility carries the lift at 100% precision over 161 fires, while baggage_allowance has only 5% precision. Negative controls (self-enforcing retail; BFCL) bound the mechanism. Suggestive unreplicated frontier evidence (gpt-5.2) is reported separately. The claim is explicitly bounded: gates deterministically block a known class of forbidden writes at the action boundary but do not guarantee task success or domain generality.","tokens_in":12356,"tokens_out":1236,"duration_ms":9822,"significance":"If the result holds, it is a useful reliability contribution for agent evaluation and deployment. It cleanly separates loud tool errors from silent wrong states, shows that action-boundary enforcement can raise final-state task success (not only bound safety cost) when tools are policy-permissive, and supplies a falsifiable admission criterion (A1–A5) for when such gates can help. Strengths include the paired bootstrap, full 50-task results, disjoint-seed replication within 0.1pp, firing vs non-firing stratification, per-gate precision audit, pass_k reliability curves, and two negative controls. The paper is careful about scope and does not overclaim generality. The main practical implication is that deterministic mediation at the write boundary remains valuable even as models improve, because it changes a system property rather than only the probability of violation.","major_comments":[{"comment":"§7(8) and §3.3 / Table 1: Gates were authored from the domain policy and selected on the same 50-task airline set used for evaluation; replication is over seeds (n=5 original, n=15 disjoint), not over held-out tasks with frozen gates. This is the primary structural risk for the reliability claim. The dominant gate’s 100% precision over 161 fires (Table 6) is partial counter-evidence against pure task-fitting, and the paper discloses the limitation, but a held-out-task freeze (or at least an explicit leave-one-task-out / frozen-predicate check for cancellation_eligibility) would make the central result substantially more robust. Without it, the free parameters of suite composition and predicate encoding remain under-constrained relative to the strength of the reliability language.","section":null},{"comment":"§5.5 / Table 6: The headline four-gate suite is not uniformly beneficial. cancellation_eligibility is load-bearing (removal Δ = −2; 100% precision); baggage_allowance has 5% precision (2 true / 40 false blocks) and removal improves pass1. The paper correctly states that gate precision must be audited, but the main text still reports the four-gate suite as the primary intervention. Either demote low-precision gates from the headline configuration, or report a precision-audited single-gate (or two-gate) primary result alongside the suite so that the load-bearing mechanism is not diluted by known false blocks.","section":null}],"minor_comments":[{"comment":"§5.3 / Table 4: Frontier stratification is point estimates only (n=5, unreplicated). The text already labels this as suggestive; ensure the abstract and conclusion never allow the +10.4pp frontier number to be read as co-equal with the replicated budget result.","section":null},{"comment":"§4.1: The evaluation-replay scrub is necessary for fair comparison but is easy to miss. A short diagram or one-sentence pseudocode of the pre-call dispatcher + scrub path would help readers reproduce the harness difference.","section":null},{"comment":"§7(7): The missing prompting / reflection baselines are correctly listed as limitations. A single sentence clarifying that deceptive task #48 cannot be fixed by prompting (user asserts false state) would further separate that case from the open non-deceptive failures.","section":null},{"comment":"Table 1 vs Table 6: basic_economy appears in the audit but not in the headline suite; a one-line note in Table 1 or §3.3 would avoid reader confusion about the candidate set.","section":null},{"comment":"§6.3: Subset inflation (+30pp on 8 tasks) is a useful caution; consider moving the exact subset size and selection criterion into the main text rather than leaving it slightly underspecified.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Fit for a workshop on evaluation and trustworthiness of agentic AI is strong; the contribution is bounded and empirical rather than a new general theory of agent safety. The same-task gate authorship is the only issue that could have justified major_revision if undisclosed; because §7(8) already flags it and the precision audit partially mitigates it, minor_revision is proportionate. No novelty or citation-pattern concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple. On τ²-bench airline, most budget-agent failures are silent wrong states—no tool error, just a forbidden write that the permissive tool executes. A thin pre-execution gate layer raises pass1 from 29.6% to 42.0% (+12.4pp, P=0.0012), the lift replicates on a disjoint 15-seed set (+12.3pp), and almost all of the gain sits in the firing stratum. That is a real reliability result, not a safety-only story.\n\nWhat is new is not the gate idea (they cite AEGIS, AgentSpec, etc. and say so). It is the failure-mode characterization, the five admission axes for when a benchmark can even see this class, the firing-share decomposition, the per-gate precision audit, and the negative controls (retail self-enforcing tools and BFCL) that bound where the mechanism helps. The load-bearing gate, cancellation_eligibility, is 100% precise over 161 fires; baggage_allowance is 5% and they report that honestly. pass_k curves show the consistency gain, not just average success. Limitations are listed without spin: one positive domain, unreplicated frontier arm, no recovery guarantee, missing prompting/reflection baselines, and same-task-set gate authorship with seed-only replication.\n\nThe soft spots are real but already disclosed and proportional. Hand-writing gates from the policy and evaluating on the same 50 tasks is the main structural risk; a frozen held-out-task test would be cleaner. Code is promised but not yet out. The frontier numbers are suggestive only. None of that overturns the central claim as scoped: in policy-permissive, state-decidable, final-state settings, deterministic blocks can recover a measurable fraction of silent corruptions.\n\nThis is for people building or evaluating tool agents who care about silent state corruption rather than loud tool errors. Workshop evaluation track is the right home. I would send it to peer review; the claim matches the evidence and the authors do not oversell it. Worth engaging once the code lands.","headline":"Bounded, carefully measured result: silent wrong-state failures are real on policy-permissive tools, and one high-precision gate recovers a replicated +12pp of final-state success without claiming more than the evidence supports.","tokens_in":12979,"tokens_out":548,"would_cite":true,"duration_ms":5613,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Deterministic pre-execution gates recover silent policy-violating writes in tool-using LLM agents, lifting airline-task success by about 12 points without relying on more model reasoning.","keywords":["LLM agents","tool use","policy compliance","deterministic verification","agent evaluation","runtime enforcement","agent reliability","silent wrong state"],"falsifier":"Freeze the four gates, evaluate them on a held-out set of new airline (or analogous) tasks never used for gate selection, and check whether the success lift and the concentration on firing tasks disappear or reverse.","tokens_in":12894,"feed_emoji":"🛡️","tokens_out":938,"duration_ms":9653,"temperature":0.7,"pith_summary":"Tool-using language-model agents can break the domain policies they are supposed to enforce and still look successful. When a write tool accepts any well-formed call, a forbidden state change (a cancelled booking, a changed passenger count) happens with no error and no warning in the transcript. In the airline domain of τ²-bench this silent wrong-state class accounts for most failures of a budget agent, and the failure rate is stable across seeds. The paper shows that a thin layer of deterministic, read-only checks placed just before each write can catch many of those violations: a four-gate suite raises full-benchmark success from 29.6 % to 42.0 % on gpt-4o-mini, and the same lift reappears on a disjoint 15-seed replication. The gain concentrates on the tasks where the gates actually fire; negative controls in a self-enforcing retail domain and in BFCL show little benefit once tools already police themselves. The claim is deliberately bounded: gates do not guarantee eventual task success, but they do give a deterministic block on a known class of silent policy-violating writes at the action boundary.","feed_headline":"Gates lift agent success 12 points by blocking silent policy breaks","feed_subtitle":"Read-only checks before writes recover forbidden airline actions that tools would otherwise execute quietly","key_machinery":"A four-gate suite of pure, read-only predicates that inspect a proposed tool call and the current database state and either allow or reject the write before mutation; the dominant high-precision gate is cancellation_eligibility.","core_discovery":"In policy-permissive tool environments whose rules are decidable from current state and call arguments, silent policy-violating writes form a recurring, reproducible failure class; lightweight deterministic pre-execution gates can recover a measurable fraction of those failures (gpt-4o-mini airline pass@1 29.6 % → 42.0 %, +12.4 pp, replicated +12.3 pp) while guaranteeing that the forbidden write is blocked whenever the gate fires.","pith_inferences":["The same pattern is likely to appear in any customer-service or enterprise agent whose tools accept well-formed writes while policy lives only in natural-language instructions.","Automatic synthesis or formal verification of gate predicates from policy documents would reduce the hand-authoring bottleneck and the risk of task-fitting.","Combining a high-precision gate with a lightweight recovery prompt after rejection could convert more blocked violations into completed tasks.","Future agent benchmarks will need deliberate violation-inducing tasks; without them the silent-policy-violation class remains invisible."],"forward_implications":["Benchmarks and deployment harnesses should separately report silent wrong-state failures versus loud tool errors.","When tools are policy-permissive and rules are state-decidable, a cheap deterministic gate can raise final-state task success, not only safety.","Gate precision itself must be audited; one high-precision gate can carry almost all of the measured lift.","Stronger models reduce but do not eliminate the silent-violation class, so action-boundary checks remain useful at the frontier.","Where tools already self-enforce their preconditions, redundant gates add little and can even introduce regressions."],"fun_headline_variants":["Deterministic gates lift agent success 12pp by blocking silent policy breaks","Pre-execution gates recover silent wrong-state failures in tool agents","Read-only gates raise airline pass rate 12pp by barring forbidden writes","Gates block silent policy violations and add 12 points of agent success","Deterministic pre-write checks recover a recurring LLM agent failure mode"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The hand-written gate predicates, chosen from the domain policy and tuned on the same airline task set used for evaluation, correctly capture the load-bearing rules without fitting idiosyncrasies of those particular tasks.","fun_headline_variants_meta":{"raw":{"variants":["Deterministic gates lift agent success 12pp by blocking silent policy breaks","Pre-execution gates recover silent wrong-state failures in tool agents","Read-only gates raise airline pass rate 12pp by barring forbidden writes","Gates block silent policy violations and add 12 points of agent success","Deterministic pre-write checks recover a recurring LLM agent failure mode"]},"model":"grok-4.5","effort":"low","cost_usd":0.004668,"raw_usage":{"total_tokens":1498,"prompt_tokens":984,"num_sources_used":0,"completion_tokens":97,"cost_in_usd_ticks":46680000,"prompt_tokens_details":{"text_tokens":984,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":417,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":984,"tokens_out":97,"duration_ms":5676,"temperature":1.0,"reasoning_tokens":417,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T15:46:07.754199+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Freeze the four gates, evaluate them on a held-out set of new airline (or analogous) tasks never used for gate selection, and check whether the success lift and the concentration on firing tasks disappear or reverse.","supporting_citations":[],"review_version":2}