{"id":"1dd7a373-b4fd-4e25-8fe4-7a32b1d291e8","arxiv_id":"2608.01619","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"State-to-draft auditing with provenance-verified transitions raises STALE strict-protocol accuracy from .686 to .736, a +5.0 point paired gain led by implicit policy adaptation and premise resistance.","lead":"This paper identifies a failure mode in personalized AI agents: a response can silently depend on an outdated fact about the user even when the agent knows the fact changed. The authors' StateAuditor audits from stored state to the draft and repairs such responses, gaining about 5 points on the STALE benchmark over a locked predecessor.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline gain is read through an LLM judge that shares a family with the system; human validation covers only the scenario-joint arm, not the strict 400-scenario claim.","rationale":"The reader pinned the soft spot on retrieval recall, which is a real deployment boundary; on the benchmark itself, however, the gain is already measured under that recall, and the matched control closes the evidence/call-budget alternative. The more direct threat to the central claim is measurement: the strict headline depends on a rubric judge from the same family as the proposer/adapter, and the only human validation was run on the scenario-joint upper bound, not on the strict 400-scenario comparison. The low human agreement (κ=0.126) on those scenario-joint labels weakens confidence in the absolute scores, even though paired human comparisons were significant. This does not mean the claim is false—the disjoint Gemini reproduction, the three-draw stability, the deterministic local re-judge, and the matched control are real supporting evidence—but it does mean the central causal attribution has not yet been checked by a blind human preference test on the exact headline protocol. Reader's CONDITIONAL verdict remains appropriate; adding a strict-arm blind human replication as an explicit condition would make the required check concrete.","tokens_in":13231,"tokens_out":6728,"duration_ms":68788,"concrete_test":"Draw a fixed random sample of 100 scenarios from the strict 400; for each, show two annotators the predecessor and VTA final responses in randomized order, with judge labels, system names, and scenario indices hidden, and ask only 'which response better reflects the user's current stated state?' rather than which is more fluent. Pre-register the paired McNemar test for a VTA-over-predecessor majority. If the human majority on the strict arm does not reproduce the LLM judges' significant ordering, the headline gain is a judge artifact; if it does, the judge-reliability concern is settled and retrieval recall remains the main external boundary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a measured score difference under a rubric-based LLM judge. In the strict arm, proposer, adapter, and judge are all gpt-5.5 (§6.1); the disjoint Gemini judge is still an LLM rubric judge scoring the same open-ended behavior. The paper's own human check was run only on the scenario-joint upper-bound arm, not on the strict 400-scenario comparison, and the two annotators' per-cell agreement is κ=0.126 (Limitations, §7). A +5.0-point macro gain is small relative to plausible judge sensitivity to surface features of repaired outputs—explicit acknowledgment sentences, length, or phrasing—which the matched control holds constant only in evidence/call budget, not in output text. If gpt-5.5 and Gemini reward the repair's form rather than its state adaptation, the causal attribution 'transition machinery, not added context/calls' is not established. This is distinct from the retrieval-recall boundary: retrieval limits how often the transition can fire, but the judge determines whether the measured gain reflects the intended behavior.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims that draft-anchored verification of generated responses misses implicit stale dependencies in open-ended personalized requests, and proposes StateAuditor, a post-generation audit-repair pipeline whose core is provenance-verified transition assembly (VTA): extract user-state premises from draft and question, build a state timeline, propose candidate old-to-new transitions from timestamped evidence, deterministically validate quotation grounding and chronology, and regenerate under typed directives. On STALE's full-history strict protocol, strict single-query VTA scores .736 versus .686 for the locked predecessor, a +5.0-point paired gain (95% CI [+2.9,+7.2], 84:37 scenarios), reproduced by the disjoint Gemini judge (.738 vs .680); a matched no-transition control scores .692 (+0.6, n.s.), which the paper uses to attribute the gain to the transition machinery. A privileged scenario-joint variant (.879) is reported explicitly as an upper bound. External transfer on HorizonBench is mostly attributable to the draft-side audit, and a harder authored set shows no gain; over-correction, cost, and latency are measured, and artifacts are released.","tokens_in":13356,"tokens_out":9347,"duration_ms":88399,"significance":"If the result holds, this is a useful and unusually well-scoped contribution: it identifies a structural blind spot in claim-based verification, introduces a deterministic provenance/chronology gate, and backs the headline with a matched control, a disjoint-judge reproduction, contamination audits, and released per-item records. The paper is also commendably explicit about its boundaries: post-development full-set use, same-family judge risk, a failed 4B distillation, a harder set with no gain, and no deployment claim. The main residual risk is that the headline effect is a small (+5.0-point) measured difference under LLM rubric judges over outputs that differ in surface form, with no human validation on the strict 400-scenario arm; this is the load-bearing uncertainty in the causal attribution.","major_comments":[{"comment":"The central claim that the transition machinery, rather than output form or added context, causes the +5.0-point gain is not yet fully supported. In the strict arm the proposer, adapter, and primary judge are all gpt-5.5, and the disjoint Gemini judge is still an LLM rubric judge; the paper's own limitations section acknowledges the same-family self-preference risk. The human evaluation covers only the scenario-joint ranking, with per-cell annotator agreement κ=0.126, and does not validate the strict 400-scenario comparison. Repaired outputs differ from predecessor outputs in surface form (for example, the one-sentence acknowledgment in correct-and-inform), and the matched no-transition control holds evidence and call budget constant but not output text. A blind human evaluation on a sample of strict-arm changed cells, or an equivalent surface-form control (for example, scoring de-acknowledged or length-matched outputs), is needed to establish that the gain reflects state adaptation rather than judge preference for repair form; without it, the causal attribution should be stated as conditional on rubric-judge validity.","section":"§6.1, Table 1; §5 Blind human protocol; §7 Limitations"},{"comment":"The strict VTA gain can only fire when the old/new evidence pair is retrieved, and the pooled strict window-level recall is .580, so on roughly 40% of the 400 scenarios the transition machinery has no verified pair to act on. The paper is honest that the audit can adjudicate only retrieved evidence, but the headline comparison currently mixes the retrieval stage with the repair stage. I request a decomposition of the strict paired gain on the subset of scenarios where the old/new pair was retrieved versus the complement, so readers can see how much of the +5.0 points is attributable to repair conditional on successful retrieval and how sensitive the result is to retrieval quality. This is a boundary analysis rather than an invalidation, but it is important for interpreting the operating envelope of the method.","section":"§4.1 and §6.1"}],"minor_comments":[{"comment":"The abstract says 'lets only these verified transitions trigger repair,' but §4.5 states that without a verified transition the system falls back to base-audit verdicts whose material STALE/UNKNOWN findings can still fire repair or verify directives. This is clarified in the body, but the abstract should be reworded to say that the gate constrains the transition channel, not the base audit, to avoid overstating the gate's authority.","section":"Abstract and §4.5"},{"comment":"Given the low per-cell agreement (κ=0.126), the paper should report the raw per-annotator confusion matrices and consider a third adjudicator or a consensus pass; as written, the per-annotator McNemar tests are suggestive but the absolute reliability of the labels is hard to assess.","section":"§5 Blind human protocol"},{"comment":"The paper reports three independent full-400 draws of strict VTA scoring .733±0.002, while the implementation note says each arm is one sampled output per item unless identified as a replicate; please clarify exactly how the three draws differ and whether judge stochasticity is included in the reported confidence intervals.","section":"§6.1 and Implementation details"},{"comment":"Table 2 lists strict VTA evidence as 'per-query+exp.' while the predecessor row is 'per-query'; since the matched control is described as having the same evidence as strict VTA, the paper should state explicitly that the control also received the expanded evidence, so readers do not confuse the evidence-expansion effect with the transition-machinery effect.","section":"Table 2 and §6.1 matched control"},{"comment":"The LoCoMo false-invalidation rate is correctly labeled an upper bound; consider also reporting a precision-oriented subset with gold conflict labels, if any exist, to complement the construction-labeled safety suite.","section":"§6.4"}],"recommendation":"major_revision","confidential_remarks":"This is a solid, honest paper whose main gap is the absence of human validation on the strict headline arm. The reviewer's skeptic concern about same-family rubric judges is legitimate and is not fully offset by the disjoint Gemini reproduction, since that is also an LLM rubric judge. I would not reject: the controlled comparisons, released artifacts, and explicit boundaries make the central claim defensible, but the causal attribution needs either additional human evidence or a more carefully scoped claim. The paper fits the journal's scope well."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper has a genuinely new verification direction: instead of decomposing the draft and checking claims, it audits from the stored state to the draft, catching stale dependencies the draft never states. Second, the headline +5.0 point gain on STALE is better supported than most results in this area: locked predecessor, frozen drafts, paired CIs, a disjoint Gemini judge, a replicate draw, and a matched no-transition control that attributes the gain to the transition machinery rather than added context or calls.\n\nWhat the paper does well: the authors are unusually honest about boundaries. They report the .879 scenario-joint result as a privileged upper bound, they report a harder authored set with no gain, they report false invalidation rates by regime, and they explicitly say \"verified\" means provenance and chronology, not semantic correctness. The deterministic validation gate is a nice engineering choice, and the released records make the numbers checkable.\n\nThe stress-test concern about the LLM judge rewarding repair form rather than state adaptation is worth taking seriously, but it does not land as a fatal objection. The disjoint Gemini judge reproduces the gain, and the matched control holds evidence and call budget fixed. The larger soft spot is human validation: only the scenario-joint arm got blind human labels, and per-cell agreement is low (κ=0.126). Both annotators still reproduce the ranking, so the paired claim survives, but the absolute .736 number should not be treated as a calibrated accuracy. Retrieval recall of .580 pooled also bounds the mechanism; the authors say themselves that on many scenarios the transition machinery has nothing verified to fire on. The repair-only operating point is post-hoc selected, which is an honest analysis but needs a pre-registered replication before deployment confidence.\n\nWho this is for: anyone building memory-augmented agents or auditing pipelines for personalized responses. It deserves a serious referee. The design is careful, the claim is bounded, and the artifacts are released. I'd suggest asking for an expanded human study or a pre-registered replication before treating the result as settled, but the core direction and evidence are solid.","headline":"A well-controlled systems paper with a genuinely new state-to-draft verification direction; the +5.0 point STALE gain is credible but the human-validation gap and retrieval recall bound keep it from being settled.","tokens_in":13975,"tokens_out":1952,"would_cite":true,"duration_ms":18382,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Auditing stored state against the draft, rather than the draft against memory, repairs stale-dependency failures that persist even when an agent knows a stored fact is outdated, lifting accuracy by 5.0 points on STALE's 400-scenario…","keywords":["memory-augmented agents","implicit policy adaptation","stale memory","state-to-draft verification","provenance-verified transition assembly","personalized assistants","response repair","STALE benchmark"],"falsifier":"Run the strict STALE protocol again with the change-scan retrieval slots disabled—the query-independent disclosure windows that recover superseding evidence—and compare strict VTA with the matched control; if the +5.0-point gain survives without that evidence being retrieved, the paper's attribution of the gain to the transition machinery rather than retrieval is wrong.","tokens_in":12939,"feed_emoji":"🧠","tokens_out":12848,"duration_ms":98737,"temperature":0.7,"pith_summary":"The paper tries to establish that a personal assistant can know a stored fact is outdated and still produce a response built around the old value, and that this failure has a fixable structural cause on the response side. The cause it identifies is draft-anchored verification: checking what a response says misses the stale dependencies the response never states. Its remedy audits in the opposite direction, from stored state to draft, and lets only provenance-verified old-to-new transitions trigger a typed repair. On STALE's full 400-scenario protocol this state-to-draft repair scores .736 against .686 for the paper's locked predecessor under the same judge, a +5.0-point paired gain concentrated in implicit policy adaptation and premise resistance, and a matched control shows the gain does not come from added evidence or calls. If the claim is right, it gives a concrete, evidence-grounded way to make personalized agents behave consistently with a user's current state without retraining the generator.","feed_headline":"Auditing stored state, not the draft, lifts stale-response accuracy","feed_subtitle":"The reversed audit beats its predecessor by 5.0 points on STALE's 400-scenario protocol.","key_machinery":"The load-bearing mechanism is provenance-verified transition assembly (VTA): an LLM proposes candidate old-to-new transitions from timestamped memory entries, and deterministic code authorizes a transition only when each evidence quotation matches at least 80% of the content tokens of a single rendered memory entry, the entry's own timestamp is later than the old entry's, both state values are present, the change is material to the response, and the causal path has at least two nodes. What is verified is provenance and chronology, not semantic supersession; the validator never decides whether the new state truly supersedes the old. Verified-only authority means a repair directive fires only through an authorized transition, while unverified verdicts cannot override the validated chronology. This state-anchored pass is what lets the system catch stale dependencies that are unstated in an open-ended response, the failure mode that draft-anchored claim extraction misses.","core_discovery":"The central claim is that the implicit policy adaptation gap—an agent that knows a stored state is outdated yet still behaves from the old value—has a structural, fixable cause on the response side. The cause is draft-anchored verification: checking what a draft says misses dependencies the draft never states. StateAuditor reverses the audit direction, and provenance-verified transition assembly (VTA) lets only transitions with matched quotations and later timestamps authorize repair. Under STALE's strict full protocol (400 scenarios, 50-session histories, one independent response per query), strict single-query VTA scores .736 versus .686 for the locked predecessor under the same judge: a +5.0-point paired gain (95% CI [+2.9, +7.2]) that comes almost entirely from implicit policy adaptation and premise resistance. The result is reproduced by the benchmark's judge from a different model family (.738 versus .680), and a matched control with identical evidence, adapter, and call budget recovers only +0.6 points, attributing the gain to the transition machinery itself.","pith_inferences":["An implication the paper leaves implicit is that evaluation of personalized agents should include counterfactual state-change probes—changing the stored state and checking whether open-ended behavior changes—because stated-content checks cannot see unsaid dependencies.","Because the validated gate is purely chronological and provenance-based, a cheap deterministic prefilter of this kind could be paired with a later semantic supersession check; the paper notes a semantic second gate reduces over-repairs on a hard set without changing accuracy, suggesting the two checks are complementary.","The T2 (implicit-conflict) premise-resistance and state-resolution regressions observed under the repair-only policy hint that some implicit conflicts are better handled by asking the user rather than repairing silently; routing by conflict type and confirmability is a testable extension.","Retrieval recall of 0.580 bounds the pipeline's ceiling, so improving lifecycle-aware retrieval—for example, detecting life-event disclosures by their event structure rather than lexical overlap—should transfer directly into further STALE gains."],"forward_implications":["On STALE's strict full protocol, the state-to-draft transition repair outscores the locked predecessor by 5.0 paired points under the same judge, with the gain concentrated in implicit policy adaptation and premise resistance.","A matched control that keeps the same evidence, adapter, and call budget but removes the transition machinery recovers only 0.6 points over the predecessor, so the measured improvement is not from extra context or extra calls.","A judge from a different model family reproduces the gain (.738 vs .680), and three independent full-400 draws land at .733 ± .002, so the headline comparison is not tied to a single judge or a single draw.","Draft-anchored verification methods, whatever their scale, are structurally blind to the relevant failure: stale-premise recall collapses to 0.06–0.38 on open-ended probes where the dependency is unstated.","The method's reach is bounded: on HorizonBench most of the external gain comes from the draft-side audit itself, and on a harder authored lifecycle set there is no gain, so the paper's claim is about the studied settings, not general-purpose agent memory."],"supporting_citations":[{"why":"Provides the STALE benchmark: the 400 expert-validated implicit-conflict scenarios, SR/PR/IPA probes, 50-session histories, and the released judge used for the headline and reproduction numbers.","marker":"Chao et al. 2026"},{"why":"Chain-of-Verification is the draft-anchored verify-and-revise baseline whose structural blind spot on unstated stale dependencies the paper demonstrates.","marker":"Dhuliawala et al. 2024"},{"why":"RARR is the other canonical draft-anchored retrieval-and-revise pipeline used as a baseline to show the same blindness.","marker":"Gao et al. 2023"},{"why":"HorizonBench supplies the independent cross-family preference-evolution benchmark used to test whether the full pipeline generalizes outside STALE.","marker":"Li et al. 2026b"},{"why":"Its deterministic freshness-resolution recipe is the ingredient the VTA validator adapts for chronological authorization.","marker":"Reddy and Challaram 2026"},{"why":"Documents the self-preference-bias risk in rubric judges, motivating the disjoint-model reproduction of the headline gain.","marker":"Pombal, Rei, and Martins 2026"}],"fun_headline_variants":["Audit stored state, not the draft, to repair stale responses","Reversed audit narrows the implicit policy adaptation gap","State-first verification beats draft-only checks by 5 points","Provenance-checked transitions fix agents' stale knowledge","From draft to state: a 5-point fix for stale-response errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The system can repair a stale dependency only when the retrieval step actually surfaces the superseding evidence, and in the official 400-scenario set pooled strict window-level recall of the old/new pair is 0.580, so if retrieval quality is lower in deployment, the measured gain would shrink or disappear.","fun_headline_variants_meta":{"raw":{"variants":["Audit stored state, not the draft, to repair stale responses","Reversed audit narrows the implicit policy adaptation gap","State-first verification beats draft-only checks by 5 points","Provenance-checked transitions fix agents' stale knowledge","From draft to state: a 5-point fix for stale-response errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000362,"raw_usage":{"total_tokens":2044,"prompt_tokens":1123,"completion_tokens":921,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":739,"completion_tokens_details":{"reasoning_tokens":836}},"tokens_in":739,"tokens_out":921,"duration_ms":7945,"temperature":1.0,"reasoning_tokens":836,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:05:44.807302+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the strict STALE protocol again with the change-scan retrieval slots disabled—the query-independent disclosure windows that recover superseding evidence—and compare strict VTA with the matched control; if the +5.0-point gain survives without that evidence being retrieved, the paper's attribution of the gain to the transition machinery rather than retrieval is wrong.","supporting_citations":[],"review_version":2}