{"id":"f9059b02-5832-4c39-8dc2-cdb8a72ab4e4","arxiv_id":"2608.12476","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A bitemporal, source-bound memory model with a verified-head release gate achieves 100% on its own frozen contract tests and sealed service arms, while explicitly limiting claims to bounded internal surfaces.","lead":"This paper defines Governed Persistent Memory, a state model that blocks retracted, superseded, or stale records from supporting an agent's outgoing claims. It provides executable contract clauses and tests them on frozen internal benchmarks, reporting perfect contract compliance on its own service evaluation surface.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 100% conformance results rest on gold labels generated by the same project, with no independent human annotation; if those labels encode implementation-shared misinterpretations, the perfect scores do not demonstrate the claimed contract.","rationale":"The reader's weakest assumption points at internally generated gold and the lack of independent human annotation, which is indeed the most load-bearing concern. I agree with that core. However, the reader frames the failure mode as 'on real-world distributions,' whereas the paper explicitly and repeatedly disclaims external validity; the sharper risk is internal: the gold may not faithfully encode the paper's own formal contract, so the perfect scores would not even demonstrate contract conformance within the constructed suite. That distinction makes my agreement partial. I examined whether a more formal gap exists—e.g., whether Eq. (10) enforces 'exact claim closure' as completeness rather than per-claim soundness—but the paper's definitions and benchmark framing treat closure as per-claim boundness, and the 'atomic match' criterion tests the full pipeline including the claim proposer. I also verified that Contract consequence 1 is a direct, correct derivation from the stated definitions; the formal model holds up. The finite verification is honestly bounded and not overclaimed. The decisive weakness is therefore evidentiary: every piece of evidence for implementation fidelity—ReleaseBench gold, Governed-QA Sealed oracle, differential tests, finite model—originates from the same project's interpretation of the contract. No independent implementation, human annotation, or third-party rescoring exists. The paper's own limitation statements acknowledge the absence of independent annotation and the exclusion of gold from the artifact, which supports keeping the verdict at CONDITIONAL: the claims are plausible and carefully scoped, but the central empirical result needs an external check before it can be relied upon.","tokens_in":17164,"tokens_out":14833,"duration_ms":142329,"concrete_test":"Publish a random 300-case subset of the GPM-ReleaseBench hidden gold and 300 per-cluster receipts from the sealed V3/V5 evaluations, then have an independent team re-derive the expected outcomes solely from the formal contract definitions in §3–§5 (Eqs. 3–10, the five clauses, and the query-family grammar), with no access to the production code. If any independently re-derived label disagrees with the sealed gold, the reported 3,600/3,600 and 2,400/2,400 counts cannot be taken as evidence of contract conformance; if all sampled labels agree, the internal-generation concern is substantially answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The formal machinery in §3.5 and Eq. (10) is internally coherent: Contract consequence 1 is a direct derivation from the bound() definition, and the paper carefully scopes it as source-relative. The load-bearing weak point is the empirical demonstration that the production system actually satisfies I1–I5. Both central evaluations use gold labels authored by the same project. §5.2 calls GPM-ReleaseBench 'an internally designed contract-conformance suite' and states that hidden gold was audited only by 'two isolated AI review lines' with 'no human blind annotation.' §5.4 says Governed-QA Sealed is 'also internally generated,' and §8 notes the artifact excludes hidden gold, evaluator keys, and per-cluster receipts. The gold generator and the system under test are two products of one codebase's interpretation of tri(), valid-interval overlap, supersession, conflict, and non-revival. If that interpretation diverges from the formal definitions in any subtle way (e.g., normalization plus multi-value registry keys, or delete barriers interacting with provenance timestamps), the generator would bake the same divergence into the 'correct' labels, and a fully conformant-but-wrong implementation could still score 3,600/3,600 and 2,400/2,400. The finite model and three-engine differential in §5.5 do not break this circularity: all stem from the same project's executable formulation of the contract. Thus the headline evidence does not yet establish that GPM's release gate enforces the paper's contract as an independent reader would read it. This is a methodological circularity concern, not a claim of bad faith.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Governed Persistent Memory (GPM), a bitemporal state-transition model for long-horizon agent memory that separates source-bound admission, derived lifecycle state, conflict isolation, retraction/deletion barriers, and a fail-closed structured release gate. The formal core is a set of five executable clauses (I1–I5) culminating in Eq. (10), which defines the release decision as requiring a fresh public view, a stable verified head, and exact binding of every structured claim to source fact identifiers. The paper reports a hash-frozen 3,600-case contract suite (GPM-ReleaseBench), a sealed end-to-end service evaluation (Governed-QA Sealed) with 2,400/2,400 correctness in the governed lane, a bounded finite model with no counterexamples, and a 100,000-trace three-engine differential. The paper repeatedly and explicitly scopes these as bounded contract and implementation results, not open-world accuracy claims.","tokens_in":17512,"tokens_out":5454,"duration_ms":48843,"significance":"If the formal contract is correctly implemented, GPM provides a coherent and useful separation between retrieval and public assertability, with a precise definition of when a structured claim may be released. The formal definitions in Section 3.5 are internally sound: Contract consequence 1 follows directly from Eq. (10) and the definition of bound(), and the paper honestly labels it as a direct consequence rather than an inductive theorem. The paper is unusually transparent about its limitations, including the internal generation of benchmarks, the failed language-judge gate, the post-freeze reducer amendment in V3, and the absence of independent human annotation. These strengths make the formal contribution a valuable reference point for systems-level agent memory. However, the empirical validation is the load-bearing support for the claim that the production system actually satisfies I1–I5, and that validation rests on gold labels and oracles authored by the same project, which substantially limits the strength of the empirical conclusions.","major_comments":[{"comment":"The central empirical claims (3,600/3,600 on GPM-ReleaseBench and 2,400/2,400 on Governed-QA Sealed) rest on gold labels designed by the same project that implemented the system. The paper states in §5.2 that GPM-ReleaseBench is 'an internally designed contract-conformance suite' and in §5.4 that Governed-QA Sealed is 'also internally generated.' Hidden gold was audited only by 'two isolated AI review lines' with no human blind annotation, and the gold sealing occurred under one OS user (§8). Because the gold generator and the system under test share the same interpretation of the contract—including normalization, conflict, supersession, and retraction/deletion barriers—a systematic misinterpretation in that shared interpretation would be reflected in both the generated 'correct' labels and the system's behavior, making the perfect scores uninformative about whether the system enforces the intended contract on any independent distribution. The paper acknowledges this, but the abstract and §6 still present these counts as the primary evidence that GPM satisfies I1–I5. To make the empirical claims load-bearing, the hidden gold should receive independent human annotation, or the results should be explicitly repositioned as self-consistency checks rather than validation of the contract.","section":"§5.2, §5.4, §8"},{"comment":"The V3 arm, which is the publicly disclosed result, required a post-freeze reducer amendment to admit only the dated form of 120 temporal questions whose date literal differed due to crossing UTC dates between candidate and baseline execution. The paper reports this transparently, but it means the V3 2,400/2,400 result is not a clean prespecified frozen evaluation; the reducer was changed after the freeze. The V5 reseal addresses this defect by pinning the generation date and reporting reducer_amended_after_freeze=false, but V5 still uses an internally generated surface. Additionally, the one-sided 95% Clopper–Pearson lower bounds (§6.2) are descriptive summaries of an all-success denominator on a deterministic generated set, not inferential estimates from a random sample; the paper notes this, but the presentation in Tables 3 and 4 may nonetheless be read as supporting a population-level success rate. Please either provide an externally authored or independently annotated evaluation, or explicitly state in the abstract and results that the 100% figures are exact counts on a constructed, internally generated suite with no sampling-frame interpretation.","section":"§5.4, Table 3"},{"comment":"The reproducibility artifact excludes hidden gold, evaluator keys, and per-cluster receipts, and has no public URL, DOI, or release date assigned. Combined with the internally generated gold, this means the empirical results cannot be independently checked or rescored by any other party. The paper's transparency is commendable, but for a paper whose contributions include empirical evaluation, this is a load-bearing limitation: the claims of 3,600/3,600 and 2,400/2,400 are not independently verifiable as presented. Please either release the hidden gold and receipts (with keys) after the evaluation, provide an independent audit protocol, or clearly state in the contribution list that the empirical results are self-reported and not externally reproducible.","section":"§8, §9"}],"minor_comments":[{"comment":"The notation is inconsistent: Eq. (1) uses 'rank k' while Eq. (2) uses 'rankk'; please harmonize to a single notation (e.g., 'rank_k').","section":"§2.1, Eq. (1)–(2)"},{"comment":"In Eq. (3), 'canon(e_i\\{h_i})' is ambiguous: clarify that it denotes the canonical encoding of the event fields excluding the chained hash h_i, and specify the canonicalization rules (field ordering, length-delimited encoding) used.","section":"§3.1, Eq. (3)"},{"comment":"The answer-bundle accuracy column in Table 2 reports percentages without explicit denominators; the text explains that violation cases whose only purpose is to diagnose unsupported free text are excluded, but the table would be clearer if it stated the denominator for each row (e.g., 1200/2300) rather than only the percentage.","section":"§5.2, Table 2"},{"comment":"When running StateFuse, the paper should cite the specific version or commit of StateFuse used and describe how the 'official state projection' was obtained, so that the shared-subset comparison is reproducible.","section":"§5.3"},{"comment":"The 'No I1' row shows 0 unsafe releases but atomic match of 3450/3600; the accompanying text explains that the mutant misses 150 required fail-closed outcomes, but a brief sentence clarifying why this does not appear as an unsafe release would help readers understand the metric design.","section":"§6.3, Table 5"},{"comment":"The abstract says 'prespecified hash-frozen 3,600-case GPM-ReleaseBench,' but §5.2 describes a public development split of 360 cases plus an evaluator-hidden split of 3,600 cases; clarify that the 3,600 figure refers to the hidden split only, or give the total count.","section":"Abstract and §5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a single-author preprint with AI-assisted writing and editing; the formal contract is a useful systems contribution, and the author's transparency about limitations is a positive signal. However, the empirical evaluation's reliance on internally generated gold labels, without independent human annotation or a public artifact, is a significant weakness for a journal-level claim of enforced fail-closed behavior. If the author can either obtain independent blind annotation for the hidden gold or substantially soften the empirical claims to 'self-consistency demonstrations on a constructed synthetic surface,' the paper could be acceptable. The current version, with its perfect scores prominently featured, overstates the external validity of the results. I recommend major revision to address the empirical validation and reproducibility concerns before reconsideration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: GPM is a serious piece of systems work. It gives long-horizon agents a concrete, checkable boundary between what may be asserted and what is merely retrievable, and it does so with unusual honesty about its own limits. The formal contract is the real contribution; the 100% scores are not yet independently established.\n\nWhat's actually new: the five executable clauses (I1–I5) and the release gate in Eq. (10). The head-reread before issuing a claim-bound record is a clean, mechanical way to make 'fresh view' precise. The derivation of Contract Consequence 1 is a direct unpacking of the definitions, and the paper openly says so. That is the right level of ambition: no inductive theorem over all reachable states, just a carefully scoped statement about what the gate checks.\n\nCredit where due: the paper is far more candid than most. It explicitly calls GPM-ReleaseBench 'internally designed,' admits there was no human blind annotation, discloses the V3 reducer amendment, and runs clean-task controls that fail their own judge gate. It consistently distinguishes deterministic contract results from language-model accuracy. That discipline is real.\n\nThe soft spot is exactly where the reader's stress-test points: the main empirical evidence is circular in a specific sense. Both the 3,600-case benchmark and the sealed 2,400-cluster evaluation use gold labels generated by the same project and audited only by two isolated AI review lines. If the label generator and the system under test share a subtle misinterpretation of the contract—say, over normalization or deletion-barrier semantics—then a fully conformant-but-wrong implementation could still produce perfect scores. The finite model and three-engine differential are bounded and also built from the same project's executable formulation, so they don't break the circularity. The paper acknowledges all of this in Section 8, which is why I call it a soft spot rather than a fatal flaw. But the headline numbers are presented with Clopper-Pearson bounds that might suggest more independence than exists. The problem is addressable: independent human annotation of a sample of gold cases, or an externally authored challenge set, would do a lot.\n\nWho should read it: anyone building memory layers for agents or working on provenance and abstention. The contract design is a useful reference even if you don't adopt the system. Does it deserve a serious referee? Yes. The formal work is coherent, the system is non-trivial, and the honesty about limitations is exactly what peer review needs. I'd send it to reviewers with a request to focus on the evaluation design rather than desk-reject it. My own verdict would be conditional: accept after the empirical claims are backed by independent gold, or at least reframed as internal demonstrations rather than evidence of contract satisfaction. For my own work, I'd cite the contract but not the 100% results.","headline":"A serious, self-aware systems paper with a sound formal contract and internally generated perfection; the empirical claims need independent gold before they can carry the headline weight.","tokens_in":18039,"tokens_out":5223,"would_cite":true,"duration_ms":40916,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces Governed Persistent Memory, a bitemporal source-bound state model in which a structured claim may be released only if it exactly matches an assertable fact in a fresh public view at a verified ledger head.","keywords":["governed persistent memory","bitemporal state model","source-bound claims","fail-closed release","retraction non-revival","structured release contract","agent memory ledger","contract conformance"],"falsifier":"Inspect the frozen GPM-ReleaseBench hidden split: if any violation-polarity case receives an unmatched release, or any valid case is not matched to an allowed outcome bundle, the central fail-closed claim collapses. Independently, a single Governed-QA Sealed cluster that returns a wrong answer, abstains on an answerable case, or regresses a baseline-correct cluster, or a full-contract counterexample inside the declared 331,776/1,990,656 finite-state bounds, would refute it.","tokens_in":16971,"feed_emoji":"🔒","tokens_out":7806,"duration_ms":63554,"temperature":0.7,"pith_summary":"The paper's pith is that retrieval is not public eligibility: a long-horizon agent may assert a fact only after the memory system has established that the fact is source-bound, not superseded or in unresolved conflict, not retracted or deleted, and closed against a freshly verified ledger head. To make that testable, it defines Governed Persistent Memory (GPM), an auditable bitemporal event-ledger model with five executable clauses covering ledger integrity, source binding, conflict isolation, non-revival after retraction or deletion, and structured release closure. On a frozen, evaluator-hidden 3,600-case suite, GPM matches every allowed outcome while the strongest simple policy matches half and releases unsafely on half of the violation cases; sealed end-to-end runs report 2,400/2,400 in the disclosed arm and again in both command-language reseal arms. The paper is careful to frame these as bounded contract and implementation results, not open-world or world-truth accuracy.","feed_headline":"Fail-closed memory gate passes 3,600/3,600 contract cases","feed_subtitle":"On a frozen 3,600-case suite and sealed end-to-end arms, all governed answers pass contract checks.","key_machinery":"The load-bearing mechanism is the release decision of equation (10), operating on the public projection $V_{\\mathrm{pub}}$ of equation (6). $V_{\\mathrm{pub}}$ is a bitemporal filtered view, where each event carries both transaction time and valid time: with respect to the verified ledger head $h_c$, it takes the transaction-time prefix $\\tau_q$, applies the latest user deletion barrier and retracted-fact set, then removes superseded and conflicted claims, so ranking only orders claims already eligible for public assertion. The gate reads the head as $h_b$, constructs $V_{\\mathrm{pub}}$, requires every proposed structured claim $a$ to satisfy $\\mathrm{bound}(a,V,u)$ — non-empty source fact identifiers all matching the normalized entity–key–value triple — digests the complete claim multiset into $D_A$, and rereads the head as $h_a$; only $h_a = h_b$ with a non-empty closed set yields a release record, and otherwise the gate fails closed or abstains. All five clauses are expressed as executable checks on this one decision boundary.","core_discovery":"Contract consequence 1 is the central claim: whenever the release gate defined by equation (10) returns release, every structured claim $a$ in the outgoing multiset $A$ is exactly bound to one or more same-user assertable facts in the fresh public view $V_{\\mathrm{pub}}(u,t,\\tau;h_b)$, its digest $D_A$ commits to the complete canonical claim multiset, and no claim is bound only to facts or episodes excluded by current retraction or user-deletion policy at the verified head $h_b$. The gate enforces this by reading the head before constructing the view, rechecking every binding against that view, and rereading the head before emitting a record; a head change raises an error and returns no record. On the paper's frozen GPM-ReleaseBench, GPM matches 3,600/3,600 complete outcomes with zero unmatched violation releases, whereas raw append, latest-first, and flat conflict-preserving policies release unsafely on 91.67%, 58.33%, and 50% of violation cases respectively. The sealed Governed-QA evaluation reports the governed lane correct on 2,400/2,400 clusters versus 600/2,400 for ungoverned local Qwen2.5-7B, with all 1,800 baseline failures repaired and no regressions.","pith_inferences":["Editorial inference: A natural extension would be to rerun the same release gate on externally authored dialogues or ontologies with independent human gold annotation; that would test whether the internally generated suite's clean results transfer to naturalistic distributions.","Editorial inference: The head-stability check suggests a general design rule for agent frameworks: emit claims only against a snapshot that is re-verified at emission time, since any delay between reading state and releasing an answer is where stale or revoked facts can slip in.","Editorial inference: The largest unclosed risk the paper itself identifies is free-form prose: I5 covers explicit structured claims only, so an agent that paraphrases or merges facts in natural language is outside the guarantee until a semantic entailment check on emitted text exists.","Editorial inference: The bounded finite model could be pushed past 331,776 semantic and 1,990,656 query states by raising the event bound or adding multi-user interference; a full-contract counterexample found there would force revision of the five clauses."],"forward_implications":["Under I1–I5, no released structured answer can be supported by a retracted or deleted fact, an unresolved conflicting fact, or a claim with no matching source identifier, as long as the trusted local commitment and verified head are intact.","The three weak policies' unsafe-release rates show the failure mode the contract removes: relevance ranking alone can surface stale or contradictory records as if they were current public facts.","A single safety rate is insufficient to audit the contract; the no-I1 mutant misses 150 required fail-closed outcomes without producing an unmatched release, so complete outcome bundles are the right evaluation unit.","The release record is local and unsigned today, so cross-process or cross-host use requires adding authentication, freshness, expiry, and key management before the same guarantee can be claimed at a distance.","The eligibility projection can be combined with different retrievers, operation languages, replicated conflict substrates, or descendant-repair planners; GPM supplies the boundary, not the ranking."],"supporting_citations":[{"why":"provides the immutable public repository commit that discloses the V3 aggregate governed-lane result used as the end-to-end evidence.","marker":"[1]"},{"why":"localizes hallucination propagation across extraction, update, and QA stages, motivating the state-level eligibility boundary GPM adds.","marker":"[4]"},{"why":"serves as a clean-task utility control in the versioned public-task evaluation.","marker":"[9]"},{"why":"adjacent typed evidence/assertion/decision memory model that GPM contrasts with its simpler source-binding contract.","marker":"[10]"},{"why":"closest contract-level neighbor; its official projection is run on the supported 2,100-case shared contract subset.","marker":"[25]"},{"why":"supplies the bitemporal contradiction-resolution operator algebra against which GPM's fixed public-policy interpretation is positioned.","marker":"[28]"},{"why":"provides the MemoryAgentBench EventQA clean-task control for deterministic substring scoring.","marker":"[29]"},{"why":"delimits the derived-artifact repair scope that GPM's I4 non-revival clause does not cover.","marker":"[31]"}],"fun_headline_variants":["Memory gate: 100% contract pass, 0 unsound releases","Source-bound memory: 3,600/3,600 contract cases clear","Fail-closed retrieval blocks stale claims, passes all tests","Governed memory: 100% exact claims, zero stale releases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The empirical case rests on gold labels and a sealed oracle generated by the same project and audited only by two isolated AI review lines with no independent human annotation; if those labels misdefine the intended contract, the all-success counts would not show fail-closed behavior on real-world data.","fun_headline_variants_meta":{"raw":{"variants":["Memory gate: 100% contract pass, 0 unsound releases","Source-bound memory: 3,600/3,600 contract cases clear","Fail-closed retrieval blocks stale claims, passes all tests","Governed memory: 100% exact claims, zero stale releases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000885,"raw_usage":{"total_tokens":3955,"prompt_tokens":1212,"completion_tokens":2743,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":828,"completion_tokens_details":{"reasoning_tokens":2667}},"tokens_in":828,"tokens_out":2743,"duration_ms":17492,"temperature":1.0,"reasoning_tokens":2667,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:07:05.796018+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the frozen GPM-ReleaseBench hidden split: if any violation-polarity case receives an unmatched release, or any valid case is not matched to an allowed outcome bundle, the central fail-closed claim collapses. Independently, a single Governed-QA Sealed cluster that returns a wrong answer, abstains on an answerable case, or regresses a baseline-correct cluster, or a full-contract counterexample inside the declared 331,776/1,990,656 finite-state bounds, would refute it.","supporting_citations":[{"cited_title":"Governed memory: Public bounded evaluation disclosure","cited_arxiv_id":null,"evidence_quote":"provides the immutable public repository commit that discloses the V3 aggregate governed-lane result used as the end-to-end evidence."},{"cited_title":"StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems","cited_arxiv_id":"2607.05844","evidence_quote":"closest contract-level neighbor; its official projection is run on the supported 2,100-case shared contract subset."},{"cited_title":"TOKI: A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory","cited_arxiv_id":"2606.06240","evidence_quote":"supplies the bitemporal contradiction-resolution operator algebra against which GPM's fixed public-policy interpretation is positioned."}],"review_version":1}