{"id":"4541854e-191e-40cb-b192-91c2d4f0fc63","arxiv_id":"2608.11392","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Compaction can leave a safety rule textually present but behaviorally dead, so presence-based audits are insufficient; rule-form text also survives better than matched facts.","lead":"After an AI agent compresses its conversation history, a safety rule can still appear in the summary text yet no longer stop the forbidden action, so checking only for its presence gives false assurance. This paper measures that gap and shows rule-like text survives compression more often than ordinary facts, which is why the failure is easy to miss.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"D-versus-W replay gap may reflect overall digest thinness rather than degraded rule form; Section 6 admits no matched-compression control, so residue-causality is not yet established.","rationale":"The reader's weakest assumption identified exactly the most load-bearing concern: the D-versus-W behavioral gap is attributed to the rule's degraded form, but D-labeled digests are disproportionately the harder-compressed ones, and replay outcomes may co-vary with overall digest length and richness. This concern is acknowledged in Section 6 but not controlled, and it directly targets the paper's headline causal claim rather than its descriptive claims. It is not fatal to the broader conclusion that presence checks give false assurance, because even textually intact W rules fail to fire on replay (23% under Qwen, 38% under Llama), and G-labeled residues also under-protect; those independent observations do not depend on the D-versus-W contrast. The concern is also addressable with a concrete controlled experiment—editing only the residue text within otherwise identical digests—so it warrants a conditional rather than a rejection. Since the reader's verdict was already CONDITIONAL on this same weakness, my stress-test does not move the verdict.","tokens_in":30570,"tokens_out":4256,"duration_ms":43633,"concrete_test":"For every D-labeled digest in the §4.4 census, construct a matched twin by replacing only the degraded residue with the full intact rule text (exact referent plus restricted action), leaving all other tokens and the overall digest otherwise unchanged; replay both twins under identical prompts and decoding seeds with Qwen and Llama. If compliance in the edited twin falls to the W baseline, the residue text is causal; if it stays near the original D rate, the gap is driven by overall digest thinness or compression and the mechanism claim fails. Optionally also match D and W digests on length and richness and re-estimate the all-case gaps as a secondary check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim—that a degraded textual residue causes the prohibited action far more often than a textually intact rule—rests on the D-versus-W replay contrast in §4.4. But D labels are assigned to digests that are disproportionately harder-compressed; the digest-validity filter and the D label both track abstraction level, and Section 6 explicitly concedes that 'part of the degraded-versus-intact gap could reflect a digest being thinner overall rather than the specific rule's degradation' and that digests matched on compression were not run. Since replay loads the entire digest as context, a D digest differs from a W digest in many correlated ways (length, richness, surviving-item count, phrase-level detail), not only in the rule's textual form. The family-stratified analysis controls for action family but not for digest-level compression, so it does not remove the confound. If overall thinness drives compliance, the mechanism-specific claim (a degraded residue lets the action through) and the design implication that operative-content checks will fix it are weakened, even though the broader 'presence is not protection' point survives via the intact-W failure rates. The authors acknowledge this unaddressed confound; it is the load-bearing gap between what §4.4 asserts and what the design establishes.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks how a standing safety rule is lost when an agent compacts its context into a model-generated summary, restricting attention to a single compaction cycle. Two rigs are used: a single-rule stress probe and a between-items deontic-versus-epistemic probe (45 conversations × 8 items × 2 marking conditions, 45k-token filler with planted tracer facts), scored under a W/G/D/X taxonomy (welded/generalized/degraded/dropped) by an LLM judge with blinded author adjudication and behavioral replay. Under confirmed compression pressure (tracer drop ≥60%), the authors find that (i) a salient rule is either welded to its exact referent or dropped wholesale in the stress probe, with degraded predicate-loss residues appearing under the tighter multi-item budget; (ii) rules are retained substantially more often than prominence-matched facts (β ≈ 2.5 with Qwen; β ≈ 2.29 with Claude under a 150-token hard cap); (iii) on behavioral replay, degraded residues and category-level survivors comply with the prohibited action far more often than intact welded rules (all-case gaps of +34 points under Qwen and +57 under Llama, with +50/+56 among headroom cases), while even textually intact rules fail to fire 23–38% of the time, so 'a presence check is not a safety check'; and (iv) an LLM judge's labels would, in two documented cases, have reversed a conclusion. The hypothesized referent-severing mode was not observed (0 of 10 generalizations non-covering, 2 unresolved).","tokens_in":30755,"tokens_out":14084,"duration_ms":111998,"significance":"If the result holds, the 'presence is not protection' finding is an important, actionable caution for context compaction in agent frameworks: audits that verify only textual presence give false assurance, and even textually intact rules fail to fire on replay (23% under Qwen, 38% under Llama in the conclusive W census). The study's concrete strengths include the tracer-based manipulation check for compression pressure, the W/G/D/X taxonomy with target-by-predicate decomposition, behavioral replay with sibling-target controls under two replay models and two label sets, dependence-aware (conversation-clustered) inference with target fixed effects and a cluster bootstrap, the reproduction of the retention premium on a second summarizer from a different provider under a hard output cap, and the shipped code and data with verification scripts. The paper is exceptionally candid: §6 flags the digest-quality confound, the single-generation design, verification bias, evaluator overlap, and replay-model capability limits.","major_comments":[{"comment":"The abstract and §7 state that degraded residues 'lead the model to perform the prohibited action far more often than an intact welded rule,' and §5 concludes that a post-compaction audit must verify operative content. This causal reading of the §4.4 replay contrast is not yet supported, because §6 concedes that 'D labels come disproportionately from harder-compressed digests' and that digests matched on compression were not run. Since replay loads the entire digest as prior context, the D-versus-W gap is confounded with overall digest thinness and richness; the family-stratified table in Appendix E controls for action family but not for digest-level compression, and the within-D split (17/26 with target present versus 16/19 absent) is too small to carry the mechanism by itself. A matched-compression control, a within-digest D-versus-W comparison, or a controlled ablation (removing the predicate from an intact W digest and re-replaying) is needed before the gap can be attributed to the rule's degraded form rather than to the thinner digest. Without such a control, the paper should present the §4.4 contrast as associational and the §5 'verify operative content' recommendation as necessary but not sufficient, consistent with the paper's own W-failure rates rather than with the causal phrasing currently in the abstract.","section":"§4.4; §6 (Digest quality as a confound); Abstract"},{"comment":"The headline quantitative claim — all-case degraded-minus-intact gaps of +34 (Qwen) and +57 (Llama) — is reported without any uncertainty estimate. The text explicitly states that 'we attach no sampling confidence intervals to these proportions, which are clustered by conversation, target, and action family,' and the only significance figure offered is the family-stratified exact conditional test of Appendix E, which the paper itself calls 'not a design-valid test' because it ignores conversation and target clustering and conditions on replay headroom. That the data are clustered is an argument for computing a clustered interval (for example, a conversation-level cluster bootstrap over the replay census, or a mixed-effects model of the replay outcome), not an argument for omitting one. The single-generation design acknowledged in §6 (each replay is one stochastic decoding at temperature 1.0 for Qwen) compounds the issue. Given the size of the effects, the direction of the conclusion is likely robust, but as written the reader cannot separate the +34/+57 point estimates from their sampling variability.","section":"§4.4 (Behavioral replay)"},{"comment":"The paper's contribution 4 (§4.6) documents two cases in which the LLM judge's labels would have reversed a conclusion, and §3.5 reports that the corrected frozen rubric reached only 85% judge–author agreement on adjudicated cases, below the paper's own 90% gate, with adjudication targeted at decisive/contested cells so that the headline numbers 'still incorporate unreviewed judge labels on non-decisive cells.' The D labels that form the numerator of the §4.4 gap are precisely the class in which judge error was observed (Case 1 of Table 4: three stress-probe digests mislabeled D). The re-bucketing under the conservative rubric (Qwen +46, Llama +51) is a re-score by the same judge, so — as the paper correctly says — it is a label-sensitivity analysis rather than an independent error bound. A small preregistered stratified human audit of the D and W replay sets, the cells that carry the central contrast, would resolve the tension between the paper's evaluation-reliability warning and its own primary labels.","section":"§3.5; §4.6"}],"minor_comments":[{"comment":"The abstract's claim that rule-form items are retained 'substantially more often than prominence-matched facts' should carry the paper's own scope caveat in the same breath: the normativity-versus-salience separation was inconclusive and the paired items differ in specificity density (§4.3, §6), so the premium is descriptive rather than attributed to rule status.","section":"Abstract; §4.3; §6"},{"comment":"The all-case floor percentages (D 43% versus W 9% under Qwen; D 88% versus W 31% under Llama) do not state their denominators in the running text; adding them (D: 33/77, W: 5/55, and the Llama analogues) would make the floors auditable at a glance.","section":"§4.4"},{"comment":"The single-generation limitation is acknowledged in §6 but is not reflected in any reported uncertainty; reporting decoding-seed variance for a subsample of replay outcomes would calibrate the reader's confidence in the §4.4 gaps.","section":"§6"},{"comment":"Panel (b) measures referent presence while panel (a) measures textual W+G survival; the caption explains this, but the figure would be much less likely to mislead if the two panels used a shared metric or an explicit visual marker of the difference.","section":"Figure 1"},{"comment":"The 'not observed' severing result is properly bounded by the Clopper–Pearson reference endpoint of ≈31% in §4.3; the main text should attach that bound when 'not observed' first appears in §4.2, since 0-of-10 with 2 unresolved is easy to over-read as 'ruled out.'","section":"§4.2; §4.3"},{"comment":"The reproducibility statement notes that the archival DOI 'will be added on posting'; since several results depend on the exact frozen label files and prompts, the paper should commit to a versioned snapshot (for example, a release tag) rather than a live repository link alone.","section":"Reproducibility statement"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the paper is unusually honest, and the §6 limitations section is a credit rather than a defect; the main revision point is aligning the abstract's causal framing with the design's associational evidence, which is achievable with a matched-compression or within-digest control. The work is incremental relative to Governance Decay (Chen 2026) but contributes a distinct morphology of loss and a worked evaluation-reliability lesson. Note that several cited works (Ding 2026, Chen 2026, Bort 2026) are very recent items that I could not independently verify; the manuscript's own claims do not depend on their veracity, since the incident is used as motivation only. If the authors add the matched-compression control and a clustered interval for the headline replay gaps, the paper would be a strong accept candidate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one for the 'presence is not protection' result, which is real and useful. The paper shows that a safety rule can survive compaction as text that looks fine but fails to fire behaviorally—intact welded rules disarmed on replay 23% (Qwen) and 38% (Llama) of the time—so any audit that checks only textual presence gives false assurance. That finding is robust to the paper's weaknesses and is the one I'd cite. The rule-versus-fact retention premium (beta ~2.5, replicated on Claude under a hard output cap) is also a solid, honestly-scoped descriptive result. The paper is unusually careful about its own measurement: manipulation checks, conservative re-labeling, dependence-aware inference, behavioral replay with two models, and a public repo with code and data.\n\nThe soft spot is the sharper causal claim—that a degraded residue lets the prohibited action through far more often than an intact rule. D-labeled digests are disproportionately the harder-compressed, thinner ones, and the authors admit in Section 6 that they did not run digests matched on compression. So part of the D-vs-W gap could be overall thinness rather than the specific rule's degradation. That is a real confound, and it is load-bearing for the mechanistic reading. The other weaknesses are more standard: each input gets a single stochastic generation, the replay proportions have no sampling intervals, adjudication was targeted rather than a preregistered random audit, and the judge, salience rater, and one summarizer all come from the same model family. The marked condition designed to separate normativity from salience came out inconclusive, so the premium stays descriptive.\n\nNone of this kills the paper. The broad 'presence is not protection' conclusion survives even if the degraded-residue mechanism is partly confounded, and the evaluator pitfalls (two judge failures that would have reversed conclusions) are documented with enough detail to be useful on their own. The fixes are addressable at revision: matched-compression digests, repeated decoding, an independent judge, and a preregistered human audit.\n\nRecommendation: yes, send it to peer review. It deserves referee time and will likely come back after a revision that tightens the causal claim. I'd use it in a reading group on agent-safety evaluation, and I'd cite it for the false-assurance point.","headline":"A careful, honest empirical paper showing 'presence is not protection' under compaction; the sharper degraded-residue causal claim is softened by a digest-quality confound the authors flag but do not control.","tokens_in":31322,"tokens_out":2721,"would_cite":true,"duration_ms":24485,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A safety rule can survive context compaction yet stop protecting.","keywords":["context compaction","agent safety","safety rule survival","presence vs behavioral protection","degraded residue","behavioral replay","LLM-as-judge evaluation","constraint degradation"],"falsifier":"Run the same degraded-versus-intact replay census with digests matched for length, output-token budget, and overall content richness; if the compliance gap collapses or reverses once the digests are equally thin, the central claim that the residue itself disarms the rule is falsified.","tokens_in":30323,"feed_emoji":"🛡️","tokens_out":7460,"duration_ms":58011,"temperature":0.7,"pith_summary":"The paper asks what happens to a standing safety rule when an agent compacts its own context into a model-generated summary, and whether the loss can be detected afterwards. It tries to establish that a presence check is not a safety check: a rule that survives only as a degraded textual residue leads the model to perform the prohibited action far more often than a textually intact welded rule does, with all-case gaps of +34 and +57 points under two replay models. If this is right, post-compaction audits that verify only that the rule text is still present give false assurance, and safety-critical agent loops need either an external ground-truth registry plus behavioral checks, or enforcement that does not depend on the summary carrying the rule. The paper also finds that rule-form items survive compaction far better than prominence-matched facts, which is precisely why presence-based checking feels adequate even though survival is not protection. All results are scoped to a single compaction cycle.","feed_headline":"A safety rule can survive context compaction yet stop protecting","feed_subtitle":"After one compaction, residual rules let the banned action through far more often; text-only audits miss it.","key_machinery":"The load-bearing instrument is the W/G/D/X survival taxonomy: W (welded) means the referent and restriction text are recoverable, G (generalized) means the restriction survives only at category level or softened, D (degraded) means only “a constraint exists” remains with the restricted action unrecoverable, and X means the rule is dropped entirely. The taxonomy separates target loss from predicate loss, and the paper couples it with behavioral replay: each summary becomes the entire prior context of a fresh session, the model is asked to perform the prohibited action on the protected target, and a sibling-target request controls for blanket refusal. The D-versus-W replay contrast carries the main argument, because a textual label alone cannot tell whether a surviving-looking rule still fires.","core_discovery":"Under a single compaction cycle, the hypothesized textual severing mode—a rule left present while its referent is silently generalized off-target—was not observed: across ten generalizations, none was non-covering. Instead, the loss is regime-dependent: a single salient rule is either welded to its referent or dropped outright, while under a tighter multi-item budget roughly a fifth of rules degrade to predicate-loss residues. The sharpest discovery is behavioral: replaying those residues in a fresh context, the model performs the prohibited action far more often than it does for intact welded rules—degraded-minus-intact gaps of +34 and +57 points across all replayed cases under two replay models, and +50 and +56 among cases with replay headroom. Even textually intact rules sometimes fail to fire, so textual presence is not behavioral protection. Separately, rule-form items are retained substantially more than prominence-matched facts (about $\\beta = 2.5$ on the log-odds scale), a descriptive advantage that reproduces on a second summarizer forced to compress, and that explains why presence-based checking feels adequate.","pith_inferences":["If the admitted confound is real—degraded digests are disproportionately the harder-compressed and thinner ones—then repairing only the predicate text may not fully restore protection; a testable next step is matching degraded and intact digests on length and richness before replaying.","The same optimistic bias that inflated the LLM judge's survival counts would plausibly infect automated audit checkers built on judges; a production safeguard should include behavioral replay or deterministic tool-call grading at least on a sample.","Because the retention premium is descriptive and the prominence-equalization condition inconclusive, an arm pairing each rule with a single-proposition, non-numeric fact would separate imperative form from content density; the paper itself suggests this extension.","Multi-cycle erosion remains the open frontier: a rule that welds through one cycle may or may not survive repeated compactions, and the likeliest test model is one that preserves content under volume pressure but is forced to compress by a hard budget."],"forward_implications":["Post-compaction audits that check only that a rule is mentioned will certify many degraded residues as safe even though behavioral replay shows the model performs the prohibited action far more often than with an intact rule.","Any audit must verify operative content—the protected target and the restricted action—rather than presence; category-level generalized survivors behave much more like degraded residues than like intact rules.","Because whole-rule omission is invisible from the summary alone, detection requires an external constraint registry held outside the lossy context; the registry establishes textual absence but not whether a surviving rule still fires.","The rule-over-fact retention premium means summarizers are not treating safety rules like incidental trivia; the risk is that a preserved-looking rule is no longer operative, so evaluation must measure behavior, not survival.","Behavioral silent disarming occurs even without textual severing: textually intact rules sometimes fail to fire on replay, so the absence of the severing text pattern does not imply the rule is being obeyed."],"supporting_citations":[{"why":"Provides the motivating behavioral premise: a constraint dropped during compaction drives violations, while a policy that survives the summary continues to be obeyed.","marker":"Chen, 2026"},{"why":"Analyzes the OpenClaw incident and proposes that compaction summarized the safety instruction away, which supplies the referent-severing hypothesis.","marker":"Ding, 2026"},{"why":"Reports that prohibition-type constraints decay under growing context, marking the constraint class this paper studies as the fragile one.","marker":"Gamage, 2026"},{"why":"Documents unstable safety and refusal behavior in long-context agents, which this paper distinguishes from compaction-specific loss.","marker":"Hadeliya et al., 2025"},{"why":"Proposes verifiable commitment-preserving compression as a mitigation counterpart to the post-compaction audit the paper motivates.","marker":"Trukhina and Vashkelis, 2026"},{"why":"Establishes that LLM safety judges are fragile, the background for the paper's worked judge-error cases.","marker":"Eiras et al., 2025"},{"why":"Provides the 'lost in the middle' baseline on how models use long contexts, which motivates isolating compaction as a distinct failure surface.","marker":"Liu et al., 2024"}],"fun_headline_variants":["Surviving text, failing behavior: AI guardrails after one compaction","Presence isn't protection: safety rules can survive yet stop working","Compaction keeps the rule text, but the ban loses its teeth","AI safety rule survives compaction, then lets the banned action through","Audit checks text, not behavior: surviving rules often fail to protect"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that a degraded residue lets the prohibited action through rests on the assumption that the residue itself—not the generally thinner, harder-compressed digest that carries it—causes the extra violations; the paper did not match digests on compression.","fun_headline_variants_meta":{"raw":{"variants":["Surviving text, failing behavior: AI guardrails after one compaction","Presence isn't protection: safety rules can survive yet stop working","Compaction keeps the rule text, but the ban loses its teeth","AI safety rule survives compaction, then lets the banned action through","Audit checks text, not behavior: surviving rules often fail to protect"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000202,"raw_usage":{"total_tokens":1436,"prompt_tokens":1055,"completion_tokens":381,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":289}},"tokens_in":671,"tokens_out":381,"duration_ms":4000,"temperature":1.0,"reasoning_tokens":289,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:13:07.974220+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same degraded-versus-intact replay census with digests matched for length, output-token budget, and overall content richness; if the compliance gap collapses or reverses once the digests are equally thin, the central claim that the residue itself disarms the rule is falsified.","supporting_citations":[],"review_version":1}