{"id":"ed3767b2-94a0-4751-8aa3-1e761a28077d","arxiv_id":"2608.12599","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Revoked constraints still shape model behavior at an 8B operating point, and compiling the net constraint state ahead of time removes the observed relapse.","lead":"A new system, ReBIND, measures when AI assistants keep obeying rules users have already withdrawn, a failure called behavioral relapse. It predicts relapse before delivery and shows that removing withdrawn rules from the prompt ahead of time eliminates the observed relapse at an 8B operating point.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The placeholder ablation in §4.2 is unvalidated against the deletion counterfactual, and the paper's own bridge data show they disagree; since all triage and prediction claims use the placeholder, the measurement/prediction contributions are not yet secured.","rationale":"The existence and restoration results are well supported: the relapse detector is blind-audited, the load contrast is pre-registered and cluster-bootstrapped, zero numerators carry denominators and rule-of-three bounds, and the compilation contrast is budget-matched. I do not see a soundness problem in those headline claims. The load-bearing soft spot is the per-clause causal measurement on which the paper's 'measure' and 'predict' contributions rest. BC is defined through a placeholder ablation, yet Section 4.2 reports that placeholder and deletion ablations agree no better than chance (κ=0.000, n=10), and Section C's placebo calibration only validates inert text at one fixed insertion position, not the placeholder-as-deletion counterfactual for real clauses. Since the five-state triage, the repair-ladder routing, and the prospective AUROC all consume BC, an invalid counterfactual would corrupt the measurement and prediction claims even though the headline relapse and compilation contrasts would survive. The reader's weakest-assumption analysis identifies the same point; I agree. The correct verdict remains CONDITIONAL: the ablation choice needs an independent check before the instrument is treated as a general measurement standard. The model-selection issue and post-hoc placebo arm are secondary and do not change this assessment.","tokens_in":19343,"tokens_out":10852,"duration_ms":116136,"concrete_test":"Re-run the Section 5.4 prospective probe on the same 134-slice horizon set with the deletion ablation substituted for the placeholder, holding the frozen stopping rule and matched budget fixed, and recompute the five-state diagnoses and the AUROC. If the deletion-based AUROC drops below the pre-registered 0.70 adequacy criterion, or if the confident diagnosis labels shift for a non-trivial fraction of clauses relative to the placeholder version, the placeholder counterfactual is load-bearing for the measurement/prediction claims.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 defines BC via an equal-length neutral-placeholder ablation and reports that this science-mode ablation and outright deletion 'agreed no better than chance (κ=0.000, raw agreement 0.6, n=10 clauses)'. The paper nevertheless builds the five-state triage (Table 1), the repair-ladder routing, and the prospective AUROC (0.897) entirely on the placeholder version. The placebo-clause calibration in Section C is not a substitute ground truth: it shows that administrative placebo text inserted into a local ledger copy at one fixed position yields zero incremental effect, but it does not show that replacing a previously in-force real clause with a placeholder preserves the counterfactual that deletion would implement. If the placeholder changes clause influence through length, position, or list structure rather than through content, BC, the triage labels, and the prediction signal can shift; the paper's own bridge data indicate such a shift is plausible. The detector-based existence result and the compilation repair contrast do not depend on BC, but the measurement and prediction contributions—the paper's first two stated gaps—do.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces \"behavioral relapse\" of revoked constraints in multi-turn LLM dialogues: models continue enacting withdrawn requirements. It presents ReBIND, a ledger-based diagnostic system that pairs each constraint with an executable checker, records revocations as tombstones, compiles the net constraint state into a single specification, measures per-clause adherence and incremental behavioral effect with a sequential ablation probe, and routes clauses through a five-state triage and a repair ladder. The empirical claims are built on RELAPSE-Code (67 HumanEval tasks, 201 verified checkers). Headline results are the load-scaling contrast (relapse at an 8B operating point rises from 0.011 at m=2 to 0.403 at m=8; pre-registered primary difference +0.392, 95% CI [+0.300,+0.483], p=1.3e-12), the prospective relapse prediction (AUROC 0.897, 95% CI [0.829,0.963]), the restoration contrast against a no-ledger verifier-retry baseline (0.192, 95% CI [0.134,0.251], p≈1e-10), zero observed relapse under ahead-of-time compilation (0/2968, rule-of-three upper bound ≤0.10%), and a placebo-controlled tombstone-note effect. The paper is unusually careful about pre-registration, matched budgets, conservative scoring semantics, zero counts with denominators, and cost reporting.","tokens_in":19610,"tokens_out":13097,"duration_ms":124478,"significance":"Behavioral relapse is a well-motivated and practically important failure mode, and the paper makes a credible case that it has not previously been isolated with executable per-clause checkers. The load-scaling and restoration results are pre-registered, use cluster-level inference, report zero counts with denominators and rule-of-three bounds, and compare arms under matched checkers, model, and token budget. The reproducibility apparatus (decision-record chain, frozen probe parameters, run identifiers, full cost accounting) is exemplary. If the measurement and prediction instruments are validated, the paper would close a real gap; as it stands, the measurement and prediction contributions rest on an unvalidated ablation choice that the paper's own bridge data call into question.","major_comments":[{"comment":"The equal-length neutral-placeholder ablation is load-bearing for the measurement and prediction contributions. Section 4.2 reports that the placeholder ablation and outright deletion \"agreed no better than chance (κ=0.000, raw agreement 0.6, n=10 clauses)\", yet the paper states that all statistics use the placeholder version. This choice feeds the five-state triage in Table 1, the sequential probe's BC classification, the diagnosis-time risk signal, and therefore the prospective AUROC of 0.897 in §5.4. The placebo-clause calibration in Section C does not validate this counterfactual: it inserts administratively inert text at a single fixed position in a local ledger copy, whereas the scientific question is whether replacing an in-force, content-bearing real clause with a length-matched placeholder preserves the causal effect that deletion would implement. The paper needs an independent ground truth for per-clause causal effect, or a sensitivity analysis reporting the triage labels and the AUROC under a deletion-based BC, or a substantially larger bridge study with a prespecified reconciliation rule. Without one of these, the first two stated gaps (measurement and prediction) are not secured, even though the detector-based existence result and the compilation restoration contrast do not depend on BC.","section":"§4.2, §3.2, §5.4"},{"comment":"The operational definition of relapse requires that a revoked clause was \"previously adopted: it demonstrably shaped earlier answers.\" Because the benchmark elicits only a single final implementation per episode and contains no assistant turns, no earlier artifact is ever observed; Section 3.3 asserts adoption \"in distribution\" on the strength of a 0.983 aggregate in-force compliance rate. The detector then labels any post-revocation marker definition as true_relapse even in an episode where that specific marker was never complied with while in force. The authors should either verify per-episode adoption (e.g., with a pre-revocation compliance checkpoint) or report the load-scaling and AUROC results restricted to episodes with verified adoption; without this, the relapse rates may include a small number of delayed first-time compliance events rather than strict relapse.","section":"§3.3, §3.4"}],"minor_comments":[{"comment":"The sentence \"over grid points 0.011, 0.238, and 0.403\" should explicitly tie those rates to loads m=2, m=5, and m=8, since the reader otherwise has to reconstruct the mapping from Figure 1a.","section":"§5.2"},{"comment":"The table rows run together in the rendered text (e.g., \"high>0adopteddeliver\"); adding explicit separators between the AC/BC columns, state names, and actions would substantially improve readability.","section":"Table 1"},{"comment":"The phrase \"adoption therefore holds in distribution rather than in-context\" is used only once and is not defined elsewhere; a one-sentence explanation of what in-context adoption would require would prevent a reader from misreading the distributional claim as per-episode evidence.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is exceptionally disciplined in its statistical reporting, pre-registration, and cost transparency, and the restoration result appears robust. My recommendation is driven almost entirely by the ablation-counterfactual issue in §4.2, which is load-bearing for the measurement and prediction contributions. If the authors can supply a convincing validation of the placeholder counterfactual, or a sensitivity analysis showing that the triage labels and AUROC are stable under a deletion-based definition of BC, I would be happy to support acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: read this one. The paper identifies a real failure mode—revoked constraints still shaping model behavior—and measures it with unusual discipline. The central existence and repair claims are solid. But the per-clause influence measurement, which drives the triage and the prediction AUROC, rests on an ablation that the paper itself shows disagrees with the natural deletion counterfactual. That needs fixing before the measurement contribution is trusted.\n\nWhat is new: behavioral relapse (revocation inertia) as the mirror image of instruction decay; a cheap API-only measurement loop (contract ledger, tombstones, compiled specification, sequential ablation probe); RELAPSE-Code with 201 verified checkers; the load-scaling result at the 8B tier; the compilation repair contrast at matched budget. The study costs under $20 in API compute, which is honestly reported with token accounting and billed prices. That reproducibility bar is real: pre-registration with decision records, leakage screens, blind human audits, zero denominators with rule-of-three bounds. The paper does not oversell—it labels exploratory contrasts, reports the nulls, and even reports the deprecated cross-family grid.\n\nWhere it is soft: Section 4.2 defines the incremental effect BC by replacing clause text with an equal-length placeholder, and explicitly reports that this 'science-mode' ablation and outright deletion 'agreed no better than chance (κ=0.000, raw agreement 0.6, n=10)'. The five-state triage, the repair-routing, and the prospective AUROC (0.897) all use the placeholder version. The paper offers a placebo-clause calibration, but that does not establish that the placeholder is the right counterfactual for a previously in-force real clause. The stress-test note has this right. The existence result (relapse scales with load) and the restoration contrast (compilation vs. verifier-retry, 0.192 CI) do not depend on BC, so those hold. The weaker tier is the measurement/prediction claim: it is conditional until the ablation choice is defended or the analysis is repeated with deletion.\n\nAlso minor: the operating point is the one model where the pilot found the phenomenon; stronger models sit at floor, so the headline is a capability gradient, not a replication. The placebo arm was added post hoc (declared as such), and the cross-family compiled cell is unusable. These are stated plainly in the paper.\n\nWho should read it: anyone building dialogue-state or contract-ledger systems, and evaluation people who care about per-clause behavioral influence rather than text-level compliance. It deserves a serious referee; a conditional accept with the ablation validity as the main revision is the right outcome. I would bring it to reading group.","headline":"Real phenomenon, unusually disciplined measurement, but the per-clause influence metric rests on an ablation the paper itself shows is unvalidated.","tokens_in":20115,"tokens_out":2391,"would_cite":true,"duration_ms":22835,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Revoked constraints in multi-turn dialogues keep shaping model behavior even after withdrawal, and this 'behavioral relapse' is measurable, predictable, and repairable through a contract-ledger intervention.","keywords":["behavioral relapse","revocation inertia","constraint influence","contract ledger","ablation probe","dialogue state","black-box LLM evaluation","multi-turn instruction following"],"falsifier":"Re-run the ablation probe on a larger matched sample comparing deletion-based and placeholder-based ablations with independent gold-standard labels of whether the revoked behavior was actually enacted; if the two modes disagree systematically on triage state, the per-clause influence measurements, the five-state diagnoses, and the prospective relapse prediction would not be tied to the clause's causal effect.","tokens_in":19085,"feed_emoji":"🧾","tokens_out":7734,"duration_ms":60830,"temperature":0.7,"pith_summary":"This paper tries to establish that revoked constraints in multi-turn LLM dialogues are not dead text: they keep shaping behavior, and this persistence is a measurable, predictable, and repairable property of dialogue state. It introduces a contract ledger that records every constraint, a tombstone for revocations, and a compiled net specification; a sequential ablation probe that estimates each clause's adherence and incremental behavioral effect; and a repair ladder run under matched token and attempt budgets. On a benchmark of 67 coding tasks with 201 executable checkers, relapse at an 8-billion-parameter model rises from 0.011 to 0.403 as constraint load grows, while stronger models show none. If true, this means revocation failures can be audited and corrected through the model API alone, and current multi-turn evaluations that ignore revoked-clause influence miss a systematic failure mode.","feed_headline":"Revoked instructions still bind: LLMs relapse at 40% under load","feed_subtitle":"A contract-ledger compile removes observed relapse at matched budgets; a tombstone note recovers a third of the effect.","key_machinery":"The central machinery is the contract ledger, a data structure that pairs every user constraint with an executable checker and a binding history. Revoking a clause writes a tombstone—the record survives but the obligation does not—and the net in-force set is compiled ahead of time into a single pseudo-single-turn specification. On top of this substrate, a sequential ablation probe measures adherence, the pass rate of the clause's checker, and incremental behavioral effect, the difference when the clause's text is replaced by an equal-length neutral placeholder, classifying clauses into five states: adopted, redundant, underpowered, inert, or adverse. A repair ladder then intervenes by editing surface text only, never clause text or checker, and re-tests under token- and attempt-matched budgets.","core_discovery":"The central claim is that behavioral relapse—continuing to enact a requirement the user has withdrawn—is not residual noise but a structured function of constraint load. At an 8-billion-parameter operating point, delayed-revocation relapse climbs from 0.011 at low load to 0.403 at high load, with a pre-registered contrast of +0.392 at p≈$10^{-12}$, while stronger models sit at zero. The paper further claims that maintaining a contract ledger and compiling net in-force constraints ahead of time removes observed relapse: 0/2968 episodes with a rule-of-three upper bound of 0.10%, against a 0.192 relapse reduction over a no-ledger verifier-retry baseline. A one-sentence tombstone note causally reduces relapse by +0.048 and survives a placebo control, so even the revocation record itself carries binding force.","pith_inferences":["A natural extension would test whether ledger-based compilation transfers to other domains, such as natural-language instructions or structured data generation, where revocations are common; the paper reports only Python code tasks.","The finding that relapse is already present at the revocation turn and does not accumulate with depth suggests the mechanism is a prompt-interpretation or retrieval failure rather than gradual memory decay; white-box studies of attention or similar in-context behavior could test this directly.","The zero observed relapse under compilation across all cells, if stable, implies that model-level instruction-following limits are not the bottleneck for this failure: the same model, given a clear compiled specification, can stop honoring withdrawn clauses.","Because the measurement pipeline uses only the model API, it could be run at deployment time as a continuous audit, flagging dialogues where a revoked clause is still shaping output; the paper stops at measurement and repair rather than recommending this deployment."],"forward_implications":["Revoked constraints should be treated as a first-class failure mode in multi-turn evaluation; current benchmarks score compliance with in-force requirements only and would miss this.","Ahead-of-time compilation of the net constraint state into a single specification can remove observed relapse at the 8-billion-parameter tier under matched budgets, with a delivery overhead of 1.49×.","Even a one-sentence tombstone note recovers about a third of the compilation effect and beats a placebo, so simply telling the model a requirement was revoked has measurable causal force.","Adaptive ladder routing adds no detectable gain beyond compiled form at this operating point; the design excludes gains of 1.3 percentage points or more.","The probe's diagnosis-time signal predicts later relapse with AUROC 0.897, so at-risk dialogues can be flagged before delivery."],"supporting_citations":[{"why":"Supplies the 67 coding tasks and checker format used to construct the RELAPSE-Code benchmark.","marker":"Chen et al., 2021"},{"why":"Provides the sequential analysis behind the probe's stopping rule.","marker":"Wald, 1945"},{"why":"Controls false discovery across clause families in the evaluation.","marker":"Benjamini & Hochberg, 1995"},{"why":"Provides the cluster bootstrap used for confidence intervals.","marker":"Efron & Tibshirani, 1994"},{"why":"Supplies the rule-of-three upper bounds for zero relapse counts.","marker":"Hanley & Lippman-Hand, 1983"},{"why":"Provides the design-by-contract notion underlying the contract ledger.","marker":"Meyer, 1992"},{"why":"Supplies the negative-control reasoning that motivates the placebo and counterfactual arms.","marker":"Lipsitch et al., 2010"}],"fun_headline_variants":["LLMs keep obeying revoked rules; ledger compile kills relapse","Behavioral relapse: why LLMs cling to withdrawn constraints","Contract ledger removes revoked-rule inertia; tombstone helps","Revoked rules still bind: load drives relapse, compile stops it","Predict and repair LLM revocation failure with contract ledgers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that swapping a clause's text for an equal-length neutral placeholder isolates that clause's causal influence on behavior, and the paper's own bridge comparison against a deletion-based ablation agreed no better than chance in 10 clauses.","fun_headline_variants_meta":{"raw":{"variants":["LLMs keep obeying revoked rules; ledger compile kills relapse","Behavioral relapse: why LLMs cling to withdrawn constraints","Contract ledger removes revoked-rule inertia; tombstone helps","Revoked rules still bind: load drives relapse, compile stops it","Predict and repair LLM revocation failure with contract ledgers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000882,"raw_usage":{"total_tokens":3876,"prompt_tokens":1074,"completion_tokens":2802,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":690,"completion_tokens_details":{"reasoning_tokens":2718}},"tokens_in":690,"tokens_out":2802,"duration_ms":16468,"temperature":1.0,"reasoning_tokens":2718,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:04:48.694015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the ablation probe on a larger matched sample comparing deletion-based and placeholder-based ablations with independent gold-standard labels of whether the revoked behavior was actually enacted; if the two modes disagree systematically on triage state, the per-clause influence measurements, the five-state diagnoses, and the prospective relapse prediction would not be tied to the clause's causal effect.","supporting_citations":[],"review_version":1}