{"id":"c3d5366a-d857-46c5-ab0e-72adc5426ff7","arxiv_id":"2607.20433","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"MOIR: estimating the preservation covariance from a model's own random-token generations reduces collapse of math/code capabilities in some knowledge-editing settings, but the claimed consistency is not supported by the paper's tables.","lead":"A new method, MOIR, builds the covariance matrix that protects a language model during knowledge editing from the model's own randomly-seeded text samples instead of from Wikipedia. In the most favorable reported case it preserves 79.9% GSM8K accuracy after 20,000 edits, versus 10.9% with the Wikipedia baseline; however, the paper's own tables show the benefit is not consistent across all tested models and editors.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CMOIR's 'consistent' preservation claim is contradicted by the paper's own tables; the effect is largely Qwen3-specific, with no gain or collapse in several OLMo-2/Llama-3 settings.","rationale":"The reader's weakest_assumption—Eq. A.1.4, that single-random-token self-generation approximates the training distribution at the activation level—is a real but secondary concern. The more decisive issue is internal inconsistency between the paper's consistency claims and its own full results. The central claim is not merely that MOIR works in one favorable configuration; it is that the model itself is a generally reliable source of the preservation distribution. That requires consistent positive effects across editors/models/regimes. The tables show several settings where CMOIR is no better than CWiki (or worse) and two models where it also collapses completely under MEMIT sequential, albeit later. This is load-bearing because it directly undermines the generality of the proposed principle, independent of whether Eq. 4 holds. The proposed concrete test—per-setting effect sizes with confidence intervals—would determine whether the apparent reversals are noise or genuine model-specific interactions. I therefore partially agree with the reader: their rationale identifies the same empirical problem, but their stated weakest_assumption does not, so my concern is a different load-bearing point. The verdict remains REJECT/UNCHANGED because the paper as written overclaims consistency and would need substantially narrowed claims and error bars to support its central conclusion.","tokens_in":28009,"tokens_out":6702,"duration_ms":71568,"concrete_test":"Run the full MEMIT/AlphaEdit battery (batch and sequential, all three models) with at least three independent MOIR generation runs per model (N=100K each) and report per-model/per-regime preservation-HM differences (CMOIR − CWiki) with 95% bootstrap confidence intervals. If the interval for OLMo-2/Llama-3 AlphaEdit includes zero or is negative, the 'consistent' claim fails. As a cheaper first check, recompute Tables 9–12 as a per-setting effect-size table and count how many settings show positive preservation differences outside Qwen3; a majority of non-positive settings would settle that the headline overgeneralizes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6 claims CMOIR 'consistently extends preservation' and, under MEMIT sequential, 'sustains 0.5–0.7 preservation throughout 10^4 edits.' The paper's own Table 10 contradicts this: CMOIR's preservation HM is 0.000 for OLMo-2 from 500 edits onward and for Llama-3 from 200 edits onward; it only delays collapse (OLMo-2 at 200 edits: 0.503 vs 0.026; Llama-3 at 100 edits: 0.460 vs 0.000). In AlphaEdit batch (Table 11), CMOIR is slightly worse than CWiki on OLMo-2 at 5K (0.600 vs 0.603) and on Llama-3 at 5K (0.662 vs 0.669). In AlphaEdit sequential (Table 12), Llama-3 is worse at 500/1K/2K (e.g., 0.670 vs 0.675 at 500). The headline result (GSM8K 79.9% vs 10.9%) is specific to Qwen3-8B with AlphaEdit. Across the full matrix, the effect is model- and regime-dependent, so the central generalization—that self-generated covariance consistently protects cross-domain capabilities—is unsupported. The mechanism may still be real for some models, but the paper's data do not establish a general principle. No confidence intervals are reported, making it impossible to distinguish real reversals from noise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the cross-domain capability collapse observed in covariance-based knowledge editors (MEMIT, AlphaEdit) stems from the choice of the preservation distribution used to estimate C. It proposes MOIR, which estimates C from the model's own decoding distribution by seeding generation with a single random vocabulary token. Claims are made that this 'self-generated manifold' consistently extends preservation across OLMo-2, Llama-3.1, and Qwen3, in both batch and sequential regimes, up to 20K edits, with a headline GSM8K recovery from 10.9% to 79.9% on Qwen3-8B under AlphaEdit batch editing. The paper also proves Proposition 1 (the parameter update depends on the input distribution only through the input covariance) and Lemma 3 (TV-Lipschitz continuity of the covariance map).","tokens_in":28368,"tokens_out":5253,"duration_ms":54863,"significance":"If the central empirical claim were true, the contribution would be significant: a data-free, drop-in replacement for the preservation covariance that consistently extends the usable edit budget across models and editors. The theoretical component is sound: Proposition 1 is correct and cleanly isolates the role of C, and Lemma 3 provides a usable bound. The method is also practical, requiring only a one-time generation pass. However, the paper's own tables contradict the 'consistently extends preservation' claim. The effect is large and clear only for Qwen3-8B under AlphaEdit; in several OLMo-2 and Llama-3 conditions, CMOIR is indistinguishable from or slightly worse than CWiki. Without confidence intervals or a pre-registered multi-seed analysis, the current evidence does not establish the proposed general principle, only a promising mechanism for a specific model/editor combination.","major_comments":[{"comment":"The claim that 'CMOIR sustains 0.5–0.7 preservation throughout 10^4 edits' under MEMIT sequential editing is directly contradicted by Table 10. OLMo-2: preservation HM is 0.503 at 200 edits, then 0.000 at 500, 1K, 2K, and 5K edits. Llama-3: preservation HM is 0.460 at 100 edits, then 0.000 at 200 edits and all later counts. Only Qwen3 retains non-zero preservation at 5K (0.657). Thus the 'across all three models' statement in Section 6 is false.","section":"Section 6, Table 10"},{"comment":"The 'consistently extends preservation' claim also fails under AlphaEdit. In Table 11 (batch), CMOIR is slightly worse than CWiki at OLMo-2 5K (0.600 vs 0.603) and Llama-3 5K (0.662 vs 0.669). In Table 12 (sequential), Llama-3 is consistently worse at 500 (0.670 vs 0.675), 1K (0.668 vs 0.676), and 2K (0.661 vs 0.672) edits. These reversals are small, but the paper claims consistency and reports no confidence intervals, so they cannot be dismissed as noise.","section":"Table 11 and Table 12"},{"comment":"The key design choice, seed length k=1, is selected on OLMo-2 using the same Se/Sp harmonic-mean scores that later serve as the headline evaluation metrics. This is selection-on-evaluation risk. The paper reports only a single run of each configuration and no variance across seeds or generated samples. A proper validation would hold out the seed-length choice or report bootstrap/interval estimates before claiming 'consistent' improvements elsewhere.","section":"Section 4, Table 2"},{"comment":"The load-bearing approximation E_{x~pθ(·|s)}[φφ^T] ≈ E_{x~Dtrain}[φφ^T] is asserted rather than demonstrated. The NLL/JSD analyses in Figure 3 and Figure 11 are suggestive about token-level distributions, but they do not directly validate equality at the activation-covariance level. The paper's own Limitations section concedes that extreme RLHF mode collapse would under-sample rare capabilities. This is an additional correctness risk, though the empirical contradictions in Tables 10–12 are already sufficient to undermine the central claim.","section":"Appendix A.1, Eq. (4)"}],"minor_comments":[{"comment":"The text says 'Please refer to Appendix C for the full mathematical formulations,' but the parameter-update formulations are in Appendix B, not Appendix C.","section":"Section 3.1"},{"comment":"The caption 'MEMIT sequential editing (cached)' uses 'cached' without explaining what is cached; either define it or remove it.","section":"Table 10 caption"},{"comment":"'The tested models sustain 79.9% accuracy on GSM8K after 20,000 edits' is inaccurate: only Qwen3-8B under AlphaEdit batch achieves this, as the abstract itself states.","section":"Section 1, final paragraph"},{"comment":"The sentence 'Figure 5 and 13 shows task-wise preservation' mixes singular/plural; should read 'Figures 5 and 13 show.'","section":"Figure 5"},{"comment":"The row 'Qwen3-8B' likely denotes the instruct variant but is not labeled as such, unlike 'OLMo2-7B-Instruct' and 'Llama-3.1-8B-Instruct.'","section":"Table 4"}],"recommendation":"reject","confidential_remarks":"The mathematical scaffolding (Proposition 1, Lemma 3) is solid and the diagnostic framing is interesting, but the paper's headline empirical claim is contradicted by its own Tables 10–12. A revision that restricted the claim to Qwen3 under AlphaEdit would be a substantially different and much narrower contribution. I see no path within the current manuscript's framing to the claimed generality, so rejection is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The novel piece here is real: sampling the preservation covariance C from the model's own random-token-seeded generations, rather than from Wikipedia or a pretraining mix, is a fresh and plausible move. Proposition 1 (G-independence) is correctly stated and proved, and the OLMo-Mix oracle ablation is a nice control. The NLL distributional analysis also supports the claim that post-training shifts the model's operative manifold away from static corpora. That part of the paper is solid and useful.\n\nThe soft spot is the empirical overclaim. The abstract and Section 6 say MOIR \"consistently extends preservation\" across editors, models, and regimes. The paper's own tables don't support that. In AlphaEdit batch, CMOIR gives essentially no gain on OLMo-2 and Llama-3.1 (Table 11: e.g., OLMo-2 5K .600 vs .603; Llama-3.1 5K .662 vs .669). In MEMIT sequential, CMOIR delays collapse on OLMo-2 and Llama-3 but still collapses to zero preservation HM by 500 edits (OLMo-2) and 200 edits (Llama-3) (Table 10), contradicting the \"sustains 0.5–0.7\" claim. The headline 79.9% GSM8K result is real but specific to Qwen3-8B with AlphaEdit. So the general principle is not established; the effect is model- and regime-dependent.\n\nA few smaller issues: no confidence intervals anywhere, so some reversals may be noise. The seed-length choice (k=1) is tuned on the same harmonic-mean score used in the main results, which is a selection-on-evaluation concern. And the load-bearing approximation in Eq. 4—that a single random token makes self-generated activations proxy the training distribution—is an assumption the authors themselves hedge in Limitations. It may hold for these instruct models but is not shown to hold generally.\n\nWho is this for? Researchers working on knowledge editing and model updating will want to know this idea. The method deserves further study, and a serious referee should see it. But the authors need to narrow their claims, add error bars, and report per-model/per-regime results without the \"consistent\" framing. The paper as written is not ready for acceptance, but it is not a desk reject either. I'd send it to review and ask for major revision.","headline":"The core idea—estimating the preservation covariance from the model's own random-seeded generations—is genuinely new and worth studying, but the paper's 'consistent' preservation claim is contradicted by its own tables; the real effect is mostly Qwen3-AlphaEdit-specific.","tokens_in":28920,"tokens_out":2126,"would_cite":true,"duration_ms":23212,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that knowledge editing destroys math and code reasoning because the preservation covariance is built from the wrong distribution, and shows that sampling the model's own continuations seeded with one random token fixes it.","keywords":["knowledge editing","model editing","covariance preservation","self-generation","capability preservation","MEMIT","AlphaEdit","distributional shift"],"falsifier":"Run MOIR on a model whose decoding distribution is pathologically mode-collapsed, for example after extreme RLHF, then apply 20,000 AlphaEdit batch edits and measure GSM8K accuracy; if math accuracy collapses despite MOIR because the self-generated covariance under-samples that specialized domain, the core approximation fails.","tokens_in":27823,"feed_emoji":"🧠","tokens_out":4328,"duration_ms":42946,"temperature":0.7,"pith_summary":"The paper claims that knowledge editing destroys math and code abilities because the covariance matrix used to protect existing knowledge is built from the wrong distribution—an encyclopedic text corpus that does not match what the model actually internalized after SFT and DPO. It proposes MOIR: build that covariance from the model's own sampled continuations, seeded by a single random vocabulary token. This change keeps GSM8K math accuracy at 79.9% after 20,000 batch edits where the encyclopedic-corpus baseline falls to 10.9%, across several 7–8B models and two editing methods. The broader claim is that the preservation space for covariance-based editors can be derived from the model itself, without any external data.","feed_headline":"A single random token keeps math intact through 20,000 edits","feed_subtitle":"Model-generated samples, not a static corpus, form the preservation space that stops edits from destroying reasoning.","key_machinery":"The uncentered activation covariance C = (1/N) Σ k_i k_i^T, computed at the input to the targeted MLP layer, is the single distributional object that determines a covariance-constrained editor's update. MOIR constructs C from autoregressive continuations seeded with one uniformly random vocabulary token, which breaks the instruction-template attractor and lets the model sample the broader subspaces it has internalized. Proposition 1 shows that the output-side K-FAC factor G cancels in the closed-form optimum, so the choice of C—self-generated versus corpus-derived—fully controls what gets preserved.","core_discovery":"MOIR's central assertion is that the self-generated covariance CMOIR samples from the model's actual internalized distribution, covering exactly the subspaces the model relies on, regardless of whether they appear in any external corpus. The paper proves that for closed-form covariance-constrained editors, the update depends on the input distribution only through the uncentered covariance C, because the output-side factor G cancels in the closed-form optimum (Proposition 1). Empirically, replacing the standard encyclopedic proxy with CMOIR delays or prevents the collapse of GSM8K and HumanEval under both MEMIT and AlphaEdit, in batch and sequential regimes, across OLMo-2, Llama-3.1, and Qwen","pith_inferences":["If the self-generated covariance truly tracks the operative distribution, the same idea could be used to audit which capability subspaces a deployed model has internalized, separate from any editing task.","A testable extension the paper leaves open is whether memory-based or hypernetwork editors, which use covariance-like operators, would also benefit from self-generated preservation spaces.","The single-random-token recipe suggests a general principle for probing internalized distributions: a minimal off-manifold perturbation escapes post-training attractors while staying within the model's learned statistics; this could inform data-free evaluation beyond editing."],"forward_implications":["Knowledge editing no longer requires access to pretraining or post-training corpora; a roughly two-hour one-time generation from the deployed model provides the preservation covariance.","Covariance-based editors can sustain tens of thousands of batch and sequential edits without collapsing math and code reasoning, as long as their preservation space is aligned with the model's operative distribution.","The model itself is the most accessible source of the shifted post-training manifold; static corpora, including the original pretraining mixture, are systematically biased proxies.","Since the editing update depends only on the input covariance, improving the distribution used to estimate C is a strict improvement across editors without architectural changes."],"fun_headline_variants":["Model's own random token beats Wikipedia for edit safety","One token from the model preserves math through 20k edits","Self-seeded covariance: edits without erasing reasoning","Model's internal distribution stops edit-induced reasoning collapse","Drop-in fix: let the model set its own preservation space"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that continuations seeded with one random token reproduce the model's true operative activation distribution closely enough that the estimated covariance covers the math and code subspaces; if the model's decoding is mode-collapsed or the random seed never enters those regions, self-generation misses them.","fun_headline_variants_meta":{"raw":{"variants":["Model's own random token beats Wikipedia for edit safety","One token from the model preserves math through 20k edits","Self-seeded covariance: edits without erasing reasoning","Model's internal distribution stops edit-induced reasoning collapse","Drop-in fix: let the model set its own preservation space"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00026,"raw_usage":{"total_tokens":1488,"prompt_tokens":865,"completion_tokens":623,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":544}},"tokens_in":609,"tokens_out":623,"duration_ms":6904,"temperature":1.0,"reasoning_tokens":544,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T14:25:18.976575+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run MOIR on a model whose decoding distribution is pathologically mode-collapsed, for example after extreme RLHF, then apply 20,000 AlphaEdit batch edits and measure GSM8K accuracy; if math accuracy collapses despite MOIR because the self-generated covariance under-samples that specialized domain, the core approximation fails.","supporting_citations":[],"review_version":1}