{"id":"ec83f471-ce14-4348-a12a-fece8e20aaf1","arxiv_id":"2607.19442","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Matching a retrained oracle on trained probes can certify models that still retain held-out forget knowledge, and oracle-free unlearning certification is only possible for counterfactual, non-inferable facts.","lead":"Machine-unlearning methods that match a retrained 'oracle' on the exact question phrasings used in training are shown, in a controlled testbed, to still retain the forgotten knowledge on different phrasings—2.82 nats below the never-learned level. The paper then builds a validated held-out screen and proves when unlearning quality can be certified at all without the oracle.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The −2.82-nat headline assumes counterfactual exchangeability (Assumption 3, §9) between held-out forget facts and never-learned probes; this identity is asserted rather than verified against the matched references, and if false the residual-knowledge interpretation is confounded by fact-set difficu","rationale":"The reader's weakest-assumption analysis identified Assumption 3 (counterfactual exchangeability) as the load-bearing premise, and I agree. The paper is unusually careful: it runs the reference through its own battery, discloses tolerance concentration, and scopes its certificate. The strongest claim is an empirical finding, but its interpretation as 'retain paraphrase-recoverable knowledge' depends on the never-learned distribution being the correct counterfactual baseline. CE is exactly that premise. The available matched references in the 45 cells make CE directly testable, so there is no need to speculate. If CE holds in the testbed, the falsification stands and the oracle-free theory is on firmer ground; if it does not, the headline number would need to be reinterpreted as a mismatch relative to the reference rather than proof of retained knowledge. The undefined 'adequate' threshold in Section 3.1 and the missing finite-query theorem are secondary: they affect reproducibility and scope, not the logical spine. The recalibrated certificate is explicitly supporting, so its tolerance sensitivity does not threaten the central claim. Therefore I keep the reader's CONDITIONAL verdict (no change) pending the CE test.","tokens_in":15995,"tokens_out":9708,"duration_ms":84582,"concrete_test":"In each of the 45 matrix cells, compute (a) the matched retraining reference's held-out forget-NLL distribution on F paraphrases and (b) M_inj's never-learned probe-NLL distribution on P paraphrases, using identical templates and scoring. Test whether the chosen functional Γ (e.g., the mean) of these two distributions is equal within the measured replication noise δ_equiv (≈0.93 nats), e.g., by a paired permutation test across seeds. Equivalently, rerun the −2.82-nats comparison using the reference's held-out forget position instead of the never-learned baseline as the yardstick; if the residual-knowledge gap is not robust to this replacement, CE is doing the work and the headline must be reframed or a direct CE validation added.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central falsification (Section 3.1) rests on comparing candidates' held-out forget NLLs to a 'never-learned level.' That level is only a valid stand-in for the retrained oracle's forget threshold if Assumption 3 (CE, §9) holds: τ*(w)=Γ(L(NLL_Mor(w) on forgotten examples))=Γ(ν_w), the Γ-functional of M_inj's NLLs on never-learned probes. The paper asserts M0 is 'symmetric' in F and P (§4), but symmetry under the base model does not imply the oracle's NLL law on forgotten examples equals M_inj's NLL law on never-learned probes. F and P are generated from disjoint seeds, so randomization is suggestive, but CE is a distributional identity that can be checked directly in the 45 cells, where the matched reference exists. The only validation offered, Ĝ on a 5-level×5-seed graded-inferability axis (Spearman 0.97), tracks headroom and does not test the full CE identity. If CE fails even in the controlled testbed, the −2.82-nats gap may reflect non-exchangeability of the fact sets rather than retained knowledge; the abstract's 'retain paraphrase-recoverable knowledge' overstates what the data show. This is not an attack on the paper's honesty — CE is labeled an assumption — but the central quantitative claim depends on it more than on any other premise. Secondary flagged omissions (the advertised finite-query impossibility theorem, the audit-pool hash deferred to a revision) do not change the logical spine but should be addressed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies machine unlearning evaluation in a controlled nonce-fact testbed with a matched retraining reference (M_oracle). It reports three main empirical findings: (1) the common trained-probe oracle-KL criterion can favor candidates that retain held-out forget knowledge, quantified as a −2.82-nat gap between candidates' held-out forget NLLs and their never-learned probe NLLs; (2) fixed absolute retain/round-trip thresholds are mis-scaled, since the injected model and even the retraining reference fail them in most of 45 model–seed cells; and (3) a base-anchored held-out rank screen is a validated necessary test with measured sensitivity on a sealed challenge panel, while a damage-relative recalibration yields a selective partial positive in 15/45 cells. The paper also presents an identifiability theorem delimiting when oracle-free selection is possible, with TOFU as the predicted boundary case, and reports that a fixed logit-suppression attack defeats the full forward battery in 12/45 cells, scoping the method as an empirical selective test rather than a formal certificate.","tokens_in":16380,"tokens_out":10850,"duration_ms":103186,"significance":"If the central falsification holds, the field's default ground-truth metric for unlearning—matching a retrained oracle on trained probes—is systematically misleading, and benchmarks built on it inherit the flaw. The paper's strengths are substantial: a controlled dose-matched testbed with a matched retraining reference, independent redraws quantifying replication noise, a sealed known-label challenge panel, frozen preregistered recalibration rules, cluster-bootstrap confidence intervals, every cell reported, and an unusually candid limitations section. The method also ships a finite-sample certifiability bound (Proposition 1) and an identifiability theorem (Theorem 1) that together clarify when oracle-free selection is even possible. The empirical apparatus is reproducible in principle, though the promised audit-pool hash is deferred. The main risk is that the headline residual-knowledge gap is not yet shown to be significantly larger than the baseline F-versus-P gap exhibited by the matched reference itself; if that baseline gap is substantial, the −2.82-nat interpretation is confounded by fact-set difficulty rather than retained knowledge.","major_comments":[{"comment":"The headline −2.82-nat claim compares trained-probe-adequate candidates' held-out forget deltas to their own never-learned probe deltas. The interpretation as retainable knowledge assumes that, had the candidate not learned F, its ΔF distribution would match its ΔP distribution. The manuscript justifies this by asserting M0 is 'symmetric' in F and P, but M0 is not the relevant counterfactual for a candidate derived from Minj; the matched reference Mor is. The reference is available in all 45 cells, yet the paper reports only that it passes the screen in 43/45 (Section 5), not the reference's own F-versus-P delta gap. A one-sided rank test with m=4, n=16 can have low power, so 'passing' does not quantify the gap. Please report the reference's median and distribution of ΔF−ΔP per cell, and show that the −2.82 gap for trained-probe-adequate candidates is significantly larger than the refere","section":"Section 3.1 / Section 4 (ΔF/ΔP definitions)"},{"comment":"The paper repeatedly advertises a 'finite-query impossibility argument' (also 'finite-query impossibility boundary' and, in Section 11, 'our finite-query argument'), and Section 6 refers to 'the finite-query impossibility argument of Section 9.' Section 9 contains Theorem 1, Lemma 1, and Proposition 2 on identifiability, and Proposition 1 in Section 4 is a finite-sample certifiability bound for the rank test; none of these is a query-complexity impossibility theorem. Either supply the promised argument/proof or revise the abstract and contributions to describe what is actually proved: identifiability limits and a measured adversarial evasion rate. This is a discrepancy between promises and content that should be resolved before publication.","section":"Abstract; §1 contribution 5; §6; §11"}],"minor_comments":[{"comment":"The appendix states that 'we publish a cryptographic hash of the audit pool with this paper,' but then says the hash and repository link 'will be added to this section in a revision.' The sealed-pool claim is therefore not currently verifiable. Include the hash or a timestamped commitment in the version of record.","section":"Appendix A"},{"comment":"The theoretical selector in Lemma 1 uses a threshold on q_a(w) against Γ(ν_w), while the implemented screen in Section 4 is a one-sided rank test on base-anchored deltas. Clarify how the rank test operationalizes the threshold condition, especially with only m=4 forget items and n=16 probes, and how the p-value floor of Proposition 1 maps to the Γ-based cutoff.","section":"Section 9 vs Section 4"},{"comment":"The statement and proof of Proposition 1 assume no ties in the rank lattice, but the implementation audit notes that ties push the code onto a normal approximation. State explicitly that the exact-lattice bound applies only in the tied-free case, and whether the 1/4845 floor was observed for all 238 candidates that attained it.","section":"Proposition 1"},{"comment":"The 0.80-vs-5.17 comparison is carefully caveated, but it is computed only in the 15 cells where the recalibrated certificate does not abstain. Please state whether the trained-probe best's 5.17-nat gap is also representative of the full 45 cells, or explain why it is not reportable there.","section":"Table 2"},{"comment":"The claim that the absolute bar fails 'its own reference' is supported by pass rates, but the reference's absolute retain-NLL and round-trip residual distributions are only summarized by medians and pass counts. Reporting per-cell values or a scatter plot would make the threshold-mis-scaling diagnosis more transparent.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest and well-controlled, and the central falsification is important if it holds. The main technical gap is the missing reference-baseline comparison for the −2.82-nat headline: the matched reference is available, and the authors should be asked to report its F-versus-P gap directly. The advertised finite-query impossibility argument also needs to be either provided or explicitly dropped from the claims. Both issues are fixable within the current experimental apparatus, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — this one is worth your time. It makes a sharp, well-controlled argument that the standard ground-truth metric for unlearning—matching a retrained oracle on trained probes—rewards methods that still retain held-out knowledge. Candidates the criterion rates adequate sit 2.82 nats below the never-learned level on held-out forget facts, with a tight cluster CI. That is a real strike against the field's default evaluation, and it is backed by an unusually careful audit: matched retraining reference, independent redraws, sealed challenge panel, and a reference that is itself run through every proposed criterion. The 45/45 rejection of the injected model and 44/45 acceptance of the reference give the screen measured sensitivity rather than hand-waving. Proposition 1's exact rank-lattice bound is a useful design rule for probe counts. The TOFU boundary case is instructive, and the identifiability theorem, though simple, is honestly scoped.\n\nSoft spots: the paper advertises a 'finite-query impossibility boundary' in the abstract and section 1, but no such theorem appears in the body—only the identifiability criterion. That should be fixed before publication, either by adding the result or cutting the claim. More importantly, code, data, and the promised audit-pool hash are absent, deferred to a revision; that makes the sealed-challenge claims uncheckable, which is awkward for an audit methodology paper. The CE assumption (counterfactual exchangeability) is doing real work: the -2.82 nat interpretation assumes the never-learned probes are a fair stand-in for the oracle's forget threshold. The paper labels it honestly, and the reference's 44/45 screen acceptance gives some empirical support, but a direct check of tau* vs Gamma(nu_w) in the 45 cells would be stronger. Tolerance concentration is disclosed: 29/38 certified candidates come from Qwen, and the count drops to 4 under the global minimum tolerance, so the 15/45 headline is fragile.\n\nNone of these are load-bearing flaws. The central falsification and the screen validation hold up internally. This is a paper for unlearning researchers and benchmark builders. It deserves a serious referee; the revision should supply the artifacts, either add or disown the finite-query theorem, and tighten the CE discussion.","headline":"A careful empirical paper that exposes a real flaw in trained-probe unlearning evaluation, but it ships without code/data and one advertised theorem is missing.","tokens_in":16904,"tokens_out":3994,"would_cite":true,"duration_ms":34244,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The standard test for machine unlearning can certify models that still hold the deleted information.","keywords":["machine unlearning","distribution restoration","retraining oracle","trained-probe criterion","held-out evaluation","counterfactual identifiability","selective certification","LLM unlearning"],"falsifier":"In any cell of the paper's testbed, find a candidate that the trained-probe oracle-KL criterion rates adequate but whose held-out forget-NLL gap from the never-learned level is within the replication tolerance (≤0.93 nats) of zero; the paper predicts a gap of roughly -2.82 nats, so such a candidate would directly contradict its central negative result. Alternatively, show that the matched retraining reference itself fails the paper's base-anchored screen in a non-trivial share of cells, which would undermine the screen's validity as a necessary test.","tokens_in":15735,"feed_emoji":"🧠","tokens_out":4958,"duration_ms":40769,"temperature":0.7,"pith_summary":"The paper claims that the standard way of evaluating machine unlearning—checking how close an unlearned model is to a retrained oracle on the same question phrasings used during training—can certify models that still retain the forgotten knowledge. In a controlled testbed with a matched retraining reference, candidates that this criterion rates as adequate score held-out paraphrases of the forget facts 2.82 nats below the never-learned level, meaning the knowledge survives rephrasing. The paper recasts good unlearning as restoration to the retrained distribution, proposes a base-anchored held-out rank screen as a necessary test, and shows that absolute certification thresholds fail even the reference model itself. It also proves an identifiability limit: an oracle-free forget threshold exists only for genuinely counterfactual facts, with the TOFU benchmark as the predicted boundary case.","feed_headline":"Standard unlearning test can certify models that still remember","feed_subtitle":"Candidates it rates adequate score held-out forget facts 2.82 nats below never-learned level—so benchmarks that rely on it inherit the flaw.","key_machinery":"The carrying instrument is a base-anchored held-out rank screen: for a candidate model, compute base-anchored deltas (the candidate's NLL minus the base model's NLL) on the forget facts and on never-learned probe facts, then run a one-sided exact Mann–Whitney rank test to exclude candidates whose forget deltas fall significantly below their own probe deltas—a sign of residual knowledge. This is paired with a round-trip residual: reacquire the forget set into the candidate and measure the KL divergence back to the injected model, used as a ranking signal. An identifiability theorem, built on a counterfactual-exchangeability assumption that the oracle's forget threshold equals a known function","core_discovery":"The central claim is that matching the retrained model on trained probes is not evidence of restoration. Candidates that the trained-probe oracle-KL criterion rates as adequate score held-out forget facts a mean of -2.82 nats (cluster CI [-3.16, -2.48]) below the never-learned level, meaning they retain paraphrase-recoverable knowledge of the forget set while matching the reference on the phrasings the criterion happens to query. Consequently, benchmarks that treat this criterion as ground truth are systematically misleading. The paper's constructive alternative is to evaluate unlearning as restoration to a matched retraining reference, using base-anchored held-out deltas and a round-trip re","pith_inferences":["The roughly 2.8-nat gap implies that the field's default metric measures phrasing-level matching rather than knowledge deletion; a direct testable extension is that unlearning methods optimized on trained probes will systematically fail on paraphrase-based adversarial probes, so rephrasing attacks should become a standard stress test.","The identifiability boundary suggests a practical deployment rule: estimate the counterfactual headroom (the paper's G-hat, which tracks true headroom at rank correlation 0.97 on a graded axis) before trusting any oracle-free selector; for inferable facts, abstain or demand the reference.","The round-trip residual—reacquire and compare—could transfer to other deletion settings beyond language models, such as image or tabular unlearning, wherever a re-acquisition protocol is feasible, and could serve as a general test of whether a method truly deleted versus merely hid.","The drift-seam detection property implies that routing-based suppression leaves a detectable distributional trace; a sophisticated adversary could plausibly attenuate that seam by also adjusting probe-side responses, which points to adaptive challenge panels as the next test."],"forward_implications":["Benchmarks that use trained-probe oracle-KL as the ground-truth unlearning metric inherit a systematic flaw: they may certify methods that merely suppress phrasing-level evidence while the knowledge survives rephrasing.","Absolute retain and round-trip thresholds should be replaced by reference-relative bars set by the retraining replication noise; otherwise even the retraining reference itself fails them in nearly every configuration.","Base-anchored held-out screens are valuable as necessary tests—they reject under-forgetting with high sensitivity—but they are not sufficiency certificates and must be paired with other axes.","Oracle-free selection is only sound for counterfactual facts; for inferable facts, where retained data can reconstruct the forgotten knowledge, the retraining reference is unavoidable.","Forward-only certification is not adversarially sound: a fixed-magnitude logit-suppression attack defeats the full forward battery in a measurable fraction of cells, so relearning axes or white-box access are required for stronger guarantees."],"fun_headline_variants":["Unlearning test can certify models that still remember","Trained-probe criterion misses retained forget facts","Unlearning benchmark flaw: adequate models leak data","Common unlearning test gives false certification"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole oracle-free framework rests on the counterfactual-exchangeability assumption that the retrained oracle's forget threshold can be read off the injected model's distribution of never-learned probe scores; if retained data can reconstruct the forgotten fact, this assumption fails and no oracle-free selector can be consistent.","fun_headline_variants_meta":{"raw":{"variants":["Unlearning test can certify models that still remember","Trained-probe criterion misses retained forget facts","Unlearning benchmark flaw: adequate models leak data","Common unlearning test gives false certification"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00025,"raw_usage":{"total_tokens":1476,"prompt_tokens":914,"completion_tokens":562,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":505}},"tokens_in":658,"tokens_out":562,"duration_ms":6113,"temperature":1.0,"reasoning_tokens":505,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T14:05:33.286888+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In any cell of the paper's testbed, find a candidate that the trained-probe oracle-KL criterion rates adequate but whose held-out forget-NLL gap from the never-learned level is within the replication tolerance (≤0.93 nats) of zero; the paper predicts a gap of roughly -2.82 nats, so such a candidate would directly contradict its central negative result. Alternatively, show that the matched retraining reference itself fails the paper's base-anchored screen in a non-trivial share of cells, which would undermine the screen's validity as a necessary test.","supporting_citations":[],"review_version":1}