{"id":"d86a7cbe-87c5-4474-837f-113ca09b0e9d","arxiv_id":"2608.11408","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Jacobian-lens audit predicts model-level relearning recovery in LLM unlearning but cannot pick which facts return and backfires when used as a training penalty.","lead":"This paper introduces J-Access, an audit that checks whether knowledge a model was trained to forget still lingers in its internal representations, and finds it can predict how easily that knowledge returns after retraining. It also shows that turning this audit into a training penalty makes models hide the signal without actually deleting the knowledge.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"J-Access is computed on a checkpoint-dependent silent-probe subset I(θ) and normalized on that same subset (Eqs.","rationale":"The paper is a well-scoped empirical study with a coherent argument and several credible controls: the full 398-checkpoint audit, the within-method Theil–Sen fits, the agreement with an independent activation-patching criterion (ρ=-0.72), and the preregistered item-level orientation all count as independent evidence. The central empirical finding that behavioral forgetting does not imply internal erasure is robust to my concern because it relies on raw access rates and gold comparisons that are not the main confounded statistic. The weaker link is the normalized, model-dependent J-Occ used for recovery prediction and optimization. The reader's CONDITIONAL verdict remains appropriate: the concern is substantive and testable, but it is not a demonstrated internal inconsistency, and the paper's own controls suggest the readout has at least some validity. I therefore recommend no change to the verdict, with the condition that the fixed-probe re-analysis and WD-Train probe-dropout statistics be reported. My agreement with the reader is partial because the reader also lists the Jacobian-lens transport assumption; I view the silent-probe selection as more load-bearing since it affects all three headline analyses and has a direct, cheap falsification.","tokens_in":12492,"tokens_out":7374,"duration_ms":72330,"concrete_test":"Recompute J-Occ(θ) on a fixed probe set—either the intersection of I(θ) across checkpoints in a method family, or a fixed random subset of probes that are silent under the original and gold anchors—and re-estimate Table 2 correlations and Table 4 WD-Train values on that fixed set. Also report |I(θ)|, its overlap with I(θ0), and how many probes drop out of I(θ) as λ increases from 0 to 10. If the model-level correlations and the 0.283→0.387 revival increase survive, the selection-bias concern is resolved; if they attenuate or reverse, the headline predictive and backfire claims need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing issue is the model-dependent conditioning set in Eqs. (4)–(6). I(θ) is defined as the probes on which the audited checkpoint does not emit a target token, and the J-Access normalization in Eq. (6) evaluates the audited model and both anchors on that same checkpoint-specific I(θ). Thus J-Occ(θ) is not measured on a fixed audit instrument; its numerator and denominator can move because probes enter or leave the silent set as unlearning changes behavior. Two failure modes follow. First, a model can lower J-Access by becoming behaviorally expressive on high-access probes, removing them from I(θ) while internal accessibility is unchanged; this is a plausible reading of Test 3, where stronger WD-Train penalties lower J-Access but increase post-attack revival, and the paper does not report whether pre-attack emission rates or |I(θ)| changed. Second, across the 398 checkpoints in Test 2, models with different silence patterns are scored on different probe sets, so a checkpoint with high recovery may be one whose silent subset happens to be easy for the original anchor (large denominator) or hard for the gold anchor (small numerator). The partial correlations controlling for FQ and MU do not control for subset composition, since I(θ) is a function of the model itself. The twin/decoy and UDS agreement controls establish that the raw readout is target-specific on a fixed subset, but they do not exercise the dynamic normalization used in the headline analyses. Because the same conditional subset enters the model-level recovery prediction and the optimization experiment, the central two findings—predictive validity and backfire—depend on this normalization being unbiased.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes J-Access, an inference-time audit that uses the Jacobian lens to decode mid-layer residual states into vocabulary space and measures whether target concept tokens appear in the top-k decoded tokens on probes where the model is behaviorally silent. The audit is normalized between the original pre-unlearning model and a retain-only gold model. The authors audit 398 public TOFU/OpenUnlearning checkpoints across eight unlearning methods, conduct cross-entity relearning attacks to define recovery outcomes, and train WD-Train variants that directly penalize J-Access. The central findings are that most unlearned checkpoints retain accessibility above the gold level; pre-attack J-Access correlates with model-level recovery (Spearman +0.35 overall, +0.45 within-method, +0.71 for the knowledge-level variant) but not with item-level recovery (AUROC near 0.5); and minimizing J-Access lowers the audit score but increases post-attack revival, indicating audit evasion rather than deletion.","tokens_in":12726,"tokens_out":7337,"duration_ms":59898,"significance":"If the results hold, J-Access would be a valuable model-level diagnostic for comparing unlearning checkpoints and for prospective risk monitoring: it is evaluated at scale on public models, uses Holm-corrected significance testing, held-out entities for recovery, partial correlations controlling for forget quality and model utility, and a convergent-validity check against the activation-patching UDS criterion. The paper's negative item-level result is important because it cautions against treating internal accessibility as a per-fact deletion certificate, and the WD-Train result is a clean demonstration that a proxy audit can be gamed. The main risks are that the model-dependent silent-probe normalization may compromise cross-checkpoint comparability and that the manuscript refers to appendices that are not included, so the specific methodology cannot currently be audited.","major_comments":[{"comment":"The silent-probe subset I(θ) is checkpoint-dependent, so J-Access is not measured on a fixed audit instrument. Eq. (4) defines I(θ) as the probes on which θ does not emit any concept token, and Eq. (6) evaluates the audited model and both anchors on that same I(θ). A checkpoint can therefore lower its J-Access either by reducing internal accessibility or by becoming behaviorally expressive on high-access probes and removing them from the audited set; conversely, across the 398 checkpoints, models with different silence patterns are scored on different probe subsets, so the anchor values in the numerator and denominator of Eq. (6) can differ due to subset composition alone. The partial correlations in Table 2 control for Forget Quality and Model Utility but not for I(θ) composition, which is a function of the model itself. Please report |I(θ)| and pre-attack emission rates per checkpoint, and re-run the Test 2 and Test 3 analyses on a fixed probe subset (for example, probes silent under all checkpoints) to show that the recovery correlations and the WD-Train backfire result are not driven by this selection effect.","section":"The J-Access Score, Eqs. (4)-(6)"},{"comment":"The method description is incomplete because the manuscript relies on appendices that are not part of the submitted text. Appendix A is said to specify the construction of concept token sets, workspace band selection, readout positions, and layer selection; Appendix B the robustness variants; Appendix C the predictor definitions; and Appendix D the WD-Train objective, hyperparameters, and seed-level results. None of these appendices are present. As a result, key audit parameters (top-k threshold k, workspace band B, readout positions P, concept-token filters, calibration entities) and the 'lens refit after unlearning' control described in Test 1 cannot be checked, and the preregistration status claimed in Table 3 is unverifiable. Please include the appendices or a code release with the full configuration before the methods can be assessed.","section":"Experimental Setup and Appendix References"},{"comment":"The central Test 3 claim is not supported with uncertainty quantification. Table 4 reports single values for WD-Train λ = 0, 5, 10 even though the text states three random seeds per setting; no standard errors, confidence intervals, or significance tests are given, and Fig. 4 shows only aggregated lines. The differences that carry the conclusion (J-Access 0.67 to 0.55, revival 0.283 to 0.387, UDS roughly flat) could plausibly be within seed noise. Please report the seed-level results and test the monotone trend, and also state whether the J-Access decrease is accompanied by changes in behavioral emission rate or |I(θ)|, which is the selection channel identified in the first major comment.","section":"Test 3, Table 4 and Fig. 4"}],"minor_comments":[{"comment":"The notation JOcc(θ) introduced in Eq. (6) is not used elsewhere; the paper consistently writes J-Access for the normalized score. Please define the relationship explicitly to avoid confusion.","section":"Notation, Eq. (6)"},{"comment":"The sentence 'This observe admits a natural interpretation' contains a typo and should read 'This observation admits a natural interpretation'.","section":"Results, Test 2"},{"comment":"The 'J-Access(preregistered)' row is not accompanied by any preregistration specification; please clarify what was fixed in advance and where this is documented.","section":"Table 3"},{"comment":"The y-axis label 'post-attack revial' should be 'post-attack revival'.","section":"Fig. 4"},{"comment":"Please state explicitly whether the Jacobian lens used in Tests 2 and 3 is the original corpus-averaged lens or the lens refitted after unlearning; the refit control is mentioned only for Test 1.","section":"Test 1 versus Tests 2 and 3"},{"comment":"Table 4 reports Model Utility as n/a for RMU; please explain why RMU utility is not reported.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the empirical resource is substantial, but I would not recommend acceptance before the missing appendices and the model-dependent silent-probe normalization are addressed. The latter is load-bearing for both the recovery correlations and the optimization backfire claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this paper earns its central claim: it shows on 398 public checkpoints that most unlearned models keep internal access to target knowledge above the gold level, and that pre-attack J-Access correlates with model-level recovery (Spearman +0.35, partial +0.33, within-method +0.45) while failing at item level (AUROC around 0.55). Second, the optimization-backfire result is the most novel and the least supported: Table 4 has no error bars or significance tests, and the lambda trend is the headline that needs the appendix most. The negative result is plausible, but as reported we only see three means.\n\nWhat is genuinely new: prior UDS and hidden-state audits are same-checkpoint diagnostics. This paper adds prospective validation—does residual accessibility forecast relearning?—and tests whether the audit survives optimization. The WD-Train finding, that minimizing the score lowers it but raises revival, is a useful cautionary result for anyone tempted to use internal audits as training objectives. The comparison with causal-criterion families (deep vs shallow GradDiff/RMU) is a good control, and the UDS agreement (rho=-0.72) gives convergent validity across two mechanistically distinct readouts.\n\nSoft spots, in order of seriousness. One: the normalization in Eq. (6) is computed on the checkpoint-dependent silent-probe subset I(theta). The stress-test note is right: probes can leave the silent set as the model changes, so the numerator and denominator are not a fixed instrument. The paper's twin/decoy and UDS controls do not exercise this dynamic normalization. This could bias both the recovery correlations and the WD-Train result. The fix is an appendix analysis conditioning on a fixed probe subset or reporting |I(theta)| and emission rates per checkpoint. Two: the knowledge-level variant's extended concept tokens overlap with answer content words used in ROUGE revival; the +0.71 correlation is the one number I would not trust until that is addressed. Three: no code or appendices shipped. Every appendix is referenced and absent, which is the main reason confidence stays moderate.\n\nI disagree with the reader on one thing: the Jacobian-lens linear-transport assumption is not the soft spot. The authors refit the lens after unlearning as a control, and prior work causally validates the lens. The dynamic normalization is the issue to press.\n\nWho this is for: unlearning evaluators and safety auditors. It deserves a serious referee, but the current version should not be accepted before the appendix, code, and Table 4 uncertainty are supplied.","headline":"J-Access is a genuinely useful empirical contribution—behavioral forgetting hides recoverable knowledge, and the audit predicts model-level recovery but fails as an optimization target—though the checkpoint-dependent normalization and missing artifacts leave the quantitative story unfinished.","tokens_in":13383,"tokens_out":2297,"would_cite":true,"duration_ms":21241,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"J-Access, an inference-time audit using the Jacobian lens, shows that residual internal accessibility predicts model-level relearning recovery but cannot certify individual facts, and that minimizing the audit score makes recovery worse.","keywords":["machine unlearning","LLM unlearning","Jacobian lens","residual knowledge accessibility","relearning attacks","behavioral forgetting","internal auditing","TOFU benchmark"],"falsifier":"Take a set of unlearned checkpoints whose pre-attack J-Access is at or below the retain-only gold level; if, after a fixed cross-entity relearning attack, a substantial fraction of them revive at rates well above the identically attacked gold model, then accessibility does not track recovery susceptibility. A more direct version is to rerun the recovery-prediction analysis with a single fixed probe set shared by all checkpoints instead of per-model silent-probe subsets; if the model-level correlation with revival disappears, the reported signal is an artifact of probe selection rather than of residual knowledge.","tokens_in":12239,"feed_emoji":"🔍","tokens_out":8119,"duration_ms":66934,"temperature":0.7,"pith_summary":"Machine-unlearned language models often stop saying forgotten facts while still holding those facts internally accessible. This paper introduces J-Access, an inference-time audit that decodes intermediate representations into the vocabulary through a Jacobian lens, and uses it to ask whether that leftover accessibility can predict future recovery when the model is fine-tuned again. Across 398 unlearned models and eight methods, the authors find that higher pre-attack J-Access predicts faster and more complete relearning at the level of the whole checkpoint, but cannot say which specific facts will come back. They also find that training a model to minimize J-Access lowers the audit score while increasing post-attack recovery, so the audit works as a diagnostic, not as an objective.","feed_headline":"Audit predicts LLM relearning risk, fails as a training target","feed_subtitle":"A Jacobian-lens audit finds residual access in 85% of unlearned models; optimizing it raises recovery risk.","key_machinery":"The central object is the Jacobian lens: a corpus-averaged Jacobian $J_\\ell = \\mathbb{E}_{x,p,p'\\ge p}[\\partial h_{T,p'}(x)/\\partial h_{\\ell,p}(x)]$ that transports residual-stream states from an intermediate layer $\\ell$ into the basis of a later target layer $T$, followed by the model's unembedding $z_{\\ell,p}(x)=W_U\\,\\mathrm{Norm}(J_\\ell h_{\\ell,p}(x))$ to produce vocabulary scores. J-Access then records, over a band of mid-to-late layers and readout positions, whether any token in the target concept set appears in the top $k$ decoded tokens, and normalizes the resulting access rate against the original model and a retain-only gold model. The lens matters because unlearning can shift representations and break the shared-basis assumption of a direct logit lens; the linear transport is what lets the audit read accessibility without relying on the model's expressed output. This machinery does the argument's work by converting residual traces into a score that can be correlated with recovery and, in the WD-Train variant, differentiated as a training penalty that augments the unlearning loss.","core_discovery":"On its own terms, the paper's central discovery is that residual internal accessibility, as measured by J-Access, is a checkpoint-level predictor of relearning vulnerability but not an item-level deletion certificate, and that it cannot be safely optimized. Most unlearned models (85%) retain target knowledge more accessible than a retain-only gold model, and the median normalized accessibility is 0.69, so behavioral forgetting does not imply internal erasure. Pre-attack J-Access correlates with excess revival at +0.35 across the pool, +0.45 within method families, and +0.71 with the knowledge-level variant, and with fewer steps to recover at -0.70. At item level, AUROC for predicting which facts revive stays near chance, and adding J-Access to behavioral and membership-inference predictors adds little. Directly minimizing J-Access via a suppression penalty lowers the score from 0.67 to 0.55 while post-attack revival rises from 0.283 to 0.387 and causal deletion depth stays flat, indicating the model hides knowledge from the audit rather than deleting it. The conclusion is that internal audits belong in unlearning evaluation as an independent diagnostic dimension, but must be validated against causal and recovery evidence before they are optimized.","pith_inferences":["If recovery is governed by a shared retrieval pathway rather than per-item traces, as the paper suggests from localization work, then monitoring J-Access on a small probe set may be enough for risk dashboards, while per-fact erasure guarantees may require causal interventions rather than audit scores.","The backfire result likely extends beyond J-Access to any differentiable white-box audit: any internal metric whose gradient can suppress the measured signal will face the same measure-optimization collapse, so new audits should be tested under optimization before being used as objectives.","A direct test of the mechanism would be to run relearning attacks on models whose J-Access is reduced by causal interventions, such as mechanistic localization edits, rather than by penalty training; the paper's account predicts these should lower both J-Access and revival, unlike WD-Train.","Because J-Access normalization conditions on each checkpoint's own behaviorally silent probe subset, cross-checkpoint comparisons could be confounded; fixing a common probe set or weighting probes by difficulty would test whether the recovery correlations hold outside the paper's protocol."],"forward_implications":["Behavioral benchmarks alone can rank unlearned checkpoints by what they output, not by how much erased knowledge remains internally reachable; J-Access supplies that missing dimension.","Checkpoint-level relearning risk can be monitored before an attack: models with higher pre-attack accessibility recover faster and more fully, even after controlling for forgetting quality and model utility and within every method family tested.","No item-level deletion certificate is achievable from this signal: facts that revive cannot be distinguished from facts that stay suppressed, so J-Access should not be treated as per-fact proof of erasure.","Optimizing the audit backfires: a model trained to lower J-Access scores lower on the audit but revives more after relearning, so audit scores should not be used as training rewards or penalties without causal validation.","Internal audits and behavioral metrics should be reported together; a checkpoint that passes both is more credible than one passing either alone."],"supporting_citations":[{"why":"Provides the Jacobian lens readout and its causal validation via steering and patching, which the paper repurposes as the audit mechanism.","marker":"Gurnee et al. 2026"},{"why":"Supplies the TOFU benchmark, the forget/retain split of fictitious authors, the gold retain-only model, and the behavioral metrics used throughout.","marker":"Maini et al. 2024"},{"why":"Releases the 398 unlearned model checkpoints spanning eight methods and the membership-inference baselines used in the item-level comparison.","marker":"Dorna et al. 2026"},{"why":"Defines UDS, the activation-patching measure used as the convergent-validity check and as the causal deletion-depth criterion distinguishing deep from shallow deletion.","marker":"Lee, Kim, and Jo 2026"},{"why":"Establishes the targeted relearning attack protocol on held-out entities that the paper uses to measure recovery.","marker":"Hu et al. 2024"},{"why":"Shows latent-space defenses can be evaded through activation obfuscation, the empirical precedent for the audit-evasion behavior WD-Train produces.","marker":"Bailey et al. 2025"},{"why":"Provides the logit lens baseline that J-Access is compared against, isolating the effect of Jacobian transport.","marker":"Belrose et al. 2023"},{"why":"Provides localization evidence that fine-tuning unlearning disables a shared retrieval pathway, used to interpret why item-level prediction fails.","marker":"Hong et al. 2024"}],"fun_headline_variants":["J-Access predicts relearning risk, not item deletion","Optimizing audit hides knowledge, raising revival risk","Audit flags unlearning risk but can't pin down facts","Measure, don't optimize: audit predicts relapse, not targets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the corpus-averaged Jacobian lens still transports intermediate residual states into the target-layer basis faithfully after unlearning, and that comparing models on their own behaviorally silent probe sets does not bias the J-Access normalization.","fun_headline_variants_meta":{"raw":{"variants":["J-Access predicts relearning risk, not item deletion","Optimizing audit hides knowledge, raising revival risk","Audit flags unlearning risk but can't pin down facts","Measure, don't optimize: audit predicts relapse, not targets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00067,"raw_usage":{"total_tokens":3124,"prompt_tokens":1083,"completion_tokens":2041,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":699,"completion_tokens_details":{"reasoning_tokens":1974}},"tokens_in":699,"tokens_out":2041,"duration_ms":14405,"temperature":1.0,"reasoning_tokens":1974,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:39.010740+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of unlearned checkpoints whose pre-attack J-Access is at or below the retain-only gold level; if, after a fixed cross-entity relearning attack, a substantial fraction of them revive at rates well above the identically attacked gold model, then accessibility does not track recovery susceptibility. A more direct version is to rerun the recovery-prediction analysis with a single fixed probe set shared by all checkpoints instead of per-model silent-probe subsets; if the model-level correlation with revival disappears, the reported signal is an artifact of probe selection rather than of residual knowledge.","supporting_citations":[],"review_version":1}