{"id":"ee4ac09c-f6f7-4df6-92df-30ed624009b0","arxiv_id":"2606.18473","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Collateral damage from LLM unlearning decays with semantic distance but persists across domains and can be predicted pre-unlearning from forget-evaluation set interaction features.","lead":"This paper measures how unlearning specific data in LLMs damages related and distant knowledge, finding a consistent decay pattern with semantic distance. It proposes auditing the forget set beforehand using data interaction features to predict that damage.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Interaction features may have been selected by correlating with post-unlearning damage on the same data, undermining the pre-unlearning prediction claim.","rationale":"The reader's weakest assumption directly matches the identified gap. Because the review was abstract-only, the concrete_test supplies the minimal experiment that would resolve whether the reported signals are usable prospectively. No other internal inconsistency is visible from the given material.","tokens_in":1691,"tokens_out":306,"duration_ms":25089,"concrete_test":"Partition the reported forget/evaluation pairs into disjoint train and test groups; fit the interaction-feature predictor only on the train group using the observed damage values from that group; evaluate its ability to rank or regress damage on the held-out test group after performing the corresponding unlearning runs; if Spearman correlation or AUC drops below 0.6, the pre-auditing utility is not supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that interaction features between forget and evaluation sets are the strongest signals for predicting collateral damage from static data geometry before any model update. For this to hold, the analysis must demonstrate that these features can be computed and used to forecast damage on unseen forget/evaluation pairs without access to the unlearning outcome. If feature importance or regression is instead fit on the full set of observed damage values (i.e., after running unlearning to obtain the labels), the reported predictive strength is retrospective rather than prospective and does not establish an early-warning capability.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper studies collateral damage during LLM unlearning from a data-centric perspective. It reports a consistent decay pattern where damage is strongest near the forget set, weakens with semantic distance, and persists across domain boundaries. It then frames forget-set auditing as a pre-unlearning prediction task and claims that interaction features between the forget set and evaluation sets are the strongest predictors of downstream damage, implying that such damage is partly encoded in static data geometry before any model update occurs.","tokens_in":1782,"tokens_out":490,"duration_ms":22110,"significance":"If the predictive relationship can be shown to hold prospectively on unseen forget/evaluation pairs without access to post-unlearning outcomes, the work would offer a practical early-warning tool for unlearning safety and could inform more robust unlearning design. The data-geometry angle is a useful complement to existing model-update analyses.","major_comments":[{"comment":"The central claim that interaction features enable pre-unlearning prediction of collateral damage is load-bearing for the contribution. The manuscript must demonstrate that feature selection, regression fitting, or importance ranking was performed without using the observed post-unlearning damage values as supervision; otherwise the reported predictive strength is retrospective rather than prospective. This directly affects whether the results support an 'early warning' capability (see abstract and the forget-set auditing formulation).","section":"Experimental setup / prediction task formulation"},{"comment":"The decay pattern and feature-predictiveness results require explicit held-out evaluation: the interaction features must be tested on forget/evaluation pairs that were never used to compute or select those features. Without such separation, the analysis risks circularity between the geometry used to define evaluation sets and the geometry used to predict damage on those sets.","section":"Results on predictive signals"}],"minor_comments":[{"comment":"The abstract states the main findings but provides no information on datasets, unlearning methods, evaluation metrics, or statistical controls; these details are essential for assessing reproducibility even in a data-centric study.","section":"Abstract"},{"comment":"Notation for 'interaction features' and 'semantic distance' should be defined precisely (e.g., via explicit formulas or embedding-based metrics) to allow independent replication.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed comments on the prospective nature of the prediction task. We address both major comments below and will revise the manuscript to provide the requested demonstrations of held-out, pre-unlearning prediction.","responses":[{"response":"We agree that a fully prospective formulation requires training the predictive model without access to post-unlearning damage on the evaluation pairs used for testing. The interaction features themselves are computed exclusively from static data geometry before any unlearning occurs. In the revision we will add an explicit cross-validation protocol: a regression model will be fit on a training partition of forget/evaluation pairs (using their post-unlearning damage only for supervision within that partition) and evaluated on completely disjoint held-out pairs whose damage values are never seen during feature selection or model fitting. This change will be documented in a new subsection on prospective auditing.","revision_made":"yes","referee_comment":"[Experimental setup / prediction task formulation] The central claim that interaction features enable pre-unlearning prediction of collateral damage is load-bearing for the contribution. The manuscript must demonstrate that feature selection, regression fitting, or importance ranking was performed without using the observed post-unlearning damage values as supervision; otherwise the reported predictive strength is retrospective rather than prospective. This directly affects whether the results support an 'early warning' capability (see abstract and the forget-set auditing formulation)."},{"response":"We acknowledge the risk of circularity when the same data geometry informs both set construction and prediction. The revised manuscript will partition all evaluation sets into disjoint training and test groups for the prediction experiments. Feature selection, importance ranking, and regression fitting will be performed only on the training partition; all reported predictive metrics (including R² and feature rankings) will be computed exclusively on the held-out test pairs. This separation will be added to the experimental setup and results sections.","revision_made":"yes","referee_comment":"[Results on predictive signals] The decay pattern and feature-predictiveness results require explicit held-out evaluation: the interaction features must be tested on forget/evaluation pairs that were never used to compute or select those features. Without such separation, the analysis risks circularity between the geometry used to define evaluation sets and the geometry used to predict damage on those sets."}],"tokens_in":1339,"tokens_out":485,"duration_ms":27353,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core idea is to audit a forget set before unlearning by checking how its data interacts with evaluation sets, using those static features to forecast how much collateral damage will occur. They also note a decay pattern where damage drops with semantic distance but lingers across domains.\n\nWhat stands out is the shift to treating auditing as a prediction problem from data geometry alone. That is a clean framing if the experiments actually test it on unseen pairs without peeking at the unlearning outcomes.\n\nThe main weakness is exactly the stress-test concern. If the interaction features were identified or the predictor trained by correlating against damage measured after unlearning runs, then the reported signals are retrospective, not an early warning. The abstract gives no indication of held-out validation or pre-specification of features, so the central claim does not yet hold. No datasets, models, metrics, or statistical tests are described either, which leaves the decay pattern and feature rankings impossible to assess.\n\nThis is narrow work aimed at the LLM unlearning community, particularly people handling privacy or safety deletions. A reader already deep in that literature might extract a useful hypothesis about data geometry, but the current version lacks the grounding to justify referee time. I would not bring it to a reading group or cite it until the prediction setup is shown to be prospective.","headline":"The paper frames unlearning collateral damage as a pre-execution prediction task from data interaction features, but the evidence for genuine prospective prediction rather than post-hoc fitting is missing from the abstract.","tokens_in":2279,"tokens_out":347,"would_cite":false,"duration_ms":17812,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Interaction features between forget sets and evaluation sets can predict collateral damage from LLM unlearning before any model updates occur.","keywords":["machine unlearning","large language models","collateral damage","data geometry","forget set","semantic distance","auditing","LLM safety"],"falsifier":"Running actual unlearning on a range of forget sets and finding that the strength of interaction features shows no correlation with the measured collateral damage on evaluation sets after the updates.","tokens_in":2569,"feed_emoji":"📊","tokens_out":614,"duration_ms":22259,"temperature":0.7,"pith_summary":"The paper studies how unlearning specified knowledge in large language models affects both nearby and distant information in the model. It documents a decay pattern in which collateral damage is strongest close to the forget set, weakens as semantic distance grows, yet continues across domain boundaries. The authors then frame forget-set auditing as a prediction task that uses pre-unlearning data features to forecast this damage. They find that interaction features between the forget set and evaluation sets give the strongest predictive signals. This positions data geometry as an early indicator that could flag risky unlearning operations without first running them.","feed_headline":"Interaction features predict unlearning collateral damage early","feed_subtitle":"Static data signals between forget and evaluation sets forecast how much related knowledge will be harmed, allowing risk checks before any u","key_machinery":"Forget-set auditing as a pre-unlearning prediction task that extracts interaction features between the forget set and evaluation sets to forecast downstream collateral damage.","core_discovery":"We find a consistent decay pattern: collateral damage is strongest near the forget set, weakens with semantic distance, but does not disappear at domain boundaries. Interaction features between the forget set and evaluation set provide the strongest signals, suggesting that collateral damage is partly reflected in data geometry before model updates occur.","pith_inferences":["The same interaction-based auditing might be tested on other model-editing operations such as targeted fine-tuning or knowledge editing.","Incorporating these features directly into unlearning algorithms could be explored as a way to adjust the forget process dynamically.","Similar pre-auditing could apply to safety interventions that modify model behavior without full retraining."],"forward_implications":["Unlearning runs can be screened in advance by computing interaction features to identify those likely to cause high collateral damage.","Semantic distance from the forget set can be used to anticipate how far damage will propagate.","Evaluation sets can be chosen or adjusted based on their geometric interaction with the forget set to reduce unexpected effects.","Unlearning procedures can incorporate pre-checks on data geometry to improve overall reliability."],"fun_headline_variants":["Pre-unlearn checks flag collateral knowledge decay","Interaction signals predict unlearning damage early","Data geometry reflects forget set collateral effects","Forget-evaluation links audit unlearning risks","Semantic decay mapped via pre-unlearning features"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Collateral damage patterns observed after unlearning can be reliably predicted in advance from static data features alone without performing any model update or training.","fun_headline_variants_meta":{"raw":{"variants":["Pre-unlearn checks flag collateral knowledge decay","Interaction signals predict unlearning damage early","Data geometry reflects forget set collateral effects","Forget-evaluation links audit unlearning risks","Semantic decay mapped via pre-unlearning features"]},"model":"grok-4.3","cost_usd":0.004112,"raw_usage":{"total_tokens":1981,"prompt_tokens":619,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":41115500,"prompt_tokens_details":{"text_tokens":619,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1301,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":619,"tokens_out":61,"duration_ms":13426,"temperature":1.0,"reasoning_tokens":1301,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T00:18:35.975288+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running actual unlearning on a range of forget sets and finding that the strength of interaction features shows no correlation with the measured collateral damage on evaluation sets after the updates.","supporting_citations":[],"review_version":1}