{"id":"2fa6dbc2-899b-46ab-a5d6-39e72d5d0e2b","arxiv_id":"2506.00848","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Existing machine unlearning methods perform poorly on speech tasks, and a SuperLoss-based structured forgetting strategy improves forgetting but is reported with almost no experimental detail.","lead":"This paper introduces speech unlearning, applying known removal methods to voice models for the first time. Tests on keyword spotting and speaker identification show that current methods either fail to forget target data or severely degrade model accuracy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract claims speech unlearning is harder than image/text, but no image/text baseline is run; the headline comparative conclusion is not supported by the reported experiments.","rationale":"The reader correctly identified the MIA metric's ambiguous direction as a weakness, but the single most load-bearing concern about the paper's central claim is the unsupported cross-modal comparison. The abstract and §1/§5.3 assert that speech unlearning is significantly more challenging than image/text unlearning, yet the experiments contain no image/text baselines and no quantitative comparison to existing results from those modalities. The absolute accuracy reductions on two speech tasks do not by themselves demonstrate greater difficulty; they could reflect dataset or model properties rather than the speech modality per se. This is an external-validity gap, not an internal inconsistency, but it is central to the paper's contribution. The MIA issue is important for the SuperLoss claim and for the 'samples remain detectable' statement, but the core empirical observation—that none of the tested methods both forgets Df and preserves Dr—is supported by the accuracy columns regardless of MIA direction. Therefore, the reader's conditional verdict remains appropriate: the paper should be revised to either add matched cross-modal experiments or temper the cross-modal claim. My read does not change the verdict, so I recommend UNCHANGED.","tokens_in":8423,"tokens_out":3797,"duration_ms":39627,"concrete_test":"Run the same five unlearning methods (GradAscent, RandLabel, Bad-T, SCRUB, SalUn) on a vision benchmark (e.g., CIFAR-10) and a text benchmark (e.g., a sentiment classification subset such as IMDB) using comparable backbone scales, the same forget-set fractions (1–10% in 1% increments), and the exact same metrics (Dt, Df, Dr, and a clearly defined MIA direction). If the average Dr retention versus Df accuracy trade-off on speech falls within the noise of the image/text trade-offs, the 'significantly more challenging' claim is refuted; if speech shows strictly larger degradation across the same methods, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in the abstract and reiterated in §5.3, is that unlearning speech data is 'significantly more challenging than unlearning image or text data.' This is a comparative, cross-modal claim, yet the experiments in §5 only evaluate unlearning on two speech tasks (keyword spotting and speaker identification). There is no image unlearning baseline, no text unlearning baseline, and no quantitative comparison to published image/text unlearning results under matched conditions. The reported evidence—accuracy drops on Df and Dr in Tables 1 and 2—shows that the tested methods struggle on these speech tasks, but it does not establish that the difficulty is greater than in other modalities. Without a matched comparison, the 'significantly more challenging' conclusion rests on an implicit assumption that existing results from vision/NLP unlearning are directly comparable despite different datasets, model architectures, training setups, and evaluation protocols. This assumption is not stated or defended, and the correctness risk is high: if a matched comparison showed similar trade-offs, the headline claim would collapse. This is the most load-bearing concern because the paper's novelty and contribution are framed around this comparative difficulty; the MIA-direction issue flagged by the reader is real but secondary, since the accuracy tables independently support the weaker claim that none of the tested methods unlearns cleanly on these speech tasks.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces machine unlearning for speech, defining sample unlearning (removing individual recordings) and class unlearning (removing a speaker or keyword class). It casts unlearning as an optimization problem in Eq. (1) with forgetting and retention losses, and it evaluates five existing methods (Gradient Ascent, Random Labeling, SalUn, SCRUB, Bad-T) on keyword spotting (Speech Commands) and speaker identification (VoxCeleb1), reporting accuracy on the forget, retain, and test sets, a membership inference attack (MIA) score, and runtime. The authors report that none of the tested methods successfully unlearns speech data without degrading retained accuracy, and they claim in the abstract and in §5.3 that speech unlearning is 'significantly more challenging than unlearning image or text data.' They also propose a SuperLoss-based structured forgetting strategy that they report reduces a forget-set score from 44.7 to 17.3, and they outline future directions.","tokens_in":8655,"tokens_out":4662,"duration_ms":46100,"significance":"If the empirical findings are properly supported, this would be a useful first systematic study of unlearning for speech, with clear task definitions, two complementary speech tasks, several baseline methods, and a time-constrained training protocol. The authors are honest about the limitations of current methods, and the observation that label-randomization and saliency methods retain substantial knowledge of the forget set while gradient-ascent methods destroy retain-set utility is a concrete, falsifiable finding. However, the paper's headline comparative claim (speech is harder than image/text), the interpretation of the MIA metric, and the reported SuperLoss improvement are not currently supported by the evidence as presented. The paper's contribution is better framed as a benchmark and difficulty analysis of existing unlearning methods on speech, rather than a demonstration that speech is inherently harder than other modalities.","major_comments":[{"comment":"The abstract states that 'unlearning speech data is significantly more challenging than unlearning image or text data,' but no image or text unlearning baseline is run in §5, and no matched comparison to published vision/NLP unlearning results is provided. This is a load-bearing comparative claim about cross-modal difficulty, and without controlled experiments or a clearly justified comparison protocol it is not supported. The authors should either remove the cross-modal claim or add matched image/text unlearning experiments under the same evaluation framework.","section":"Abstract and §5.3"},{"comment":"The MIA metric is internally inconsistent as reported. In §4.4 the text says that if an MIA model classifies most Df samples as non-members, this indicates the unlearned model has limited knowledge of Df, which suggests the reported MIA(↑) is the rate of correct non-member classification. Under that reading, the original model's low MIA of 17.3 is sensible (it has trained on Df and therefore classifies those samples as members), but the unlearned models' MIA values around 50 percent indicate chance-level membership discrimination, which would mean the samples are not clearly detectable. In contrast, §5.3 interprets higher MIA as 'unlearned samples remain detectable,' which would imply that the original model is the least detectable, an implausible outcome. The direction of the metric and the threshold for 'detectable' must be defined precisely, and the conclusion in §5.3 must be re-derived from the chosen definition.","section":"§4.4 and Tables 1-2"},{"comment":"The SuperLoss result is not verifiable from the manuscript. The text says 'the overall unlearning performance on Df effectively decreased from 44.7 to 17.3,' but no SuperLoss row appears in Tables 1 or 2, and the baseline value 44.7 coincides with the MIA value for RandomLabel class unlearning in Table 1 while 17.3 is the original model's MIA in the same table. The authors do not state which metric (Df accuracy, MIA, or another score) is being reported, which method SuperLoss is applied to, or under which setup. Since this is presented as a promising solution, full experimental details and a table row or figure are required.","section":"§5.4"},{"comment":"The reported results are aggregated averages across several backbone architectures (Whisper, Wav2Vec 2.0, HuBERT) and across forget-set fractions from 1% to 10%, but the tables contain only point estimates with no standard deviations, no per-architecture breakdown, and no description of how the averaging is performed. The claim that all methods fail is probably robust, but specific quantitative comparisons (for example, Df values of 12.3 vs. 11.2 for GradAscent and SCRUB) cannot be evaluated without variance information. The authors state that all experiments were repeated five times, so these standard deviations should be reported.","section":"§5.2 and Tables 1-2"}],"minor_comments":[{"comment":"The figure caption contains a typo: 'Utterence' should be 'Utterance.'","section":"Figure 1"},{"comment":"There is a typo in the first paragraph: 'suggeting' should be 'suggesting.'","section":"§5.3"},{"comment":"The conclusion contains a duplicated article: 'the the high-dimensional, sequential, and speaker-dependent nature' should read 'the high-dimensional, sequential, and speaker-dependent nature.'","section":"§7"},{"comment":"The dataset name is typeset as 'V oxCeleb1' with an extra space; it should be 'VoxCeleb1.'","section":"§5.1"},{"comment":"The phrase 'If a MIA model classifies most samples of Df as non-members' should specify that the subject is a membership-inference classifier, not a model, to avoid confusion with the speech model being evaluated.","section":"§4.4"},{"comment":"The authors do not state the random-chance accuracy levels for the forget set. For speaker identification with 1,211 speakers, an accuracy of 33.4 is far above chance, while for keyword spotting with 12 classes the interpretation is different; stating chance levels would help readers interpret Df values.","section":"§4.2 and Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is very short relative to the scope of its claims. The cross-modal difficulty claim and the SuperLoss result are both presented as headline contributions, but neither is backed by the reported experiments; both are fixable, but they require either new experiments or a substantial reframing of the paper's contribution. The high number of self-citations in the related work is noticeable but not itself a reason to reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what you should know: this is the first systematic benchmark of machine unlearning on speech tasks, and that part is real. The paper defines sample and class unlearning for keyword spotting and speaker identification, runs five standard methods across several speech backbones, and reports the predictable failure mode: either the forget set isn't forgotten (RandLabel, SalUn) or retained accuracy collapses (GradAscent, Bad-T, SCRUB). That's a useful service to the unlearning community, and the time-constraint check (unlearning must beat full retraining) is a good practical touch. I'd trust the accuracy tables as evidence that off-the-shelf unlearning does not transfer cleanly to speech.\n\nThe soft spots, in order of severity. First, the abstract's claim that speech unlearning is 'significantly more challenging than unlearning image or text data' is not supported by anything in the paper. There is no image or text baseline, no matched comparison, no quantitative reference to published numbers. The experiments show these methods struggle on these two speech tasks; they don't show speech is harder than vision or NLP. This is the load-bearing framing sentence, and it rests on an assumption the authors never defend. The stress-test note is right about this. Second, the MIA metric as reported is confusing: the original model has MIA 17.3, the text says higher MIA means samples remain detectable, which would imply the original model has low detectability—odd for a model that trained on those samples. Either the direction is reversed or the metric is something else. The accuracy tables independently support the weaker conclusion that unlearning is incomplete, so this is fixable by clarification, not fatal. Third, the one positive result—SuperLoss dropping Df from 44.7 to 17.3—is under-specified. We aren't told which task, which method it was combined with, which hyperparameters, or which rows in Tables 1 and 2 it corresponds to. As written it's an unreviewable anecdote.\n\nThe citation pattern is fine; self-citations are contextual. The paper is honest in its limitations section, which counts for something. I'd send it to review, but with a flag on the cross-modal claim and a request for the SuperLoss details. The MIA clarification is mandatory but mechanical.\n\nWho it's for: people working on machine unlearning as an application area, and anyone thinking about privacy for voice AI. It's not a methods paper; it's a benchmark-with-a-warning. Worth a serious referee, conditional on the fixes.","headline":"First speech-unlearning benchmark, but the headline cross-modal claim outruns the experiments.","tokens_in":9169,"tokens_out":1628,"would_cite":false,"duration_ms":14315,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper defines speech unlearning — erasing one recording or an entire speaker from a trained model without retraining — and reports that all five standard methods fail: each either leaves the targeted data recognizable or destroys…","keywords":["speech unlearning","machine unlearning","keyword spotting","speaker identification","membership inference attack","right to be forgotten","curriculum learning","sample reweighting"],"falsifier":"Re-measure membership inference as a plain attack-success rate: the fraction of forget-set samples that an attack model labels as training members, reported for the original model and each unlearned model. The original model trained on the forget set, so it should show the highest success rate; whichever quantity in Tables 1 and 2 lines up with that fact settles the metric's direction, and with it whether the SuperLoss reduction from 44.7 to 17.3 is an improvement or a return to original-level memorization.","tokens_in":8206,"feed_emoji":"🎙️","tokens_out":20418,"duration_ms":167464,"temperature":0.7,"pith_summary":"Speech unlearning is the task of removing a chosen voice recording (sample unlearning) or an entire speaker or keyword (class unlearning) from an already trained speech model without retraining from scratch. The paper argues this matters for privacy law compliance, for deleting outdated or sensitive content such as wake words, and for correcting biased or noisy training data. Its central empirical claim is negative: across keyword spotting and speaker identification, every one of five standard unlearning methods adopted from vision and text research either leaves the targeted data recognizable or destroys accuracy on retained and test data. The paper's positive result is that a structured forgetting objective based on curriculum-style sample reweighting (SuperLoss) cuts the forget-set score from 44.7 to 17.3, pointing to a research direction rather than a finished solution.","feed_headline":"Five unlearning methods all fail to erase one voice","feed_subtitle":"Each approach either leaves the erased recording detectable or wrecks accuracy on the data that must be kept.","key_machinery":"The central machinery is a unified unlearning objective, $f' = \\arg\\min_{\\theta'} \\mathcal{L}_f(\\mathcal{D}_f) + \\lambda \\mathcal{L}_r(\\mathcal{D}_r)$, which casts speech unlearning as fine-tuning the original model to minimize a forgetting loss on the forget set and a retention loss on the retain set; each tested method is one choice of these two terms, and $\\lambda$ sets the trade-off. The two tasks are sample unlearning (erase one recording) and class unlearning (erase a speaker or keyword), evaluated by accuracy on the forget, retain, and original test sets plus a membership inference attack on the forget set. The paper's own promising result runs through SuperLoss, $\\mathcal{L}_\\lambda = (l_i - \\tau)\\sigma_i + \\lambda(\\log\\sigma_i)^2$, where $\\tau$ is a moving-average loss threshold and each sample weight $\\sigma_i$ is given in closed form; this turns flat fine-tuning into a curriculum that reweights hard-to-forget samples. Experiments average across seven speech backbones — Whisper Tiny and Base, Wav2Vec 2.0 Base and Large, HuBERT Base, Large, and X-Large — on the Speech Commands and VoxCeleb1 datasets.","core_discovery":"The paper claims that speech unlearning is genuinely harder than unlearning in vision or text, and that off-the-shelf unlearning methods fail on it. On keyword spotting and speaker identification, five methods — gradient ascent, random labeling, saliency-based unlearning (SalUn), SCRUB, and bad-teacher training — are tested across Whisper, Wav2Vec 2.0, and HuBERT backbones. The failure pattern is consistent: gradient ascent, SCRUB, and bad-teacher forget the target data but collapse accuracy on everything else, while random labeling and SalUn preserve accuracy but leave forget-set samples recognizable, as judged by subset accuracy and by a black-box membership inference attack. Class unlearning, removing an entire keyword or speaker, proves harder than sample unlearning. The paper also reports that a structured forgetting objective using SuperLoss — a curriculum-style loss that reweights each sample by a closed-form function of its own loss — improves the forget-set score from 44.7 to 17.3, and frames speech unlearning as an open area with feature-level, certifiable, and adversarial directions to explore.","pith_inferences":["My inference: the membership-inference numbers in the paper are internally ambiguous — the original model that trained on the forget set has the lowest score in both tables, while every unlearned model scores higher, and the metric's definition in section 4.4 and its use in section 5.3 point in opposite directions; the metric direction needs one explicit statement to decide which reading is right.","A testable extension the paper does not run: order the forget set by per-sample loss and vary the SuperLoss threshold schedule, to see whether class unlearning improves monotonically with curriculum strength; the closed-form weights make this a minor change to the reported setup.","If the entanglement claim is right, feature-level unlearning — removing pitch, timbre, or accent while keeping linguistic content — is not a niche add-on but the core of the problem: successful speech unlearning must first separate speaker identity from phonetic content in the representation space.","A natural next benchmark is a relearning attack: after supposedly forgetting a speaker, fine-tune the unlearned model briefly and measure how many steps re-identify that speaker; the paper cites such attacks but does not run them, and they would quantify exactly how much residual knowledge each failure mode leaves."],"forward_implications":["Unlearning methods built for images and text do not transfer to speech; the uniform failure pattern in Tables 1 and 2 means the forgetting and retention objectives have to be redesigned for temporal, speaker-entangled representations.","Class unlearning is the harder task: in every method, erasing an entire keyword or speaker drags down recognition of the remaining classes more than sample unlearning does.","Forgetting-first methods (gradient ascent, SCRUB, bad teacher) and retention-first methods (random labeling, SalUn) fail in opposite, predictable ways, so a working method needs both mechanisms rather than either extreme.","The SuperLoss experiment, cutting the forget-set score from 44.7 to 17.3, indicates that how samples are weighted and scheduled during unlearning matters as much as the gradient update itself.","If speech unlearning works, concrete applications open up: retiring wake words, removing a deactivated user's voice from authentication systems, deleting outdated commands, and de-identifying speakers in medical or legal speech records."],"supporting_citations":[{"why":"Defines the gradient-ascent and random-labeling unlearning methods that the paper instantiates on speech tasks.","marker":"[2]"},{"why":"Supplies the SCRUB baseline that the paper tests and finds sacrifices retained and test accuracy for forgetting.","marker":"[29]"},{"why":"Supplies the SalUn saliency baseline, which retains accuracy but leaves forget-set knowledge in the model.","marker":"[30]"},{"why":"Supplies the bad-teacher (Bad-T) baseline evaluated for speech unlearning.","marker":"[31]"},{"why":"Defines the keyword-spotting and speaker-identification task setups the experiments build on.","marker":"[32]"},{"why":"Provides the Speech Commands dataset used for keyword-spotting unlearning.","marker":"[33]"},{"why":"Provides the VoxCeleb1 dataset used for speaker-identification unlearning.","marker":"[34]"},{"why":"Provides the SuperLoss objective whose closed-form sample weights drive the structured-forgetting result (44.7 to 17.3).","marker":"[38]"}],"fun_headline_variants":["Five unlearning methods fail to erase one voice recording","Speech unlearning: harder than vision, five methods flop","Erasing a voice is harder than erasing a face from AI","Speech models resist unlearning: five attacks all fail","Removing a speaker from a speech model remains unsolved"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions depend on one reading of the membership-inference score (a test of whether the model still recognizes supposedly erased recordings): the paper treats a higher score as meaning the erased data stays detectable, yet the original trained model reports the lowest score in both tables, so the direction of this metric is the load-bearing assumption.","fun_headline_variants_meta":{"raw":{"variants":["Five unlearning methods fail to erase one voice recording","Speech unlearning: harder than vision, five methods flop","Erasing a voice is harder than erasing a face from AI","Speech models resist unlearning: five attacks all fail","Removing a speaker from a speech model remains unsolved"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000388,"raw_usage":{"total_tokens":2043,"prompt_tokens":936,"completion_tokens":1107,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":1025}},"tokens_in":552,"tokens_out":1107,"duration_ms":11504,"temperature":1.0,"reasoning_tokens":1025,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:56:03.762267+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-measure membership inference as a plain attack-success rate: the fraction of forget-set samples that an attack model labels as training members, reported for the original model and each unlearned model. The original model trained on the forget set, so it should show the highest success rate; whichever quantity in Tables 1 and 2 lines up with that fact settles the metric's direction, and with it whether the SuperLoss reduction from 44.7 to 17.3 is an improvement or a return to original-level memorization.","supporting_citations":[{"cited_title":"Ma- chine unlearning of federated clusters,","cited_arxiv_id":null,"evidence_quote":"Supplies the SalUn saliency baseline, which retains accuracy but leaves forget-set knowledge in the model."},{"cited_title":"Continual learning and private unlearning,","cited_arxiv_id":null,"evidence_quote":"Supplies the bad-teacher (Bad-T) baseline evaluated for speech unlearning."},{"cited_title":"Robust speech recognition via large-scale weak supervision,","cited_arxiv_id":null,"evidence_quote":"Provides the Speech Commands dataset used for keyword-spotting unlearning."},{"cited_title":"Cognivoice: Multi- modal and multilingual fusion networks for mild cognitive impair- ment assessment from spontaneous speech,","cited_arxiv_id":null,"evidence_quote":"Provides the VoxCeleb1 dataset used for speaker-identification unlearning."},{"cited_title":"Can bad teaching induce forgetting? unlearning in deep networks using an incompetent teacher,","cited_arxiv_id":null,"evidence_quote":"Provides the SuperLoss objective whose closed-form sample weights drive the structured-forgetting result (44.7 to 17.3)."}],"review_version":1}