{"id":"aaf6e95e-2ac6-48c1-b4aa-501d8eaac350","arxiv_id":"2411.15113","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Post-training magnitude pruning of Stable Diffusion 2 preserves image quality at 38.5% sparsity and outperforms Wanda on the text encoder.","lead":"This paper tests what happens when you delete weights from Stable Diffusion 2, a text-to-image model, without retraining it. It finds that simple magnitude pruning keeps image quality at 38.5% total sparsity and beats a more advanced method, but too much pruning makes images collapse.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 38.5% minimal-loss claim rests on unreplicated FID/CLIP numbers whose own Table 6 shows pruned 33.5-36% configurations beating the unpruned baseline, so the 0.08 FID gap is likely inside run-to-run noise.","rationale":"The reader's weakest assumption is that single-run FID/CLIP values without seeds or intervals cannot resolve the small differences that carry the paper's central claim. My reading agrees and sharpens it with internal evidence: Table 6 already contains configurations at 33.5% and 36% sparsity whose FID is lower than the unpruned baseline (16.53 and 16.66 vs 18.07), and Table 10 shows the identical 38.5% configuration with a different pruning variant at FID 17.97, below baseline. A deterministic pruning method should not be expected to improve FID a priori; for one variant to beat baseline while the other does not is exactly what one expects from sampling noise or seed variation. This makes the headline 0.08 FID gap and the post hoc choice of 38.5% unverifiable from the reported tables. The correct disposition is therefore conditional pending replicated evaluation, not rejection: the qualitative findings (magnitude competitive or better than Wanda; text encoder collapses at high sparsity) are plausible and consistently reported, and the code is released, so the missing evidence is obtainable. I did not find internal inconsistency or circular derivation; the concern is about evidential strength, matching the reader's verdict. A secondary overclaim about SparseGPT (Section 6) is noted but is less load-bearing than the unreplicated quantitative claims. Verdict remains CONDITIONAL, so no adjustment is needed.","tokens_in":12194,"tokens_out":4781,"duration_ms":41681,"concrete_test":"Using the released code, rerun the Table 6 configurations (unpruned baseline; text/image sparsities 42.5/30, 45/32.5, 47.5/35, 50/37.5, 52.5/40) with at least 5 independent sampler seeds over the same 10,000 MSCOCO captions, and also rerun the text-encoder-only 60% and 62.5% magnitude configurations with the same seeds. Report mean and standard deviation (or 95% CI) for FID and CLIP. If the CI for the 38.5% configuration overlaps the baseline CI, or if the lower-sparsity 33.5%/36% configurations remain better than baseline in most seeds, the central minimal-loss claim and the optimal-configuration recommendation are unsubstantiated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 4.4, Tables 5-6) that Stable Diffusion 2 can be pruned to 38.5% sparsity with minimal quality loss depends on a single FID (18.15 vs 18.07) and CLIP score (0.311 vs 0.314) per configuration. No seeds, repeated runs, or confidence intervals are reported (Section 3.4). The same table undermines the assumption that these numbers resolve 0.08 FID: with magnitude pruning, 36% sparsity gives FID 16.66 and 33.5% gives 16.53, both better than the 18.07 unpruned baseline; with OWL (Table 10), the same 38.5% configuration gives FID 17.97, below baseline. Because diffusion sampling is stochastic and FID is estimated from 10,000 generated images, a pruned model should not be expected to beat the baseline deterministically. This pattern indicates either seed-dependent sampling noise greater than the claimed loss, or that the 'optimal' 38.5% configuration was chosen post hoc from a monotone curve without a stated selection rule (Section 4.4, second approach). If measurement noise exceeds roughly 0.1-0.5 FID, the 'minimal quality loss' and the specific 47.5%/35% recommendation are not established, and the sharp text-encoder threshold near 60-62.5% is also not quantified. The failure to test SparseGPT despite claiming it does not extend (Section 6) is a secondary overclaim, but the primary load-bearing weakness is the unreplicated evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a post-training pruning study of Stable Diffusion 2, comparing magnitude pruning, Wanda, SparseGPT (claimed but not tested), and OWL-based variants across text-encoder-only, U-Net-only, and full-model pruning. The authors report that magnitude pruning outperforms Wanda, that Stable Diffusion 2 can be pruned to 38.5% total sparsity (47.5% text encoder, 35% U-Net) with minimal quality loss, and that both components exhibit sharp performance drop-off thresholds. The central claims are evaluated with FID and CLIP Score on 10,000 MSCOCO captions, with qualitative image examples supporting the quantitative results.","tokens_in":12486,"tokens_out":2835,"duration_ms":27233,"significance":"If the findings hold, the paper would provide a practically useful result: for at least one text-to-image model, simple magnitude pruning is competitive or superior to activation-aware pruning, and substantial sparsity can be achieved with little degradation. The study is also valuable for separately examining text-encoder and U-Net pruning, and for releasing code. The paper has no fitted free parameters or circular derivation; the main methodological risk is the unreplicated evaluation, since the headline differences (e.g., 0.08 FID) are small relative to known stochastic variation in diffusion sampling and FID estimation. The sharp-threshold observation is interesting but similarly rests on single-run numbers.","major_comments":[{"comment":"The central quantitative claims are based on a single FID and a single CLIP Score per configuration, with no repeated runs, seeds, or confidence intervals. The headline result of 'minimal quality loss' at 38.5% sparsity rests on FID 18.15 vs. 18.07 and CLIP 0.311 vs. 0.314. Table 6 itself shows that magnitude-pruned configurations at 36% and 33.5% sparsity achieve FID 16.66 and 16.53, both below the unpruned baseline of 18.07; Table 10 shows the same 38.5% configuration with OWL at FID 17.97. This pattern strongly suggests that run-to-run or seed-dependent variation is comparable to or larger than the reported degradation. The authors should provide repeated evaluations with error bars, or at minimum a formal statistical test, before the 'minimal quality loss' and sharp-threshold claims can be considered established.","section":"§3.4, Tables 5 and 6"},{"comment":"The 'optimal' configuration of 47.5% text sparsity and 35% image sparsity is selected post hoc from a sweep without a stated selection criterion. The text says the authors 'recommend' this configuration based on qualitative evaluations and improvements as sparsity is reduced, but Table 6 shows that lower-sparsity configurations (36% and 33.5%) have better or equal FID and CLIP Score than the 38.5% configuration. If the goal is maximal sparsity with minimal loss, the paper needs a predefined acceptable-degradation threshold; if the goal is simply the best quality-sparsity trade-off, the 36% or 33.5% configurations appear superior. A formal selection rule is needed to justify calling 38.5% 'optimal.'","section":"§4.4, Second Approach and Table 6"},{"comment":"The paper states that 'both Wanda and SparseGPT do not effectively extend to these models,' but SparseGPT is never tested in the experiments. The experimental comparison is limited to magnitude pruning, Wanda, and their OWL variants. The unsupported claim about SparseGPT weakens the general conclusion that magnitude pruning outperforms 'more advanced techniques.' Either implement SparseGPT, or restrict the claim to the methods actually evaluated.","section":"§6, Future Work"},{"comment":"The U-Net-only pruning threshold used in the second full-model approach is identified as 50% sparsity, but the text says performance 'steadily decreases as sparsity exceeded 40%,' and Figure 11 states that performance 'greatly suffers beyond 40%.' The quantitative basis for choosing 50% as the threshold is not reported, and the inconsistency between 40% and 50% should be resolved. Since the second approach starts from this threshold, a clear and quantitative justification is needed.","section":"§4.2, Figures 10 and 11"}],"minor_comments":[{"comment":"The text says 'a lower CLIP Score indicates a higher correlation,' but the CLIP Score is a similarity measure where higher values indicate better text-image alignment; this is inconsistent with the tables, where better configurations have higher CLIP scores.","section":"§3.4"},{"comment":"The entries listed as '0' for FID and CLIP Score in presumably invalid configurations should be marked as 'N/A' rather than 0, since a FID of 0 would indicate perfect similarity and a CLIP Score of 0 is not meaningful.","section":"Appendix A, Tables 7 and 8"},{"comment":"References [32] and [33] appear to cite the same latent diffusion paper twice, once as an arXiv preprint and once as the CVPR version; this duplication should be consolidated.","section":"References"},{"comment":"The abstract mentions 'interoperability' in the context of future research; this appears to be a typo for 'interpretability.'","section":"Abstract"},{"comment":"The claim that Wanda 'consistently outperforms' magnitude at every sparsity level is clear from Figure 9, but the figure is not accompanied by the underlying numeric values; adding a table of text-encoder-only results would make the comparison reproducible.","section":"§4.1 and Figure 9"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and practical question, and the authors have been transparent about releasing code. However, the unreplicated evaluation makes the headline quantitative claims unverifiable as stated. The revision should focus on adding repeated runs or confidence intervals, formalizing the configuration-selection rule, and either testing SparseGPT or softening the claim. With these changes, the paper could become publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful empirical report, not a definitive compression recipe. The core signal—magnitude pruning beats Wanda on the CLIP text encoder, and both components have sharp collapse thresholds—is probably right. The headline \"38.5% with minimal quality loss\" is not established by the reported numbers, because every FID and CLIP score is a single run and the optimal configuration is chosen from a sweep after the fact.\n\nWhat is actually new: a systematic per-component pruning study of Stable Diffusion 2, comparing magnitude, Wanda, and OWL variants across text encoder, U-Net, and full-model sparsity allocations, with code released. The finding that magnitude beats Wanda on the text encoder is a genuinely useful surprise and it is consistently supported across figures and tables. The sharp text-encoder collapse near 60–62.5% sparsity is also striking and visually corroborated. That empirical map is worth having on record.\n\nThe main soft spot is evaluation methodology. Section 3.4 reports one FID and one CLIP score per configuration from 10,000 generated images, with no seeds, repeated runs, or confidence intervals. Table 6 itself undercuts the 0.08 FID gap at the recommended 38.5% sparsity: the 36% and 33.5% configurations report FID 16.66 and 16.53, below the unpruned baseline of 18.07. A pruned model beating its own baseline by 1.5 FID is not evidence of measurement precision; it is evidence of run-to-run noise or a favorable seed. Table 10 shows the same pattern with OWL. So the specific quantitative claims—minimal loss at 38.5%, the 47.5%/35% recommendation, and the precise threshold values—are not supported by the reported evidence. The qualitative sample plus the overall trend suggest something in that sparsity neighborhood is usable, but that is a weaker claim than the abstract makes.\n\nSecond, the \"optimal\" configuration is selected post hoc without a formal criterion. The second approach starts from the individual component thresholds and reduces both sparsities in 2.5% steps, stopping when visual inspection looks acceptable. A stated decision rule, such as smallest sparsity with FID within X of baseline, would fix this. Third, Section 6 says Wanda and SparseGPT do not extend to these models, but SparseGPT was never run; only magnitude, Wanda, and OWL appear in the experiments. That sentence should be deleted or backed with experiments. Minor issues: the CLIP Score direction is stated backwards in Section 3.4 (\"lower indicates higher correlation\") while the analysis treats lower as worse, and the \"no prior work\" framing needs a more careful sentence relative to the cited LD-Pruner work on latent diffusion pruning. The acknowledged CLIP-evaluation bias in Section 4.3 is handled honestly.\n\nWho this is for: people building compressed text-to-image pipelines or studying how pruning methods transfer across modalities. They will get a practical map of where SD2 breaks and a reproducible baseline comparison. The paper deserves a serious referee, but the review should ask for repeated seeded runs and a stated selection rule before the quantitative claims are accepted.","headline":"Useful empirical report on pruning Stable Diffusion 2, but the headline 38.5% claim rests on unreplicated FID/CLIP numbers and a post hoc optimum.","tokens_in":13049,"tokens_out":2545,"would_cite":false,"duration_ms":24565,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Simple magnitude pruning outperforms Wanda and SparseGPT for Stable Diffusion 2, which can be pruned to 38.5% sparsity with minimal quality loss.","keywords":["Computer Vision","Natural Language Processing","Model Compression","Pruning","Text-to-Image Generation","Stable Diffusion","Post-training Pruning","Magnitude Pruning"],"falsifier":"Re-run the recommended 38.5% sparsity configuration (text 47.5%, image 35%) and the unpruned baseline multiple times with different random seeds, computing FID and CLIP Score over the same 10,000 MSCOCO prompts; if the spread across runs exceeds roughly 0.1 FID, the minimal-quality-loss claim is not established. For the magnitude-wins claim, a decisive test is to hold constant the evaluation protocol and swap only the pruning method at 50% text-encoder sparsity, comparing the resulting FID distributions rather than point estimates.","tokens_in":11952,"feed_emoji":"✂️","tokens_out":8021,"duration_ms":64843,"temperature":0.7,"pith_summary":"This paper asks whether post-training pruning, a compression technique proven on large language models, transfers to text-to-image generation models like Stable Diffusion 2. It finds that the simple baseline, magnitude pruning, beats the more sophisticated Wanda and SparseGPT methods in this multimodal setting. The paper's main quantitative claim is that Stable Diffusion 2 can be pruned to 38.5% total sparsity with minimal quality loss, dropping FID from 18.07 to 18.15 and CLIP score from 0.314 to 0.311. It also reports sharp collapse thresholds in each component: the CLIP text encoder degrades suddenly beyond roughly 60-62.5% sparsity, while the U-Net generator declines gradually. If correct, this means a straightforward magnitude-based pruning recipe can make billion-parameter text-to-image models substantially smaller for edge deployment.","feed_headline":"Simple pruning cuts Stable Diffusion 2 to 38.5% with no quality loss","feed_subtitle":"Wanda and SparseGPT, top methods for language models, lose to baseline magnitude pruning on text-to-image models.","key_machinery":"The central objects are the two sub-models of Stable Diffusion 2: the CLIP text encoder and the U-Net diffusion generator, evaluated separately and together at sparsity levels stepping by 10% (and finer near thresholds). The load-bearing mechanism is magnitude pruning, which zeroes weights with the smallest absolute values, compared against Wanda, which weights magnitudes by activation norms, and OWL, which allocates layerwise sparsity based on outlier density. The argument is carried by the FID and CLIP Score metrics computed on 10,000 generated images, with the observed sharp thresholds used to set the optimal per-component sparsities.","core_discovery":"Stable Diffusion 2, a 1.2-billion-parameter text-to-image model with a 340M-parameter CLIP text encoder and an 860M-parameter U-Net, can be post-training pruned with magnitude pruning to 38.5% total sparsity (47.5% text, 35% image) with negligible degradation: FID 18.15 vs 18.07 and CLIP Score 0.311 vs 0.314 on 10,000 MSCOCO captions. Contrary to the established trend in large language model pruning, activation-aware Wanda pruning and outlier-weighted OWL are consistently worse than or equal to plain magnitude pruning in this context. The paper interprets this as evidence that the information encoding in text-to-image models is different: both the text encoder and diffusion generator show sharp drop-off thresholds rather than graceful degradation, implying that specific weights carry critical semantic information whose removal collapses generation quality.","pith_inferences":["A direct consequence the paper leaves implicit: single-run FID on 10,000 images likely has non-negligible stochastic variance, so the advertised 0.08 FID gap between the pruned and unpruned model should be re-measured with repeated seeds and confidence intervals to be sure the 'minimal quality loss' conclusion is not measurement noise.","The claim that specific weights encode critical semantics could be tested directly: prune the weights just below and just above the collapse threshold, and probe how prompts alter failure modes; if failure is prompt-specific, it would support a semantic-locality interpretation, not just a generic capacity effect.","If the magnitude-over-Wanda result generalizes to other text-to-image models, such as SDXL or Stable Diffusion 3, then pruning research for generative vision should focus on conditional layers like cross-attention rather than on activation outlier statistics; a testable extension is to run the same component-wise sparsity sweep on those models.","The paper's optimal 47.5/35 split is computed for SD2's parameter mix; for architectures with a different text-encoder fraction, the same threshold-relative reasoning would shift the optimal split, and a simple algebraic rescaling of the thresholds could be checked on another model."],"forward_implications":["If the central claim holds, a trivial magnitude-based pruning script, with no activation statistics and no retraining, is the practical choice for compressing Stable Diffusion 2-class models, and Wanda-style methods should not be transferred from LLMs without adaptation.","The 38.5% sparsity configuration would cut roughly 460 million parameters from the 1.2B model while keeping image quality on par with the full model, easing memory and latency constraints for on-device deployment.","Because the diffusion generator holds 72% of the parameters, allocating the majority of sparsity to the U-Net yields the best full-model quality, a direct corollary of the component-wise results.","Pruning either component beyond its threshold induces catastrophic degradation, so any practical pruning scheme must respect the text-encoder threshold near 60-62.5% and the generator's roughly 50% qualitative threshold.","The sharp thresholds suggest that saliency analysis of weights just below the collapse point could reveal which parameters encode prompt semantics, opening a window into interpretability and potential bias identification."],"supporting_citations":[{"why":"Defines magnitude pruning, the baseline method that the paper finds to be the best-performing technique.","marker":"[12]"},{"why":"Introduces Wanda, the activation-aware LLM pruning method that the paper tests and finds underperforms magnitude pruning on the text encoder.","marker":"[37]"},{"why":"Introduces SparseGPT, the one-shot LLM pruning algorithm referenced as the approach that does not transfer to text-to-image models.","marker":"[8]"},{"why":"Presents the latent diffusion architecture of Stable Diffusion, the model being pruned.","marker":"[33]"},{"why":"Provides the CLIP text encoder used in Stable Diffusion 2 and the CLIP Score evaluation metric.","marker":"[30]"},{"why":"Defines the Fréchet Inception Distance (FID), the primary image-quality metric used in all evaluations.","marker":"[16]"},{"why":"Defines CLIP Score, the semantic alignment metric used to evaluate prompt-image consistency.","marker":"[15]"},{"why":"Supplies the MSCOCO 2017 dataset of 10,000 captions used to generate evaluation images.","marker":"[23]"},{"why":"Introduces OWL, the outlier-weighted layerwise sparsity method tested on top of magnitude and Wanda pruning.","marker":"[40]"}],"fun_headline_variants":["Magnitude pruning wins for text-to-image, slims SD2 to 38.5%","Simple pruning beats Wanda and SparseGPT on Stable Diffusion","Pruning SD2: magnitude beats advanced methods, 38.5% sparsity","Sharp pruning thresholds reveal critical weights in SD2","No quality loss: magnitude pruning cuts SD2 by 38.5%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central quantitative claims depend on the stability of one FID and one CLIP Score measurement per configuration, with an observed gap as small as 0.08 FID and no repeated runs, seeds, or confidence intervals.","fun_headline_variants_meta":{"raw":{"variants":["Magnitude pruning wins for text-to-image, slims SD2 to 38.5%","Simple pruning beats Wanda and SparseGPT on Stable Diffusion","Pruning SD2: magnitude beats advanced methods, 38.5% sparsity","Sharp pruning thresholds reveal critical weights in SD2","No quality loss: magnitude pruning cuts SD2 by 38.5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000826,"raw_usage":{"total_tokens":3656,"prompt_tokens":1037,"completion_tokens":2619,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":2531}},"tokens_in":653,"tokens_out":2619,"duration_ms":17171,"temperature":1.0,"reasoning_tokens":2531,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:28:25.800734+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the recommended 38.5% sparsity configuration (text 47.5%, image 35%) and the unpruned baseline multiple times with different random seeds, computing FID and CLIP Score over the same 10,000 MSCOCO prompts; if the spread across runs exceeds roughly 0.1 FID, the minimal-quality-loss claim is not established. For the magnitude-wins claim, a decisive test is to hold constant the evaluation protocol and swap only the pruning method at 50% text-encoder sparsity, comparing the resulting FID distributions rather than point estimates.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines magnitude pruning, the baseline method that the paper finds to be the best-performing technique."},{"cited_title":"Zico Kolter","cited_arxiv_id":null,"evidence_quote":"Introduces Wanda, the activation-aware LLM pruning method that the paper tests and finds underperforms magnitude pruning on the text encoder."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces SparseGPT, the one-shot LLM pruning algorithm referenced as the approach that does not transfer to text-to-image models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CLIP text encoder used in Stable Diffusion 2 and the CLIP Score evaluation metric."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines CLIP Score, the semantic alignment metric used to evaluate prompt-image consistency."},{"cited_title":"Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C","cited_arxiv_id":null,"evidence_quote":"Supplies the MSCOCO 2017 dataset of 10,000 captions used to generate evaluation images."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces OWL, the outlier-weighted layerwise sparsity method tested on top of magnitude and Wanda pruning."}],"review_version":1}