{"id":"903c0fa1-eb97-4238-9bba-011fb1781c42","arxiv_id":"2511.00382","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Benign PEFT fine-tuning changes LLM safety and fairness: adapter-based methods (LoRA, IA3) preserve alignment better than prompt-based methods, and the base model strongly moderates outcomes.","lead":"Fine-tuning LLMs with parameter-efficient methods like LoRA or prompt tuning can shift safety and fairness even when the training data is benign. Across four 7–8B models and 235 fine-tuned variants, adapter methods generally preserved alignment better than prompt-based methods, with base model choice a major factor.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Utility-outlier removal concentrated in LLaMA-Prompt-Tuning (13/19) may drive the reported 'LLaMA stable' and prompt-vs-adapter contrast; re-analysis including failures is needed.","rationale":"The reader's weakest assumption was cross-family judge bias in LLaMA-Guard-2 and BBQ-Lite. That is a legitimate concern, but the utility-outlier removal is more load-bearing because it directly removes a large fraction of one method-model cell (13 of 19 outliers are Prompt-Tuning on LLaMA) that anchors the 'LLaMA is stable' moderator claim. Judge bias would affect cross-model absolute scores but is less likely to reverse the within-model adapter-vs-prompt ordering, since the same judge evaluates all methods on the same base model. Outlier removal, by contrast, changes which data points exist for one cell, and the correlation between utility and safety means the filtering is likely informative about safety. The paper's own threat-to-validity statement acknowledges the possibility but treats it as a worst-case underestimation; the concentration in a single cell makes it a structural bias. I still think the central qualitative claim may survive—excluding degenerate models could plausibly attenuate prompt-based degradation, so including them would strengthen the method contrast—but the 'LLaMA stable' claim could easily flip, and the failure rate itself is practically important. Thus the verdict remains CONDITIONAL, pending the re-analysis; I do not see grounds to move to ACCEPT or REJECT without the sensitivity check.","tokens_in":42538,"tokens_out":6126,"duration_ms":64473,"concrete_test":"Re-run the main safety and fairness analyses (Figs. 1–2, Tables X/XII) with the 19 utility outliers included as a separate 'degenerate fine-tuning' category rather than excluded. Specifically, report the Prompt-Tuning-LLaMA cell's mean safety/fairness change and failure rate before and after exclusion. If LLaMA's Prompt-Tuning mean shifts from near-zero to a significant negative change when the 13 excluded runs are counted (e.g., assigned worst-case safety or analyzed separately), then the LLaMA-stability and adapter-vs-prompt conclusions are artifacts of the filtering rule.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—adapter-based PEFT is safer/fairer than prompt-based PEFT, with LLaMA comparatively stable—depends on the filtering step in Section IV and Appendix C. The paper removes 19 utility outliers via Tukey's fences; Appendix C reports that 13 of these are Prompt-Tuning applied to LLaMA. If LLaMA had 18 Prompt-Tuning runs (6 settings × 3 repeats), this removes 72% of that cell. The remaining Prompt-Tuning-LLaMA runs are a non-representative survivor set, and the paper's own correlation analysis (Section IV.C) shows utility and safety positively correlated, so dropping low-utility runs likely drops low-safety runs. This can inflate LLaMA's apparent safety stability and attenuate the prompt-based degradation, directly shaping the moderation claim. The paper acknowledges this in VI.A with a brief 'worst-case degradation may be underestimated' caveat, but that undersells the problem: the removal is not random, it is concentrated in one method-model pair that anchors the 'LLaMA stable' result. A method that collapses utility 72% of the time is itself a safety/fairness-relevant outcome and should be reported, not filtered away. The cross-family judge-bias concern raised by the reader is also valid, but it affects cross-model comparisons less directly than this cell-specific filtering, which biases even within-model method comparisons.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a large-scale empirical study of how four PEFT methods (LoRA, IA3, Prompt-Tuning, P-Tuning) affect the safety and fairness of four instruction-tuned 7–8B LLMs. The authors fine-tune models on benign conversational data under SFT/DPO, varying dataset, learning rate, and epoch count; evaluate safety with HEx-PHI prompts scored by Llama-Guard-2, fairness with a manually modified BBQ-Lite, and utility with GPT-4o-based MT-Bench scoring. After removing inference failures and utility outliers, 235 models are analyzed with paired non-parametric tests. The central claim is that adapter-based methods generally preserve or improve safety/fairness, while prompt-based methods more often degrade them, with base-model choice as a strong moderator: LLaMA stable, Qwen modest gains, Gemma steepest safety decline, Mistral most variable.","tokens_in":42933,"tokens_out":4257,"duration_ms":43791,"significance":"If the central claim holds, this is a valuable and practically relevant contribution: it provides the first systematic, multi-model comparison of PEFT-specific alignment risks, and the practical guidance (prefer adapters, audit category-level metrics, start from a well-aligned base) is actionable. The study's strengths include the large experimental matrix (264 initial fine-tunes), the use of external benchmarks and an independent guard model for measurement, the detailed statistical appendix, and the provision of a replication package with the modified BBQ-Lite. The paper also transparently reports threats to validity. However, the robustness of the headline comparisons is currently undermined by a non-random filtering step and by some inconsistent reporting of sample sizes and significance.","major_comments":[{"comment":"The removal of utility outliers is concentrated in one cell: 13 of the 19 removed models are Prompt-Tuning runs of LLaMA (Appendix C). Because §IV.C reports positive correlations between utility and safety/accuracy, removing these low-utility runs likely removes disproportionately low-safety, low-fairness runs. This directly inflates LLaMA's apparent safety stability and attenuates the prompt-vs-adapter contrast, both of which anchor the paper's central claim. The §VI.A caveat that 'worst-case degradation may be underestimated' does not address the cell-specific, non-random nature of the removal. Please reanalyze with the 19 outlier models included (treating severe utility collapse as an outcome, not a missing value) and/or provide a sensitivity analysis excluding each cell in turn.","section":"§IV (Data Filtering) and Appendix C"},{"comment":"The reported numbers do not reconcile. §III.C states '24 models were fine-tuned for each of LLaMA, Mistral, and Qwen, and 16 for Gemma, bringing the total number of fine-tuned models to 264'; 24+24+24+16=88, not 264. The abstract also mentions a coding-task extension with 96 additional fine-tuned models, but the main text contains no section describing or analyzing this extension. These inconsistencies matter because filtering fractions, cell sizes for paired tests, and the claimed scope of the study depend on the exact counts. Please correct the counts and either integrate the coding-extension results or remove that claim.","section":"§III.C and Abstract"},{"comment":"The text says safety 'increases significantly with LoRA (p = 0.059)' while the stated significance level is α = 0.05. Appendix Table X lists the LoRA-vs-base comparison as p = 0.0587, which is not significant at the chosen threshold. This is not merely a typo: the findings summary in §IV.A claims 'Adapter-based techniques yield statistically significant safety gains,' which is supported only by IA3 under the stated α. Please reword the LoRA claim and adjust any summary statements that rely on it.","section":"§IV.A.1 and Table X/Appendix E"},{"comment":"Safety scores are produced by Llama-Guard-2, an 8B model from the Llama-3 family, and are used to rank all fine-tuned models, including Llama-3-8B-Instruct itself. A systematic same-family or judge-model bias could distort the LLaMA-stable versus Gemma-declining comparison, which is the paper's second headline finding. The threat is acknowledged but not quantified. Please report agreement on a subsample with a second independent judge (e.g., a different guard model or human annotations), or at least discuss the direction of the likely bias and why the cross-model moderation finding survives it.","section":"§III.D.2 / §VI.A"}],"minor_comments":[{"comment":"Typographical errors: 'Poytechnique Montreal' should be 'Polytechnique Montreal'.","section":"Affiliation block"},{"comment":"The bias-score equations are correct but the explanation of n_biased_ans and n_non-UNKNOWN_outputs is spread across the main text and Appendix B; consider moving the full derivation into the main text for readability.","section":"§III.D.3, Eq. (1)–(2)"},{"comment":"Table VI headers 'Accuracy AMB' etc. are clear, but the text switches between 'Bias AMB' and 'BiasScore AMB' without consistency; unify notation.","section":"§IV.B"},{"comment":"The manual modifications to BBQ-Lite are substantial (200 examples excluded, 32 corrected). The appendix is helpful, but the paper should state explicitly whether the modified benchmark is released and whether the base-model fairness scores in Table VI are computed on the original or modified set.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a promising and unusually systematic empirical study, but the concentrated exclusion of LLaMA–Prompt-Tuning utility outliers is a load-bearing methodological threat that the current caveat does not cover. The numeric inconsistencies (88 vs. 264, abstract coding extension absent from the text) also need to be resolved before the claims can be trusted. The topic is squarely within the journal's scope and the replication package is a clear strength. I would be willing to see a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the first systematic comparison of four PEFT methods (LoRA, IA3, Prompt-Tuning, P-Tuning) across four 7-8B instruction-tuned families on both safety and fairness, using external benchmarks and 235 surviving fine-tunes. That scope alone is the contribution: prior work mostly studies LoRA or full fine-tuning, and fairness effects of PEFT were largely unmeasured. The paper is also careful with paired non-parametric tests, effect sizes, category-level breakdowns, and it ships a corrected BBQ-Lite and a replication package. The central qualitative claim—adapter methods disturb alignment less than prompt-based methods—is plausible and reasonably supported by the aggregated data, especially for Mistral and Gemma. I would not desk-reject this.\n\nThe soft spots are real, though. The utility-outlier removal is load-bearing. Nineteen outliers were removed, and thirteen of them are Prompt-Tuning on LLaMA. That is roughly 72% of that cell gone. Since the paper's own correlation analysis shows utility and safety move together, removing the lowest-utility runs almost certainly removes the lowest-safety runs. The paper's VI.A caveat that 'worst-case degradation may be underestimated' undersells this: the removal is not random noise, it is concentrated in exactly the method-model pair that anchors the 'LLaMA is stable' conclusion and the prompt-vs-adapter contrast. Re-analysis including the failures, or at least reporting results with and without the filtered runs, is needed before I'd trust the specific rankings. Second, LoRA's p=0.059 is described as significant in Section IV.A.1; that's just wrong at alpha=0.05, even if the paper elsewhere reports the actual value. Third, the abstract promises a coding-task extension with 96 additional models, but the body I received has no such section—either missing content or a mismatch that needs fixing. The Llama-Guard-2 judge bias across non-Llama families is a lesser concern; it exists, but it doesn't affect within-family method comparisons as much as the filtering does.\n\nWho should read this: practitioners choosing PEFT for safety-critical deployments, and researchers working on alignment effects of fine-tuning. It deserves a serious referee, but only with the understanding that the analysis must be re-run transparently around the outlier removal. I'd bring it to a reading group for a method discussion, but I wouldn't cite the specific LLaMA-stability numbers until that re-analysis is done.","headline":"Useful first broad PEFT safety/fairness map, but the concentrated utility-outlier removal likely biases the 'LLaMA stable' result and the abstract's coding extension is missing from the body; needs re-analysis before I'd trust the specific rankings.","tokens_in":43360,"tokens_out":1944,"would_cite":false,"duration_ms":21457,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Benign parameter-efficient fine-tuning can significantly shift LLM safety and fairness, with adapter methods preserving alignment and prompt-based methods degrading it.","keywords":["Large Language Models","Parameter-Efficient Fine-Tuning","Safety alignment","Fairness","LoRA","IA3","Prompt-Tuning","BBQ"],"falsifier":"Re-score a sample of the fine-tuned variants' prompt-response pairs with human annotators (or a second guard model from a different training lineage) and check whether the adapter-vs-prompt ranking and the base-model ordering (LLaMA stable, Gemma declining) reproduce; a systematic disagreement that varies by model family would indicate the reported differences are measurement artifacts rather than alignment shifts.","tokens_in":42480,"feed_emoji":"⚖️","tokens_out":5227,"duration_ms":45703,"temperature":0.7,"pith_summary":"This paper asks whether parameter-efficient fine-tuning (PEFT) on benign, everyday data can quietly change an LLM's safety and fairness. It fine-tunes four instruction-tuned 7-8B models (LLaMA, Qwen, Mistral, Gemma) with four PEFT methods -- LoRA, IA3, Prompt-Tuning, P-Tuning -- across six training settings, producing 235 usable variants evaluated on 11 safety hazard categories and 9 fairness dimensions. The central finding: adapter-based methods (LoRA, IA3) generally preserve or even improve safety and fairness, while prompt-based methods (Prompt-Tuning, P-Tuning) more often degrade both. The base model strongly moderates the effect: LLaMA stays stable, Qwen improves modestly, Gemma shows the steepest safety decline, and Mistral varies most. The authors conclude that benign intent does not guarantee safe behavior, and recommend starting from a well-aligned base model, favoring adapters, and auditing at category level.","feed_headline":"Prompt-tuning erodes LLM safety; adapters protect it","feed_subtitle":"A 235-model study shows benign fine-tuning shifts alignment -- base model choice and category-level audits matter most.","key_machinery":"The central instrument is a controlled experimental grid: four instruction-tuned 7-8B base models (LLaMA-3-8B, Qwen2.5-7B, Mistral-7B, Gemma-7B) are each fine-tuned with four PEFT methods (LoRA, IA3, Prompt-Tuning, P-Tuning) under six training settings (SFT/DPO x two datasets x one or five epochs x two learning rates), yielding 264 variants of which 235 pass validity filters. Safety is scored with an automated guard model (LLaMA-Guard-2) on the 330-prompt HEx-PHI benchmark across 11 hazard categories; fairness is scored with a manually cleaned version of the BBQ-Lite multiple-choice benchmark (15,876 questions after corrections) across 9 demographic categories, using accuracy and bias scores","core_discovery":"On the paper's own terms, the discovery is that even benign, task-appropriate fine-tuning is not alignment-neutral: the choice of PEFT method and base model can shift measured safety and fairness by large amounts. Across 235 fine-tuned variants, adapter-based methods (LoRA, IA3) tend to raise or preserve safety scores and keep fairness accuracy higher with lower bias, whereas prompt-based methods (Prompt-Tuning, P-Tuning) significantly reduce safety in most hazard categories and depress fairness accuracy, especially in ambiguous contexts. The base model is a strong moderator -- LLaMA is comparatively robust, Qwen shows modest gains, Gemma exhibits the steepest and most consistent safety decl","pith_inferences":["If the mechanism (adapters leave core weights intact, prompt methods rewrite the input path) is causal, then any PEFT variant that modifies input embeddings or activation distributions -- such as other soft-prompt or prefix methods -- may carry similar alignment risk; this is a testable extension.","The steepest fairness drops occurring in categories with the highest base accuracy suggest a ceiling or regression-to-the-mean pattern; a direct test would compare fine-tuning effects on held-out category variants with matched base accuracy.","The findings imply an audit protocol: before deploying any PEFT-adapted model, rerun a category-level guard model and bias benchmark even when the tuning dataset was benign; this could become a standard pre-deployment check on model hubs.","Because the base model effect dominates the fine-tuning config effect, organizations could screen candidate base models once per use case and then fix a conservative adapter setting, substantially reducing the cost of per-configuration safety and fairness audits."],"forward_implications":["Favoring adapter-based PEFT (LoRA, IA3) over prompt-based methods should reduce the risk of benign fine-tuning harming safety or fairness.","Base-model selection is a risk decision: starting from a well-aligned model (e.g., Qwen for fairness, LLaMA for safety stability) is more decisive than hyperparameter tuning.","Safety and fairness do not move together; a configuration that improves one can worsen the other, so both must be audited separately at category level.","Specific categories are early-warning indicators: Child Abuse Content and Adult Content for safety, Sexual Orientation and Nationality for fairness -- aggregate scores should not be used alone.","Fine-tuning hyperparameters (learning rate, epochs, dataset, SFT vs DPO) have limited and sporadic effects on alignment compared to method and base model choice."],"fun_headline_variants":["Benign PEFT still shifts LLM safety and fairness","Adapter-based tuning safer than prompt-tuning for LLMs","LLM fine-tuning method and base model alter safety","Prompt-tuning harms safety; adapters preserve it","Even benign tuning can erode LLM alignment"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The rankings of methods and base models assume the automated safety judge (LLaMA-Guard-2) and the manually edited BBQ-Lite fairness benchmark measure safety and fairness equally across all four model families; if the judge is systematically biased toward or against one family, the central comparisons could be distorted.","fun_headline_variants_meta":{"raw":{"variants":["Benign PEFT still shifts LLM safety and fairness","Adapter-based tuning safer than prompt-tuning for LLMs","LLM fine-tuning method and base model alter safety","Prompt-tuning harms safety; adapters preserve it","Even benign tuning can erode LLM alignment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001001,"raw_usage":{"total_tokens":4138,"prompt_tokens":873,"completion_tokens":3265,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":3190}},"tokens_in":617,"tokens_out":3265,"duration_ms":20981,"temperature":1.0,"reasoning_tokens":3190,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T06:50:25.492522+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score a sample of the fine-tuned variants' prompt-response pairs with human annotators (or a second guard model from a different training lineage) and check whether the adapter-vs-prompt ranking and the base-model ordering (LLaMA stable, Gemma declining) reproduce; a systematic disagreement that varies by model family would indicate the reported differences are measurement artifacts rather than alignment shifts.","supporting_citations":[],"review_version":1}