Pith. sign in

REVIEW 2 cited by

The effect of fine-tuning on language model toxicity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.15821 v1 pith:73CMOHP6 submitted 2024-10-21 cs.AI

classification cs.AI
keywords fine-tuningmodelstoxicitymodelassessimpactlanguageopen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Fine-tuning language models has become increasingly popular following the proliferation of open models and improvements in cost-effective parameter efficient fine-tuning. However, fine-tuning can influence model properties such as safety. We assess how fine-tuning can impact different open models' propensity to output toxic content. We assess the impacts of fine-tuning Gemma, Llama, and Phi models on toxicity through three experiments. We compare how toxicity is reduced by model developers during instruction-tuning. We show that small amounts of parameter-efficient fine-tuning on developer-tuned models via low-rank adaptation on a non-adversarial dataset can significantly alter these results across models. Finally, we highlight the impact of this in the wild, demonstrating how toxicity rates of models fine-tuned by community contributors can deviate in hard-to-predict ways.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ConSens: Assessing context grounding in open-book question answering

    cs.CL 2025-04 conditional novelty 5.0 of 10

    ConSens measures context grounding in open-book QA as a sigmoid-transformed log ratio of answer perplexity without context to perplexity with context, reaching ROC AUC 0.88 to 0.93.

  2. Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation

    cs.SE 2025-04 reject novelty 4.0 of 10

    Fine-tuned open-source LLMs can restructure bug reports into standard templates, with Qwen 2.5 reaching 77% CTQRS, comparable to ChatGPT-4o.

Pith tools