Pith. sign in

REVIEW 2 cited by

NoisyTune: A Little Noise Can Help You Finetune Pretrained Language Models Better

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.12024 v2 pith:IYXJA4BP submitted 2022-02-24 cs.CL

classification cs.CL
keywords plmsdifferenttasksdownstreamfinetuningnoisytunebenchmarkbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Effectively finetuning pretrained language models (PLMs) is critical for their success in downstream tasks. However, PLMs may have risks in overfitting the pretraining tasks and data, which usually have gap with the target downstream tasks. Such gap may be difficult for existing PLM finetuning methods to overcome and lead to suboptimal performance. In this paper, we propose a very simple yet effective method named NoisyTune to help better finetune PLMs on downstream tasks by adding some noise to the parameters of PLMs before fine-tuning. More specifically, we propose a matrix-wise perturbing method which adds different uniform noises to different parameter matrices based on their standard deviations. In this way, the varied characteristics of different types of parameters in PLMs can be considered. Extensive experiments on both GLUE English benchmark and XTREME multilingual benchmark show NoisyTune can consistently empower the finetuning of different PLMs on different downstream tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bi-Level Optimization for Self-Supervised AI-Generated Face Detection

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Bi-level reweighting of EXIF and manipulation pretext tasks yields a photo-only encoder that beats prior detectors on cross-generator AI-face detection.

  2. Put Teacher in Student's Shoes: Cross-Distillation for Ultra-compact Model Compression Framework

    cs.CL 2025-07 conditional novelty 5.0 of 10

    EI-BERT compresses a Chinese NLU model to 1.91 MB with competitive accuracy using attention-based vocabulary pruning, cross-distillation, and module-wise INT8 quantization, and reports deployment at Alipay.

Pith tools