Pith. sign in

REVIEW 1 cited by

SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03437 v5 pith:JUPSADAL submitted 2019-11-08 cs.CL cs.LGmath.OC

classification cs.CLcs.LGmath.OC
keywords pre-trainedmodelsdownstreamfine-tuninglanguagemodeltaskscapacity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transfer learning has fundamentally changed the landscape of natural language processing (NLP) research. Many existing state-of-the-art models are first pre-trained on a large text corpus and then fine-tuned on downstream tasks. However, due to limited data resources from downstream tasks and the extremely large capacity of pre-trained models, aggressive fine-tuning often causes the adapted model to overfit the data of downstream tasks and forget the knowledge of the pre-trained model. To address the above issue in a more principled manner, we propose a new computational framework for robust and efficient fine-tuning for pre-trained language models. Specifically, our proposed framework contains two important ingredients: 1. Smoothness-inducing regularization, which effectively manages the capacity of the model; 2. Bregman proximal point optimization, which is a class of trust-region methods and can prevent knowledge forgetting. Our experiments demonstrate that our proposed method achieves the state-of-the-art performance on multiple NLP benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Generalization in Chain of Thought Reasoning for Smaller Models

    cs.LG 2025-01 reject novelty 4.0 of 10

    PRADA combines P-Tuning and domain-adversarial training with CoT distillation and claims improved cross-domain reasoning in small models, though the evaluation is confounded by target-data access.

Pith tools