Pith. sign in

REVIEW 10 cited by

How to Fine-Tune BERT for Text Classification?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.05583 v3 pith:VN6PAQ3R submitted 2019-05-14 cs.CL

How to Fine-Tune BERT for Text Classification?

classification cs.CL
keywords bertlanguageclassificationmodeltextfine-tuningpre-trainingrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Language model pre-training has proven to be useful in learning universal language representations. As a state-of-the-art language model pre-training model, BERT (Bidirectional Encoder Representations from Transformers) has achieved amazing results in many language understanding tasks. In this paper, we conduct exhaustive experiments to investigate different fine-tuning methods of BERT on text classification task and provide a general solution for BERT fine-tuning. Finally, the proposed solution obtains new state-of-the-art results on eight widely-studied text classification datasets.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Stay Focused: Problem Drift in Multi-Agent Debate

    cs.CL 2025-02 unverdicted novelty 7.0

    The paper defines and measures 'problem drift' in multi-agent LLM debates across tasks and proposes DRIFTJudge and DRIFTPolicy as baselines to detect and reduce it.

  2. OPT: Open Pre-trained Transformer Language Models

    cs.CL 2022-05 unverdicted novelty 7.0

    OPT releases open decoder-only transformers up to 175B parameters that match GPT-3 performance at one-seventh the carbon cost, along with code and training logs.

  3. FinBERT: Financial Sentiment Analysis with Pre-trained Language Models

    cs.CL 2019-08 unverdicted novelty 7.0

    FinBERT adapts BERT to the financial domain and outperforms prior state-of-the-art methods on financial sentiment analysis tasks.

  4. Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning

    cs.CL 2026-05 conditional novelty 5.0

    Fine-tuned transformers with multi-task learning recover substantial wording-derived signal for item difficulty at small sample sizes typical in applied testing.

  5. Exploring Data Augmentation and Resampling Strategies for Transformer-Based Models to Address Class Imbalance in AI Scoring of Scientific Explanations in NGSS Classroom

    cs.AI 2026-03 unverdicted novelty 4.0

    Targeted data augmentation with GPT-4 synthetic responses and ALP phrase-level extraction substantially improves SciBERT performance on severely imbalanced rubric categories for NGSS scientific explanations, achieving...

  6. ADMEDTAGGER: an annotation framework for distillation of expert knowledge for the Polish medical language

    cs.CL 2025-12 unverdicted novelty 4.0

    Llama 3.1 annotates Polish medical texts to train DistilBERT classifiers achieving F1 scores above 0.80 that are 500 times smaller than the teacher model.

  7. Towards the Anonymization of the Language Modeling

    cs.CL 2025-01 unverdicted novelty 4.0

    Authors introduce MLM and CLM specialization methods that avoid memorizing identifiers in sensitive training data while aiming for a privacy-utility tradeoff on medical datasets.

  8. PortBERT: Navigating the Depths of Portuguese Language Models

    cs.CL 2026-06 unverdicted novelty 3.0

    PortBERT releases two RoBERTa models for Portuguese that match or beat prior monolingual and multilingual models on translated GLUE/SuperGLUE tasks while reporting training and inference times.

  9. Enhancing Construction Worker Safety in Extreme Heat: A Machine Learning Approach Utilizing Wearable Technology for Predictive Health Analytics

    cs.AI 2026-04 unverdicted novelty 3.0

    An attention-based LSTM model predicts heat stress from heart rate, HRV, and SpO2 data with 95.4% accuracy and 0.982 F1 score on a 19-worker dataset.

  10. To Tune or Not To Tune? How About the Best of Both Worlds?

    cs.CL 2019-07 unverdicted novelty 3.0

    A sequential fine-tuning strategy for pre-trained language models reports modest accuracy gains of 4.7%, 0.99%, and 0.72% on semantic similarity, sequence labeling, and text classification tasks.