Pith. sign in

REVIEW 11 cited by

Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08660 v2 pith:6UWFNICG submitted 2024-06-12 cs.CL cs.AI

Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification

classification cs.CL cs.AI
keywords classificationllmstextfine-tunedgenerativemodelstrainingchatgpt
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Generative AI offers a simple, prompt-based alternative to fine-tuning smaller BERT-style LLMs for text classification tasks. This promises to eliminate the need for manually labeled training data and task-specific model training. However, it remains an open question whether tools like ChatGPT can deliver on this promise. In this paper, we show that smaller, fine-tuned LLMs (still) consistently and significantly outperform larger, zero-shot prompted models in text classification. We compare three major generative AI models (ChatGPT with GPT-3.5/GPT-4 and Claude Opus) with several fine-tuned LLMs across a diverse set of classification tasks (sentiment, approval/disapproval, emotions, party positions) and text categories (news, tweets, speeches). We find that fine-tuning with application-specific training data achieves superior performance in all cases. To make this approach more accessible to a broader audience, we provide an easy-to-use toolkit alongside this paper. Our toolkit, accompanied by non-technical step-by-step guidance, enables users to select and fine-tune BERT-like LLMs for any classification task with minimal technical and computational effort.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training

    cs.LG 2026-06 unverdicted novelty 7.0

    ART optimizes visual pixel inputs to frozen MLLMs to achieve LoRA-competitive accuracy on math and structured tool-use benchmarks without modifying computational graphs.

  2. RePrompT: Recurrent Prompt Tuning for Integrating Structured EHR Encoders with Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0

    RePrompT uses recurrent prompt tuning to inject prior-visit latent states and cohort-derived population prompt tokens into LLMs, yielding better performance than pure EHR or pure LLM baselines on MIMIC clinical predic...

  3. Just Use XML: Revisiting Joint Translation and Label Projection

    cs.CL 2026-03 unverdicted novelty 6.0

    LabelPigeon jointly performs translation and label projection via XML tags, improving translation quality in 11 languages and cross-lingual transfer by up to +40.2 F1 on NER across 27 languages.

  4. AnnotateThis: Analyzing a human-LLM system for annotating social media data with the concept of climate change mitigation pessimism

    cs.CY 2026-06 unverdicted novelty 5.0

    AnnotateThis lets users improve LLM annotations for climate change mitigation pessimism on social media, yielding 0.15 higher F-Measure and 0.23 higher accuracy than automated prompt refinement when ground truth label...

  5. Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents

    cs.CR 2026-04 unverdicted novelty 5.0

    Targeted prompting and system interventions enable local LLMs such as Llama 3.1 70B to exploit 83% of tested Linux privilege escalation vulnerabilities.

  6. AIx4Soccer: A Unified Platform Architecture for Football Club Management and Structured Athlete Development

    cs.CY 2026-07 conditional novelty 4.0

    A conceptual multi-tenant football-club SaaS design embeds a PDI development cycle and a 75/25 certified video-analyst marketplace on a proposed event-sourced knowledge-graph substrate.

  7. Comparing BERT Sentence-Pair Classification and Few-Shot LLM Prompting for Detecting Threat and Solution Framing in German Climate News

    cs.CL 2026-06 unverdicted novelty 4.0

    Fine-tuned BERT sentence-pair classifiers reach F1 0.83 while few-shot LLM prompting reaches F1 0.78 on threat and solution framing detection in 440 manually coded German climate news articles.

  8. Supervision versus Demonstration-Based In-Context Learning for Multiword Expression Classification

    cs.CL 2026-06 unverdicted novelty 4.0

    On a controlled Turkish dataset of 147 examples, few-shot prompting lets some LLMs match or beat a supervised BERT baseline for LVC detection, though results are highly sensitive to prompt design.

  9. Poodle: Seamlessly Scaling Down Large Language Models with Just-in-Time Model Replacement

    cs.DB 2025-12 unverdicted novelty 4.0

    Poodle shows that LLMs can be automatically replaced with cheaper models for recurring tasks to save significant cost and energy without extra user effort.

  10. Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

    cs.AI 2026-07 unverdicted novelty 3.0

    SPG-Layout combines statistical object priors with hierarchical large-object-first placement to produce physically plausible text-driven 3D scenes in non-Manhattan rooms and outperforms baselines on a new 500-scene benchmark.

  11. Short-form Text Rewriting with Phi Silica

    cs.CL 2026-05 unverdicted novelty 3.0

    Finetuning Phi Silica on curated short presentation text improves semantic fidelity, reduces hallucinations, and raises preference win rates over GPT-5-chat rewrites.