Pith. sign in

REVIEW 6 cited by

Open, Closed, or Small Language Models for Text Classification?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.10092 v1 pith:IM6Y64RA submitted 2023-08-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelstasksclosedperformanceacrossclassificationdatasetslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in large language models have demonstrated remarkable capabilities across various NLP tasks. But many questions remain, including whether open-source models match closed ones, why these models excel or struggle with certain tasks, and what types of practical procedures can improve performance. We address these questions in the context of classification by evaluating three classes of models using eight datasets across three distinct tasks: named entity recognition, political party prediction, and misinformation detection. While larger LLMs often lead to improved performance, open-source models can rival their closed-source counterparts by fine-tuning. Moreover, supervised smaller models, like RoBERTa, can achieve similar or even greater performance in many datasets compared to generative LLMs. On the other hand, closed models maintain an advantage in hard tasks that demand the most generalizability. This study underscores the importance of model selection based on task requirements

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    PiFi adds one frozen LLM layer to an SLM and fine-tunes it, reporting consistent but modest gains across NLU and NLG tasks, with larger gains when the LLM matches the target language.

  2. IYKYK: Using language models to decode extremist cryptolects

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs struggle with extremist in-group jargon, but prompting with example posts and domain-adapting encoders substantially improves detection and decoding.

  3. Large Language Models in the Task of Automatic Validation of Text Classifier Predictions

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLM-based annotators using token-probability thresholds, RAG, and reasoning fine-tuning matched or exceeded human annotator quality on a proprietary 250-class intent-validation task.

  4. SOI Matters: Analyzing Multi-Setting Training Dynamics in Pretrained Language Models via Subsets of Interest

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A new fine-grained taxonomy of example-level learning dynamics (SOI) is applied to multi-task, multi-source, and multi-lingual fine-tuning, showing robust OOD gains for multi-source training and small gains from SOI-g...

  5. Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A single BERT-scale model with a game-context token and LLM-assisted label transfer achieves toxicity detection comparable to per-game models while extending to seven languages.

  6. Rethinking the Understanding Ability across LLMs through Mutual Information

    cs.CL 2025-05 conditional novelty 4.0 of 10

    The paper uses token-level recoverability as a computable lower bound on mutual information to compare LLMs and to fine-tune them, finding encoder-only models preserve information better than decoder-only models.

Pith tools