Pith. sign in

REVIEW 2 cited by

Large Language Models For Text Classification: Case Study And Comprehensive Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.08457 v1 pith:XKQWDFCO submitted 2025-01-14 cs.CL cs.LG

classification cs.CLcs.LG
keywords classificationmodelslanguagellmsmodelbinarydifferentevaluate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Unlocking the potential of Large Language Models (LLMs) in data classification represents a promising frontier in natural language processing. In this work, we evaluate the performance of different LLMs in comparison with state-of-the-art deep-learning and machine-learning models, in two different classification scenarios: i) the classification of employees' working locations based on job reviews posted online (multiclass classification), and 2) the classification of news articles as fake or not (binary classification). Our analysis encompasses a diverse range of language models differentiating in size, quantization, and architecture. We explore the impact of alternative prompting techniques and evaluate the models based on the weighted F1-score. Also, we examine the trade-off between performance (F1-score) and time (inference response time) for each language model to provide a more nuanced understanding of each model's practical applicability. Our work reveals significant variations in model responses based on the prompting strategies. We find that LLMs, particularly Llama3 and GPT-4, can outperform traditional methods in complex classification tasks, such as multiclass classification, though at the cost of longer inference times. In contrast, simpler ML models offer better performance-to-time trade-offs in simpler binary classification tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Binary Moderation: Identifying Fine-Grained Sexist and Misogynistic Behavior on GitHub with Large Language Models

    cs.SE 2025-07 conditional novelty 6.0 of 10

    An instruction-tuned GPT-4o prompt achieves an MCC of 0.501 on 12-category sexism/misogyny classification of GitHub comments, but the evaluation was tuned on the same test set.

  2. Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn

    cs.AI 2026-06 conditional novelty 5.0 of 10

    A GPT-4-distilled small LM plus grouped LoRA adapters improves LinkedIn's job-attribute classification over legacy models.

Pith tools