Pith. sign in

REVIEW 9 cited by

Pushing The Limit of LLM Capacity for Text Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07470 v2 pith:LWBJD3ZC submitted 2024-02-12 cs.CL

classification cs.CL
keywords classificationtextlearnersllmsbasergptlanguagequestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The value of text classification's future research has encountered challenges and uncertainties, due to the extraordinary efficacy demonstrated by large language models (LLMs) across numerous downstream NLP tasks. In this era of open-ended language modeling, where task boundaries are gradually fading, an urgent question emerges: have we made significant advances in text classification under the full benefit of LLMs? To answer this question, we propose RGPT, an adaptive boosting framework tailored to produce a specialized text classification LLM by recurrently ensembling a pool of strong base learners. The base learners are constructed by adaptively adjusting the distribution of training samples and iteratively fine-tuning LLMs with them. Such base learners are then ensembled to be a specialized text classification LLM, by recurrently incorporating the historical predictions from the previous learners. Through a comprehensive empirical comparison, we show that RGPT significantly outperforms 8 SOTA PLMs and 7 SOTA LLMs on four benchmarks by 1.36% on average. Further evaluation experiments show a clear surpassing of RGPT over human classification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Transformer-based Autoregressive Decoder Architecture for Hierarchical Text Classification

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A decoder that generates child-to-parent symbolic label sequences, with no graph encoder and no label semantics, matches state-of-the-art hierarchical text classification on three benchmarks at roughly half the infere...

  2. Regulation of Language Models With Interpretability Will Likely Result In A Performance Trade-Off

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Forcing an LLM to classify using only human-specified legal concepts costs about 7.34% accuracy, but can speed up human decision-making despite the loss.

  3. RAMIE: Retrieval-Augmented Multi-task Information Extraction with Large Language Models on Dietary Supplements

    cs.CL 2024-11 conditional novelty 6.0 of 10

    RAMIE, a retrieval-augmented multi-task instruction-tuned framework, improves LLM information extraction for dietary supplements from clinical records, with RAG recovering accuracy lost in multi-task training.

  4. Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support

    cs.CY 2025-08 unverdicted novelty 5.0 of 10

    A multi-agent reliability check and an LLM-as-a-judge quality check both reduce hallucinations in AI-generated study scaffolds, with the multi-agent check matching human expert judgments almost perfectly.

  5. How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting

    cs.CL 2025-07 conditional novelty 5.0 of 10

    For multilingual RAG intent classification, the best translation strategy depends on the model and language; translating instructions into the user's language helps some models, while making the model answer in low-re...

  6. Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuned BERT-like models outperform zero-shot and internal-state LLM methods on four of six challenging text classification datasets.

  7. Data Quality Enhancement on the Basis of Diversity with Large Language Models for Text Classification: Uncovered, Difficult, and Noisy

    cs.CL 2024-12 conditional novelty 5.0 of 10

    DQE selects about half of a training set via greedy sampling plus similarity-based categorization of uncovered, difficult, and noisy samples, and reports improved text classification accuracy over full-data fine-tuning.

  8. Multiple Abstraction Level Retrieve Augment Generation

    cs.CL 2025-01 conditional novelty 4.0 of 10

    MAL-RAG retrieves document, section, paragraph, and multi-sentence chunks together and claims a 25.7% improvement in AI-judged answer correctness on glycoscience questions over single-level RAG.

  9. A Survey on Large Language Models for Communication, Network, and Service Management: Application Insights, Challenges, and Future Directions

    cs.NI 2024-12 conditional novelty 4.0 of 10

    A systematic survey of 108 papers classifies how large language models are used for communication network and service management across four network domains.

Pith tools