Pith. sign in

REVIEW 2 cited by

A Survey on Text Classification: From Shallow to Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.00364 v6 pith:PGUZPTQC submitted 2020-08-02 cs.CL

classification cs.CL
keywords classificationtextdeeplearningmodelsresearchsurveyarea
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text classification is the most fundamental and essential task in natural language processing. The last decade has seen a surge of research in this area due to the unprecedented success of deep learning. Numerous methods, datasets, and evaluation metrics have been proposed in the literature, raising the need for a comprehensive and updated survey. This paper fills the gap by reviewing the state-of-the-art approaches from 1961 to 2021, focusing on models from traditional models to deep learning. We create a taxonomy for text classification according to the text involved and the models used for feature extraction and classification. We then discuss each of these categories in detail, dealing with both the technical developments and benchmark datasets that support tests of predictions. A comprehensive comparison between different techniques, as well as identifying the pros and cons of various evaluation metrics are also provided in this survey. Finally, we conclude by summarizing key implications, future research directions, and the challenges facing the research area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Large Language Models for Multimodal Sentiment Analysis: Challenges, Benchmarks, and Future Directions

    cs.CL 2024-11 conditional novelty 5.0 of 10

    On two Twitter multimodal aspect-based sentiment benchmarks, zero-shot/few-shot LLMs like Llama2, LLaVA, and ChatGPT score 7 to 16 F1 points below supervised baselines and take orders of magnitude longer to run.

  2. The Text Classification Pipeline: Starting Shallow going Deeper

    cs.CL 2024-12 conditional novelty 1.0 of 10

    A survey monograph on text classification pipeline stages, with proposed preprocessing acronyms and a reproduced word-embedding case study, but no new results.

Pith tools