Pith. sign in

REVIEW 5 cited by

Is ChatGPT a Good Sentiment Analyzer? A Preliminary Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.04339 v2 pith:4JIDGC37 submitted 2023-04-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords emphevaluationchatgptsentimentanalysisanalyzerconductpreliminary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, ChatGPT has drawn great attention from both the research community and the public. We are particularly interested in whether it can serve as a universal sentiment analyzer. To this end, in this work, we provide a preliminary evaluation of ChatGPT on the understanding of \emph{opinions}, \emph{sentiments}, and \emph{emotions} contained in the text. Specifically, we evaluate it in three settings, including \emph{standard} evaluation, \emph{polarity shift} evaluation and \emph{open-domain} evaluation. We conduct an evaluation on 7 representative sentiment analysis tasks covering 17 benchmark datasets and compare ChatGPT with fine-tuned BERT and corresponding state-of-the-art (SOTA) models on them. We also attempt several popular prompting techniques to elicit the ability further. Moreover, we conduct human evaluation and present some qualitative case studies to gain a deep comprehension of its sentiment analysis capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rank, Don't Generate: Statement-level Ranking for Explainable Recommendation

    cs.IR 2026-04 unverdicted novelty 7.0 of 10

    The work reframes explainable recommendation as statement-level ranking, introduces the StaR benchmark from Amazon reviews, and finds popularity baselines outperforming SOTA models in item-level personalized ranking.

  2. Visibility vs. Engagement: How Two Indian News Websites Reported on LGBTQ+ Individuals and Communities during the Pandemic

    cs.HC 2025-07 conditional novelty 6.0 of 10

    Indian news websites gave LGBTQ+ communities visibility during the pandemic but often with little depth, and Times of India's language was at times transphobic.

  3. MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction Tuning for Emotion-Cause Pair Extraction

    cs.CL 2025-07 conditional novelty 5.0 of 10

    MEKiT improves LLM emotion-cause pair extraction by adding emotional label knowledge to instruction prompts and mixing causal examples into training data, achieving 61.49 F1 on NTCIR-13.

  4. FCKT: Fine-Grained Cross-Task Knowledge Transfer with Semantic Contrastive Learning for Targeted Sentiment Analysis

    cs.CL 2025-05 conditional novelty 5.0 of 10

    FCKT improves targeted sentiment analysis through fine-grained transfer of aspect boundary knowledge into sentiment prediction, using token-level contrastive learning and an alternating real/predicted training strategy.

  5. Can ChatGPT Perform Image Splicing Detection? A Preliminary Study

    cs.CV 2025-05 conditional novelty 5.0 of 10

    GPT-4.1 detects image splicing on a curated CASIA v2.0 subset with over 85% zero-shot accuracy, and chain-of-thought prompting gives the most balanced performance across authentic and spliced images.

Pith tools