Pith. sign in

REVIEW 2 cited by

A Wide Evaluation of ChatGPT on Affective Computing Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.13911 v1 pith:ZUTNL6NL submitted 2023-08-26 cs.AI cs.CL

classification cs.AIcs.CL
keywords modelsproblemschatgptdetectionrankingaffectivecomputingintensity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

With the rise of foundation models, a new artificial intelligence paradigm has emerged, by simply using general purpose foundation models with prompting to solve problems instead of training a separate machine learning model for each problem. Such models have been shown to have emergent properties of solving problems that they were not initially trained on. The studies for the effectiveness of such models are still quite limited. In this work, we widely study the capabilities of the ChatGPT models, namely GPT-4 and GPT-3.5, on 13 affective computing problems, namely aspect extraction, aspect polarity classification, opinion extraction, sentiment analysis, sentiment intensity ranking, emotions intensity ranking, suicide tendency detection, toxicity detection, well-being assessment, engagement measurement, personality assessment, sarcasm detection, and subjectivity detection. We introduce a framework to evaluate the ChatGPT models on regression-based problems, such as intensity ranking problems, by modelling them as pairwise ranking classification. We compare ChatGPT against more traditional NLP methods, such as end-to-end recurrent neural networks and transformers. The results demonstrate the emergent abilities of the ChatGPT models on a wide range of affective computing problems, where GPT-3.5 and especially GPT-4 have shown strong performance on many problems, particularly the ones related to sentiment, emotions, or toxicity. The ChatGPT models fell short for problems with implicit signals, such as engagement measurement and subjectivity detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comprehensive Evaluation of Large Language Models on Aspect-Based Sentiment Analysis

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Across 13 datasets and 8 ABSA subtasks, efficiently fine-tuned LLMs outperform cited fine-tuned SLM baselines, and retrieval-based demonstration selection improves in-context learning for API models.

  2. Emotions in the Loop: A Survey of Affective Computing for Emotional Support

    cs.HC 2025-05 conditional novelty 1.0 of 10

    A 20-paper survey of affective computing for emotional support, organized into four domains, with a focus on ChatGPT's emotion recognition strengths and weaknesses.

Pith tools