Pith. sign in

REVIEW 4 cited by

Testing the Reliability of ChatGPT for Text Annotation and Classification: A Cautionary Remark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.11085 v1 pith:R7F4HWJN submitted 2023-04-17 cs.CL

classification cs.CL
keywords chatgptclassificationannotationtextreliabilityidenticaloutputsconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies have demonstrated promising potential of ChatGPT for various text annotation and classification tasks. However, ChatGPT is non-deterministic which means that, as with human coders, identical input can lead to different outputs. Given this, it seems appropriate to test the reliability of ChatGPT. Therefore, this study investigates the consistency of ChatGPT's zero-shot capabilities for text annotation and classification, focusing on different model parameters, prompt variations, and repetitions of identical inputs. Based on the real-world classification task of differentiating website texts into news and not news, results show that consistency in ChatGPT's classification output can fall short of scientific thresholds for reliability. For example, even minor wording alterations in prompts or repeating the identical input can lead to varying outputs. Although pooling outputs from multiple repetitions can improve reliability, this study advises caution when using ChatGPT for zero-shot text annotation and underscores the need for thorough validation, such as comparison against human-annotated data. The unsupervised application of ChatGPT for text annotation and classification is not recommended.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 31 citations worldwide. Full citation record

  1. Language Models Agree With Each Other, Not With Readers

    cs.IR 2026-07 accept novelty 7.0 of 10

    Across 18 model arms, model-model excess agreement (+0.093 median) is 2.3x human-human agreement (+0.040), against a naturalistic uninstructed reader baseline.

  2. Auditing Differential Visibility of Political Content on TikTok

    cs.SI 2026-07 conditional novelty 6.0 of 10

    Account-level analysis finds no evidence of moderate-to-large reach suppression on TikTok for three political topics; an apparent pooled gap is a statistical artifact, while oppositional content earns more engagement ...

  3. A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol

    cs.CL 2026-07 accept novelty 5.0 of 10

    The teaching-feedback classification protocol remains durable across three representation generations and transfers to English sentiment, so model choice is a deployment decision.

  4. Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications

    cs.IR 2025-07 reject novelty 4.0 of 10

    LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.

Pith tools