Pith. sign in

REVIEW 16 cited by

ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.06588 v1 pith:TWIGRFN7 submitted 2023-04-13 cs.CL cs.AIcs.SI

ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning

classification cs.CL cs.AIcs.SI
keywords accuracychatgpt-4messagestwitterbiasclassifierscrowdhigher
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This paper assesses the accuracy, reliability and bias of the Large Language Model (LLM) ChatGPT-4 on the text analysis task of classifying the political affiliation of a Twitter poster based on the content of a tweet. The LLM is compared to manual annotation by both expert classifiers and crowd workers, generally considered the gold standard for such tasks. We use Twitter messages from United States politicians during the 2020 election, providing a ground truth against which to measure accuracy. The paper finds that ChatGPT-4 has achieves higher accuracy, higher reliability, and equal or lower bias than the human classifiers. The LLM is able to correctly annotate messages that require reasoning on the basis of contextual knowledge, and inferences around the author's intentions - traditionally seen as uniquely human abilities. These findings suggest that LLM will have substantial impact on the use of textual data in the social sciences, by enabling interpretive research at a scale.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

    cs.CL 2026-06 unverdicted novelty 7.0

    Introduces GenAI agent framework for auditing personalization algorithms via synthetic accounts with fixed personas, applied to X post-2024 election showing amplification of toxic and right-leaning content varying by ...

  2. Structure Before Collapse: Transient semantic geometry in next-token prediction

    cs.LG 2026-06 unverdicted novelty 7.0

    Semantic geometry emerges transiently early in next-token prediction training before collapsing to Neural Collapse symmetry in synthetic settings with latent semantic factors.

  3. The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel

    cs.CY 2026-06 unverdicted novelty 7.0

    A panel of nine LLMs achieves Krippendorff's alpha of 0.86 for political position measurement in data-sparse regions, with added axis definitions improving agreement and disagreements revealing interpretive issues.

  4. SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models

    cs.CL 2026-04 unverdicted novelty 7.0

    SPAGBias reveals that LLMs form nuanced gender associations with specific urban micro-spaces that exceed real-world distributions and produce failures in planning and descriptive tasks.

  5. The Shrinking Lifespan of LLMs in Science

    cs.DL 2026-04 unverdicted novelty 7.0

    LLM adoption in science follows a compressing inverted-U trajectory where release year predicts time-to-peak and lifespan better than model attributes.

  6. Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement

    cs.CL 2026-07 conditional novelty 6.5

    Across five subjective tasks and five open-source LLMs, demographic prompting improves human agreement only for 1–3 high-signal, directionally coherent attributes and degrades under the full attribute set.

  7. When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering

    cs.AI 2026-07 conditional novelty 6.0

    LLM answers to HDL questions are often redundant and verbose; a task-aware multi-agent framework cuts redundancy by 37% and padding by 31% while raising judge-based quality scores.

  8. What Prediction Markets Can See: Market Formation, Settlement Legibility, and the Geography of Tradable Uncertainty in Africa and Latin America

    econ.GN 2026-06 unverdicted novelty 6.0

    Prediction market inventories for Africa and Latin America topics are shaped more by settlement legibility than by public salience, with sports and elections favored over conflicts.

  9. Interpretable Discriminative Text Representations via Agreement and Label Disentanglement

    cs.CL 2026-05 unverdicted novelty 6.0

    LFD discovers predictive text features via LLM contrastive proposals, cross-LLM Cohen's kappa screening, and residual held-out gain selection, matching baseline accuracy while achieving higher human agreement and lowe...

  10. Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest

    cs.CL 2026-04 unverdicted novelty 6.0

    LLMs show mixed results on authorship verification, post generation, and attribute inference from Twitter data, with new frameworks and user studies establishing benchmarks for these analytics tasks.

  11. Evaluating LLMs as Human Surrogates in Controlled Experiments

    cs.HC 2026-03 unverdicted novelty 6.0

    LLMs reproduce several directional effects from a human accuracy perception experiment but show inconsistent effect magnitudes and moderation patterns across models.

  12. Characterizing initial human-AI proof formalization workflows

    cs.AI 2026-06 unverdicted novelty 5.0

    A controlled user study and qualitative survey find that AI assistance raises formalization accuracy for math proofs, with users flexibly combining multiple tools while retaining oversight.

  13. Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation

    cs.AI 2025-09 conditional novelty 5.0

    Introduces PAS and FAS task abstractions plus the LLM-S^3 benchmark to evaluate LLMs on generating sociodemographic survey responses across 11 real datasets and multiple models.

  14. VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents

    cs.CL 2025-06 unverdicted novelty 5.0

    VIDEE introduces a human-in-the-loop system using Monte-Carlo Tree Search for task decomposition, executable pipeline generation, and LLM-based evaluation with visualizations to support non-expert text analytics.

  15. Improving Medical Communication using Rubric-Guided Counterfactual Recommendations

    cs.CL 2026-06 unverdicted novelty 4.0

    An LM-guided counterfactual pipeline recommends minimal ordinal changes to communication features like tone and actionability, yielding a mean +6.41% gain in predicted positive feedback under independent auditor models.

  16. LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

    cs.CL 2024-12 accept novelty 3.0

    A survey that organizes LLMs-as-judges research into functionality, methodology, applications, meta-evaluation, and limitations.