Pith. sign in

REVIEW 1 cited by

Investigation of the Inter-Rater Reliability between Large Language Models and Human Raters in Qualitative Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.14764 v2 pith:HP3OGVSN submitted 2025-08-20 physics.ed-ph

Investigation of the Inter-Rater Reliability between Large Language Models and Human Raters in Qualitative Analysis

classification physics.ed-ph
keywords humanreliabilityanalysisinter-raterqualitativeagreementdemonstratefindings
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Qualitative analysis is typically limited to small datasets because it is time-intensive. Moreover, a second human rater is required to ensure reliable findings. Artificial intelligence tools may replace human raters if we demonstrate high reliability compared to human ratings. We investigated the inter-rater reliability of state-of-the-art Large Language Models (LLMs), ChatGPT-4o and ChatGPT-4.5-preview, in rating audio transcripts coded manually. We explored prompts and hyperparameters to optimize model performance. The participants were 14 undergraduate student groups from a university in the midwestern United States who discussed problem-solving strategies for a project. We prompted an LLM to replicate manual coding, and calculated Cohen's Kappa for inter-rater reliability. After optimizing model hyperparameters and prompts, the results showed substantial agreement (${\kappa}>0.6$) for three themes and moderate agreement on one. Our findings demonstrate the potential of GPT-4o and GPT-4.5 for efficient, scalable qualitative analysis in physics education and identify their limitations in rating domain-general constructs.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Me and My Bot: What Users Talk About in AI Companion Communities on Reddit

    cs.HC 2026-08 conditional novelty 5.0

    Across eight AI companion subreddits, only about 45% of retrieved posts concerned the user's own bot, own-bot posts were markedly more emotional than general-bot posts, and companionship and romance dominated relation...