REVIEW 4 cited by
LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding
read the original abstract
Deductive coding is a widely used qualitative research method for determining the prevalence of themes across documents. While useful, deductive coding is often burdensome and time consuming since it requires researchers to read, interpret, and reliably categorize a large body of unstructured text documents. Large language models (LLMs), like ChatGPT, are a class of quickly evolving AI tools that can perform a range of natural language processing and reasoning tasks. In this study, we explore the use of LLMs to reduce the time it takes for deductive coding while retaining the flexibility of a traditional content analysis. We outline the proposed approach, called LLM-assisted content analysis (LACA), along with an in-depth case study using GPT-3.5 for LACA on a publicly available deductive coding data set. Additionally, we conduct an empirical benchmark using LACA on 4 publicly available data sets to assess the broader question of how well GPT-3.5 performs across a range of deductive coding tasks. Overall, we find that GPT-3.5 can often perform deductive coding at levels of agreement comparable to human coders. Additionally, we demonstrate that LACA can help refine prompts for deductive coding, identify codes for which an LLM is randomly guessing, and help assess when to use LLMs vs. human coders for deductive coding. We conclude with several implications for future practice of deductive coding and related research methods.
Forward citations
Cited by 4 Pith papers
-
Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications
In four years of Copilot tasks, students wrote mostly natural-language What comments, used more How comments for procedural constructs, and focused effort on verifying output rather than rewriting comments.
-
Uncovering the Internet's Hidden Values: An Empirical Study of Desirable Behavior Using Highly-Upvoted Content on Reddit
LLM analysis of highly-upvoted Reddit comments yields 64-72 macro/meso/micro values per year; existing prosocial measures capture only 18% on average while the method also recovers and extends prior qualitative taxonomies.
-
Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding
Temperature and persona variations shape consensus speed in LLM multi-agent coding but produce no robust accuracy gains over single agents on human-annotated tutoring transcripts.
-
Effects of Collaboration on the Performance of Interactive Theme Discovery Systems
The study introduces a framework and reports differences in consistency, cohesiveness, and correctness of themes produced under synchronous versus asynchronous collaboration across three interactive NLP tools.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.