Pith. sign in

REVIEW 23 cited by

LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.14924 v1 pith:BYXUGBTD submitted 2023-06-23 cs.CL cs.AIcs.LGstat.AP

classification cs.CLcs.AIcs.LGstat.AP
keywords codingdeductivelacaanalysiscontentgpt-3languagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deductive coding is a widely used qualitative research method for determining the prevalence of themes across documents. While useful, deductive coding is often burdensome and time consuming since it requires researchers to read, interpret, and reliably categorize a large body of unstructured text documents. Large language models (LLMs), like ChatGPT, are a class of quickly evolving AI tools that can perform a range of natural language processing and reasoning tasks. In this study, we explore the use of LLMs to reduce the time it takes for deductive coding while retaining the flexibility of a traditional content analysis. We outline the proposed approach, called LLM-assisted content analysis (LACA), along with an in-depth case study using GPT-3.5 for LACA on a publicly available deductive coding data set. Additionally, we conduct an empirical benchmark using LACA on 4 publicly available data sets to assess the broader question of how well GPT-3.5 performs across a range of deductive coding tasks. Overall, we find that GPT-3.5 can often perform deductive coding at levels of agreement comparable to human coders. Additionally, we demonstrate that LACA can help refine prompts for deductive coding, identify codes for which an LLM is randomly guessing, and help assess when to use LLMs vs. human coders for deductive coding. We conclude with several implications for future practice of deductive coding and related research methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. My Favorite Streamer is an LLM: Discovering, Bonding, and Co-Creating in AI VTuber Fandom

    cs.HC 2025-09 conditional novelty 7.0 of 10

    Fans of AI VTuber Neuro-sama treat her as a consistent persona, not a human mimic, and use paid SuperChats to co-create stream content, creating a distinct participatory fan economy.

  2. Design-Based Supervised Learning with Noisy Human Labels

    stat.ML 2026-07 conditional novelty 6.0 of 10

    A nested design-based correction that combines noisy human audit labels with partial expert adjudication yields unbiased downstream estimates and better efficiency than using only adjudicated labels.

  3. Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications

    cs.SE 2026-07 conditional novelty 6.0 of 10

    In four years of Copilot tasks, students wrote mostly natural-language What comments, used more How comments for procedural constructs, and focused effort on verifying output rather than rewriting comments.

  4. The Role of Partisan Culture in Mental Health Language Online

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Partisan culture is associated with measurable differences in how Republican and Democrat Reddit users express distress, with Democrats using more clinical and polarization-related language and Republicans using more ...

  5. SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.

  6. Simulating Ethics: Using LLM Debate Panels to Model Deliberation on Medical Dilemmas

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Two AI ethics debates with differently composed panels reached the same policy recommendation but through different arguments and coalitions, showing that panel membership can shift reasoning even when the facts are fixed.

  7. ChatCollab: Exploring Collaboration Between Humans and AI Agents in Software Teams

    cs.HC 2024-12 conditional novelty 6.0 of 10

    ChatCollab lets humans and AI agents work side-by-side in software teams and offers a way to measure how those teams collaborate.

  8. An Empirical Examination of the Evaluative AI Framework

    cs.HC 2024-11 conditional novelty 6.0 of 10

    A pre-registered experiment found that an AI providing only pro and con evidence, without recommendations, did not improve decision performance and was used shallowly by participants.

  9. A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health

    cs.LO 2026-07 conditional novelty 5.0 of 10

    Ethical rules for financial digital phenotyping can be written as deontic temporal constraints whose violations Z3 proves unsatisfiable inside the formal model.

  10. Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

    cs.CY 2026-06 conditional novelty 5.0 of 10

    Across 25,000 stories from five LLMs, an LLM judge rated stories mentioning intellectual disabilities as more infantile, paternalistic, dependent, and inspirational than stories without the label.

  11. EchoAid: Enhancing Livestream Shopping Accessibility for the DHH Community

    cs.HC 2025-08 conditional novelty 5.0 of 10

    EchoAid combines speech-to-text, LLM summarization, and rapid serial visual presentation to help deaf and hard of hearing users follow livestream shopping shows; user studies show lower cognitive load and modestly bet...

  12. From Assistance to Autonomy -- A Researcher Study on the Potential of AI Support for Qualitative Data Analysis

    cs.CY 2025-01 conditional novelty 5.0 of 10

    Interviews with 15 HCI researchers show openness to AI in qualitative data analysis under conditions of privacy, control, and reliability, leading to a framework of AI involvement levels from minimal to high.

  13. Optimizing Code Runtime Performance through Context-Aware Retrieval-Augmented Generation

    cs.SE 2025-01 conditional novelty 5.0 of 10

    An LLM code optimizer using control-flow-graph differences and retrieved examples reports 7.3% average runtime reduction on 116 C++ programs versus zero-shot GPT-4o.

  14. LeMo: Enabling LEss Token Involvement for MOre Context Fine-tuning

    cs.CL 2025-01 conditional novelty 5.0 of 10

    LeMo reduces long-context fine-tuning memory by eliminating low-informativeness tokens, predicting sparsity patterns, and optimizing kernels, while keeping perplexity close to LoRA.

  15. Interpreting Language Reward Models via Contrastive Explanations

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Reward model preferences can be explained by generating counterfactual and semifactual answer variations along 15 hand-picked evaluation attributes and measuring which attribute changes flip the model's preference.

  16. Conversational AI for Rapid Scientific Prototyping: A Case Study on ESA's ELOPE Competition

    cs.AI 2026-01 conditional novelty 4.0 of 10

    One engineer paired with ChatGPT and reached second place in ESA's ELOPE competition in about one week of work; the paper draws best-practice lessons from that experience.

  17. A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks

    cs.LG 2025-08 reject novelty 4.0 of 10

    A risk score built from confidence and vote entropy is proposed for triaging LLM qualitative coding, but the central R-squared=0.979 result is inflated by the definitions of agreement and diversity.

  18. From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering

    cs.SE 2025-04 conditional novelty 4.0 of 10

    GPT-4 labels software requirements with substantial agreement to human analysts (Cohen's Kappa up to 0.738) when given detailed few-shot prompts, while zero-shot performance is only moderate.

  19. Concept Navigation and Classification via Open-Source Large Language Model Processing

    cs.CL 2025-02 reject novelty 4.0 of 10

    A pipeline combining LLM summarization, iterative category generation, and human-in-the-loop refinement achieves frame and topic classification accuracy comparable to human coders on three text corpora.

  20. A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.

  21. Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery

    cs.CR 2025-07 reject novelty 3.0 of 10

    Scout applies off-the-shelf LLMs and vision models to triage digital evidence, but only anecdotal examples are shown and accuracy is withheld.

  22. Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI

    cs.SE 2025-05 conditional novelty 3.0 of 10

    A qualitative taxonomy positions vibe coding and agentic coding as complementary paradigms rather than rivals in AI-assisted software development.

  23. A Framework for LLM-powered Design Assistants

    cs.HC 2025-02 conditional novelty 2.0 of 10

    LLMs are organized into a three-modality framework for design assistance, but no evidence is provided that the framework works.

Pith tools