Pith. sign in

REVIEW 3 major objections 6 minor 44 references

Human vs. LLM-Based Thematic Analysis for Digital Mental Health Research: Proof-of-Concept Comparative Study

T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper claims GPT-4o can match humans on parent-code development and saturate far faster, but cannot match human depth in child codes, excerpts, or themes.

desk verdict Useful empirical benchmark of LLM vs human thematic analysis, but its headline kappa and saturation numbers are misreported and must be corrected before the comparisons can be trusted. read the letter →

arxiv 2507.08002 v1 pith:X4OY6XAN submitted 2025-05-02 cs.HC cs.AI

classification cs.HCcs.AI
keywords largelanguagemodelsthematicanalysisqualitativecodingdigitalmentalhealthGPT-4osaturationhuman-AIcollaborationpromptengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tests whether GPT-4o, steered by the RISEN prompt framework, can do the work of a human qualitative researcher on interview transcripts from a digital mental health trial. It claims that the LLM produces deductive parent codes comparable to human codes and, in its out-of-the-box form, identifies about as many excerpts (417 versus 428) with strong inter-rater reliability (K = 0.84). The knowledge-based LLM reaches coding saturation after 10–15 transcripts and the out-of-the-box model after 15–20, whereas human researchers needed 90–99 transcripts to finalize their full coding structure. The same speed does not carry over to depth: humans generated 65 child codes and nine themes against the LLMs' 22 and six, and human excerpts were longer and more frequently multi-coded. The paper's conclusion is that LLM-based thematic analysis is dramatically cheaper and faster but less deep, and that a hybrid workflow — LLMs for initial coding and humans for refinement — is the recommended path.

What carries the argument

The machinery is a four-step thematic-analysis pipeline — code development, operational definition of codes, excerpt extraction and code application, and theme synthesis — executed by GPT-4o under the RISEN prompt framework (Role, Instructions, Steps, End-Goal, Narrowing). One variant runs the model out of the box; the other injects the standard six-phase thematic analysis methodology through retrieval-augmented generation, which retrieves relevant methodological guidance without retraining the model. The comparison is anchored by the saturation metric: the number of transcripts needed until no new codes appear, tracked by the LLMs in batches of five transcripts. This setup lets a single postdoctoral researcher complete the entire analysis in about 40 hours at $12.10 of API cost, against 110 hours and $3,537 of personnel cost for the human team.

What would settle it

Recompute both saturation points with the same stopping rule: have a human coder review transcripts in the same batches of five and report the first batch with no new child codes, and have the LLM review all 99 transcripts before declaring its final coding structure; if the gaps shrink to within 10–20 transcripts, the paper's saturation comparison is an artifact of the metric.

Watch

Extended reading notes

Core claim

The central discovery is a mixed result. With the RISEN prompt structure, GPT-4o can develop deductive parent codes that align with human-derived codes (four of its ten parent codes matched human code labels directly), and the out-of-the-box model extracts a comparable number of excerpts to human coders, with 73% of its excerpts fully embedded within longer human excerpts and an inter-rater reliability of K = 0.84 against human coding. Coding saturation arrives far earlier for the LLMs: 10–15 transcripts for the knowledge-based variant and 15–20 for out-of-the-box, versus 90–99 for humans. But the LLMs' codes and themes are systematically shallower — 22 child codes versus 65, six themes versus nine, mostly single-coded excerpts — and 44% of human-identified excerpts are missed by both LLMs. The knowledge-based LLM is more efficient on code development but markedly worse at excerpt extraction (251 excerpts versus 417 for out-of-the-box and 428 for humans). The paper frames the result as a proof that LLM thematic analysis is a viable, low-cost first pass, but not a substitute for human interpretation.

Load-bearing premise

The study's headline saturation gap rests on assuming that the human saturation point — the number of transcripts needed to finalize the full coding structure after reviewing all 99 — can be directly compared with the LLM saturation point — the batch size at which no new codes appear.

Editorial extensions

If this is right

  • The out-of-the-box LLM can cut the cost of coding a 20-transcript dataset from about $3,537 to roughly $1,272 in personnel plus $12 in API fees, making large-scale qualitative analysis affordable.
  • A knowledge-based LLM that reaches saturation in 10–15 transcripts could support rapid, iterative analysis in pilot feasibility trials where participant feedback must be incorporated quickly.
  • Human oversight should remain in the loop for child-code development and theme synthesis, since the LLMs miss 44% of human-selected excerpts and one-third of human themes.
  • The strong excerpt-level agreement (K = 0.84) suggests LLMs can serve as a reliable pre-coder, flagging candidate excerpts for human review rather than replacing the analyst.
  • The cost asymmetry means qualitative studies no longer need to cap sample sizes purely because of coding labor, as long as researchers accept coarser initial codes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The saturation comparison may not be apples-to-apples: human saturation was defined as the point at which the full coding structure across all 99 transcripts was finalized, whereas LLM saturation was the point at which no new codes appeared in successive five-transcript batches; these are different stopping rules, and the 10-15 vs 90-99 gap could partly reflect that difference.
  • LLM saturation may be artificially early because the model generates a coarser codebook (10 parent codes and 22 child codes vs 7 and 65); with a smaller codebook, new codes naturally stop appearing sooner, so a fairer comparison would track the number of new codes per transcript rather than the batch cutoff.
  • The knowledge-based LLM's inferior excerpt extraction (251 vs 417 excerpts) suggests that injecting methodological knowledge via RAG can over-constrain the model; a hybrid might use out-of-the-box extraction with knowledge-based code development.
  • A direct test would be to run the same thematic analysis on a separate corpus with human coders reporting saturation in five-transcript batches and with LLM codebooks expanded to match human granularity; if the gap persists, the efficiency claim is solid.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper reports a proof-of-concept comparison of human reflexive thematic analysis with two GPT-4o pipelines (an out-of-the-box model and a knowledge-base/RAG version incorporating Braun and Clarke's framework) on 99 semi-structured VR debrief interview transcripts from a healthcare-worker stress trial, with a random subset of 20 transcripts used for excerpt extraction and coding. The study compares code development, coding saturation, excerpt identification, theme synthesis, and cost. The authors find that the LLMs produce parent codes comparable to human-derived deductive codes, reach coding saturation in 10--20 transcripts versus 90--99 for humans, and are substantially cheaper ($12.10 API cost versus $3,537 personnel cost), but that human analysis produces richer child codes, longer multi-coded excerpts, and more nuanced themes; they recommend a hybrid human--AI approach.

Significance. If the empirical claims held, this would be a valuable practical benchmark for using LLMs in qualitative digital-mental-health research, using real clinical trial data and a transparent multi-stage comparison. The paper is commendable for reporting detailed frequency counts, a concrete cost ledger, explicit code-level overlap analyses, and an explicit limitations section, and for addressing a timely methodological question. However, the headline saturation comparison and the abstract's reliability statement are not supported by the methods as written, and the absence of run-to-run variability information limits the precision of the quantitative claims. With those issues corrected, the study would be a useful contribution to the emerging literature on human--AI collaboration in qualitative health research.

major comments (3)
  1. [Abstract; II.D; II.E.2; III.B] The headline claim that knowledge-based LLMs reached coding saturation with 10--15 transcripts versus 90--99 for humans compares two different quantities. For humans (II.D), saturation is indexed by the number of transcripts needed to finalize the entire coding structure across all 99 transcripts, and for distributed codes this required all 99 transcripts. For LLMs (II.E.2), saturation is defined as the point at which no new codes appear across batches of five transcripts, measured after Step 1 had already exposed the model to all 99 transcripts. The LLM number is a stability point in a post-hoc batch scan, not the number of transcripts needed to build and finalize the coding scheme. The abstract and conclusions present the two numbers as directly comparable. The paper should either re-analyze the human data under the same batch-stability rule or explicitly present these as distinct quantities and remove the direct 'versus' framing.
  2. [Abstract; II.D; III.C] The abstract attributes 'strong inter-rater reliability (K = 0.84)' to the out-of-the-box LLM, but II.D reports Cohen's κ = 0.84 as the agreement between two human coders, and III.C states that the LLMs 'could not replicate reliability measures.' No human--LLM inter-rater reliability is computed. The abstract and results should be corrected so that κ = 0.84 is not presented as validation of LLM outputs, and the absence of a human--LLM agreement metric should be explicitly listed as a limitation.
  3. [II.E; III.B] The quantitative saturation and excerpt-count comparisons rest on a single execution of each GPT-4o pipeline; the paper reports no repeated runs, temperature settings, or run-to-run variability. Because GPT-4o outputs are stochastic, the saturation boundaries (15--20 vs 10--15 transcripts) and the specific excerpt sets could change upon re-running. At minimum, the manuscript should acknowledge this; ideally it should report repeated-run agreement or variance. As written, the proof-of-concept cannot support the apparent precision of the saturation numbers.
minor comments (6)
  1. [II.E] The text says 'GPT-4o, am OpenAI transformer-based model'; 'am' should be 'an'.
  2. [III.A] The reported human labor hours are 92 (code development) + 10 (excerpt coding) + 10 (theme synthesis) = 112 hours, but the text states 110 hours; please reconcile the sum.
  3. [III.C] The sentence '73% were fully embedded within longer, more context-rich human excerpts (81%)' uses the parenthetical 81% without explaining its denominator; clarify whether it means 81% of human excerpts contained LLM excerpt text or some other base.
  4. [Abstract; III.B] The human saturation range '90--99' is not derived in the body text; III.B describes saturation within 20--30 transcripts for deductive codes and all 99 transcripts for distributed codes. Please state how the 90--99 composite range is computed or adjust the abstract.
  5. [References] Reference [6] misspells 'research' as 'reserach' and reference [28] misspells 'language' as 'langauge'; please correct the bibliographic entries.
  6. [IV.A] The phrase 'The GPT-4o structure enhanced the LLMs’ abilities' is imprecise; the intervention being compared is the RISEN prompt framework, not GPT-4o's architecture, so the sentence should be reworded accordingly.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper is a direct empirical benchmark against external human thematic analysis.

full rationale

This paper is a comparative empirical study, not a derivation. It runs three independent pipelines (human analysis in Dedoose, out-of-the-box GPT-4o, and knowledge-base GPT-4o with RAG) on the same interview transcripts and directly compares outputs: parent/child codes, saturation points, excerpts, themes, reliability, and cost. The comparison target, human thematic analysis, is external to the LLM pipeline, so the LLM outputs are not defined in terms of the study's conclusions. The knowledge-base variant retrieves Braun and Clarke's published framework via RAG; that is a methodological input, not a restatement of the paper's own findings, and no fitted parameter is renamed as a prediction. The inter-rater reliability value (Cohen's κ = 0.84) is measured against human coding and is not manufactured by the LLM design. The only substantive concern is that 'saturation' is operationalized differently for humans (number of transcripts needed to finalize the entire coding structure across all 99 transcripts; Methods D) versus LLMs (batch-stability after the model had already reviewed all 99 transcripts; Methods E.2 and Results B). That is a measurement-validity or comparability flaw in the headline '10-15 vs 90-99' contrast, but it is not circularity: each pipeline's saturation number is computed from its own stated stopping rule, and neither number is defined in terms of the other or in terms of the paper's conclusions. There are self-citations to the authors' prior DHMI-S trial as the data source, but they are not load-bearing derivational premises, and no uniqueness theorem, ansatz smuggled in by citation, or renaming of a known result is present. The paper's central claims therefore rest on independent empirical comparison, so the circularity burden is minimal.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

No mathematical free parameters are fitted; the only hand-chosen parameter that directly shapes a headline result is the five-transcript batch size used to detect LLM saturation. The central claims rest on the methodological assumptions listed above. No invented entities are introduced.

free parameters (1)
  • LLM saturation batch size = 5 transcripts
    LLMs reviewed transcripts in batches of five to detect when no new codes appeared (Methods E.2). The reported saturation points (10-15 and 15-20) are multiples of this batch size, so a different batch size could change the result. It is a design choice, not fitted to a target.
assumptions (3)
  • domain assumption Human reflexive thematic analysis, as operationalized by a single qualitative expert and two coders, is a valid gold-standard benchmark for LLM outputs.
    The study treats human coding and theme synthesis as the reference; no independent validation of the human coding itself is provided. Invoked throughout Results and Discussion.
  • domain assumption Braun and Clarke's four-step thematic analysis can be operationalized through the RISEN prompt framework with GPT-4o, including via RAG.
    Methods E assumes RISEN prompts plus Braun and Clarke text retrieved via RAG reproduce the intended qualitative methodology; no ablation or validation of this mapping is given.
  • ad hoc to paper A single run of GPT-4o is representative of the method's performance.
    Saturation and excerpt metrics are reported as point values from one pass per LLM variant; stochasticity and prompt sensitivity are acknowledged in Limitations but not quantified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Human vs. LLM-Based Thematic Analysis for Digital Mental Health Research: Proof-of-Concept Comparative Study." pith.science (2026). https://pith.science/paper/X4OY6XAN

@misc{pith2026250708002,
  author       = {Pith},
  title        = {Pith review of: Human vs. LLM-Based Thematic Analysis for Digital Mental Health Research: Proof-of-Concept Comparative Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X4OY6XAN}},
  note         = {Machine review of arXiv:2507.08002}
}
read the original abstract

Thematic analysis provides valuable insights into participants' experiences through coding and theme development, but its resource-intensive nature limits its use in large healthcare studies. Large language models (LLMs) can analyze text at scale and identify key content automatically, potentially addressing these challenges. However, their application in mental health interviews needs comparison with traditional human analysis. This study evaluates out-of-the-box and knowledge-base LLM-based thematic analysis against traditional methods using transcripts from a stress-reduction trial with healthcare workers. OpenAI's GPT-4o model was used along with the Role, Instructions, Steps, End-Goal, Narrowing (RISEN) prompt engineering framework and compared to human analysis in Dedoose. Each approach developed codes, noted saturation points, applied codes to excerpts for a subset of participants (n = 20), and synthesized data into themes. Outputs and performance metrics were compared directly. LLMs using the RISEN framework developed deductive parent codes similar to human codes, but humans excelled in inductive child code development and theme synthesis. Knowledge-based LLMs reached coding saturation with fewer transcripts (10-15) than the out-of-the-box model (15-20) and humans (90-99). The out-of-the-box LLM identified a comparable number of excerpts to human researchers, showing strong inter-rater reliability (K = 0.84), though the knowledge-based LLM produced fewer excerpts. Human excerpts were longer and involved multiple codes per excerpt, while LLMs typically applied one code. Overall, LLM-based thematic analysis proved more cost-effective but lacked the depth of human analysis. LLMs can transform qualitative analysis in mental healthcare and clinical research when combined with human oversight to balance participant perspectives and research resources.

Figures

Figures reproduced from arXiv: 2507.08002 by the authors.

Figure 1
Figure 1. Methodological process for the proof-of-concept comparison between traditional (human-based; red left-hand panel) and LLM approaches (out-of-the-box and knowledge-base; blue right-hand panel) to thematic analysis. Created in BioRender. Bhat, V. (2025) https://BioRender.com/3zeeh2d B. Data Source Semi-structured interview transcripts from the VR debrief component of the DHMI-S trial in healthcare workers [37, 38] wer… view at source ↗
Figure 4
Figure 4. Distribution of code outputs for human (red) and LLM (blue) approaches across dataset-specific research questions. Created in BioRender. Bhat, V. (2025) https://BioRender.com/ghyv39m [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Histograms depicting frequency counts of human, out-of-the￾box GPT-4o LLM, and knowledge-base GPT-4o LLMs. Panel A: Frequency comparison of codes and themes. Panel B: Frequency comparison of excerpts. Created in BioRender. Bhat, V. (2025) https://BioRender.com/tolmebo Overall, 238 human excerpts (56%) were also identified by the out-of-the-box LLM, while the knowledge-based LLM matched 124 human excerpts (29%). Out-… view at source ↗
Figures from the paper (1 more)
Figure 6
Figure 6. Figure 6: Themes generated by human researchers, out-of-the-box GPT-4o LLM, and a knowledge-base GPT-4o LLM with conceptual overlap indicated with bidirectional arrows. Created in BioRender. Bhat, V. (2025) https://BioRender.com/dzfaore IV. DISCUSSION This study compared the eff…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 22 canonical work pages

  1. [1]

    LLM-Based Thematic Analysis for Digital Mental Health Research: Proof-of-Concept Comparative Study (2025) Karisa Parkington*, Bazen G

    1 Human vs. LLM-Based Thematic Analysis for Digital Mental Health Research: Proof-of-Concept Comparative Study (2025) Karisa Parkington*, Bazen G. Teferra*, Marianne Rouleau-Tang, Argyrios Perivolaris, Alice Rueda, IEEE Member, Adam Dubrowski, Bill Kapralos, Reza Samavi, Andrew Greenshaw, Yanbo Zhang, Bo Cao, Yuqi Wu, Sirisha Rambhatla, Sridhar Krishnan, ...

  2. [2]

    Created in BioRender

    Methodological process for the proof-of-concept comparison between traditional (human-based; red left-hand panel) and LLM approaches (out-of-the-box and knowledge-base; blue right-hand panel) to thematic analysis. Created in BioRender. Bhat, V. (2025) https://BioRender.com/3zeeh2d B. Data Source Semi-structured interview transcripts from the VR debrief co...

  3. [3]

    Created in BioRender

    Description of the RISEN prompt engineering framework. Created in BioRender. Bhat, V. (2025) https://BioRender.com/pd1cnzr RISEN offers a structured yet flexible approach tailored to qualitative research, enabling targeted, goal-driven outputs. Compared to alternative strategies—few-shot prompting [42], which lacks adaptability; Chain-of-Thought [43], whi...

  4. [4]

    Created in BioRender

    Distribution of code outputs for human (red) and LLM (blue) approaches across dataset-specific research questions. Created in BioRender. Bhat, V. (2025) https://BioRender.com/ghyv39m Fig

  5. [5]

    Panel A: Frequency comparison of codes and themes

    Histograms depicting frequency counts of human, out-of-the-box GPT-4o LLM, and knowledge-base GPT-4o LLMs. Panel A: Frequency comparison of codes and themes. Panel B: Frequency comparison of excerpts. Created in BioRender. Bhat, V. (2025) https://BioRender.com/tolmebo Overall, 238 human excerpts (56%) were also identified by the out-of-the-box LLM, while ...

  6. [6]

    Created in BioRender

    Themes generated by human researchers, out-of-the-box GPT-4o LLM, and a knowledge-base GPT-4o LLM with conceptual overlap indicated with bidirectional arrows. Created in BioRender. Bhat, V. (2025) https://BioRender.com/dzfaore IV. DISCUSSION This study compared the effectiveness of out-of-the-box and knowledge-base GPT-4o LLM thematic analysis processes i...

  7. [7]

    Using thematic analysis in psychology,

    V. Braun and V. Clarke, "Using thematic analysis in psychology," Qualitative Research in Psychology, vol. 3, no. 2, pp. 77-101, 2006, doi: 10.1191/1478088706qp063oa

  8. [8]

    Pragmatic measures: what they are and why we need them,

    R. E. Glasgow and W. T. Riley, "Pragmatic measures: what they are and why we need them," American Journal of Prevention Medicine, vol. 45, no. 2, pp. 237-43, Aug 2013, doi: 10.1016/j.amepre.2013.03.010

Show all 44 references
  1. [9]

    Studying complexity in health services research: desperately seeking an overdue paradigm shift,

    T. Greenhalgh and C. Papoutsi, "Studying complexity in health services research: desperately seeking an overdue paradigm shift," BMC Medicine, vol. 16, no. 1, p. 95, Jun 20 2018, doi: 10.1186/s12916-018-1089-4

  2. [10]

    Qualitative methods in implementation research: An introduction,

    A. B. Hamilton and E. P. Finley, "Qualitative methods in implementation research: An introduction," Psychiatry Research, vol. 280, p. 112516, Oct 2019, doi: 10.1016/j.psychres.2019.112516

  3. [12]

    What Is Qualitative Research? An Overview and Guidelines,

    W. M. Lim, "What Is Qualitative Research? An Overview and Guidelines," Australasian Marketing Journal, pp. 1-31, 2024, doi: 10.1177/14413582241264619

  4. [14]

    Large Language Models Can Enable Inductive Thematic Analysis of a Social Media Corpus in a Single Prompt: Human Validation Study,

    M. S. Deiner, V. Honcharov, J. Li, T. K. Mackey, T. C. Porco, and U. Sarkar, "Large Language Models Can Enable Inductive Thematic Analysis of a Social Media Corpus in a Single Prompt: Human Validation Study," JMIR Infodemiology, vol. 4, p. e59641, Aug 29 2024, doi: 10.2196/59641

  5. [15]

    Best practices for implementing ChatGPT, large language models, and artificial intelligence in qualitative and survey-based research,

    J. Kantor, "Best practices for implementing ChatGPT, large language models, and artificial intelligence in qualitative and survey-based research," JAAD Int, vol. 14, pp. 22-23, Mar 2024, doi: 10.1016/j.jdin.2023.10.001

  6. [16]

    Exploring the Use of AI in Qualitative Analysis: A Comparative Study of Guaranteed Income Data,

    L. Hamilton, D. Elliott, A. Quick, S. Smith, and V. Choplin, "Exploring the Use of AI in Qualitative Analysis: A Comparative Study of Guaranteed Income Data," International Journal of Qualitative Methods, vol. 22, 2023, doi: 10.1177/16094069231201504

  7. [17]

    Qualitative and mixed methods in mental health services and implementation research,

    L. A. Palinkas, "Qualitative and mixed methods in mental health services and implementation research," Journal of Clinical Child and Adolescent Psychology, vol. 43, no. 6, pp. 851-61, 2014, doi: 10.1080/15374416.2014.910791

  8. [18]

    Available: https://arxiv.org/abs/2310.15100

    [Online]. Available: https://arxiv.org/abs/2310.15100

  9. [19]

    Comparing GPT-4 and Human Researchers in Health Care Data Analysis: Qualitative Description Study,

    K. D. Li et al., "Comparing GPT-4 and Human Researchers in Health Care Data Analysis: Qualitative Description Study," J Med Internet Res, vol. 26, p. e56500, Aug 21 2024, doi: 10.2196/56500

  10. [20]

    A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges,

    M. A. K. Raiaan et al., "A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges," IEEE Access, vol. 12, pp. 26839-26874, 2024, doi: 10.1109/access.2024.3365742

  11. [21]

    Job-Related Problems Prior to Nurse Suicide, 2003-2017: A Mixed Methods Analysis Using Natural Language Processing and Thematic Analysis,

    J. E. Davidson et al., "Job-Related Problems Prior to Nurse Suicide, 2003-2017: A Mixed Methods Analysis Using Natural Language Processing and Thematic Analysis," Journal of Nursing Regulation, vol. 12, no. 1, pp. 28-39, 2021, doi: 10.1016/s2155-8256(21)00017-x

  12. [22]

    Thematic analysis and natural language processing of job-related problems prior to physician suicide in 2003-2018,

    K. Kim, G. Y. Ye, A. M. Haddad, N. Kos, S. Zisook, and J. E. Davidson, "Thematic analysis and natural language processing of job-related problems prior to physician suicide in 2003-2018," Suicide and life-threatening behavior, vol. 52, no. 5, pp. 1002-1011, 2022, doi: 10.1111/...

  13. [23]

    Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?,

    W. S. Mathis, S. Zhao, N. Pratt, J. Weleff, and S. De Paoli, "Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?," Comput Methods Programs Biomed, vol. 255, p. 108356, Oct 2024, ...

  14. [24]

    Performing an Inductive Thematic Analysis of Semi-Structured Interviews With a Large Language Model: An Exploration and Provocation on the Limits of the Approach,

    S. De Paoli, "Performing an Inductive Thematic Analysis of Semi-Structured Interviews With a Large Language Model: An Exploration and Provocation on the Limits of the Approach," Social Science Computer Review, vol. 42, no. 4, pp. 997-1019, 2023, doi: 10.1177/08944393231220483

  15. [26]

    Can large language models understand context?,

    Y. Zhu et al., "Can large language models understand context?," arXiv, vol. 2402.00858 2024, doi: https://doi.org/10.48550/arXiv.2402.00858

  16. [27]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural Computation, vol. 9, pp. 1735-1780, 1997, doi: https://doi.org/10.1162/neco.1997.9.8.1735

  17. [29]

    Current applications and challenges in large language models for patient care: A systematic review,

    F. Busch et al., "Current applications and challenges in large language models for patient care: A systematic review," Communications in Medicine, vol. 5, no. 1, p. 26, Jan 21 2025, doi: 10.1038/s43856-024-00717-2

  18. [30]

    Qualitative data analysis in the age of artificial general intelligence,

    M. S. Abdüsselam, "Qualitative data analysis in the age of artificial general intelligence," International Journal of Advanced Natural Sciences and Engineering, vol. 7, no. 4, pp. 1-5, 2023, doi: 10.59287/ijanser.2023.7.4.454

  19. [31]

    Clinician voices on ethics of LLM integration in healthcare: a 10 thematic analysis of ethical concerns and implications,

    T. Mirzaei, L. Amini, and P. Esmaeilzadeh, "Clinician voices on ethics of LLM integration in healthcare: a 10 thematic analysis of ethical concerns and implications," BMC Med Inform Decis Mak, vol. 24, no. 1, p. 250, Sep 9 2024, doi: 10.1186/s12911-024-02656-3

  20. [33]

    Available: https://arxiv.org/abs/2311.14693

    [Online]. Available: https://arxiv.org/abs/2311.14693

  21. [34]

    AI–Human Hybrids for Marketing Research: Leveraging Large Language Models (LLMs) as Collaborators,

    N. Arora, I. Chakraborty, and Y. Nishimura, "AI–Human Hybrids for Marketing Research: Leveraging Large Language Models (LLMs) as Collaborators," Journal of Marketing, vol. 89, no. 2, pp. 43-70, 2025, doi: 10.1177/00222429241276529

  22. [35]

    Applications of large language models in psychiatry: A systematic review,

    M. Omar, S. Soffer, A. W. Charney, I. Landi, G. N. Nadkarni, and E. Klang, "Applications of large language models in psychiatry: A systematic review," Frontiers in Psychiatry, vol. 15, p. 1422807, 2024, doi: 10.3389/fpsyt.2024.1422807

  23. [36]

    Analyzing patient perspectives with large language models: a cross-sectional study of sentiment and thematic classification on exception from informed consent,

    A. E. Kornblith et al., "Analyzing patient perspectives with large language models: a cross-sectional study of sentiment and thematic classification on exception from informed consent," Sci Rep, vol. 15, no. 1, p. 6179, Feb 20 2025, doi: 10.1038/s41598-025-89996-w

  24. [37]

    Digital Interventions to Understand and Mitigate Stress Response: Protocol for Process and Content Evaluation of a Cohort Study,

    J. Martin et al., "Digital Interventions to Understand and Mitigate Stress Response: Protocol for Process and Content Evaluation of a Cohort Study," JMIR Res Protoc, vol. 13, p. e54180, May 6 2024, doi: 10.2196/54180

  25. [38]

    From bard to Gemini: An investigative exploration journey through Google’s evolution in conversational AI and generative AI,

    Z. Bin Akhtar, "From bard to Gemini: An investigative exploration journey through Google’s evolution in conversational AI and generative AI," Computing and Artificial Intelligence, vol. 2, no. 1, 2024, doi: 10.59400/cai.v2i1.1378

  26. [39]

    Created in BioRender

    and anonymization conventions. Created in BioRender. Bhat, V. (2025) https://BioRender.com/srhvxog C. Use-Case Research Questions VR debrief interviews from the DHMI-S trial [37, 38] were selected as use-case examples for this study because they captured detailed participant e...

  27. [40]

    Prompts, Pearls, Imperfections: Comparing ChatGPT and a Human Researcher in Qualitative Data Analysis,

    J. Wachinger, K. Barnighausen, L. N. Schafer, K. Scott, and S. A. McMahon, "Prompts, Pearls, Imperfections: Comparing ChatGPT and a Human Researcher in Qualitative Data Analysis," Qual Health Res, p. 10497323241244669, May 22 2024, doi: 10.1177/10497323241244669

  28. [41]

    The art of prompting: Unleashing the power of large language models,

    A. Thakur, "The art of prompting: Unleashing the power of large language models," preprint, doi: 10.13140/RG.2.2.18470.54089

  29. [42]

    Cloud application for managing, analyzing, and presenting qualitative and mixed method reserch data. (2024). SocioCultural Research Consultants, LLC, Los Angeles, CA, USA. [Online]. Available: https://www.dedoose.com

  30. [47]

    Available: https://arxiv.org/pdf/2305.10601

    [Online]. Available: https://arxiv.org/pdf/2305.10601

  31. [2015]

    Available: https://www.audiotranskription.de/wp-content/uploads/2020/11/manual-on-transcription.pdf

    [Online]. Available: https://www.audiotranskription.de/wp-content/uploads/2020/11/manual-on-transcription.pdf

  32. [2017]

    Available: https://arxiv.org/abs/1706.03762

    [Online]. Available: https://arxiv.org/abs/1706.03762

  33. [2020]

    Available: https://arxiv.org/pdf/2005.14165

    [Online]. Available: https://arxiv.org/pdf/2005.14165

  34. [2021]

    Available: https://arxiv.org/abs/2005.11401

    [Online]. Available: https://arxiv.org/abs/2005.11401

  35. [2023]

    Available: https://arxiv.org/abs/2310.18729

    [Online]. Available: https://arxiv.org/abs/2310.18729

  36. [2024]

    Available: https://arxiv.org/pdf/2405.08828

    [Online]. Available: https://arxiv.org/pdf/2405.08828

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.