{"id":"9747ccd5-8b95-46d4-9d43-6db11f935669","arxiv_id":"2506.06083","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A human-in-the-loop computational grounded theory framework that adds data exploration, validation, and theoretical sampling to topic-model based coding is introduced and illustrated on Reddit tutoring posts.","lead":"This paper proposes a three-phase framework for applying grounded theory to large social media datasets, combining human coding, topic models, and annotation checks. It demonstrates the framework on 52,000 Reddit posts from online tutors in the gig economy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing assumption is that LDA topics from the full 52K corpus can be validated by 160 posts from two subreddits; the Phase One comparison excludes two codes, uses no quantitative agreement measure, and is not representative, so the framework's trustworthiness claim is not yet established.","rationale":"The paper is strongest as a process contribution: it spells out three phases, gives detailed annotation guidance (Appendix A), provides term extraction tables, and documents an end-to-end case study. Those are real, and they support a conditional acceptance rather than rejection. My concern is not that the framework is incoherent; it is that the central trustworthiness claim rests on a single validation step whose design cannot support the weight placed on it. The reader identified essentially the same point; I agree. The concrete check would settle it: repeat the validation on a representative sample with all codes included, independent coders, and a quantitative agreement measure. If that check fails, the framework is still usable as a structured way to combine TM with GT, but the claim that it 'maintains the rigour of established GT methodologies' over big data would need to be narrowed. If it passes, the concern is resolved. I therefore recommend no change to the reader's CONDITIONAL verdict: the burden is on the authors to provide the missing validation evidence.","tokens_in":27451,"tokens_out":4690,"duration_ms":50465,"concrete_test":"Re-run the Phase One validation with independent researchers: (a) draw a stratified random sample of posts from all 18 subreddits (not just the two small ones), (b) have coders blind to LDA outputs produce GT codes using the same open-coding protocol, including the two abstract codes, and (c) map all resulting codes to the existing 13- and 17-topic LDA models using a pre-registered matching rule. Report precision, recall, and F1 per code and an overall inter-coder agreement measure for the mapping. If the abstract codes cannot be mapped, or if F1 is not substantially above the chance baseline, the reported 12/13 overlap is insufficient to validate LDA as a substitute for human reading on the full corpus.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the framework lets researchers analyze large qualitative datasets while preserving GT rigor—depends on treating LDA topics computed on the full 52K corpus as validated substitutes for human GT codes. In Phase One, this validation is performed by comparing LDA topics (13 and 17 topics) with 15 GT codes derived from ~160 posts in two randomly selected subreddits. That evidence is weaker than the framework requires for three reasons. (1) Two codes are excluded because 'LDA is not expected to model them' (Phase One, Codes 14 and 15); this post hoc exclusion makes the reported overlap a selected measurement. (2) The comparison is a qualitative count ('12 topics detected') with no confusion matrix, no inter-rater reliability, and no chance baseline; because the same researchers produced the GT codes and judged the overlap, the 12/13 figure is an interpretation rather than a measured agreement. (3) Two subreddits are not shown to be representative of the 18-subreddit corpus (e.g., platform-specific issues may dominate), so even a perfect overlap on this sample would not license full-corpus validation. The rest of the pipeline inherits this assumption: QDTM queries are built from these 'validated' topics, annotator agreement is fair (Fleiss' kappa 0.21–0.38), 27% of QDTM topics are excluded, and Phase Three hand-codes only top-10 posts per topic. If the LDA–GT correspondence is unreliable, the claim that computational coding substitutes for human reading over the full dataset is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a three-phase human-in-the-loop computational grounded theory (CGT) framework for large qualitative datasets. Phase One combines grounded-theory (GT) coding of a random subset of data with LDA topic modelling on the full corpus, using the comparison as concurrent validation. Phase Two applies a query-driven topic model (QDTM) with term expansion, followed by human annotation of topic coherence, labelling, and main-subtopic relatedness. Phase Three involves line-by-line hand-coding of representative documents, construction of higher-level categories, and computational theoretical sampling (sentiment analysis) to saturate core categories. The framework is illustrated through a Reddit case study of gig-economy tutors (52K posts, 18 subreddits), producing a substantive grounded theory centred on 'staying financially afloat' and 'persisting'. The paper argues that the framework maintains GT rigour while enabling analysis of data too large for manual coding alone.","tokens_in":27787,"tokens_out":4095,"duration_ms":39498,"significance":"The paper addresses a genuine methodological gap: existing CGT frameworks often omit explicit validation of unsupervised topic models and rarely implement theoretical sampling. The proposed framework is described in unusual operational detail, including annotation guidelines (Appendix A), term extraction tables (Appendix B), and extensive coding outputs (Appendices C and D), which strengthen reproducibility. The case study contributes substantive findings about an under-studied gig-economy population. However, the paper's central trustworthiness claim rests on the Phase One concurrent validation, and that validation is currently too weak to establish the claim. If the authors can strengthen or reframe this validation, the framework would be a valuable addition to the CGT literature.","major_comments":[{"comment":"The concurrent validation of LDA topics against GT codes is the load-bearing step for the paper's trustworthiness claim, but the reported 12-of-13 overlap is a selected and unquantified comparison. Two of the 15 GT codes (Codes 14 and 15) are excluded from comparison because 'LDA is not expected to model them', which is a post hoc selection; no confusion matrix, chance-adjusted agreement, or reliability statistic is reported; and the same researchers who produced the GT codes judged the overlap. Because QDTM queries in Phase Two are built from these 'validated' topics, the rest of the pipeline inherits this unmeasured agreement.","section":"Phase One (Data Exploration), Table 3"},{"comment":"The validation sample is not shown to represent the corpus. GT coding is performed on approximately 160 posts randomly drawn from only two of the 18 subreddits (GoGoKidTeach and Palfish), while LDA is run on the full 52K posts. The manuscript provides no analysis of whether these two subreddits are typical of the remaining 16, and platform-specific issues could dominate the codes. Without a representativeness argument or a multi-sample validation, the Phase One result cannot license full-corpus generalization.","section":"Phase One (Data Exploration), Case study data"},{"comment":"The reported inter-annotator agreement (Fleiss' kappa 0.21-0.38) is 'fair' at best, and the paper's fallback to 97% majority agreement does not address the fact that chance-adjusted agreement is low. Additionally, 27% of QDTM topics are excluded based on annotator disagreement or quality; the manuscript should explain how this exclusion could bias the remaining 55 topics and whether the excluded topics were distributed evenly across main topics. This matters because Phase Three hand-codes only the top-10 posts of the surviving topics.","section":"Phase Two (Human Evaluation of QDTM Topics), Table 4"},{"comment":"The theoretical sampling step selects 'Gratitude' and 'Realization' after inspecting the emotion frequencies in Figure 2, and then randomly samples 50 posts per emotion. While theoretical sampling is legitimately iterative and data-driven, the manuscript should report the full set of emotions considered and justify why only these two were sampled, in order to distinguish gap-driven sampling from cherry-picking of emotions that happen to confirm the emerging 'Redditing' category.","section":"Phase Three (Supporting Theoretical Sampling), Figure 2"}],"minor_comments":[{"comment":"The sentence 'Sometimes, these models even outperform humans' is vague; specify the task types and benchmarks being referenced.","section":"Introduction and Background"},{"comment":"The decision rule for excluding Codes 14 and 15 is stated only as a researcher judgment; a more principled criterion (e.g., based on whether a code is a communicative function rather than a topic) would improve replicability.","section":"Phase One (Data Exploration)"},{"comment":"The term extraction table lists terms from the 13-topic LDA even though the 17-topic model was eventually selected; clarify the role of the 13-topic model in the term curation process.","section":"Appendix B, Table 8"},{"comment":"The denominators for Stage One differ across tasks (12 topics for coherence and issue identification, 10 for relatedness to main topic); these should be stated explicitly in the table.","section":"Phase Two, Table 4"},{"comment":"The manuscript references Figure 1 and Appendices A-D; verify that all are included in the final version and that Figure 1 clearly labels the three phases.","section":"General presentation"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for Big Data & Society and the case study is genuinely useful. The main concern is the Phase One validation, which is load-bearing for the trustworthiness claim and is currently a qualitative, self-assessed overlap with post hoc exclusions. I would encourage the editor to request that the authors either add a quantitative agreement measure (e.g., a confusion matrix with chance baseline) or substantially reframe the claim so that Phase One is presented as 'exploratory topic discovery' rather than 'validation'. The low kappa values and the 27% topic exclusion rate also deserve more careful discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, well-documented human-in-the-loop framework for computational grounded theory, and the case study shows the whole pipeline can be run end-to-end. The catch is that the framework's main trust claim rests on a comparison between LDA topics on 52K posts and GT codes from ~160 posts in two subreddits, and that comparison is weaker than the paper acknowledges.\n\nWhat's actually new: the three-phase structure—random-subset open coding, LDA validation, QDTM hierarchical modeling with four-task annotator evaluation, and computational theoretical sampling—is a genuine integration not present in the four prior frameworks they compare. The paper also does a good job grounding each design choice in GT principles (constant comparison, memo-writing, saturation), and the annotation guidelines in Appendix A are specific enough for someone else to reuse. The case study yields a substantive, plausible theory of tutor experiences in the gig economy, and the theoretical sampling via sentiment analysis is a creative way to follow up on the 'Redditing' category.\n\nWhere it goes soft: the Phase One validation load-bears more than it can carry. Two codes are excluded because 'LDA is not expected to model them,' the overlap is reported as a qualitative 12-of-13 with no confusion matrix or chance baseline, and two subreddits are not shown to be representative of the 18-subreddit corpus. That makes the 12/13 figure an interpreted agreement, not a measured one. The internal loop is also real: the same research team chose the LDA topics, wrote the QDTM queries from those topics, and designed the annotator tasks around their own codes. Annotator agreement is fair (kappa 0.21–0.38), 27% of topics are dropped, and the sampling emotions are selected after seeing the sentiment results. None of these are fatal for a qualitative methods paper, but they do mean the 'maintaining GT rigour' claim runs ahead of the evidence.\n\nAlso, there are no executable artifacts—no code, no topic key, no data release—so another team cannot independently reproduce the QDTM runs or the annotation counts. The terms table helps, but it is not enough for full reproducibility.\n\nWho this is for: qualitative methodologists and computational social scientists who want a concrete, well-documented template for doing HITL topic modeling on large text corpora. It deserves serious peer review, not desk rejection. A revision should either strengthen the validation evidence (larger validation sample, a quantitative agreement measure, a sensitivity analysis on the excluded codes) or explicitly reframe the trust argument as resting on the whole HITL process rather than the single LDA–GT comparison.\n\nRecommendation: send to review with requests for that validation strengthening. I would not cite it in its current form, but I would revisit after revision.","headline":"A clearly written HITL integration for computational grounded theory, but the central LDA-vs-GT validation is too thin to carry the trustworthiness claim; worth reviewing with revisions.","tokens_in":28335,"tokens_out":2294,"would_cite":false,"duration_ms":24083,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a three-phase, human-in-the-loop framework that scales Computational Grounded Theory to big social data, demonstrated on 52,000 Reddit posts about gig-economy tutoring.","keywords":["computational grounded theory","human-in-the-loop","topic modelling","LDA","query-driven topic model","big social data","grounded theory","gig economy"],"falsifier":"Take a new random sample of posts from the same 18 subreddits, have independent researchers code them, and measure the agreement between those codes and the LDA topics quantitatively; if the overlap on this second sample falls well below the reported 12 of 13 on the first, or if two independent coding teams produce materially different codes, then the validation step does not establish that the topic model reliably substitutes for human reading.","tokens_in":27193,"feed_emoji":"🧵","tokens_out":6515,"duration_ms":56519,"temperature":0.7,"pith_summary":"This paper proposes a three-phase framework for Computational Grounded Theory (CGT) that lets social scientists analyze big qualitative datasets while keeping the core principles of traditional Grounded Theory. The framework begins with hand-coding a small random subset of the data, uses those codes to validate topic models run on the full dataset, then applies a query-driven hierarchical topic model to organise the corpus into main topics and subtopics, and finishes with interpretive line-by-line coding, constant comparison, and theory building by researchers. The authors test it on roughly 52,000 Reddit posts about gig-economy tutors, producing a substantive theory centred on tutors' persistence in staying financially afloat. The central contribution is a practical, trust-preserving route from unstructured text at scale to a grounded theory.","feed_headline":"Human-in-the-loop pipeline makes Grounded Theory scale to big data","feed_subtitle":"A three-phase CGT framework validates topic models against manual codes, then hand-codes only representative posts to build theory.","key_machinery":"The load-bearing mechanism is Query-Driven Topic Modelling (QDTM), a semi-supervised hierarchical topic model built on a Hierarchical Dirichlet Process. The researcher inputs curated query terms per topic; QDTM expands them using frequency-based extraction, KL-divergence-based extraction, and a relevance model with word embeddings, then organises the corpus into main topics and automatically inferred subtopics. Its tree-like structure is what lets the framework map main topics to focused codes and subtopics to sub-codes, supporting constant comparison and partially automating theoretical sampling. The framework's validation step, comparing hand-derived Grounded Theory codes with LDA topics, is the trust mechanism that carries the claim that the machine outputs can substitute for human reading at scale.","core_discovery":"The claim is that a human-in-the-loop pipeline can automate enough of Grounded Theory's coding work to make it scalable without sacrificing the methodology's rigour. Concretely, the paper argues that treating topic models as provisional codes, validating them against independently derived human codes, expanding them through a semi-supervised hierarchical model, and then reserving hand-coding for representative documents preserves the analytical steps that make a theory grounded. The pipeline is shown to work end to end on a large real-world corpus, yielding a theory in which tutors' core concern is financial survival and their core behaviour is persistence through Redditing, solving, strategising, and sometimes leaving.","pith_inferences":["A natural testable extension would measure the GT-LDA overlap quantitatively on several random samples of different sizes to establish how small a subset still yields stable validation, rather than relying on a single roughly 160-post sample.","The exclusion of abstract codes from validation suggests that LDA may systematically miss emotional and interactional dimensions of the data; future work could ask whether the framework's claims hold for studies whose core concerns are primarily affective.","The framework equates saturation with computational coverage of terms and subtopics; one could compare this with traditional theoretical sampling by having a researcher conduct purposive follow-up analysis after coding to see whether new categories still emerge.","If the validation step is the sole gate for trusting the topic model, its reliability across different platforms, languages, and corpus sizes is an open empirical question that the paper does not address."],"forward_implications":["Social scientists can now apply Grounded Theory's open, focused, and theoretical coding to corpora of tens of thousands of posts, with hand-coding confined to representative subsets.","The framework's use of automatically inferred subtopics provides a concrete way to implement constant comparison and theoretical saturation in big-data settings where returning to fieldwork is impossible.","Human evaluation of topic quality, issue identification, labelling, and main-subtopic relatedness, combined with majority voting and explicit exclusion criteria, offers a template for trust in computational coding.","The case study demonstrates that the framework produces domain-level findings, here a substantive theory of gig-economy tutor persistence, rather than only a methodological demonstration.","Because the framework is described as NLP-technique agnostic, it can in principle accommodate newer language models without changing the three-phase structure."],"supporting_citations":[{"why":"Defines Grounded Theory's open coding, constant comparison, core categories, and theoretical sampling, which the framework must preserve.","marker":"Glaser and Strauss (1967)"},{"why":"The prominent CGT framework this paper extends and critiques, providing the baseline that lacks validation and uses a single-layered topic model.","marker":"Nelson (2020)"},{"why":"Introduces QDTM, the query-driven hierarchical topic model that is the framework's main computational machinery.","marker":"Fang et al. (2021)"},{"why":"Shows statistical topic quality measures correlate poorly with human judgments, motivating the human-in-the-loop evaluation approach.","marker":"Chang et al. (2009)"},{"why":"Provides the empirical method used to compare LDA models with different numbers of topics in the case study.","marker":"Quinn et al. (2010)"},{"why":"Supplies the complementary empirical approach for selecting the LDA K value.","marker":"Grimmer (2010)"},{"why":"Provides the Hierarchical Dirichlet Process that lets QDTM automatically infer the number of subtopics.","marker":"Teh et al. (2004)"},{"why":"Offers the term-based topic coherence evaluation alternative that the framework sets aside in favour of document-based human evaluation.","marker":"Mimno et al. (2011)"}],"fun_headline_variants":["Human-validated topic models scale Grounded Theory","Grounded Theory at big data scale with human checks","CGT framework: human-in-the-loop for rigorous big data analysis","Scalable Grounded Theory via human-validated topic models","Automated coding, human validation: Grounded Theory at scale"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework's trustworthiness rests on the assumption that a randomly selected roughly 160-post subset of the corpus, hand-coded by the researchers, is representative enough that LDA topics matching those codes across the full 52,000 posts can be taken as proof that the topic model substitutes for human reading.","fun_headline_variants_meta":{"raw":{"variants":["Human-validated topic models scale Grounded Theory","Grounded Theory at big data scale with human checks","CGT framework: human-in-the-loop for rigorous big data analysis","Scalable Grounded Theory via human-validated topic models","Automated coding, human validation: Grounded Theory at scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001171,"raw_usage":{"total_tokens":4811,"prompt_tokens":879,"completion_tokens":3932,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":3849}},"tokens_in":495,"tokens_out":3932,"duration_ms":25913,"temperature":1.0,"reasoning_tokens":3849,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T06:00:01.132816+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a new random sample of posts from the same 18 subreddits, have independent researchers code them, and measure the agreement between those codes and the LDA topics quantitatively; if the overlap on this second sample falls well below the reported 12 of 13 on the first, or if two independent coding teams produce materially different codes, then the validation step does not establish that the topic model reliably substitutes for human reading.","supporting_citations":[],"review_version":1}