{"id":"592df962-8156-40f5-bc32-7a513e542ac9","arxiv_id":"2507.00547","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper presents an STM-based illustration and eight rigour guidelines for topic modelling, but does not empirically validate the guidelines.","lead":"This paper proposes eight guidelines for applying topic modelling algorithms with methodological rigour, illustrated through a structural topic model analysis of 1,520 blockchain research abstracts. The paper is aimed at researchers, reviewers, and editors who need to judge whether computationally intensive text analysis is done rigorously.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Model selection rests on a small human evaluation with no inter-coder reliability; K=35's better topic log odds means the K=20 justification is not robust.","rationale":"Read in good faith, the paper's central claim is not an empirical discovery but a methodological illustration: it aims to show how to achieve rigour in topic modelling and to offer generalisable guidelines. For that claim to hold, the illustrative application itself must be defensible. The weakest point is therefore the human evaluation in Section 3.4 that selects the final model. The evaluation uses two coders, one of whom is the author, only 10 topic intrusion cases per model, and no inter-coder reliability statistic. The reported metrics are also mutually inconsistent: K=35 has better topic log odds (-1.01 vs -1.09) while K=20 has better precision (0.68 vs 0.50), so the claimed justification for choosing K=20 is not transparently quantitative. Since the paper explicitly advises balancing accuracy and interpretability and designing context-specific evaluation tasks, the illustrative evaluation should meet a higher bar; it currently does not. This concern is the same one the reader identified, and it supports a CONDITIONAL verdict rather than ACCEPT. I do not see an internally inconsistent or fatal flaw in the guidelines themselves, so no stronger verdict change is warranted.","tokens_in":14043,"tokens_out":4646,"duration_ms":55290,"concrete_test":"Run a larger blind evaluation using the same word and topic intrusion tasks, with at least 50 topic-intrusion cases per model and two coders blind to model identity; report per-coder agreement (e.g., Cohen's kappa) and exact binomial 95% confidence intervals for model precision. If K=35 ties or beats K=20 on either metric, or if inter-coder agreement is low, the stated basis for selecting K=20 is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central demonstration of rigour depends on the model-selection evaluation in Section 3.4. Table 1 reports word intrusion precision and topic log odds from two coders (one the author) with only 10 topic-intrusion cases per model and no reported inter-coder reliability. The chosen model K=20 has the highest precision (0.68) but not the best topic log odds (K=35 scores -1.01 vs -1.09). No uncertainty intervals are given; with roughly 20 topic-intrusion observations per model, a log-odds gap of 0.08 is within noise, and the precision gap (0.68 vs 0.50) may also overlap once binomial uncertainty is accounted for. Because the paper offers this STM application as the worked illustration of rigour and builds its guidelines on it, a selection step that is not demonstrably robust undermines the central illustrative claim, even if the guidelines themselves are reasonable. The generalisation to other ML algorithms is a further assertion, but it is the weaker link only through this illustration.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes that methodological rigour in topic modelling can be established through a three-stage iterative process—data preparation, model building, and pattern evaluation—and presents eight guidelines (Table 3) for achieving that rigour. To illustrate the approach, the author applies the Structural Topic Model (STM) to 1520 academic abstracts about blockchain, describing decisions at each stage: preprocessing choices, hyperparameter selection via diagnostic metrics (held-out likelihood, residuals, semantic coherence, exclusivity), human evaluation using word and topic intrusion tasks, and topic labelling. The paper further claims that the guidelines can generalize to other machine learning algorithms with context-specific adjustments. The intended contribution is methodological guidance for novice researchers, editors, and reviewers, rather than an empirical finding about blockchain or topic models.","tokens_in":14248,"tokens_out":4607,"duration_ms":57057,"significance":"If the illustration is accepted as rigorous, the paper offers a useful pedagogical checklist and a worked example that consolidates scattered practices from existing tutorials (e.g., Debortoli et al. 2016; Schmiedel et al. 2019) and from the STM literature. Its stage-linked guidelines and emphasis on transparency are well aligned with current discussions of computational rigour in IS research. However, the load-bearing demonstration of rigour rests on a model-selection evaluation that is currently too thin: two coders (one the author), ten topic-intrusion cases per model, no inter-coder reliability, and no uncertainty quantification. The chosen model (K=20) is not clearly superior on the reported metrics. The generalizability claim to other ML algorithms is asserted rather than argued. These weaknesses do not invalidate the guideline content, but they prevent the paper from fully making its central claim that the illustration itself exemplifies rigour.","major_comments":[{"comment":"The model selection decision is the crucial step in the illustration, yet it rests on a small and potentially unreliable human evaluation. Two coders (one being the author), ten topic-intrusion cases per model, no inter-coder reliability statistic, and no confidence intervals are reported. The selected model, K=20, has the highest model precision (0.68), but K=35 has better topic log odds (-1.01 vs -1.09). With roughly 20 topic-intrusion judgments per model, the precision gap between 0.68 and 0.50 is plausibly within binomial sampling error, and the 0.08 log-odds difference is negligible. As reported, the preference for K=20 is not robust. To support the paper's claim that the illustration demonstrates methodological rigour, please report inter-coder agreement (e.g., Cohen's kappa), quantify uncertainty (e.g., confidence intervals or a binomial test), increase the number of intrusion cases or justify the sample size, and/or pre-specify a decision rule for selecting K. Without this, the central illustrative example does not convincingly establish the rigour it advocates.","section":"Section 3.4, Table 1"},{"comment":"The claim that the guidelines 'can be applied to other algorithms with context-specific adjustments' is not substantiated. The paper provides a single illustration using STM on one corpus of blockchain abstracts, and the Discussion and Conclusion acknowledge that future work is needed to develop guidelines for other algorithm families. While it is plausible that some guidelines (e.g., 'report computational procedures') are algorithm-agnostic, the transfer of the data-preparation and model-building guidelines to supervised ML or deep learning requires argument, not just assertion. Either soften the generalizability claim to a hypothesis or add a short rationale explaining which principles are algorithm-independent and why.","section":"Abstract; Section 4; Section 5"},{"comment":"The shortlisting of candidate models is described qualitatively: the text states that models 10, 15, and 20 have the highest held-out likelihood, that residuals are lower from 35 to 60, and that models 35 and 45 were included because they have 'relatively higher' held-out likelihood and semantic coherence. No explicit decision rule or table of all twelve diagnostic values is provided. For a paper whose central theme is rigour, the model-selection criteria should be transparent and reproducible: for example, a table of all models with their held-out likelihood, residual, semantic coherence, and exclusivity values, and a clear statement of the threshold or ranking procedure used to discard K=5 and to retain K=35 and K=45.","section":"Section 3.3"}],"minor_comments":[{"comment":"The sentence 'For the topic labelling task, an external coder with the five most probable words and the ten most probable article abstracts corresponding to each of the top 10 topics' lacks a main verb; it should read 'was provided with' or 'was presented with'.","section":"Section 3.4"},{"comment":"There is inconsistent spelling between 'Rigor' (Table 3 header) and 'Rigour' (title and text). Choose one spelling convention and apply it consistently.","section":"Table 3 and throughout"},{"comment":"In the version under review, Figures 2 and 3 are not rendered properly; the axis labels and tick values appear as raw text interspersed with plotting instructions, making it impossible to verify the claims about held-out likelihood, semantic coherence, and exclusivity. Please ensure the figures are embedded as high-resolution images.","section":"Figure 2 and Figure 3"},{"comment":"The phrase 'I did not use the lower bound metric' should briefly define what the lower bound is (the variational lower bound) and why it is relevant to model accuracy, since many readers will not be familiar with STM diagnostics.","section":"Section 3.3"},{"comment":"The sentence 'There exists a plethora of algorithms to date' is awkward; consider revising to 'There is a plethora of algorithms available today' or similar.","section":"Section 4.4.2"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is an ACIS 2024 conference proceedings paper, and its contribution is primarily pedagogical/methodological rather than an empirical discovery. The guideline content is reasonable and could be useful to a practitioner audience, but the worked example—which is the paper's central evidence for its claims—needs stricter empirical reporting (inter-coder reliability, uncertainty, and transparent selection criteria). The generalizability claim also overreaches relative to the evidence. These issues are fixable within the scope of the paper, hence major_revision rather than reject. The fit with cs.CL is somewhat borderline; the paper is more in the IS methods tradition than in computational linguistics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a methods/teaching paper, not a research paper. The actual contribution is a worked STM application with eight guidelines for rigorous topic modelling. The guidelines aren't new—justify preprocessing, understand hyperparameters, balance accuracy and interpretability, use human evaluation, report transparently—all appear in the cited tutorials and in Chang et al. But it does a real service in making those explicit for IS researchers and showing them in one coherent workflow. The data preparation narrative is honest: the author discusses why stemming wasn't used, how collocations were handled, and how searchK diagnostics were interpreted. That kind of concrete transparency has pedagogical value.\n\nThe soft spot is exactly where the stress-test hits: the selection of K=20 over K=35 rests on a human evaluation of 10 topic-intrusion cases per model, two coders, one being the author, and no inter-coder reliability. The topic log-odds are actually slightly worse for K=20 (-1.09 vs -1.01 for K=35), and with 20 observations the gap is noise. The model precision difference (0.68 vs 0.50) is also within binomial uncertainty. The paper doesn't report uncertainty or sensitivity. Because the whole illustration is meant to demonstrate rigour, a selection step that isn't demonstrably robust undermines the central claim. Also, no code or data are provided, despite the \"report to help build the community\" guideline—an irony the author doesn't address.\n\nMinor: the generalisation to other ML algorithms is asserted, not argued. And the novelty claim about filling a gap in \"explicit methodological rigour\" is weak; the cited tutorials already cover much of this. The paper is honest about its scope limitations, though it doesn't flag the evaluation issue.\n\nVerdict: solid teaching material, not a substantive methodological advance. A serious editor should send it to review before publishing—it needs a major revision on the evaluation step, or a re-framing of the contribution as a tutorial rather than a new rigour framework. But it deserves referee time; it's not a desk reject.","headline":"A readable, honest methods tutorial whose model-selection step doesn't live up to its own rigour message.","tokens_in":14712,"tokens_out":1925,"would_cite":false,"duration_ms":22423,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that topic modelling studies become trustworthy when researchers follow eight guidelines spanning data preparation, model building, and pattern evaluation, and it demonstrates these guidelines on a corpus of blockchain…","keywords":["topic modelling","methodological rigour","structural topic model (STM)","computationally intensive theory construction","interpretability","transparency","human evaluation","guidelines"],"falsifier":"A replication with more coders and reported inter-coder reliability that found the 20-topic model no more interpretable than the 35-topic model (which had better topic log odds), or found the intrusion-task scores unstable, would weaken the paper's claim that its demonstration establishes rigour, even though the guidelines themselves could still stand.","tokens_in":13837,"feed_emoji":"📋","tokens_out":4143,"duration_ms":40240,"temperature":0.7,"pith_summary":"The paper tries to establish that methodological rigour in topic modelling is something a researcher can deliberately do, not merely aspire to. It offers eight practical guidelines organised around three stages—data preparation, model building, and pattern evaluation—and demonstrates them on a structural topic model built from 1,520 blockchain article abstracts. The author argues that algorithmic opacity and unexamined defaults undermine trust in research, and that explicit justification of every decision restores that trust. A sympathetic reader would care because the guidelines give novice researchers, editors, and reviewers a concrete checklist for judging whether topic-modelling patterns are plausible enough to support theory. The paper's contribution is guidance, not a new empirical discovery.","feed_headline":"Eight guidelines make topic-modelling research trustworthy","feed_subtitle":"A worked STM study of 1,520 blockchain abstracts shows how to justify every step.","key_machinery":"The load-bearing object is the three-stage iterative process shown in Figure 1—data preparation, model building, and pattern evaluation—together with the eight-guideline checklist in Table 3. The machinery that carries the argument includes STM's diagnostic metrics (held-out log-likelihood, residuals, semantic coherence, and exclusivity) used to shortlist candidate topic counts, and the word- and topic-intrusion evaluation tasks that inject an intruder word or intruder topic for human coders to detect. The worked illustration on 1,520 blockchain abstracts is what turns the abstract idea of rigour into concrete, reviewable steps.","core_discovery":"The central claim is that rigour in topic modelling comes from systematically justifying decisions at every stage of an iterative three-stage process, rather than from following standard procedures as fixed recipes. The paper illustrates this by applying the structural topic modelling (STM) algorithm to blockchain abstracts: it justifies data-preparation choices, compares twelve candidate models through diagnostic metrics such as held-out likelihood, residuals, semantic coherence, and exclusivity, and then uses word- and topic-intrusion tasks plus a label-quality evaluation by two coders to settle on a 20-topic model. The final contribution is a set of eight guidelines, six tied to the three stages and two about reporting computational procedures, which the author argues can generalise to other machine learning algorithms with context-specific adjustments.","pith_inferences":["The eight guidelines could be converted into a structured reporting checklist or preregistration template for topic-modelling studies, making adherence easier to verify.","The human-evaluation step would be stronger with more coders and reported inter-coder reliability; without those, the demonstration shows the method but not its stability.","The same three-stage logic could be adapted to supervised learning or large-language-model research by substituting evaluation tasks tailored to those outputs, such as calibration checks for predictions.","A testable extension would be to compare studies that follow these guidelines against studies that do not, measuring whether readers or reviewers judge the former as more trustworthy."],"forward_implications":["If the guidelines are right, researchers can structure a topic-modelling study as a series of justified choices, and reviewers can demand those justifications.","The reporting guidelines imply that a topic-modelling paper is not complete until its computational procedure is transparent enough for replication.","The selection of the 20-topic model shows that combining statistical diagnostics with human evaluation can make the number-of-topics decision defensible.","Because the guidelines are not tied to one algorithm, they can be adapted to other machine learning algorithms with context-specific adjustments.","Novice researchers gain a starting template that replaces copying prior studies' defaults with context-aware decisions."],"supporting_citations":[{"why":"Provides the stm R package, hyperparameter guidance, and the spectral initialization recommendation that ground Stage II of the illustration.","marker":"(Roberts et al. 2019)"},{"why":"Establishes that STM is suited to short texts such as abstracts and open-ended survey responses, justifying the algorithm choice.","marker":"(Roberts et al. 2014)"},{"why":"Supplies the word- and topic-intrusion tasks and the model precision and topic log-odds metrics used for the human evaluation in Stage III.","marker":"(Chang et al. 2009)"},{"why":"Documents the lack of rigour and context-blind reuse of practices in topic-modelling studies that motivates the need for the guidelines.","marker":"(Günther and Joshi 2020)"},{"why":"Provides the process schematic that the paper simplifies into its three-stage iterative process.","marker":"(Shmueli and Koppius 2011)"},{"why":"Defines probabilistic topic models and the tension between predictive accuracy and interpretability that shapes the guidelines.","marker":"(Blei 2012)"},{"why":"Represents an existing tutorial that the paper positions as helpful but lacking explicit discussion of methodological rigour.","marker":"(Debortoli et al. 2016)"}],"fun_headline_variants":["Eight guidelines make topic modeling research trustworthy","Justify every step: eight guidelines for topic modeling","A worked STM example yields eight rigor guidelines","Keep topic modeling transparent with eight clear rules","One study, eight rules: rigorous topic modeling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the human evaluation of interpretability—two coders on ten topic-intrusion cases per model, with the author among the coders and no reported inter-coder reliability—is sound enough to choose the final model and to demonstrate rigour across the whole process.","fun_headline_variants_meta":{"raw":{"variants":["Eight guidelines make topic modeling research trustworthy","Justify every step: eight guidelines for topic modeling","A worked STM example yields eight rigor guidelines","Keep topic modeling transparent with eight clear rules","One study, eight rules: rigorous topic modeling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000244,"raw_usage":{"total_tokens":1478,"prompt_tokens":840,"completion_tokens":638,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":569}},"tokens_in":456,"tokens_out":638,"duration_ms":6759,"temperature":1.0,"reasoning_tokens":569,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:12:28.813629+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication with more coders and reported inter-coder reliability that found the 20-topic model no more interpretable than the 35-topic model (which had better topic log odds), or found the intrusion-task scores unstable, would weaken the paper's claim that its demonstration establishes rigour, even though the guidelines themselves could still stand.","supporting_citations":[{"cited_title":"whether the association between a document and a topic makes sense","cited_arxiv_id":null,"evidence_quote":"Supplies the word- and topic-intrusion tasks and the model precision and topic log-odds metrics used for the human evaluation in Stage III."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the lack of rigour and context-blind reuse of practices in topic-modelling studies that motivates the need for the guidelines."}],"review_version":1}