{"id":"7f0b1b33-7708-4fe1-8c16-92449e3ec301","arxiv_id":"2412.10139","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A structured prompt framework with role, task, procedure, context, and output-format elements improves LLM performance scores on three corpus discourse analysis tasks, but the evaluation relies on a self-defined human rubric.","lead":"This paper introduces TACOMORE, a structured prompting protocol for using large language models in corpus-based discourse analysis, and tests it on keyword, collocate, and concordance tasks in a COVID-19 research abstract corpus. The authors report improved accuracy and reproducibility scores with the framework, though hallucination in model outputs persists.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-hoc exclusion of the only unstable concordance line (line 7) from the evaluation inflates reported scores; re-including it may overturn the concordance task's 'satisfying results' claim.","rationale":"We focused on the post-hoc exclusion rather than the reader's concern about rater independence because the exclusion is an explicit, documented procedural choice that demonstrably affects the reported numbers, whereas the lack of blinding is an omitted detail that might be addressed in a revision. The exclusion also undermines the integrity of the evaluation: excluding data because it conflicts with the expected pattern is a classic form of cherry-picking, and it directly impacts the central claim that TACOMORE produces satisfying, reproducible results. A simple re-analysis can resolve this, so the appropriate verdict remains CONDITIONAL: the paper requires this re-analysis before the concordance-task claims can be accepted.","tokens_in":51824,"tokens_out":9783,"duration_ms":97213,"concrete_test":"Restore the excluded concordance line 7 and rerun the evaluation for GPT-4o and Gemini-1.5-Pro using the same five-point rubric (Accuracy, Ethicality, Reasoning, Reproducibility) with the original raters or independent ones, then recompute the mean scores for Table 3. Compare the new totals with the published 17 and 17; if either task's total drops by more than 2 points or any sub-score (e.g., Reproducibility) decreases, the claim of 'satisfying results' in concordance analysis is not supported by the full dataset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 4.2 (Concordance Analysis), the authors state that 'Only concordance line 7 is confusing...' and that 'this case is not taken into consideration in evaluation.' This is a post-hoc exclusion of the single data point where the models' judgments were unstable. The central claim of achieving 'satisfying results' and the measured Accuracy and Reproducibility scores are directly derived from evaluations that omit this case. Since Reproducibility is defined as stable output across trials, excluding the one line with unstable answers guarantees a higher Reproducibility score. Similarly, the claim that 'both models could successfully detect whether a concordance is biased' is only true for the 19 retained lines, not the full set. This is outcome-based data deletion, which biases the evaluation in favor of the framework in one of the three tasks that support the central claim. The issue is not merely cosmetic: the concordance task is a central demonstration of TACOMORE's utility, and the numerical scores in Table 3 would be lower if line 7 were included, potentially changing the qualitative conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes TACOMORE, a prompting framework for LLM-assisted corpus-based discourse analysis, built on four principles (Task, Context, Model, Reproducibility) and five prompt elements (Role Description, Task Definition, Task Procedures, Contextual Information, Output Format). The authors apply it to keyword, collocate, and concordance analysis on a corpus of COVID-19 research abstracts using GPT-4o, Gemini-1.5-Pro, and Gemini-1.5-Flash, and score the outputs on a self-defined 5-point Likert rubric (Accuracy, Ethicality, Reasoning, Reproducibility). They report that TACOMORE improves accuracy and replicability, with a monotonic improvement in an ablation study, while acknowledging persistent hallucination in model outputs.","tokens_in":52055,"tokens_out":4575,"duration_ms":49118,"significance":"The paper addresses a timely and practical problem: making LLM outputs more replicable and interpretable for qualitative corpus analysis. Its concrete strengths are the open corpus, the detailed TACOMORE design, and the unusually complete appendix with full prompts and raw outputs for all models and ablation conditions, which is a genuine reproducibility resource. If the evaluation were adequately controlled, this would be a useful methodological contribution for corpus linguistics. However, the current evidence is not conclusive: the outcome measure is a non-validated in-house rubric applied without reported blinding or statistical analysis, and one task excludes the only ambiguous data point from evaluation.","major_comments":[{"comment":"The manuscript excludes the only unstable concordance line from the evaluation. Section 4.2 states that \"Only concordance line 7 is confusing\" and \"this case is not taken into consideration in evaluation.\" Because Reproducibility is defined as stable output across trials (Section 3.4), dropping the one line in which the two large models disagreed (GPT-4o marked [Yes], Gemini-1.5-Pro marked [No]) mechanically raises the reported Reproducibility and Accuracy scores in Table 3. The claim that \"both models could successfully detect whether a concordance is biased\" is therefore supported only for 19 of 20 lines. Re-including line 7 may lower the scores and could change the qualitative conclusion; the authors should report full-set results and treat line 7 as an object of analysis rather than an exclusion.","section":"4.2, Concordance Analysis"},{"comment":"The evaluation rests entirely on a self-defined 5-point Likert rubric with no scoring anchors, no evidence of blinding, no report of per-rater or per-item scores, and no statistical tests or confidence intervals. The two raters are described only as \"two human experts,\" and the single reported Krippendorff's alpha does not establish that the scores are valid measures of analytical quality rather than prompt-format compliance. Because every central claim—\"satisfying results,\" the effectiveness of TACOMORE, and the monotonicity of the ablation—is derived from these scores, the paper needs a pre-specified coding protocol, blind and independent rating, and at least per-item agreement and error bars; ideally, the rubric should be validated against expert human analysis of the same corpus.","section":"3.4 and 4.1"},{"comment":"Table 4 is presented as evidence that \"the score climbs up every time a new element is added to the prompt,\" but the design is a single fixed-order cumulative ablation on one model with one run per condition. The observed monotonic increase does not identify which element is responsible, and the baseline is an extremely terse prompt that may simply benefit from longer instructions. Furthermore, since the Reproducibility metric rewards stable output and the Output Format element explicitly constrains output structure, part of the final score gain is built into the prompt by construction. A factorial ablation or length-matched control prompts, together with variance estimates across repeated runs, are needed to support the framework-attribution claim.","section":"4.3, Ablation Study"}],"minor_comments":[{"comment":"There are several typos and infelicities: \"Gemimi-1.5-Pro\" in Section 1, \"air and ethical\" in Section 3.4, \"Bellow\" in the Output Format prompt, \"attach attach\" in Section 4.1, and \"Inte-rater Reliability\" in Section 4.1.","section":"Throughout"},{"comment":"One ablation condition is labeled \"B. + R. D. + T. D. + T. D + C. I.\"; this should presumably be \"B. + R. D. + T. D. + T. P. + C. I.\" to match Table 4.","section":"Appendix A.2"},{"comment":"The concordance experiments were run on a third-party platform while the other tasks used APIs; the possible effect of the platform on output variability is not discussed.","section":"4.1, Target LLMs"},{"comment":"Figure 3 is referenced but its content is not described in the text, so readers cannot see the rubric details beyond the four one-sentence definitions in Section 3.4.","section":"3.4, Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a methods-oriented application study with a useful open-data appendix. I do not see grounds for rejection, because the framework and the released materials are reusable, but the current evaluation does not yet support the strong claims about effectiveness and reproducibility. The concordance-line exclusion and the unvalidated rubric should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has a genuinely useful core idea: TACOMORE, a structured prompting framework for LLM-assisted corpus analysis, built from five elements (role, task definition, procedures, contextual information, output format). The most distinctive piece is the contextual co-text principle—giving the LLM the same concordance lines a human analyst would see. That is a simple, actionable insight that could help many corpus linguists. The ablation study, though small, shows monotone improvement as elements are added, which is a nice practical demonstration. The paper also reports high inter-rater reliability (Krippendorff's alpha 0.927) and includes full prompts and many raw outputs in the appendix, which supports replication.\n\nThe soft spots are real but not fatal. The evaluation rests on a self-defined 5-point Likert rubric applied by two raters with no reported blinding. That is a moderate concern; the high inter-rater agreement helps, but it still leaves room for expectation effects. More serious is the post-hoc exclusion of concordance line 7 in Section 4.2. The authors say the line is confusing and \"not taken into consideration.\" That is outcome-based data deletion. It directly inflates the Accuracy and Reproducibility scores for the concordance task and underpins the claim that both models \"successfully detect\" bias. Re-including that line could change the qualitative conclusion for that task, and the paper should report both the unstable answers and a clear rule for exclusions. The \"first time ever\" claim in the introduction is unsupported and should be removed. And for a paper centered on reproducibility, the lack of a released dataset or code is a gap, though the appendix partially compensates.\n\nOverall, the framework is plausible and worth testing, but the current evidence is suggestive rather than conclusive. The central idea—especially the co-text principle—deserves independent validation.\n\nRecommendation: send it to peer review, but expect heavy revision. The manuscript should fix the line-7 exclusion, describe the rating procedure in enough detail to allow independent replication, tone down the overclaims, and ideally release the prompts and a minimal corpus subset. who is this for: corpus linguists, social scientists using LLMs for qualitative coding, and prompt-engineering researchers. A serious referee can help this become a useful paper.","headline":"A useful domain-specific prompting protocol, but the evaluation has a post-hoc exclusion problem and overclaims.","tokens_in":52551,"tokens_out":1867,"would_cite":false,"duration_ms":21487,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A five-element prompting protocol called TACOMORE moves LLM corpus analysis from baseline failure to human-rated accuracy, the authors claim.","keywords":["prompt engineering","large language models","corpus linguistics","discourse analysis","keyword analysis","collocation","concordance","replicability"],"falsifier":"Have two independent teams, blind to the paper's scores, re-score the recorded model outputs and run the same prompts again against a pre-registered rubric that separates 'followed the format' from 'interpreted the corpus correctly'; if format-compliant outputs still score high while interpretive quality (for example, concordance bias judgments against an expert gold standard) does not improve over baseline, the central claim fails.","tokens_in":51604,"feed_emoji":"🤖","tokens_out":6376,"duration_ms":59831,"temperature":0.7,"pith_summary":"LLMs are good at counting words but not at interpreting them, and earlier attempts to use them for corpus-based discourse analysis produced unreliable and unreproducible results. The paper argues that the failure is largely a prompting problem: when the prompt specifies a role, a task definition, step-by-step procedures, the co-text around each target word, and a strict output format, the same models produce acceptable keyword, collocate, and concordance analyses. The authors test this on a public corpus of COVID-19 research abstracts with three LLMs and report human-rated gains in accuracy, ethicality, reasoning, and reproducibility, with an ablation study showing scores climbing as each prompt element is added. If the claim holds, structured prompting protocols could give corpus linguists a standard, transparent way to delegate parts of qualitative analysis to LLMs while keeping human oversight.","feed_headline":"Prompt recipe lifts LLM corpus analysis to human-grade scores","feed_subtitle":"How a five-element protocol turns unreliable LLMs into auditable research assistants.","key_machinery":"TACOMORE is the paper's central object: a prompting protocol that encodes the standard operating procedure of corpus-based discourse analysis into a fixed prompt template. Its five compulsory elements are Role Description (assign the model the persona of a corpus linguist), Task Definition (state the goal, data, and background), Task Procedures (break the task into numbered human-like steps, drawing on chain-of-thought prompting), Contextual Information (attach the same co-text—concordance lines and original text—that human analysts would see), and Output Format (compel exhaustive, machine-checkable output through delimiters and an example). The framework's work is to replace ad-hoc single-shot instructions with a replicable protocol, and the ablation isolates the contribution of each element by adding them one at a time.","core_discovery":"The paper's central claim is that TACOMORE—a framework built on four principles (Task, Context, Model, Reproducibility) and five prompt elements (Role Description, Task Definition, Task Procedures, Contextual Information, Output Format)—turns LLMs from unreliable statistical predictors into effective assistants for corpus-based discourse analysis. On three tasks (keyword theme grouping, collocate analysis of the word china, and bias detection in concordances of 'China virus' and 'Chinese virus'), GPT-4o and Gemini-1.5-Pro receive roughly 4-out-of-5 human Likert scores across accuracy, ethicality, reasoning, and reproducibility, while the smaller Gemini-1.5-Flash scores lower. The keyword ablation shows the total score rising monotonically from 5/20 for a baseline prompt to 16/20 when all five elements are present, which the authors read as direct evidence that each element contributes. They describe this as the first time LLM outputs in these tasks were satisfying. The paper also reports that hallucination—fabricated or mis-cited corpus context—persists even with the full framework, so the method is proposed as a complement to, not a replacement for, human validation.","pith_inferences":["The authors do not test whether two independent teams prompting without seeing the paper's exact templates would converge on the same TACOMORE-style prompts; a multi-lab replication would clarify whether the framework's four principles alone reproduce the reported gains.","The claimed improvement may partly reflect formatting compliance: the Output Format element forces exhaustive coverage (all 83 keywords), which mechanical completeness checks would reward on Accuracy and Reproducibility even if interpretive depth is unchanged, and the paper's Likert scale cannot fully separate these.","The concordance bias task could become a quantitative benchmark by measuring the two large models' judgments against an expert gold-standard annotation set rather than the two raters' consensus, giving a testable extension beyond the current design.","If the framework generalizes, the same five elements could be tuned for other qualitative corpus tasks such as stance detection or metaphor analysis, with the COVID-19 abstract corpus serving as a reproducibility benchmark."],"forward_implications":["If the protocol works as claimed, corpus linguists could adopt TACOMORE as a documented default for LLM-assisted keyword, collocate, and concordance analysis, making AI involvement in qualitative work auditable step by step.","The monotonic ablation suggests each prompt element earns its place, so omitting co-text or output format forfeits measurable accuracy and completeness, arguing for always pairing LLMs with the human-in-the-loop context.","The persistence of hallucination implies that even with strong prompts, model outputs must be checked against the original corpus, so TACOMORE would complement rather than eliminate manual verification in discourse studies.","The reproducibility principle—stable and reasonable, not identical, outputs—gives the field a practical target for replicated LLM use, since identical outputs across runs are not guaranteed.","The evaluation rubric of Accuracy, Ethicality, Reasoning, and Reproducibility may transfer to other qualitative analysis tasks, giving researchers a common scoring language."],"supporting_citations":[{"why":"Supplies the baseline prompt and the prior negative finding that LLMs are unreliable in corpus discourse analysis, which TACOMORE is built to improve.","marker":"(Curry et al., 2024)"},{"why":"Provides the standard operating procedure for corpus-based discourse analysis that the framework translates into step-by-step prompts.","marker":"(Baker, 2023)"},{"why":"AntConc 4.2.4, the tool used to generate the keyword, collocate, and concordance data for all three tasks.","marker":"(Anthony, 2023)"},{"why":"Chain-of-thought prompting, the technique behind the Task Procedures element that breaks analysis into human-like steps.","marker":"(Wei et al., 2022)"},{"why":"Survey establishing prompt engineering as essential to LLM performance, used to justify the need for structured prompts.","marker":"(Schulhoff et al., 2024)"},{"why":"Defines keyword analysis and keyness statistics that set up the first task's expected output.","marker":"(Baker, 2004)"},{"why":"Defines collocation analysis, the methodological basis for the collocate task's L5-R5 span and content-word focus.","marker":"(Gries and Stefanowitsch, 2004)"},{"why":"Source cited for the four evaluation metrics used to score outputs on the 5-point Likert scale.","marker":"(Hu and Zhou, 2024)"}],"fun_headline_variants":["TACOMORE protocol boosts LLM corpus analysis, but hallucination remains","Structured prompts turn LLMs into auditable corpus research assistants","Five-prompt recipe yields 4/5 human scores in LLM corpus tasks","TACOMORE makes LLM corpus analysis replicable, not perfect","TACOMORE's prompt framework hits 4/5 human scores, still hallucinates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the researchers' own 5-point Likert rubric, scored by two human raters without reported blinding or independence, captures genuine analytical quality rather than rewarding format compliance or matching rater expectations.","fun_headline_variants_meta":{"raw":{"variants":["TACOMORE protocol boosts LLM corpus analysis, but hallucination remains","Structured prompts turn LLMs into auditable corpus research assistants","Five-prompt recipe yields 4/5 human scores in LLM corpus tasks","TACOMORE makes LLM corpus analysis replicable, not perfect","TACOMORE's prompt framework hits 4/5 human scores, still hallucinates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000619,"raw_usage":{"total_tokens":2900,"prompt_tokens":1000,"completion_tokens":1900,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":1799}},"tokens_in":616,"tokens_out":1900,"duration_ms":16841,"temperature":1.0,"reasoning_tokens":1799,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:18:09.238348+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent teams, blind to the paper's scores, re-score the recorded model outputs and run the same prompts again against a pre-registered rubric that separates 'followed the format' from 'interpreted the corpus correctly'; if format-compliant outputs still score high while interpretive quality (for example, concordance bias judgments against an expert gold standard) does not improve over baseline, the central claim fails.","supporting_citations":[],"review_version":1}