{"id":"0bc9d85c-6598-4a72-b510-416d53fc59da","arxiv_id":"2506.14371","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A pipeline pairing Llama 3.1 8B as a question generator with Gemma 2 9B as a judge won the CQs-Gen 2025 critical question generation shared task.","lead":"This paper describes a two-step system that generates and selects critical thinking questions from debate speeches, and reports that it ranked first in the CQs-Gen 2025 shared task. It matters because it shows that small, open-weight language models can beat a large proprietary model on a targeted reasoning benchmark, pointing to deployable, private educational tools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported shared-task win is not statistically substantiated: Table 5's automatic scores point against the chosen submission, and the manual re-annotation that produced 67.6 is not broken down.","rationale":"The reader's weakest assumption correctly identifies the NoEval-to-useful upgrade and the low automatic test score of the winning submission. My concern is the same issue sharpened: the paper's internal automatic experiments (Tables 1-4) are plausible and include one significance test (Table 4, McNemar p<0.05 for Judge vs random), but they do not validate the final submission choice under the official metric because Table 5 contradicts the selection. The external shared-task ranking is the strongest evidence, yet the paper provides insufficient evaluation metadata to adjudicate the ranking's reliability. Credit is due for releasing code, using small open-weight models, and including a limitation statement about the noisy metric. However, the central scientific claim that this small-model Questioner-Judge pipeline outperforms a GPT-4o pipeline is currently supported only by a single, undocumented 67.6 manual score. The conditional verdict should stand, requiring the organizers' official evaluation details, per-submission manual labels, and a significance check on the top-two gap before the effectiveness claim can be fully accepted.","tokens_in":10261,"tokens_out":3886,"duration_ms":43269,"concrete_test":"Request from the CQs-Gen 2025 organizers the per-intervention official human labels for all three final submissions on the 34 test interventions, and recompute the official score while conservatively scoring every Not-able-to-evaluate question as Not Useful. If Submission 1 no longer ranks first under this worst-case treatment, the reported win depends entirely on the manual upgrade of otherwise unevaluable questions.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Table 5 is the crux. Submission 1 (Llama 3.1 8B Questioner, Gemma 2 9B Judge) has the lowest automatic test usefulness (36.3%) of the three final systems, while Submission 3 (GPT-4o for both roles) reaches 50.0%. The paper's only evidence that Submission 1 is actually better is one sentence: \"After the manual annotation of the questions by the organizers, the score of the best performing submission rose to 67.6, ranking first in the task.\" No breakdown is given of how many Not-able-to-evaluate questions were upgraded by manual annotation, how many annotators judged them, what the inter-annotator agreement was, or what the official score formula is (for example, whether 67.6 is the proportion of questions judged Useful after manual re-annotation). The automatic metric was used in Sections 3.2 and 4.2 to select the final configuration, yet on the held-out test set that metric does not favor the chosen system, so the manual re-annotation step is doing all the work in the reported win. The win could be real, but it could also be a small-sample artifact: 34 test interventions, three questions each, and the NoEval reclassification likely affects a large fraction (36% for Submission 1). Without the official per-question labels or at least the evaluation details from the shared-task overview paper, the paper does not establish that the Questioner-Judge architecture with small open models outperforms the GPT-4o baseline; it only reports an external ranking whose statistical robustness is unknown. The limitations section acknowledges the noisy automatic metric, but the central claim still rests on the undocumented manual evaluation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a two-stage pipeline for critical question generation in debate interventions: a Questioner LLM generates multiple candidate questions, and a Judge LLM selects the three most useful ones. Using small open-weight models (Llama 3.1 8B as Questioner, Gemma 2 9B as Judge) and including argumentation schemes selectively in prompts, the system was entered into the CQs-Gen 2025 shared task. The authors report that after manual annotation by the organizers, the best submission reached a score of 67.6 and ranked first. The paper also presents internal experiments on model choice, number of candidates, scheme prompting, and Judge selection.","tokens_in":10553,"tokens_out":5166,"duration_ms":54840,"significance":"If the reported shared-task result is fully substantiated, the paper has clear practical significance: a small, open-weight, untuned LLM pipeline outperformed a GPT-4o-based pipeline on a human-evaluated task, which supports locally deployable educational tools that preserve privacy. The paper gives credit for releasing code, running experiments on commodity hardware, and honestly listing limitations, including the mismatch between automatic and human evaluation. However, the internal evidence for the specific architectural choices is noisy, and the central claim of winning rests on a manual re-annotation whose details are not reported in the manuscript.","major_comments":[{"comment":"The central claim that the proposed system won the shared task is not substantiated within the paper. Submission 1 has the lowest automatic test usefulness (36.3%) among the three submissions, while Submission 3 reaches 50.0%; the only evidence for the win is the sentence 'After the manual annotation of the questions by the organizers, the score of the best performing submission rose to 67.6, ranking first in the task.' The paper should report the official scoring formula, the number of annotators and their agreement, how many Not-able-to-evaluate questions were reclassified into each label, and the manual scores for all three submissions. If these details are in the shared-task overview paper, cite the specific table or section; otherwise include them directly.","section":"§4.5, Table 5"},{"comment":"Configuration choices that motivated the final submission are based on single-run automatic scores on Dtest with no variance estimates. The differences among the top rows of Table 1 (57.6 vs. 57.1 vs. 56.8) and between 'Both' and 'Without' in Table 2 (62.4 vs. 57.7) are small, and Table 3 shows overlapping confidence intervals for 4, 6, and 8 candidate questions (59.3±3.36, 57.2±0.88, 57.3±0.76), yet §4.3 concludes that four candidates are best. Report multiple runs with confidence intervals or perform statistical tests for the key comparisons that drive the design decisions.","section":"§3.2, Tables 1–3"},{"comment":"The claim that the Judge improves over random selection by 3.4 percentage points with p < 0.05 (McNemar's test) lacks essential procedural details. State the number of paired items, whether the test was applied to pooled questions or per intervention, how the three runs were aggregated, and the test statistic. Without this information, the statistical claim cannot be verified.","section":"§4.4, Table 4"},{"comment":"The paper acknowledges that Not-able-to-evaluate questions may still be useful and that the authors intentionally avoided overfitting to the automatic metric. This assumption is load-bearing because the final submission was chosen despite its lowest automatic test usefulness. A quantitative characterization of the Not-able-to-evaluate category—for example, a manual sample annotation with agreement figures, or a breakdown of the official re-annotation—would directly support this assumption and is currently missing.","section":"Limitations and §A.2.1"}],"minor_comments":[{"comment":"The table header contains a typo, 'Valiadation' for 'Validation,' and the surrounding text says 'DShared train and DShared test' while the table columns are labeled 'Valiadation' and 'Test'; clarify which split each column reports.","section":"§4.5, Table 5"},{"comment":"The axis labels contain '/glyph1197umber' instead of 'Number,' indicating a LaTeX rendering issue that should be fixed.","section":"Figures 2 and 3"},{"comment":"The text refers to 'the column No in Tables 1, 2, and 3,' but the corresponding column is labeled 'NoEval' in the tables; align the terminology.","section":"§A.2.1"},{"comment":"The caption should explicitly state that rows with '—' for LLMJ correspond to direct generation by LLMQ without a Judge, rather than leaving that interpretation to the reader.","section":"Table 1 caption"},{"comment":"The meaning of '# quest.' is ambiguous: clarify whether the numbers denote candidate questions per prompt, total candidates including both scheme and non-scheme prompts, or something else, because §4.3 mentions 'four candidate questions per prompt (eight in total).'","section":"Tables 3 and §4.3"},{"comment":"The shared-task overview (Figueras et al., 2025) is cited in §2.2 but with no page numbers or URL in the reference list; add the complete proceedings information.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-style shared-task system description. The headline result is externally grounded in the organizers' manual evaluation, so the paper is not fatally flawed, but the reported win depends on a manual re-annotation that is not documented here. The authors should be able to address the major comments by incorporating the official evaluation details from the shared-task overview paper and by adding variance estimates or statistical tests for the internal comparisons."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this as a competent shared-task system description, not as a research claim with much statistical bite. The pipeline is straightforward: a small open LLM (Llama 3.1 8B) generates candidate critical questions, and another (Gemma 2 9B) selects the best three. That generate-then-select architecture is not new, but applying it to the CQs-Gen 2025 task with small models and getting the top official rank is a genuine, if modest, result. Credit where due: the paper is clearly written, the ablations are thoughtful, and it honestly reports negative results on fine-tuning, BERT, and data augmentation. Code is linked. That is more than many shared-task papers give you.\n\nThe soft spots are real and concentrated in the evidence for the headline claim. Table 5 shows that their final submission had the lowest automatic usefulness on the test set (36.3%) compared with the GPT-4o baseline (50.0%). The only evidence that it actually won is one sentence telling you the organizers' manual re-annotation raised the score to 67.6. No breakdown of that re-annotation appears: no annotator count, no agreement, no per-question labels, no statement of what the 67.6 means. This is the load-bearing fact, and it is a black box. It may well be true—the organizers did run a manual evaluation—but the paper does not let a reader check it. Internal ablations have their own issues: Tables 1 and 2 lack error bars, and Table 3 shows overlapping confidence intervals for different candidate counts, despite the claim that four is best.\n\nThat said, I would not call this fatal. Shared-task papers routinely report organizer-provided results, and the authors cite the overview paper that presumably contains the evaluation details. The limitations section also acknowledges the automatic metric's unreliability. The result is plausible; it is just not substantiated within the paper. A revision should add either the official evaluation details or a serious caveat that the win is reported as received and not independently verified.\n\nWho is this for? People working on argument mining or educational NLP who want a concrete example of a small-model LLM-as-judge pipeline. It is not going to change anyone's research direction. A serious referee should see it, but the bar should be: does the paper give enough for the reader to trust the shared-task result? Right now, it does not. Recommend acceptance only if the authors add the missing evaluation details or explicitly frame the result as an organizer-reported outcome with the automatic score as the only internal evidence.","headline":"A solid shared-task system paper whose reported win rests on an under-documented manual evaluation and automatic metrics that point the other way.","tokens_in":11167,"tokens_out":2391,"would_cite":false,"duration_ms":26513,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Small open-weight LLMs win critical-question generation task","keywords":["critical question generation","LLM-as-a-judge","Questioner-Judge pipeline","argumentation schemes","debate interventions","shared task","open-weight LLMs","critical thinking"],"falsifier":"If an independent set of human judges re-annotated the test questions and the winning submission's manual usefulness score fell substantially below 67.6, or if a different random split of the 189 interventions reversed the ranking between the small-model pipeline and the GPT-4o baseline, the paper's central claim of superiority would be undercut. More directly, rerunning the final submission with the same prompts and models but with the 'Both' scheme configuration removed should drop the manual score if the scheme mixture is truly load-bearing.","tokens_in":10053,"feed_emoji":"❓","tokens_out":4494,"duration_ms":36369,"temperature":0.7,"pith_summary":"This paper claims that a two-step pipeline using two small, open-weight language models—a Questioner that generates candidate critical questions and a Judge that selects the best three—can produce critical questions about debate interventions that human judges rate as useful. The authors report that this configuration, Llama 3.1 8B as Questioner and Gemma 2 9B as Judge, ranked first in the CQs-Gen 2025 shared task with a post-annotation score of 67.6. The result matters because it suggests that locally deployable, untuned models can support critical-thinking tools without relying on large proprietary systems, and that separating generation from selection improves question quality.","feed_headline":"Small open-weight LLMs win critical-question generation task","feed_subtitle":"A Llama 3.1 questioner plus Gemma 2 judge scored 67.6 after human annotation, top of CQs-Gen 2025.","key_machinery":"The pipeline is the central mechanism: a Questioner LLM generates N candidate questions, then a Judge LLM ranks and selects three. The generation prompt optionally includes argumentation scheme definitions and critical-question templates from Walton et al. (2008); the Judge's prompt instructs it to prefer redundant-but-important questions. The paper's key design choice is the 'Both' configuration, where the Questioner is prompted once with and once without schemes and the candidates are merged before selection.","core_discovery":"The central claim is that combining a generative Questioner LLM with an evaluative Judge LLM, both in the 7B–14B parameter range and without fine-tuning, outperforms other submissions to the CQs-Gen 2025 shared task on critical question generation. In the final submission, Llama 3.1 8B generates four candidate questions without argumentation schemes and four with schemes in a single prompt; Gemma 2 9B then selects the three most relevant. After organizers' manual annotation, this submission scored 67.6 and ranked first. The paper also reports that selectively mixing scheme-based and scheme-free prompts ('Both') beats either alone, and that the Judge adds a small but statistically significant gain over random selection.","pith_inferences":["If the automatic metric's 'Not able to evaluate' labels hide many useful questions, then leaderboard rankings on this task hinge on human annotation; future shared tasks should consider designing metrics that better credit novel-but-valid critical questions.","The Questioner–Judge architecture could be applied to other argumentative domains (e.g., scientific claims, policy documents) with minimal adaptation, since it only needs an intervention text and optional scheme definitions.","The near-oracle gap in Judge performance (93.5% oracle vs 59.3% Gemma 2) suggests headroom for better selection methods, e.g., training a small classifier on human preferences or using uncertainty-based selection."],"forward_implications":["Separating generation from selection improves question quality over direct generation.","Selectively adding argumentation schemes to prompts yields better questions than strict enforcement, which reduces diversity.","Small, open-weight models in the 7B–14B range can compete with or beat a GPT-4o-based pipeline on this task.","The automatic similarity-based evaluation underestimates useful questions, so manual annotation can substantially change rankings."],"supporting_citations":[{"why":"Defines the shared task, the dataset of debate interventions, and the usefulness labels used for evaluation.","marker":"(Figueras et al., 2025)"},{"why":"Supplies the argumentation schemes and critical-question templates integrated into the prompts.","marker":"(Walton et al., 2008)"},{"why":"Motivates the LLM-as-a-Judge approach used for selecting the best candidate questions.","marker":"(Li et al., 2024)"},{"why":"Basis for the hypothesis that redundant questions are important, encoded in the Judge's instructions.","marker":"(Guo et al., 2023)"},{"why":"Prior work on critical question generation that the 'Both' configuration builds on.","marker":"(Figueras and Agerri, 2024)"},{"why":"Provides the Llama 3.1 8B model used as the Questioner.","marker":"(Dubey et al., 2024)"},{"why":"Provides the Gemma 2 9B model used as the Judge.","marker":"(Team et al., 2024)"},{"why":"Grounds the analytic, creative, and evaluative framing of the two-step architecture.","marker":"(Elder and Paul, 2020)"}],"fun_headline_variants":["Small LLM pair wins critical-question generation task","Questioner and Judge LLMs rank first in CQs-Gen","Open models, no tuning: first for critical questions","Two open LLMs beat all in critical question task"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's winning configuration was chosen assuming that questions the automatic metric cannot evaluate may still be genuinely useful, so a lower automatic score does not mean worse questions.","fun_headline_variants_meta":{"raw":{"variants":["Small LLM pair wins critical-question generation task","Questioner and Judge LLMs rank first in CQs-Gen","Open models, no tuning: first for critical questions","Two open LLMs beat all in critical question task"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1309,"prompt_tokens":852,"completion_tokens":457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":391}},"tokens_in":468,"tokens_out":457,"duration_ms":4765,"temperature":1.0,"reasoning_tokens":391,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:17:19.890530+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If an independent set of human judges re-annotated the test questions and the winning submission's manual usefulness score fell substantially below 67.6, or if a different random split of the 189 interventions reversed the ranking between the small-model pipeline and the GPT-4o baseline, the paper's central claim of superiority would be undercut. More directly, rerunning the final submission with the same prompts and models but with the 'Both' scheme configuration removed should drop the manual score if the scheme mixture is truly load-bearing.","supporting_citations":[],"review_version":1}