{"id":"2d3b8c2b-5fa5-4c8d-b963-1bdf7e415b9f","arxiv_id":"2412.07923","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Repeating a question 3 or 5 times in a prompt does not significantly improve LLM reading comprehension accuracy across five models, three datasets, and four prompt configurations.","lead":"This paper tests whether repeating a question inside a single prompt changes how well five large language models answer reading comprehension questions. Across three datasets and four prompt styles, repetition did not produce a statistically significant improvement in accuracy, though some individual models gained up to 6% in specific settings.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The blanket 'no meaningful impact' claim in §3.3 rests on a single pooled Friedman test that is likely underpowered to detect the 3–6% per-cell differences visible in Table 1; per-cell significance and effect sizes are needed.","rationale":"The reader's weakest-assumption correctly identifies pooling as a threat to the global null result: Table 1 contains heterogeneous per-cell changes, and a single Friedman test across all 60 blocks could easily mask real effects that point in different directions. My stress test sharpens this further. Even if repetition effects were directionally consistent, the test may be underpowered because it uses only 60 rank-based blocks and cannot detect accuracy shifts of only a few percentage points. A 6-point change (as in Llama-3.1 closed-book HotPotQA) is substantial relative to the n=500 per cell, so at least some per-cell differences are worth formal testing. The manuscript does not provide those tests, confidence intervals, or a power analysis, yet it concludes 'no meaningful impact' in §3.3 and the Conclusion. This is a genuine inferential gap, not a mere preference for different statistics. A cell-level McNemar analysis with FDR control would directly adjudicate whether any per-setting effect is real; a power calculation would calibrate how much absence-of-significance is worth. If per-cell effects survive, the conclusion must be narrowed to 'no consistent effect' or to specific settings. If none survive and the power analysis shows a small minimal detectable effect, the blanket claim becomes credible. Since the paper's data table is sufficiently detailed for an external reanalysis (though per-question outputs are not released), conditional acceptance with a request for this reanalysis is the appropriate outcome. My concern does not shift the reader's verdict; it reinforces the same CONDITIONAL recommendation, hence UNCHANGED.","tokens_in":10577,"tokens_out":8181,"duration_ms":79682,"concrete_test":"Obtain the per-question predictions from the authors (or regenerate them via the described API and local setups). For each of the 60 model×configuration×dataset cells, run McNemar's test on the 500 paired questions comparing Qx1 vs Qx5 and Qx1 vs Qx3, then apply Benjamini-Hochberg false-discovery-rate control across all 120 tests. Separately, compute the smallest accuracy-point effect the Appendix C.1 Friedman design can detect at 80% power with 60 blocks. If any per-cell difference survives FDR, or if the minimal detectable effect is above ~3 points, the §3.3 blanket conclusion is unsupported; if not, the null claim is credible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in §3.3 ('question repetition within a prompt neither significantly improves nor significantly degrades the performance of LLMs across all models, datasets, and settings') is supported by a single Friedman test on aggregated cell accuracies (Appendix C.1: test statistic 0.7118, p=0.70). The load-bearing assumption is that this pooled test has adequate power to detect any practically meaningful repetition effect. That assumption is insecure. Table 1 shows per-cell differences of 2–6 percentage points (e.g., Llama-3.1 closed-book HotPotQA improves from 0.24 to 0.30; Mistral open-book HotPotQA drops from 0.51 to 0.48). With n=500 questions per cell, a 6-point difference is roughly two standard errors and could be real for that cell. The Friedman test only ranks within the 60 model×configuration×dataset blocks; if effects point in opposite directions across blocks, as Table 1 suggests, the test has little power. The paper reports no per-cell significance tests, confidence intervals, or power analysis. Interpreting p=0.70 as 'no meaningful impact' (§3.3, Conclusion) therefore conflates absence of evidence with evidence of absence. The universal negative claim depends on a test that cannot reliably detect the very effects the data display.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether repeating a question 1, 3, or 5 times within a single prompt affects the accuracy of five LLMs (GPT-4o-mini, DeepSeek-V3, Llama-3.1 8B, Mistral 7B, Phi-4 14B) on SQuAD, HotPotQA, and Natural Questions, under open-book, closed-book, question-context-question, and paraphrasing configurations. Accuracy is measured by substring matching on 500 sampled questions per condition, for a total of 90,000 prompted evaluations. A Friedman test across the pooled condition blocks yields p=0.70, and the authors conclude that question repetition neither significantly improves nor significantly degrades performance across all models, datasets, and settings. The paper also reports descriptive per-configuration trends, including a 6% accuracy gain for Llama-3.1 on closed-book HotPotQA.","tokens_in":10781,"tokens_out":4418,"duration_ms":39865,"significance":"This is a carefully executed empirical study with a clearly stated null result. The scope is substantial: five recent LLMs, three QA benchmarks, four prompt configurations, and 500 questions per condition. The main positive finding—that literal question repetition produces at most small accuracy changes on reading-comprehension tasks—is a useful data point for prompt engineering, and the paper correctly contrasts with EchoPrompt's restatement-based gains. However, the statistical support for the strong universal claim is currently inadequate, as detailed in the major comments.","major_comments":[{"comment":"The universal negative claim—that question repetition 'neither significantly improves nor significantly degrades the performance of LLMs across all models, datasets, and settings'—is supported only by a single Friedman test on pooled block ranks. This test is not sensitive to effects that are large in a few cells but opposite in direction across the 60 model×configuration×dataset blocks. Table 1 shows several 3–6 percentage point changes (e.g., Llama-3.1 closed-book HotPotQA Qx1=0.24 vs. Qx5=0.30; Mistral 7B open-book HotPotQA Qx1=0.51 vs. Qx5=0.48). With n=500 per cell, a 6-point difference is roughly two standard errors and may be statistically significant for that cell. The paper reports no per-cell significance tests, confidence intervals, effect sizes, or power analysis. Interpreting p=0.70 as 'no meaningful impact' therefore conflates absence of evidence with evidence of absence. I request per-cell paired tests (e.g., McNemar) with multiple-comparison correction, plus a power analysis showing the detectable effect size at n=500.","section":"§3.3, Appendix C.1, Table 1"},{"comment":"The Paraphrasing configuration is pooled together with the identical-repetition conditions in the Friedman test, but it is a different manipulation: it introduces model-generated paraphrases rather than literal repetition. If repetition effects are masked by paraphrase noise, pooling reduces power. The paper should either analyze Paraphrasing separately or provide a justification for pooling. Also, the definition of repetition levels for Paraphrasing is unclear: Qx3 appears to correspond to one original plus two paraphrases, not three repetitions, so the column labels are misleading.","section":"§2.1, Appendix A.4"},{"comment":"The abstract states that repetition 'can increase models’ accuracy by up to 6%' while §3.3 concludes there is 'no meaningful impact.' These statements are not contradictory only if the 6% increase is shown to be statistically indistinguishable from noise. The paper does not provide a confidence interval for that 6% cell (Llama-3.1 closed-book HotPotQA), so the reader cannot evaluate whether the largest observed effect is real. Please report the uncertainty around the largest per-cell differences, or soften the universal claim to 'no consistent statistically significant effect across the pooled analysis.'","section":"Abstract, §3.1"}],"minor_comments":[{"comment":"There are subject-verb agreement errors, e.g., 'performance remain stable' (§3.1) and 'NQ show the smallest' (§3.2); the manuscript should be proofread for such grammatical issues.","section":"§3.1, §3.2"},{"comment":"The phrase 'total of 90,000 questions' is ambiguous; since the same 500 questions are reused across all 180 settings, the correct description is 90,000 prompted evaluations, not 90,000 unique questions.","section":"§2.3"},{"comment":"The prompt templates are only described with placeholders; for reproducibility, please provide the exact full prompts for Qx1, Qx3, and Qx5 in each configuration.","section":"Appendix A"},{"comment":"No code, data, or model-output release is mentioned; given the use of API models, releasing the prompt templates, sampled question IDs, and inference configuration would substantially improve reproducibility.","section":"General"},{"comment":"Figures 2–5 are referenced in the text but appear after the appendix; please ensure all figures are called out in numerical order in the final version and that captions are self-contained.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":"The main statistical concern is fixable with additional analyses. The paper is otherwise a straightforward empirical contribution that fits the scope of an NLP venue. I would not reject; the authors need to either add per-cell tests and effect-size reporting or revise the universal claim so that it does not overstate the pooled null result. The paraphrasing confound should also be resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something genuinely under-tested: it varies how many times the same question appears in a prompt (1, 3, or 5 times) across open-book, closed-book, QCQ, and paraphrase configurations, using five models and three QA datasets. That is a systematic look at a practical prompt-design question, and I am not aware of prior work that directly measures this manipulation. The 90k-query scale and the honest reporting of the null result — including the up-to-6% per-cell improvements that do not survive the global test — are real strengths. I believe the authors are reporting what they actually saw.\n\nThe soft spot is the statistics, and it is load-bearing. The blanket claim in Section 3.3 — that repetition neither significantly improves nor degrades performance across all models, datasets, and settings — rests on a single Friedman test over the aggregated rankings. That test can miss effects that point in opposite directions across the 60 model×dataset×configuration blocks, and Table 1 shows several such candidates: Llama-3.1 closed-book HotPotQA rises from 0.24 to 0.30, Phi-4 closed-book SQuAD rises from 0.37 to 0.41, Mistral open-book HotPotQA drops from 0.51 to 0.48. With 500 questions per cell, differences of 4–6 percentage points are roughly two to three standard errors. None of those cells are tested individually, and the paper reports no confidence intervals, no per-configuration significance tests, and no power analysis. So the strong conclusion conflates absence of evidence with evidence of absence for the per-setting claims. It also leaves out how the 500 questions per dataset were sampled, and no code or data are released, which hurts reproducibility.\n\nThat said, the central practical takeaway — duplicating a question in a prompt will not meaningfully change reading-comprehension accuracy — is plausible and consistent with the data as a pooled statement. The paper's own limitations section is honest, and the authors are not overclaiming in the abstract. The main fix is either to add per-setting analyses (with appropriate multiple-testing corrections) or to temper the universal conclusion. As written, I would call this a solid negative result that needs a revision, not a desk reject.\n\nWho gets value: anyone working on prompt sensitivity, input robustness, or efficient prompt design. It is a modest finding, not a landmark. I would send it to a serious referee, expecting heavy or at least moderate revision on the statistical framing.","headline":"A clean, honest null-result study on question repetition that overreaches its pooled statistics when claiming 'no meaningful impact' per setting.","tokens_in":11347,"tokens_out":2603,"would_cite":false,"duration_ms":27308,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Repeating a question 1, 3, or 5 times within a prompt does not significantly change the reading-comprehension accuracy of large language models, across five models, three datasets, and four prompt settings.","keywords":["large language models","question repetition","prompt engineering","reading comprehension","LLM robustness","statistical significance","Friedman test","prompt design"],"falsifier":"Run the same 90,000-question experiment but perform separate Friedman tests or paired comparisons for each model × dataset × configuration cell (or at least for cells with large swings, such as Llama-3.1 closed-book HotPotQA, which rises from 0.24 to 0.30, or Mistral open-book HotPotQA, which falls from 0.51 to 0.48). If any cell shows a statistically significant difference across Qx1, Qx3, and Qx5 after multiple-comparison correction, the blanket 'no meaningful impact' claim would be false. A simpler check is a permutation test that respects block structure to see whether the observed 6% gains exceed what chance would produce, or a larger-sample replication of one high-swing cell to determine whether the difference is real.","tokens_in":10338,"feed_emoji":"🤖","tokens_out":4504,"duration_ms":39700,"temperature":0.7,"pith_summary":"This paper asks whether repeating a question inside a single prompt helps or hurts large language models. Across five models (GPT-4o-mini, DeepSeek-V3, Llama-3.1 8B, Mistral 7B, Phi-4 14B), three reading-comprehension datasets, and four prompt configurations, the authors compare accuracy when the question appears 1, 3, or 5 times. They find that repetition can move accuracy by up to 6% in specific cells, but a pooled Friedman test yields p = 0.70, meaning the overall effect is not statistically significant. The paper concludes that verbatim or paraphrased question repetition alone is not a reliable prompt-engineering lever for reading comprehension, and that the tested models are robust to redundant input structures.","feed_headline":"Repeating a question in an LLM prompt changes almost nothing","feed_subtitle":"A 90,000-question test across five models finds no significant accuracy gain from asking 1, 3, or 5 times.","key_machinery":"The key machinery is a controlled prompt-repetition experiment with four configurations: open-book (context before question), closed-book (no context), question-context-question (QCQ, question both before and after context), and paraphrasing (original question plus model-generated paraphrases appended to context). Repetition levels are Qx1, Qx3, and Qx5, and accuracy is measured by substring matching of the gold answer. The statistical workhorse is the Friedman test, a non-parametric repeated-measures test that ranks the three repetition conditions within each experimental block; here it produced a p-value of 0.70, which carries the conclusion that repetition has no significant aggregate effect.","core_discovery":"The central claim is that question repetition within a prompt neither significantly improves nor significantly degrades LLM performance across all models, datasets, and settings tested. The evidence comes from 90,000 questions (500 sampled per dataset × 3 repetition levels × 4 configurations × 3 datasets × 5 models), with accuracy measured by substring matching. A Shapiro-Wilk test showed non-normality, so the authors used the non-parametric Friedman test on the pooled accuracy scores across repetition levels; the test statistic was 0.7118 with p = 0.70. Individual configurations show contrasting movements—for example, Llama-3.1 closed-book HotPotQA rises from 0.24 to 0.30 from Qx1 to Qx5, while Mistral open-book HotPotQA falls from 0.51 to 0.48—but the aggregate analysis treats these as noise. The authors interpret this as evidence that LLMs process the question effectively regardless of how many times it is repeated, and that repetition does not encourage the model to focus more on the repeated information.","pith_inferences":["The single pooled Friedman test may cancel out opposing effects: if repetition helps smaller models in closed-book settings but hurts others, the aggregate null result could hide real, heterogeneous behavior that a per-cell analysis would reveal.","The 500-question subsamples may be underpowered for detecting small but genuine accuracy differences; a power analysis could show whether the experiment could have detected a 1–2% effect, and if not, the 'no meaningful impact' conclusion is weaker than it appears.","The paper's own limitation about causal-only models suggests an untested extension: masked or encoder-only language models might show a different sensitivity to question repetition, since they process input dependencies differently.","The interaction between repetition and explicit instructions (e.g., 'focus on the repeated question') is not tested here; combining repetition with such directives could still produce a measurable effect, contrary to the paper's broad null framing."],"forward_implications":["Question repetition is not a dependable prompt-engineering technique for reading-comprehension tasks; practitioners should not expect verbatim or paraphrased repetition to improve accuracy.","The tested LLMs are robust to redundant phrasing, suggesting that simply showing a question multiple times does not strengthen the model's attention to it.","The null result contrasts with prior work showing benefits from instructing models to restate questions, so the mechanism behind such gains is not mere repeated presentation.","The conclusion holds across a wide range of model sizes (7B to 685B) and context conditions, indicating a general robustness rather than a quirk of a single model.","Any true repetition effects that exist are likely small and setting-specific, and would require larger samples or per-condition analyses to detect reliably."],"supporting_citations":[{"why":"Provides the SQuAD dataset, one of the three reading-comprehension benchmarks used to evaluate the models.","marker":"Rajpurkar et al., 2016"},{"why":"Provides the HotPotQA dataset, the multi-hop benchmark that shows the largest per-cell accuracy swings under repetition.","marker":"Yang et al., 2018"},{"why":"Provides the Natural Questions dataset, the third benchmark, with long answers used as context and short answers as gold labels.","marker":"Kwiatkowski et al., 2019"},{"why":"Supplies the finding that context-before-question ordering improves performance, which the open-book prompt design follows.","marker":"Shaier et al., 2024a"},{"why":"Serves as the contrast case: instructing models to restate questions helps in their work, which the paper's null result for mere repetition is measured against.","marker":"Mekala et al., 2024"},{"why":"Establishes the standard practice of concatenating question and context into a single input string, which the experimental setup adopts.","marker":"Brown et al., 2020"},{"why":"One of the references cited for the substring-matching accuracy metric used to score all model outputs.","marker":"Liu et al., 2023"}],"fun_headline_variants":["Repetition won't boost LLM answers, 90k test shows","Question repetition yields no significant LLM gain","Repeating queries doesn't move LLM accuracy","LLM answers stay put when you repeat the question"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The blanket conclusion rests on the assumption that pooling all models, datasets, and configurations into one Friedman test is a valid way to detect repetition effects, which requires that meaningful repetition effects, if any, are consistent enough in direction to show up in the global rank comparison.","fun_headline_variants_meta":{"raw":{"variants":["Repetition won't boost LLM answers, 90k test shows","Question repetition yields no significant LLM gain","Repeating queries doesn't move LLM accuracy","LLM answers stay put when you repeat the question"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000848,"raw_usage":{"total_tokens":3668,"prompt_tokens":901,"completion_tokens":2767,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":2702}},"tokens_in":517,"tokens_out":2767,"duration_ms":19064,"temperature":1.0,"reasoning_tokens":2702,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:25:06.066541+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 90,000-question experiment but perform separate Friedman tests or paired comparisons for each model × dataset × configuration cell (or at least for cells with large swings, such as Llama-3.1 closed-book HotPotQA, which rises from 0.24 to 0.30, or Mistral open-book HotPotQA, which falls from 0.51 to 0.48). If any cell shows a statistically significant difference across Qx1, Qx3, and Qx5 after multiple-comparison correction, the blanket 'no meaningful impact' claim would be false. A simpler check is a permutation test that respects block structure to see whether the observed 6% gains exceed what chance would produce, or a larger-sample replication of one high-swing cell to determine whether the difference is real.","supporting_citations":[],"review_version":1}