{"id":"56c79993-bc21-4651-84bc-835ba48b51d5","arxiv_id":"2506.02302","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Feeding an LLM-generated grammar explanation back to a model before a grammaticality judgment improves minimal-pair accuracy, with the largest gains for smaller models.","lead":"This paper tests a trick called grammar prompting: a large language model writes a short grammar rule, then that rule is fed back to a model before it chooses the grammatical sentence from a pair. The method improves accuracy on English, Chinese, and Russian grammar benchmarks and shrinks the gap between large and small models, though the headline numbers come from hand-picked hard problems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 56% gap-reduction claim is computed on a post-hoc subset of hard paradigms; full-benchmark gaps are much smaller, so the equalization result may not generalize.","rationale":"The reader's weakest assumption accurately identifies the main soft spot: the headline gap-reduction statistic is computed over a post-hoc selected subset of hard paradigms, and the paper's own full-benchmark baselines show a much smaller LLM-SLM gap (e.g., 7.2 pp on full BLiMP vs 12.9 pp on the selected subset). This undermines the generalizability of the 56% figure and the 'within 5.8 pp' claim, especially since the per-language results are uneven and lack error bars. The direction of the paper is plausible and the controlled comparisons (control and textbook conditions) provide some support for the mechanism, so rejection is not warranted. A conditional verdict — accept pending full-benchmark evaluation and artifact release — remains the appropriate outcome, which matches the reader's assessment.","tokens_in":21788,"tokens_out":9573,"duration_ms":86966,"concrete_test":"Run the Base, CoT, GP, and GP+CoT conditions on all 67 BLiMP, 38 SLING, and 45 RuBLiMP paradigms (or a random sample of paradigms not filtered by gpt-4o accuracy), using the same 50-item-per-paradigm protocol and the same o1-generated beginner prompts, and recompute the per-language and cross-language LLM-SLM gaps exactly as in Table 7. If the full-benchmark cross-language gap reduction falls materially below 56% (e.g., below 30%), or if the English gap does not improve over base, the headline claim should be revised to reference the selected challenging paradigms. Also compute bootstrap confidence intervals over paradigm categories to assess whether the gap reduction is statistically distinguishable from noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim — that GP+CoT brings SLMs within 5.8 pp of GPT-4-class models, cutting the LLM-SLM gap by 56% — is computed only on paradigms selected because gpt-4o scored at or below 90% (BLiMP), 6 SLING categories excluding easier ones, and 7 RuBLiMP categories at or below 96% (Section 4.2). On the full BLiMP benchmark (Table 8), gpt-3.5 and gpt-4o score 84.0 and 91.2, a 7.2 pp gap, versus 12.9 pp on the selected English subset. Easy paradigms leave little room for gap reduction, so excluding them inflates both the baseline gap and the relative improvement. The per-language gaps in Table 7 are also heterogeneous: English GP alone increases the gap from 12.9 to 17.4 pp, and even GP+CoT leaves English at 9.4 pp, far above the 5.8 pp cross-language average. No confidence intervals or significance tests are reported, and several paradigms regress (e.g., gpt-3.5 ellipsis drops from 76.0 base to 72.0 GP+CoT; Chinese wh-fronting drops from 98.3 to 88.7 with GP+CoT). Thus the claim that grammar prompting 'cuts the capacity gap by 56%' is not established for the full benchmarks; it is established only for a hand-picked challenging subset.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes 'grammar prompting' (GP), an explain-then-process paradigm in which a large LLM first generates a concise metalinguistic explanation of a syntactic paradigm, and this explanation is then fed back as context to the target model (an LLM or a smaller SLM) before it chooses the grammatical sentence in a minimal pair. The method is evaluated on English BLiMP, Chinese SLING, and Russian RuBLiMP with five models (GPT-4o, GPT-3.5, Claude 3.5 Sonnet, Claude Haiku, Llama 3.3 9B), across base, CoT, control, textbook, and three-shot conditions. The central reported result is that on SLMs, GP alone reduces the average LLM-SLM accuracy gap by about 20%, and GP+CoT reduces it by 56% (13.0 pp to 5.8 pp), with the abstract and conclusion presenting this as bringing small models close to frontier-LLM performance.","tokens_in":22079,"tokens_out":4578,"duration_ms":38779,"significance":"If the central claim were established on full benchmarks, this would be a practically useful and inexpensive finding: the method requires only a single generated explanation per paradigm, works across three typologically different languages, and the experimental design includes valuable controls (an irrelevant-explanation control, a textbook-style condition, and three-shot baselines). The paper also documents concrete failure cases, which is a strength. However, the headline quantitative claim is computed on a post-hoc subset of 'challenging' paradigms, not on the full benchmarks, and the cross-language average hides substantial heterogeneity. The contribution is therefore better characterized as a promising prompt-engineering result on a selected subset than as the general equalizer claimed in the abstract.","major_comments":[{"comment":"The headline '56% reduction' (13.0 pp to 5.8 pp) is computed only over a post-hoc selection of challenging paradigms, not over the full benchmarks. The English gap on the full BLiMP (Table 8: gpt-3.5 84.0 vs. gpt-4o 91.2) is 7.2 pp, whereas the selected English subset in Table 7 has a base gap of 12.9 pp. Because the selection thresholds are based on gpt-4o performance (BLiMP categories at or below 90%, SLING categories excluding easier ones, RuBLiMP categories at or below 96%), the baseline gap is inflated and the general claim that grammar prompting 'cuts the capacity gap by 56%' is not supported for the full benchmarks. The authors should either report the gap analysis on all paradigms or explicitly and prominently reframe the claim as applying to the selected challenging subset.","section":"§4.2, §5.7, Tables 7–8"},{"comment":"The cross-language average conceals strong per-language heterogeneity. For English, GP alone increases the LLM-SLM gap from 12.9 to 17.4 pp, and GP+CoT leaves it at 9.4 pp, while Chinese and Russian gaps under GP+CoT drop to 3.4 and 4.4 pp. The statement in §5.7 and §7 that three SLMs finish within 5.8 pp 'across English, Chinese, and Russian' is therefore misleading: the 5.8 pp figure is an average, and English remains roughly twice that. The paper should present the per-language gap results as a central result and temper the equalizer claim accordingly.","section":"Table 7, §5.7"},{"comment":"No uncertainty quantification is provided for the main effect. Accuracies are averages over the first 50 items per paradigm and three A/B order trials, so item-level bootstrap confidence intervals or per-paradigm standard errors are feasible. Without them, the precise 13.0-to-5.8 pp reduction and the derived 20% and 56% relative reductions cannot be distinguished from sampling noise, especially given only three SLMs and small numbers of selected categories per language.","section":"§4.2–§4.3, §5.7"},{"comment":"The claim that grammar prompting 'consistently outperforms' base and CoT conditions is contradicted by systematic regressions: gpt-3.5 English ellipsis drops from 76.0 (base) to 72.0 (GP+CoT, Sonnet prompt); gpt-4o English ellipsis falls from 70.3 to 59.3 (GP+CoT, Sonnet); and gpt-3.5 Chinese wh-fronting falls from 98.3 to 91.7/88.7 under GP+CoT, a case the paper itself documents in Appendix C.2. The discussion should acknowledge these paradigms and state the scope of the improvement claim rather than presenting the gains as uniform.","section":"§5.3, Tables 3–4, Appendix C.2"}],"minor_comments":[{"comment":"There is a typo in 'aninstruction template' near the description of the instruction template; it should read 'an instruction template.'","section":"§3"},{"comment":"'The grammar prompting conditions consistently outperforms basic and chain-of-thought conditions' has a subject-verb agreement error and should be rephrased.","section":"§5.3"},{"comment":"The notation 'GP(b)-o1' is used without an explicit caption definition; the caption should state that this is GP with the o1-generated beginner prompt.","section":"Table 6"},{"comment":"The choice to use the first 50 minimal pairs per paradigm is not justified; if item ordering affects difficulty, this could bias the per-paradigm estimates, so a sentence on why this truncation is appropriate (or a robustness check) would help.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the control and three-shot comparisons are genuine strengths. The main risk is the mismatch between the advertised headline and the selected-subset analysis: the abstract and conclusion claim a general 'capacity gap' reduction, while the evidence supports a more limited claim about hand-picked challenging paradigms. I would accept a revision that reframes the claims, adds full-benchmark gap numbers, and reports uncertainty. I do not see a need for rejection, since the core method and controlled comparisons are defensible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jim -\n\nQuick read on 2506.02302. The core idea is simple and the execution is mostly careful: have a strong LLM write a short metalinguistic explanation of the target paradigm, then feed that explanation to the model (usually a small one) before a minimal-pair acceptability judgment. That is a reasonable application of explanation-in-context prompting to syntactic judgments, and the paper does several things right. The control condition (an irrelevant explanation) and the textbook condition (a pile of explanations) are exactly the right checks, and the pattern is consistent with what you'd want: relevant explanations help, irrelevant ones don't, and too much information hurts. The three-language design and the beginner/expert split are also genuinely informative. The appendix's failure analysis—models misparsing constituents or losing track of their own reasoning—is honest and useful.\n\nThe soft spots are real but mostly fixable. The headline '56% gap reduction' is computed on a hand-picked set of challenging paradigms (BLiMP categories where gpt-4o is at or below 90%, six SLING categories, seven RuBLiMP categories). On full BLiMP the base LLM-SLM gap is about 7 points, not 13, so the equalization story is much weaker on the full benchmark. Per-language, English is the weak case: GP alone widens the gap (12.9 to 17.4) and GP+CoT still leaves 9.4 pp, far from the 5.8 pp average. No error bars or significance tests on the main effect, and several paradigms regress. There is also no code or data release, which matters for a prompting study where prompt templates are the artifact. The explanation source (Sonnet or o1) is in the same model family as the evaluation targets, which isn't circular but does make the 'large LLM writes the rule' story a little less clean.\n\nStill, the central direction holds. The controlled comparisons show the method works on many hard paradigms, especially for small models, and the mechanism story is plausible. The paper just oversells the full-benchmark equalization. With the subset analysis reframed as 'hard-paradigm improvements' and full-benchmark results given equal weight, this would be a solid ACL-type contribution.\n\nMy take: worth a serious referee. I'd send it to review but with a clear request for full-benchmark reporting, confidence intervals, and artifact release. I'd bring it to the reading group if people care about prompting or grammaticality evaluation.","headline":"An honest, controlled study of feeding LLM-written grammar rules back to SLMs, whose headline 56% gap reduction is a subset artifact but whose method and controls deserve a careful referee.","tokens_in":22613,"tokens_out":2450,"would_cite":true,"duration_ms":23000,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Grammar prompting — feeding a model's own explanation of a syntactic rule back to it before a minimal-pair judgment — closes most of the accuracy gap between small and large language models on English, Chinese, and Russian grammaticality…","keywords":["grammar prompting","explain-then-process","grammatical acceptability judgments","minimal pairs","chain-of-thought reasoning","small language models","multilingual evaluation","syntactic reasoning"],"falsifier":"Run the identical GP+CoT protocol on the full BLiMP (67 paradigms), SLING (38), and RuBLiMP (45) benchmarks, averaging over all paradigms with the same three trials per pair. If the small-model gains concentrate in the selected hard categories and the full-benchmark LLM-SLM gap remains near the 13-point baseline (or the 56% reduction shrinks to insignificance), the central claim — that one rule explanation closes most of the capacity gap — would be an artifact of paradigm selection, not a property of grammar prompting.","tokens_in":21582,"feed_emoji":"⚖️","tokens_out":6904,"duration_ms":56082,"temperature":0.7,"pith_summary":"The paper proposes \"grammar prompting,\" a two-step recipe: first a large language model writes a concise explanation of a syntactic phenomenon, then that explanation is inserted into the prompt of the model that must decide which sentence of a minimal pair is grammatical. The authors' central claim is that this \"explain-then-process\" loop turns latent rule knowledge into applied rule use, and that it helps smaller models most. On selected challenging paradigms from BLiMP (English), SLING (Chinese), and RuBLiMP (Russian), grammar prompting alone cuts the average LLM-SLM accuracy gap from 13.0 to 10.4 percentage points, and grammar prompting combined with chain-of-thought reasoning cuts it to 5.8 points, a 56% reduction. The authors argue this lets low-cost smaller models approach frontier-LLM performance on structured linguistic judgments without any training or task-specific tuning. The headline gap numbers are computed on manually chosen hard paradigm subsets, not on full benchmarks.","feed_headline":"A grammar rule in the prompt cuts the LLM-SLM gap by 56%","feed_subtitle":"Feeding smaller models an LLM's rule explanation nearly closes the grammar-judgment gap across three languages.","key_machinery":"The load-bearing mechanism is the \"grammar prompt\": a few-hundred-word, model-generated explanation of the target syntactic phenomenon, written for a novice learner, with full example sentences deliberately excluded. It is combined with the minimal-pair judgment prompt, optionally with chain-of-thought reasoning (GP+CoT). The explanation supplies the linguistic categories and constraints (for example, that \"only\" licenses the negative-polarity item \"ever\" only when it scopes over the whole phrase), redirecting the model's reasoning from semantic paraphrase — the failure mode the paper diagnoses — to structural rule application. Beginner-oriented explanations outperform expert-oriented ones, and a single relevant explanation beats both an irrelevant control explanation and a multi-explanation \"textbook\" condition.","core_discovery":"The paper's discovery is that the failure of language models to judge grammatical acceptability despite being able to explain grammar reflects a knowing-versus-using gap, and that the gap can be bridged with the model's own metalinguistic output. The method, grammar prompting, first elicits a beginner-oriented explanation of a target phenomenon (e.g., negative polarity licensing, aspect marking) from a large model, then feeds that explanation back as context to the deciding model before it chooses between two minimal-pair sentences. Across English, Chinese, and Russian, this consistently improves accuracy over base and chain-of-thought conditions. The key quantitative claim is that with a single LLM-generated rule explanation plus chain-of-thought, three smaller models (GPT-3.5, Claude Haiku, and Llama 3.3 9B) finish on average within 5.8 percentage points of GPT-4o and Claude Sonnet, reducing the average LLM-SLM gap from 13.0 to 5.8 percentage points (a 56% relative reduction) on the selected challenging paradigm categories.","pith_inferences":["An implied extension, not tested in the paper, is that the same explain-then-process loop could improve other structural tasks where models know a rule but fail to apply it, such as syntax-aware translation or code generation.","The 'meaning-first' failure mode identified here suggests the benefit of grammar prompting may be partly to block paraphrase-based shortcuts; a direct testable corollary is that removing the explanation should revert accuracy to baseline, which the control condition already supports.","A reader should be cautious that the 13.0 to 5.8 percentage-point gap figure relies on hard paradigm subsets only; a fair comparison on all paradigms of each benchmark would reveal whether the method's average benefit is smaller than the headline number.","For languages where models cannot generate reliable explanations, one could test translating grammar prompts from high-resource languages or using a multilingual explanation generated by a frontier model."],"forward_implications":["Small models can be brought within a few points of frontier-LLM accuracy on grammatical acceptability judgments using only a short generated rule explanation, with no fine-tuning.","Grammar prompting works best for phenomena with clear distributional constraints or functional morphemes (NPI licensing, aspect, alternative questions) and helps least for constituent-sensitive island effects and lexical classifier-noun agreement.","A single relevant explanation is worth more than a compiled set of explanations: the \"textbook\" condition performed like the irrelevant control, implying models cannot reliably select the right rule from a larger body of grammar.","Because explanations are self-generated by the model rather than drawn from curated resources, the paradigm is portable to languages and domains where no grammar textbook exists, though the paper explicitly leaves low-resource languages untested.","Combined with chain-of-thought, the LLM-generated explanation lets SLMs reach roughly 90% accuracy or higher in Chinese and Russian where baselines were 74–82%."],"supporting_citations":[{"why":"Supplies the English BLiMP minimal-pair dataset that anchors the BLiMP experiments.","marker":"Warstadt et al., 2020"},{"why":"Supplies the Chinese SLING benchmark and its curated paradigms.","marker":"Song et al., 2022"},{"why":"Supplies the Russian RuBLiMP benchmark and paradigm categories.","marker":"Taktasheva et al., 2024"},{"why":"The MTOB grammar-book method motivates eliciting grammatical explanations from the model.","marker":"Tanzer et al., 2024"},{"why":"Provides the formal-versus-functional competence distinction used to frame the knowing-using gap.","marker":"Mahowald et al., 2024"},{"why":"Grounds the grammatical acceptability judgment task in psycholinguistic methodology.","marker":"Schütze, 2016"},{"why":"Evidence that examples can mislead in grammar-book prompting, cited to justify excluding full example sentences from grammar prompts.","marker":"Aycock et al., 2024"},{"why":"Shows that explanations in context can improve LLM task performance, the prior result that grammar prompting extends.","marker":"Lampinen et al., 2022"}],"fun_headline_variants":["Grammar prompting: one explanation shrinks LLM-SLM gap by 56%","Explain-then-process: 56% smaller grammar-judgment gap for smaller models","Feed a rule explanation, cut the LLM-SLM accuracy gap 56%","Know the rule, use it: grammar prompting closes 56% of the model gap","LLM's grammar rule feedback trims small-model gap 56%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The gap-reduction numbers are computed only over manually selected 'challenging' paradigm categories — BLiMP categories where gpt-4o scored at or below 90%, six SLING categories chosen by hand, and RuBLiMP categories at or below 96% — and the paper assumes that subset represents the general LLM-SLM gap.","fun_headline_variants_meta":{"raw":{"variants":["Grammar prompting: one explanation shrinks LLM-SLM gap by 56%","Explain-then-process: 56% smaller grammar-judgment gap for smaller models","Feed a rule explanation, cut the LLM-SLM accuracy gap 56%","Know the rule, use it: grammar prompting closes 56% of the model gap","LLM's grammar rule feedback trims small-model gap 56%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1392,"prompt_tokens":972,"completion_tokens":420,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":315}},"tokens_in":588,"tokens_out":420,"duration_ms":4324,"temperature":1.0,"reasoning_tokens":315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:26:32.879764+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical GP+CoT protocol on the full BLiMP (67 paradigms), SLING (38), and RuBLiMP (45) benchmarks, averaging over all paradigms with the same three trials per pair. If the small-model gains concentrate in the selected hard categories and the full-benchmark LLM-SLM gap remains near the 13-point baseline (or the 56% reduction shrinks to insignificance), the central claim — that one rule explanation closes most of the capacity gap — would be an artifact of paradigm selection, not a property of grammar prompting.","supporting_citations":[{"cited_title":"RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs","cited_arxiv_id":"2406.19232","evidence_quote":"Supplies the Russian RuBLiMP benchmark and paradigm categories."},{"cited_title":"Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?","cited_arxiv_id":"2409.19151","evidence_quote":"Evidence that examples can mislead in grammar-book prompting, cited to justify excluding full example sentences from grammar prompts."}],"review_version":1}