{"id":"a62e4fd6-e392-413b-bf99-4dfa2e7cd76f","arxiv_id":"2505.17037","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Synonymizing prompts with more specific vocabulary does not generally improve LLM accuracy and can hurt reasoning when verbs are made more specific.","lead":"This thesis tests whether using more specific synonyms in prompts changes how well large language models answer STEM, law, and medicine questions. It finds that specificity mostly does not matter, except for verbs in math reasoning, where it hurts, and points to a mid-level specificity range where models do best.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed optimal specificity range is the span of the experimental grid (Tab. 8/9), not a performance peak; the paper's own controlled experiments show null or negative effects of specificity, so the design rule is unsupported.","rationale":"The reader's CONDITIONAL verdict is appropriate, and my concern supports it, though through a different route than the stated weakest_assumption. The reader's weakest_assumption focused on WSD misclassification biasing synonym categories; that is a genuine limitation of the Q2 manipulation, but the headline optimal-range claim (Q3) is built from original, unmodified prompts and never passes through WSD. The load-bearing problem is that Table 8's 'ranges' match the min/max/median of the nine experimental prompt-specificity values in Table 9, not any performance comparison; the paper's own controlled experiments show null or negative correlations (Table 7), so moving prompts into the claimed range is never demonstrated to improve accuracy. The Q3 histograms (Fig. 10) are unadjusted marginal associations, confounded by dataset composition and difficulty. This is an internal inconsistency: the evidence adduced for the central claim fails to support it. Credit is due for the synonymization framework and the replicated negative verb effect in GSM8K, and for the paper's honest reporting of null results; but the abstract and conclusion overstate the range as a design rule. A reanalysis of existing tables or a prospective manipulation test would settle the matter. I therefore keep the verdict UNCHANGED and partially agree with the reader's weakest_assumption.","tokens_in":26124,"tokens_out":9660,"duration_ms":89014,"concrete_test":"Reanalyze the existing data: for each model and part of speech, plot the mean accuracy of the nine (specificity level x replacement ratio) conditions, already encoded in Figs. 8–9, against the nine prompt-specificity values in Table 9, and locate the performance maximum. If the maximum sits at the lowest specificity condition or accuracy is monotone decreasing, then Table 8's 'optimal range' is an artifact of the experimental grid, not a performance peak. A decisive follow-up would be a prospective manipulation: select original prompts below and above the claimed range, synonymize nouns or verbs to move their specificity into the range, and compare accuracy against the original prompts; no accuracy gain would refute the design rule.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central design rule — noun specificity ~17.7–19.7 and verb specificity ~8.1–10.6 yields best performance — rests on Section 4, Q3, but that analysis does not support it. Table 8's 'optimal ranges' are simply the minimum-to-maximum span of the nine prompt-specificity values produced by the synonymization grid (Table 9, 'All' rows); they carry no performance information. The abstract's 17.7–19.7 mixes the Q1 KDE boundary (17.74) with one model's median (19.70), and the Fig. 10 histogram modes (17.72–18.79) are a third number. More importantly, the controlled manipulation in Q2 (Table 7, Figs. 8–9) shows null effects for nouns and significantly negative effects for verbs in GSM8K (e.g., Llama-3.1-70B: rho = -0.89, p = 5.4e-4); no condition shows an interior performance peak inside the claimed range. The Q3 histogram analysis uses only original, unmodified prompts, so the concentration of correct answers near specificity 18 is a confounded marginal association: original questions with higher specificity may come from different domains or difficulty levels. The paper frames the range as a design rule ('manipulating prompts within this range could maximize LLM performance'), which would require a prospective test that is never run. The WSD limitation (Section 6) is real but secondary: it affects the Q2 manipulation, not the Q3 original-prompt analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a synonymization framework that replaces nouns, verbs, and adjectives in domain-specific prompts with synonyms of varying WordNet-based specificity, at replacement ratios of 33%, 67%, and 100%. It evaluates four LLMs (Llama-3.1-70B-Instruct, Granite-13B-Instruct-V2, Flan-T5-XL, and Mistral-Large 2) on MMLU, GPQA, and GSM8K tasks. The main reported findings are that increasing prompt specificity generally has little or negative impact on accuracy, and that there exists an 'optimal specificity range' (noun specificity roughly 17.7–19.7, verb specificity roughly 8.1–10.6) within which LLMs perform best. The paper also introduces an adjective-specificity measure and presents a GPT-4o-based validation of that measure.","tokens_in":26466,"tokens_out":3738,"duration_ms":39301,"significance":"If the optimal-range claim were supported, it would provide a practical, actionable rule for prompt design. The controlled manipulation experiments are a useful negative result: they suggest that blanket increases in specificity do not help and can hurt, especially for verbs in reasoning tasks. The framework is transparent, spans four models and three task formats, and includes a WSD validation step; this is a solid basis for follow-up work. However, the central design rule is not supported by the controlled experiments as reported, and the current evidence supports only a descriptive, exploratory statement about where correct answers concentrate on original prompts.","major_comments":[{"comment":"The claimed optimal specificity range is derived entirely from original, unmodified prompts, not from the specificity-manipulation experiments. The histograms in Fig. 10 are marginal associations: original questions with higher specificity may differ in domain, topic, or difficulty, so the concentration of correct answers near a given specificity value is confounded. The abstract's design rule ('manipulating prompts within this range could maximize LLM performance') would require a prospective test in which prompts are independently varied to fall inside versus outside the range; no such test is presented. At minimum, the claim should be reframed as an exploratory hypothesis, not a design rule.","section":"Section 4, Q3; Fig. 10; Table 8"},{"comment":"The numeric 'optimal ranges' in Table 8 appear to be the minimum-to-maximum span of the nine prompt-specificity values produced by the synonymization grid in Table 9, not performance peaks estimated from accuracy. Table 9 reports specificity values per replacement ratio and specificity level but does not report any accuracy associated with those values, and its caption calls them 'Optimal Specificities' without a performance criterion. The abstract's 17.7–19.7 noun range mixes the Q1 KDE boundary (17.74) with one model's median (19.70), and Fig. 10's histogram modes are yet another set of numbers; the paper should specify exactly which statistic defines the optimal range and justify why that statistic identifies an optimum.","section":"Tables 8 and 9"},{"comment":"The controlled manipulation experiments directly undercut the optimal-range claim. Table 7 shows that only 5 of 24 correlations are significant, and the significant correlations are negative, e.g., verb specificity in GSM8K for Llama-3.1-70B-Instruct (rho = -0.89, p = 5.4e-4) and Mistral-Large 2 (rho = -0.87, p = 0.001). No condition in the controlled experiments shows an interior performance peak inside the claimed optimal ranges. If the optimal-range claim is to be retained, the authors must explain why the controlled manipulation does not produce the predicted non-monotonic pattern, or they must soften the conclusion to a descriptive observation about original prompts.","section":"Section 4, Q2; Table 7; Figs. 8–9"},{"comment":"The WSD accuracy of 0.79 means that roughly one in five synonym choices may use the wrong sense, and the paper assumes this adds only noise. That assumption is load-bearing for Q2: if misidentified senses concentrate in particular specificity categories, the observed accuracy differences could reflect semantic corruption rather than specificity per se. The limitation is acknowledged in Section 6, but the manuscript should provide a sensitivity analysis, for instance by re-running the key correlations on a human-verified subset of synonym substitutions, before the null effect for nouns and the negative effect for verbs can be attributed to specificity.","section":"Section 6 and Section 3.3"}],"minor_comments":[{"comment":"The text says 'Since all p-values are smaller than 0.05, we cannot reject the null hypothesis and therefore the distributions cannot be considered as normal distributions.' This is backwards: a p-value below 0.05 leads to rejection of the null hypothesis of normality, which is the intended conclusion. The wording should be corrected.","section":"Section 4, Q1"},{"comment":"The phrase 'precentral changes in accuracy' appears to be a typo; it should read 'percentage changes in accuracy'.","section":"Section 2, Q2 (categorical approach)"},{"comment":"The dataset name 'MMMUL' appears once in Table 1; this should be 'MMLU'.","section":"Section 3.1; Table 1"},{"comment":"Figure 10's dashed line is described as marking 'the prompt specificity associated with the highest number of correct answers,' but Table 8 reports medians. The paper should clarify whether the optimal specificity is defined by the histogram mode, the median, or some other statistic, and use that definition consistently.","section":"Figure 10 and Table 8"},{"comment":"The adjective specificity equation is based on four equal-weight additive assumptions, and the GPT-4o validation reports a median Spearman correlation of 0.50 with 32 of 91 negative correlations. The paper should state more explicitly that this is a preliminary measure, and that the small adjective sample sizes (e.g., 4 GSM8K samples) preclude any conclusions about adjective effects on performance.","section":"Section 3.3, adjective specificity"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a master's thesis with a transparent experimental pipeline and a useful negative result, but the headline 'optimal specificity range' is not supported by the controlled experiments and appears to be a post hoc description of original-prompt data. The revision should either add a prospective test of the range or substantially weaken the claim, and should fix the statistical reporting issues. I do not see grounds for rejection, because the core negative-result direction is defensible and the methodological framework has value, but the current framing overstates the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you spend time on this. First, the controlled experiments are the honest part: synonym-based specificity changes mostly do nothing, and for verbs in GSM8K the effect is significantly negative in three of the four models. That is a useful calibration result for prompt designers. Second, the advertised 'optimal specificity range' (nouns ~17.7–19.7, verbs ~8.1–10.6) is not supported by the paper's own data. Table 8's ranges are just the min-to-max span of the nine prompt-specificity values produced by the synonymization grid (Table 9); they carry no performance information. The abstract appears to mix the Q1 KDE boundary with one model's median, and Figure 10's histogram modes are yet another number. The Q3 analysis only uses original, unmodified prompts, so any association between specificity and correct answers is confounded with domain and difficulty. No prospective manipulation of the claimed range is ever run, so the design rule in the abstract and conclusion is overreach.\n\nWhat the paper does well: the synonymization framework is systematic and reusable—three specificity levels, three replacement ratios, four models, three domains—and the authors are upfront about the null result and the tiny adjective sample (4 GSM8K instructions). The WSD accuracy check (Table 3) is a reasonable sanity check, even if 0.79 accuracy means a real error rate in the pipeline. The adjective specificity equation is a genuine first step, but the validation against GPT-4o shows only moderate median correlation (0.50) with negative correlations in 32 of 91 cases, so it should be treated as exploratory.\n\nSoft spots, in proportion: the missing prospective test is the load-bearing one. The lack of released code/data and the absence of confidence intervals on accuracies are correctable and should be fixed before any publication. The WSD limitation is real but secondary—it affects the Q2 manipulation more than the Q3 original-prompt analysis. None of this is fatal to the controlled null/negative finding; that part holds up.\n\nWho this is for: people working on prompt sensitivity or lexical variation in LLMs will want the framework and the verb-reasoning negative result, but they should ignore the optimal-range claim. I would send this to peer review with heavy revision expected, mainly to force a prospective test or a much more careful framing of Q3. The reader's conditional verdict is fair, and the stress-test note lands.\n\nMy recommendation: engage with the controlled results, discard or heavily rework the optimal-range claim, and ask for artifacts before accepting.","headline":"The paper's real finding is a null/negative controlled result about prompt specificity; its headline 'optimal specificity range' is a post-hoc descriptive artifact, not a tested design rule.","tokens_in":26947,"tokens_out":1284,"would_cite":false,"duration_ms":15342,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that LLM performance peaks inside a mid-range of prompt vocabulary specificity, not at maximum specificity.","keywords":["prompt specificity","prompt engineering","synonymization","WordNet","large language models","domain knowledge","question answering","reasoning tasks"],"falsifier":"Take the same prompts, correct the word senses by hand before synonymizing, and re-measure the accuracy-versus-specificity curves: if the reported optimal bands for nouns and verbs shift or disappear, the ranges were artifacts of sense errors rather than effects of specificity. A quicker check is to tabulate the disambiguation error rate per specificity category (low, intermediate, high) across the three datasets.","tokens_in":25883,"feed_emoji":"🎯","tokens_out":5964,"duration_ms":55728,"temperature":0.7,"pith_summary":"The thesis asks whether making the vocabulary of prompts more specific improves how large language models answer expert questions in STEM, law, and medicine. It builds a synonymization framework that swaps nouns, verbs, and adjectives for synonyms of measured low, intermediate, or high specificity, applied at three replacement rates to samples from MMLU, GPQA, and GSM8K. The central finding, consistent across all four tested models, is that performance is best inside a middle band of prompt specificity, roughly 17.7 to 19.7 for nouns and 8.1 to 10.6 for verbs, and that pushing beyond the band does not help nouns and significantly hurts verbs in reasoning tasks. If correct, the practical rule for prompt design is to tune specificity into the band rather than maximize it.","feed_headline":"LLMs answer best at mid-range prompt specificity","feed_subtitle":"Across STEM, law, and medicine, four models peak in the same bands; over-specific verbs hurt reasoning most.","key_machinery":"The load-bearing mechanism is a five-step synonymization pipeline: part-of-speech tagging, word sense disambiguation (performed by Llama-3.1-70B-Instruct, measured at 0.79 accuracy), crawling of synonyms through WordNet hypernym and hyponym relations, assignment of continuous specificity scores, and replacement of words at 33, 67, and 100 percent rates. For nouns and verbs, specificity is computed as distance from the taxonomy root plus a normalized hyponym count minus a sense-count penalty; for adjectives, a new fractional measure is introduced based on counts of similar words, synonyms, antonyms, and senses. These scores convert coarse low/intermediate/high categories into a continuous prompt-level specificity value, which is what lets the paper locate the performance peaks as specific numerical bands.","core_discovery":"Prompt specificity has a non-monotonic effect on LLM performance: for nouns and verbs there is an optimal specificity interval, shared across four differently sized and differently trained models, within which question-answering and reasoning accuracy peaks, with performance degrading or plateauing outside it. The claim is supported by a stepwise synonymization of sampled instructions, where prompts are assigned continuous specificity scores and correlated with accuracy. Beyond the optimal range, raising noun specificity leaves performance largely unchanged, while raising verb specificity in the reasoning benchmark (GSM8K) yields strong negative correlations, for example -0.89 for Llama-3.1-70B-Instruct and -0.87 for Mistral-Large 2. Only 5 of 24 model-dataset evaluations reach statistical significance, so the paper's headline negative result is concentrated in verbs on reasoning tasks; the paper presents the optimal band as the actionable discovery, correcting the intuition that maximally specific expert wording is always best.","pith_inferences":["A natural testable extension is variance: within the optimal band, model outputs may be more stable across seeds and paraphrases; if the band is where performance peaks and jitter drops, it becomes a genuine operating point rather than a statistical artifact.","The adjective-specificity equation, validated only in passing (median Spearman 0.50 against an LLM judge), may be the most portable piece: it could be tested on sentiment or concreteness benchmarks where adjectives carry the semantics, and if it survives there it becomes a general lexical tool.","Because the noun optimum (about 17.7 to 19.7) sits below the corpus mean specificity (21.37) while verbs cluster near their own mean, the band may track the lexical density of pretraining text; that predicts the band shifts for other languages or heavily technical corpora.","Read together with the mostly flat noun results, the findings suggest specificity is a cost curve with a flat middle: a prompt-selection heuristic that holds specificity inside the band while minimizing verb perturbation could be evaluated on held-out benchmarks as a practical optimization rule."],"forward_implications":["Prompt designers in specialized domains should target the mid-specificity band instead of maximizing term precision, since over-specific nouns do not help and over-specific verbs measurably degrade reasoning.","The optimal band's consistency across four models of different families and sizes suggests a shared property of how LLMs process prompt wording rather than a quirk of any single model.","Verb specificity is the more consequential lever: the significant negative correlations concentrate in verbs on the chain-of-thought math task, implying that rewording actions and relations disrupts stepwise reasoning more than rewording entities.","The WordNet-based specificity score offers a cheap, parameter-free diagnostic that could complement existing prompt-sensitivity methods in deciding when a prompt rewrite is worth testing."],"supporting_citations":[{"why":"Supplies the base specificity formula for nouns and verbs that the thesis extends with a sense-count penalty.","marker":"Bolognesi, Burgers, and Caselli (2020)"},{"why":"WordNet, the lexical database whose hypernym/hyponym taxonomies define the specificity geometry for nouns and verbs.","marker":"Fellbaum (1998)"},{"why":"Provides the word sense disambiguation approach the framework depends on; the thesis substitutes Llama-3.1-70B-Instruct after benchmarking it against the original fine-tuned T5.","marker":"Wahle, Ruas, Meuschke, et al. (2021)"},{"why":"MMLU, one of the three benchmark suites whose STEM, law, and medicine subsets carry the evaluation.","marker":"Hendrycks et al. (2021)"},{"why":"GPQA, the PhD-level question-answering benchmark used as the hardest test of specificity effects.","marker":"Rein et al. (2023)"},{"why":"GSM8K, the grade-school math reasoning benchmark where the significant negative verb-specificity effects appear.","marker":"Cobbe et al. (2021b)"},{"why":"Prior evidence that domain vocabulary in prompts can improve LLM performance; the thesis extends it with a stepwise specificity analysis.","marker":"Zheng et al. (2023)"},{"why":"Evidence that synonym substitution shifts prompt performance; the thesis categorizes those synonyms by specificity level to test where the effect lives.","marker":"Leidinger, van Rooij, and Shutova (2023)"}],"fun_headline_variants":["Mid-range prompt specificity: LLMs' hidden sweet spot","Optimal prompt specificity is a band, not a ceiling","Four LLMs converge on a mid-range prompt specificity optimum","LLMs hit accuracy sweet spot at mid-range prompt specificity","Prompt specificity: too vague or expert? Try the middle"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes that the 21 percent of word senses the disambiguation step gets wrong only add random noise and do not systematically push wrong-sense synonyms into particular specificity categories.","fun_headline_variants_meta":{"raw":{"variants":["Mid-range prompt specificity: LLMs' hidden sweet spot","Optimal prompt specificity is a band, not a ceiling","Four LLMs converge on a mid-range prompt specificity optimum","LLMs hit accuracy sweet spot at mid-range prompt specificity","Prompt specificity: too vague or expert? Try the middle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000981,"raw_usage":{"total_tokens":4163,"prompt_tokens":939,"completion_tokens":3224,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":3143}},"tokens_in":555,"tokens_out":3224,"duration_ms":22269,"temperature":1.0,"reasoning_tokens":3143,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:39:10.248100+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same prompts, correct the word senses by hand before synonymizing, and re-measure the accuracy-versus-specificity curves: if the reported optimal bands for nouns and verbs shift or disappear, the ranges were artifacts of sense errors rather than effects of specificity. A quicker check is to tabulate the disambiguation error rate per specificity category (low, intermediate, high) across the three datasets.","supporting_citations":[],"review_version":1}