{"id":"74c15841-9b06-48d9-9a73-996d1132f25f","arxiv_id":"2508.13365","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Reasoning-tuned DeepSeek models give only small, inconsistent gains over base models on idiomaticity detection, and supplying definitions from large models can improve smaller models on some datasets.","lead":"This paper tests whether 'reasoning' versions of AI language models detect idioms better than their non-reasoning counterparts, across four datasets and five model sizes. The gains are small and inconsistent, and feeding small models definitions generated by large models helps on some tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reasoning-vs-base effect sizes in Table 2 have no uncertainty or significance tests; 5-seed means are too thin to support the size-dependence claim.","rationale":"The reader's weakest assumption (no inter-annotator agreement, author-only manual labels) is valid but mostly limits RQ2 ('understanding') and the distillation interpretation; it does not bear directly on the headline effect-of-reasoning claim. The load-bearing condition for that claim is that the Table 2 differences are signal rather than seed-level noise. The paper reports no standard deviations or tests, and several key contrasts are tiny (14B +0.024, 32B +0.017, 70B +0.035 mean macro F1). With only five seeds and small evaluation sets, these could easily be noise. If a paired permutation/bootstrap check shows CIs crossing zero, the 'modest improvements for larger models' and the size-dependence pattern collapse; the remaining defensible conclusion is 'no clear effect.' I therefore keep the verdict conditional (UNCHANGED) but for a different reason than the reader's, and would require the per-seed scores or a re-run as a condition.","tokens_in":10913,"tokens_out":9846,"duration_ms":106984,"concrete_test":"Request the per-seed macro-F1 scores for Tables 1 and 5, or rerun with at least 20 seeds. For each base-vs-reasoning pair and dataset, compute a paired permutation test over seeds and bootstrap 95% confidence intervals on the mean difference; apply Benjamini-Hochberg correction across the 20+ hypotheses. If the 14B/32B/70B positive mean differences (0.024/0.017/0.035) have CIs crossing zero or p>0.05 after correction, the abstract's size-dependent 'modest improvements' claim should be weakened to 'no significant effect,' and the strongest_claim should be revised accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline conclusion rests on Table 2's point estimates: mean macro-F1 differences over 5 seeds, e.g., +0.024 (14B), +0.017 (32B), +0.035 (70B), and -0.167 (7B), with no standard deviations, confidence intervals, or significance tests reported anywhere. The datasets are relatively small (FLUTE test=250; DICE=2066; SemEval has 2342 sentences but only 150 PIEs), and five random seeds can easily produce mean shifts of 0.02-0.03 macro F1. The claim that larger models show modest improvements while smaller models show decreases depends on contrasts of exactly that magnitude and direction; if those differences are within run-to-run noise, the 'varies across model size' finding is not established, and the 'small effect' finding collapses to 'no detectable effect.' The Limitations section acknowledges inconsistency within Llama-70B but does not quantify uncertainty, and Table 5's 'significantly improved' claims are not accompanied by test statistics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates the suite of DeepSeek-R1 distilled reasoning models (1.5B–70B) on four idiomaticity detection datasets (FLUTE, SemEval 2022 Task 2a, MAGPIE, DICE), comparing them with their non-reasoning base models and, for the two smallest sizes, with intermediate math-tuned variants. Macro-F1 scores are averaged over five seeds. The main empirical claim is that the effect of reasoning on idiomaticity detection is small and varies with model size: larger models (14B, 32B, 70B) show modest improvements, while smaller models often underperform their base counterparts. The paper also reports a manual analysis of chain-of-thought outputs, arguing that larger models produce accurate idiom definitions while smaller models do not, and an experiment in which definitions generated by the 32B model are appended to the prompts of smaller models, yielding gains on FLUTE but not DICE.","tokens_in":11147,"tokens_out":5702,"duration_ms":64329,"significance":"If the headline results are reliable, the paper provides a useful empirical datapoint on reasoning models for idiomaticity detection and introduces a plausible knowledge-distillation idea: using definitions from a larger model to prompt smaller models. The work uses public datasets, open-source checkpoints, and a straightforward evaluation protocol, with no fitted models or circular derivations. The manual CoT analysis is a thoughtful attempt to separate definitional understanding from contextual disambiguation. However, the central size-dependence claim and the claimed significant distillation gains are not supported by the reported statistics: no standard deviations, confidence intervals, or significance tests are given for the differences that drive the conclusions. As a result, the contribution is currently exploratory and requires substantial strengthening before the claims can be accepted.","major_comments":[{"comment":"The headline conclusion—reasoning has a small, size-dependent effect—rests entirely on 5-seed mean macro-F1 differences without any reported variance or test statistics. For example, Table 2 reports +0.024 (14B), +0.017 (32B), +0.035 (70B), and −0.167 (7B) as bare means. With FLUTE’s test set of 250 examples, MAGPIE’s test split of roughly 1% of 4,840 sentences, and only five seeds, these differences are well within plausible run-to-run noise. The Limitations section acknowledges Llama-70B’s inconsistency (worse by 0.037 on SemEval English) but does not quantify uncertainty. Please report per-seed standard deviations and confidence intervals, and perform paired significance tests (e.g., bootstrap across seeds or McNemar/permutation over items). If the contrasts are not significant, the ‘small effect’ finding reduces to ‘no detectable effect,’ which is a materially weaker claim.","section":"§3, Tables 1–2"},{"comment":"The caption and text claim ‘significantly improved’ results for the definition-prompting experiment, and ‘no significant difference’ for DICE and for the 14B model, yet no significance test, p-value, confidence interval, or standard deviation is reported. State the test used (paired bootstrap across the five seeds, McNemar over items, or similar) and report its outcomes. For the null results on DICE, report the smallest effect size the design could plausibly detect; otherwise the conclusion that definition prompting ‘can be applied generally without risk of regression’ is unsupported.","section":"§6.1, Table 5"},{"comment":"The manual analysis uses 30 responses per model per dataset (15 correct, 15 incorrect), annotated by the three authors with no inter-annotator agreement reported. Table 3’s percentages are therefore based on 15 ‘incorrect’ predictions per row (e.g., 40% = 6 examples). The qualitative conclusion that larger models have better idiomatic understanding—and the subsequent choice of the 32B model as definition generator—rests on these scores. Please report per-label agreement (e.g., Cohen’s kappa), provide full label distributions, and ideally have annotation performed blind to model identity and correctness. At minimum, present Table 3 as illustrative and state its sample-size limitations explicitly.","section":"§4.1–§4.2, Figure 1, Table 3"}],"minor_comments":[{"comment":"The GPT-4o label extraction step is used without reporting extraction accuracy or the frequency of unparseable outputs. If a large fraction of outputs require extraction, the reported scores mix model behavior with extractor behavior; please report both quantities.","section":"§2.2"},{"comment":"The correlation table reports p-values below 0.05 without describing the test or applying any multiple-testing correction. The pseudo-R² values mentioned in the text are not tabulated; please include them or remove the claim.","section":"Table 4 / §5"},{"comment":"MAGPIE’s test split is very small (1% of 4,840 sentences), which makes per-dataset macro-F1 scores noisy. This should be stated wherever MAGPIE results are interpreted, and it strengthens the need for uncertainty estimates.","section":"§2.1"},{"comment":"Definition quality is illustrated by five example definitions; consider reporting automatic or manual quality statistics on the full generated set so the reader can assess the distillation source.","section":"§6, Table 6"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern from the reader is well-founded: the central size-dependence and distillation claims are not statistically grounded. The paper is otherwise methodologically straightforward and likely suitable for the journal after the uncertainty analysis is added. I would also ask the authors to double-check the MAGPIE test-set size and the interpretation of the SemEval combined few-/zero-shot test sets, as both affect the precision of the reported differences."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Useful paper. It does something straightforward and valuable: evaluates the DeepSeek-R1 distilled family (1.5B to 70B) against their base and math-tuned variants on four idiom detection benchmarks. The main finding—reasoning tuning gives at best small, inconsistent gains—is honestly presented and likely robust. The definition-injection distillation idea is genuinely new and cheap: taking definitions from a strong 32B model and appending them to smaller models' prompts improved FLUTE by 0.069 macro F1 for the 1.5B model, with no significant regression elsewhere. Credit also for evaluating public checkpoints and datasets with clear prompting details, and for the Limitations section explicitly acknowledging the Llama-70B inconsistency. That is the kind of transparency that makes a paper useful even when its conclusions are modest.\n\nThe soft spots are real but proportionate. The stress-test note lands: Table 2's mean differences (e.g., -0.167 for 7B, +0.024 to +0.035 for 14B-70B) are presented without standard deviations, confidence intervals, or significance tests. Five seeds on datasets sized 250 to a few thousand examples can easily shift macro F1 by 0.02-0.03, so the 'varies across model size' claim is not statistically established. The authors should either report per-seed results with CIs or soften the size-dependence language. The manual analysis (300 examples, three co-author annotators, no inter-annotator agreement) supports the understanding-vs-reasoning distinction only suggestively; I would not hang strong conclusions on it. Also, Table 5 says 'significantly improved' but gives no test statistics. These are fixable with a revision.\n\nThat said, the central qualitative conclusion—reasoning tuning does not dramatically help idiomaticity detection—survives the noise concern, because even if the differences are within run-to-run variation, the effect sizes are small, not dramatic. The paper's cautious tone is appropriate. The definition distillation result is less affected by the statistical issue since the FLUTE improvements are sizable.\n\nWho is this for? Anyone working on idiom processing or evaluating open reasoning-tuned models in a practical setting. It gives concrete numbers and a cheap distillation method worth trying. It deserves a serious referee—send it to peer review, and require the revisions above on uncertainty quantification and the manual annotation.","headline":"A useful, honest empirical benchmark of reasoning models for idiom detection, but the headline size-dependent conclusions rest on 5-seed point estimates with no uncertainty.","tokens_in":11606,"tokens_out":1612,"would_cite":true,"duration_ms":19605,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Chain-of-thought reasoning yields only small, inconsistent gains for idiomaticity detection.","keywords":["reasoning models","chain-of-thought","idiomaticity detection","multiword expressions","DeepSeek-R1 distillation","definition prompting","LLM evaluation","figurative language"],"falsifier":"Take a held-out set of at least 100 MAGPIE and DICE instances per model, have multiple annotators independently label the CoTs for definition accuracy and reasoning quality, check inter-annotator agreement, and rerun the definition-prompt experiment with human-written definitions instead of the 32B model's; if agreement is low or the gains vanish, the reported understanding gap and distillation effect would not replicate.","tokens_in":10837,"feed_emoji":"🧠","tokens_out":5055,"duration_ms":48510,"temperature":0.7,"pith_summary":"This paper investigates whether adding chain-of-thought reasoning to open-source LLMs improves idiomaticity detection—deciding whether a phrase like \"play with fire\" is used literally or idiomatically in context. Evaluating DeepSeek-R1 distilled models from 1.5B to 70B parameters on four datasets, it finds the reasoning effect is small and inconsistent: small models gain little or lose, while 14B–70B models show modest gains. The central reason, the authors argue, is that smaller models fail to produce accurate definitions of the idioms, whereas larger models do; the remaining hard part is using context to disambiguate, which even the largest models get wrong on subtle cases. The paper then shows that feeding the larger models' definitions to smaller models can improve performance on one dataset (FLUTE) without hurting another (DICE), suggesting a practical distillation path.","feed_headline":"Reasoning barely helps LLMs spot idioms","feed_subtitle":"Across four benchmarks, small models lose or gain little from reasoning; definitions supplied by larger models help more.","key_machinery":"The central object is the DeepSeek-R1 distillation suite, an open-source family of reasoning models fine-tuned to emit chain-of-thought (CoT) text before a final answer, evaluated against their untuned base counterparts (Qwen2.5 and Llama-3.3) and, for the small sizes, the intermediate math-tuned checkpoints. The evaluation uses four idiomaticity detection datasets (FLUTE, SemEval 2022 Task 2a, MAGPIE, DICE) and a manual labelling rubric that scores each CoT separately for 'understanding' (whether the model gives a valid idiomatic definition) and 'reasoning' (whether it correctly disambiguates the expression in context). The same rubric underlies the paper's final experiment, where definitio","core_discovery":"The paper's central claim is that reasoning—generating a chain of thought before answering—has a smaller and more varied effect on idiomaticity detection than the recent enthusiasm for reasoning models would suggest. On four benchmarks (FLUTE, SemEval 2022 Task 2a, MAGPIE, DICE), moving from base models to DeepSeek-R1 reasoning variants improves macro F1 for 14B, 32B, and 70B models by small amounts, while the 1.5B and 7B models improve only relative to the math-tuned intermediate checkpoints and still underperform their base models. Manual inspection of reasoning traces shows the divide in capability: larger models usually produce accurate idiomatic definitions, smaller models often cannot,","pith_inferences":["Editorial extension: the definition-distillation result could generalize into a low-cost pipeline for under-resourced languages or idiom inventories, provided the teacher model's definitions are validated by human speakers first.","Editorial extension: the manual finding that many 'incorrect' larger-model predictions are judged valid against ambiguous gold labels suggests benchmark noise, not only model error, is at play; cleaning the datasets could change the reported ordering.","Editorial extension: a direct test of the paper's explanation would be to give small models human-written definitions instead of model-generated ones; if gains disappear, the definitions' quality, not their source, is the operative ingredient."],"forward_implications":["If the finding holds, adding CoT reasoning to small open-source models is not a reliable route to better idiom detection: the 1.5B and 7B reasoning models stay below their base variants on average.","Larger reasoning models gain modestly (up to ~0.095 macro F1 on MAGPIE), so reasoning helps most when the model already has enough linguistic knowledge.","Because reasoning quality—not definition knowledge—is the deciding factor for correctness, better context-disambiguation methods are the next bottleneck, not more idiom definitions.","Definition prompting is a safe, sometimes helpful intervention: it improved FLUTE scores for 1.5B and 7B models and did not significantly hurt DICE, so it can be applied generally without regression.","CoT length does not predict correctness, so users should not treat longer reasoning as a sign of better idiomaticity judgements."],"supporting_citations":[{"why":"Establishes the evaluation setup and the four-dataset comparison standard that this paper extends to reasoning models.","marker":"Phelps et al., 2024"},{"why":"Supplies the SemEval 2022 Task 2a multilingual idiomaticity detection test sets used in the evaluation.","marker":"Madabushi et al., 2022"},{"why":"Supplies the FLUTE dataset and its idiom subset, where definition prompting produces the reported gains.","marker":"Chakrabarty et al., 2022"},{"why":"Supplies the MAGPIE corpus, the main dataset for the manual CoT analysis and definition questions.","marker":"Haagsma et al., 2020"},{"why":"Supplies the DICE dataset, which controls for surface form and drives the claim that context disambiguation remains hard; also cited for why definitions do not help DICE.","marker":"Mi et al., 2024"},{"why":"Supplies the DeepSeek-R1 reasoning models and the distillation/CoT training scheme whose effect is measured.","marker":"DeepSeek-AI et al., 2025"},{"why":"Supplies the base Qwen2.5 models and the math-tuned intermediates that serve as the non-reasoning baselines.","marker":"Qwen, 2024; Qwen et al., 2025"},{"why":"Supplies the Llama-3.3 70B base model paired with the DeepSeek-R1 70B variant.","marker":"Llama Team, 2024"}],"fun_headline_variants":["Reasoning barely moves the needle on idiom detection","Chain-of-thought fails to boost small LLM idiom skills","Idiom spotting: reasoning helps big models, not small ones","Definitions, not reasoning, rescue small models on idioms","Reasoning on idioms: less impact than expected"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The manual analysis in Section 4 rests on 30 responses per model (15 correct, 15 incorrect) annotated only by the paper's authors, without reported inter-annotator agreement; the conclusion that larger models understand idioms while smaller ones do not depends on these labels being representative and unbiased.","fun_headline_variants_meta":{"raw":{"variants":["Reasoning barely moves the needle on idiom detection","Chain-of-thought fails to boost small LLM idiom skills","Idiom spotting: reasoning helps big models, not small ones","Definitions, not reasoning, rescue small models on idioms","Reasoning on idioms: less impact than expected"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000158,"raw_usage":{"total_tokens":1077,"prompt_tokens":771,"completion_tokens":306,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":227}},"tokens_in":515,"tokens_out":306,"duration_ms":4315,"temperature":1.0,"reasoning_tokens":227,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:04:03.366446+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out set of at least 100 MAGPIE and DICE instances per model, have multiple annotators independently label the CoTs for definition accuracy and reasoning quality, check inter-annotator agreement, and rerun the definition-prompt experiment with human-written definitions instead of the 32B model's; if agreement is low or the gains vanish, the reported understanding gap and distillation effect would not replicate.","supporting_citations":[],"review_version":1}