Pith. sign in

REVIEW 4 major objections 11 minor 69 references

A large accuracy drop from rare to brand-new Chinese xiehouyu riddles marks memorization in frontier Chinese LLMs, even as some models beat humans on novel items and still lag at creating them.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Δacc between low-frequency and novel xiehouyu is ~23.6% for Chinese frontier LLMs vs ~5.1% for English-centric models and ~2.9% for humans, while LLM-created xiehouyu rate below human creations.

T0 review reviewed 2026-07-30 challenge →

load-bearing objection Solid contamination-aware Chinese riddle benchmark; the Δacc story is useful but only partly identified, and the authors mostly own that. the 4 major comments →

arxiv 2607.23440 v1 pith:MFLLYBSX submitted 2026-07-26 cs.CL cs.AI

Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

classification cs.CL cs.AI
keywords xiehouyuLLM reasoningmemorizationdata contaminationChinese wordplayhomophonic punscreative generationmultiple-choice evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models truly reason about non-literal Chinese wordplay or mostly retrieve answers seen in training. The authors build X-Riddles, mixing dictionary xiehouyu with hundreds of new homophonic riddles written by linguists so the novel set cannot have leaked into pretraining. On multiple-choice matching, native speakers score almost the same on low-familiarity existing items and on novel ones, which the authors take as shared reasoning. Frontier Chinese models instead lose on average about 24 percentage points from low-familiarity to novel items, while English-centric models lose only about 5; the authors read that split as heavier Chinese-data memorization. At the same time, the best model reaches 92.6% on novel items—above human accuracy—yet model-written xiehouyu are rated less reasonable and less funny than human ones. The work argues that reasoning claims need contamination-aware tests, and that creative language play in this Chinese form still favors human experts.

Core claim

Using linguist-written novel xiehouyu to block contamination, the authors show that frontier Chinese models have a large mean accuracy drop (about 23.6 points) from low-familiarity existing homophonic items to novel ones, versus near-zero drop for native speakers and about 5 points for English-centric models—evidence they treat as memorization from larger Chinese training data—while Gemini 3.1 Pro still scores 92.6% on novel items (above human MCQ accuracy) and LLM-created xiehouyu receive worse reasonableness and funniness ratings than human creations.

What carries the argument

Δacc (accuracy on low-familiarity existing homophonic xiehouyu minus accuracy on novel ones): a lower-bound memorization index, justified because humans show almost no gap and are therefore argued to use the same reasoning process on both sets.

Load-bearing premise

The score gap between rare dictionary riddles and linguist-written new ones mainly reflects training-set memorization, not leftover difficulty, style, or distribution shift between those two sets—and the Chinese-versus-English contrast is driven mainly by Chinese data volume even though the model groups also differ in openness and scale.

What would settle it

Build a new matched novel set with the same human difficulty profile, or measure how often the low-familiarity dictionary items actually appear in public Chinese pretraining crawls: if the Low–New gap for Chinese models vanishes under tighter difficulty matching, or fails to track documented Chinese-data exposure, the memorization reading of Δacc weakens.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Benchmarks built from culturally circulated material need paired novel items of matched difficulty, or they will mix retrieval with reasoning.
  • Claims of strong LLM reasoning on Chinese figurative language should be re-checked on uncontaminated, expert-written items.
  • Homophonic Chinese wordplay remains a harder test than polysemy or direct proverb types for current models.
  • Even models that match or beat humans on novel multiple-choice xiehouyu still underperform human linguists at creating reasonable, funny ones.
  • Heavier Chinese pretraining can inflate scores on rare existing items without equal gains on truly new riddles.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Phonology-aware training or explicit pronunciation modules may be needed before models close the remaining gap on strict homophonic links.
  • The same Low-versus-New template could be ported to other language-specific games (puns, two-part allegories, dialect wordplay) to audit memorization outside Chinese.
  • Creation tasks may stay a stricter creativity filter than multiple-choice understanding even after contamination is controlled.
  • Open reporting of Chinese pretraining mix would let future work test whether Δacc scales with documented Chinese token volume.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 11 minor

Summary. The paper introduces X-Riddles, a benchmark of 1,143 Chinese xiehouyu: 900 dictionary-sampled items balanced across homophonic/polysemous/direct types and high/low familiarity, plus 243 novel homophonic items written by linguistics students and validated by the authors. Three experiments probe LLMs: MCQ riddle–answer matching (Exp 1), free-form explanation (Exp 2), and xiehouyu creation (Exp 3). The central index is Δacc = Acc(low-familiarity) − Acc(novel) on homophonic items: native speakers show Δ=2.9 (71.5→68.6), frontier Chinese models average 23.6 (e.g., DeepSeek-V3.2 95.8→64.9), and English-centric models 5.1. The authors interpret the Chinese-model gap as memorization from larger Chinese training data, while noting Gemini 3.1 Pro reaches 92.6% on novel items (above the human 68.6). Exp 2 reports degraded explanation quality and hallucinated homophones on novel items; Exp 3 finds DS-R1/o1 creations rated less reasonable and funny than human ones.

Significance. If the memorization reading holds, this is a valuable contribution to the retrieval-vs-reasoning debate and to evaluation methodology for culturally specific language. Named strengths: (i) expert-authored novel items designed to sidestep contamination, with an explicit validation protocol; (ii) a 100+ native-speaker MCQ baseline providing an external anchor for item difficulty (human Δ≈2.9); (iii) a matched-difficulty existing-vs-novel pairing that is a reusable decontamination template; (iv) broad model coverage with condition ablations (context, shots, CoT, literal vs figurative target); (v) a token-effort analysis yielding a falsifiable dissociation (Chinese models spend +40.6% tokens from Low to New while losing 23.6 points); (vi) full prompt transparency in appendices. The paper also reports a result against its own headline (Gemini 3.1 Pro, Δ=−0.2) and concedes the openness/scale confound in §4.3.4, which increases credibility. The creation experiment, if its evaluation is sound, is a rare controlled comparison of expert vs LLM linguistic creativity.

major comments (4)
  1. [§3, §4.3.3, Table 4] §3 and §4.3.3 (Table 4): 256 of 312 candidate novel items—hence the large majority of the retained 243—were seed-based, reusing the exact homophone pair of an existing dictionary xiehouyu (e.g., 舅/旧 from 外甥打灯笼——照旧). Exp 2 itself shows homophone identification is precisely where models fail; a model that memorized the seed gets the hardest link of the reasoning chain for free, so only the riddle→literal mapping is genuinely novel for ~4/5 of the 'uncontaminated' set. The abstract's 'to avoid data contamination' is therefore too strong, and Δacc's magnitude is not cleanly interpretable. A cheap, decisive check exists: report per-model accuracy and Δacc separately for the 56 free-created items vs the seed-based ones. If the Chinese/English contrast replicates on the free-created subset, the central claim is much better supported; either way, foreground the construction detail in the abstrac
  2. [§4.1, Table 4] All MCQ distractors are randomly sampled dictionary answers. Thus in the Low condition the correct option is an attested answer among attested answers, while in the New condition the correct option is a novel string among attested distractors. A model that prefers familiar/attested strings—plausible precisely for models trained on more Chinese text—would be helped on Low and hurt on New, producing a Δacc that reflects option-familiarity bias rather than, or in addition to, memorization of the items. The human Δ=2.9 controls intrinsic item difficulty but not model-specific weighting of this cue, since humans and models need not use surface familiarity the same way. A free-form answer-generation condition (no options) or a novel-distractor control on a subsample would discriminate the accounts; at minimum, this alternative should be analyzed and discussed.
  3. [§4.3.3–4.3.4, Table 4] The group contrast in Table 4 decomposes into two effects: Chinese models are higher on Low (90.6 vs 82.7) and lower on New (67.0 vs 77.5). Only the first component bears on memorization; the second is a reasoning gap (addressed separately via the overthinking discussion, Table 5). Because Δacc = Low − New sums both, the headline 'memorization' framing attributes the full ~18.5-point group difference in Δ to one mechanism, when roughly half of it is the New-side deficit that memorization does not explain. Please present the (Low, New) plane directly, state how much of the group Δ difference comes from each side, and temper the abstract sentence ('likely trained with much larger Chinese data, thus memorizing more') accordingly—particularly given the conceded openness/scale confound.
  4. [§6.1, Table 10] The abstract-level claim that LLM creations are less reasonable and funny than humans' rests on three of the authors rating 60 items per model against 312 human items, and the text does not state whether ratings were blind to source. Since the authors also validated (and supervised the writing of) the human items, non-blind rating would be a serious bias. Please state the blinding and item-interleaving procedure, and ideally add independent raters. Relatedly, Exp 3 evaluates only o1 and DS-R1—dated relative to the Exp 1 frontier set (Gemini 3.1 Pro, GPT-5.2, etc.)—so the generalization should be scoped or a current model added.
minor comments (11)
  1. [Table 5, §4.3.3] Text says 'from High to Low (166.3% vs. 38.8%)' but the column is N/H (High→New); the caption's direction for L/H ('increase from Low to High') is also confusing relative to the column headers.
  2. [Tables 3–5, 7, 10] No confidence intervals or statistical tests are reported anywhere. With n≈243–288 per split, per-model Δacc CIs are roughly ±6–8 points, so individual rankings (e.g., Doubao 13.7 vs GLM 16.6) are not reliable even though the group contrast is. Please add at least bootstrap CIs for Table 4.
  3. [Table 3] Total is 1,142 here vs 1,143 stated in §3/§8. High-familiarity cells are tiny (n=12/25/61), so the human 99.0% and models' 100.0 on High rest on a handful of items; the familiarity cutoff ≥2.5 would benefit from a sensitivity check (e.g., tertile split).
  4. [Table 4] The † on Grok-4 is unexplained in this table (the footnote appears only under Table 5).
  5. [Table 8, §5.2] The pattern is mixed—Q-max scores higher on New (67.5) than Low (45.4)—and per-cell n is only 12–40, so the text claim that homophone identification is worse 'especially for novel xiehouyu' should be qualified.
  6. [§5.1] Inter-rater agreement for the 56 explanation raters is not reported. Model versions also differ across experiments: Exp 2 used DS-R1/Qwen-max/Kimi-k2 (July 2025) while Exp 1 used V3.2/Qwen3.5-Plus/Kimi-k2.5 (Feb–Mar 2026); please harmonize or justify, and report access dates for all models.
  7. [Table 10] Human N=312 is the pre-filtering count (243 survived validation); state explicitly that the comparison uses raw human output against raw model output, and whether human and model items were interleaved during rating.
  8. [Appendix E] Character-set Jaccard overlap will be systematically inflated for Chinese given the small active character inventory; τ=0.5 is unjustified. Consider character-bigram or embedding-based novelty, or at least a justification of the threshold.
  9. [§3, §8] Release of X-Riddles (with licensing terms for the dictionary-derived items) and the scoring scripts is not stated; for a benchmark paper this would substantially increase utility and reproducibility.
  10. [passim] Typos/wording: 'homophonoic' (§4.2), 'noticable' (§4.3.2), 'are able of complex reasoning' (§1), 'close-sourced' (§4.1), Table 8 header 'accurately in identifying'. §4.2 mentions 'DeepSeek-R1' although Table 4 reports DeepSeek-V3.2.
  11. [§4.3.3] As convergent evidence for the indirect Δacc index, consider a direct contamination probe—e.g., greedy completion of the answer given the riddle, or n-gram overlap with known corpora for open-weight models—even on a subsample.

Circularity Check

0 steps flagged

Empirical benchmark with external human anchors; Δacc is a proposed proxy, not a result forced by definition or self-citation.

full rationale

This paper does not present a first-principles derivation whose outputs reduce to its inputs. The central quantity Δacc is the observed accuracy gap between low-familiarity dictionary items and linguist-written novel items; humans provide an external near-zero baseline (2.9), and model groups are compared on the same held-out splits. Calling Δacc an “index for memorization” is an interpretive framing with a validity assumption (Low≈New difficulty for pure reasoners; gap ≈ retrieval), not algebraic self-definition or a fitted parameter renamed as a prediction. No load-bearing uniqueness theorem, ansatz, or self-citation chain forces the Chinese-vs-English contrast or the Gemini result on novel items. Creation ratings are independent human judgments. Concerns about MCQ option statistics, seed-reused homophone pairs, or openness/scale confounds are identification/correctness issues, not circularity. Score 0; steps empty.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 2 invented entities

Load-bearing commitments are methodological: familiarity thresholding, treating Low-vs-New accuracy gap as a memorization index, equating human-validated novel items with uncontaminated test material, and interpreting cross-lab model differences as largely data-mixture effects. No physical constants or fitted scientific laws; free choices are annotation and split rules.

free parameters (3)
  • High-familiarity cutoff (mean human rating ≥ 2.5 on 1–4 scale) = ≥2.5
    Defines the High vs Low split used throughout accuracy and token analyses; different cutoffs would reallocate items and change reported High/Low gaps.
  • Novel-item validation rule (remove if ≥2 of 3 authors judge unreasonable) = majority-of-3 author veto
    Determines which of 312 drafts enter the 243-item New split that anchors Δacc; committee threshold is a design choice affecting measured novelty difficulty.
  • Creation novelty overlap threshold τ=0.5 (Jaccard on character sets) = 0.5
    Used to label model/human creations as original vs possible duplicates; threshold chosen in implementation (Appendix E).
axioms (5)
  • domain assumption Native-speaker near-equal accuracy on Low and New homophonic items implies those sets engage the same reasoning mechanism and are difficulty-comparable for interpreting model Δacc as memorization.
    Stated in §4.2–4.3.4; without this, Δacc could be difficulty shift rather than training contamination.
  • domain assumption Expert-created xiehouyu that passed author filters did not exist in any pretraining corpus before the study.
    Core contamination-avoidance premise in Abstract/§3; plausible but not cryptographically audited against model training sets.
  • domain assumption Homophonic, polysemous, and direct categories plus literal/figurative answer distinction capture the main reasoning routes in xiehouyu.
    Typology in §2 and Table 2 structures sampling and later restriction of frontier eval to homophonic items.
  • domain assumption Standard multiple-choice accuracy, Likert explanation quality, and author reasonableness/funniness ratings are valid proxies for understanding and creative quality.
    Evaluation design across Experiments 1–3; common in NLP but still an operationalization choice.
  • ad hoc to paper MCQ option sampling and prompt formats in Appendix C do not systematically favor memorized surface forms beyond the intended task.
    Distractors are randomly sampled answers from other items; no extensive adversarial option-control study is reported.
invented entities (2)
  • X-Riddles benchmark (900 dictionary + 243 novel annotated xiehouyu) no independent evidence
    purpose: Provide balanced existing items and uncontaminated novel items for matching, explanation, and creation tests.
    Primary artifact enabling all experiments; independent usefulness depends on public release and external re-annotation.
  • Δacc memorization index (Acc_low − Acc_new on homophonic items) no independent evidence
    purpose: Scalar proxy separating retrieval of attested rare items from generalization to unseen riddles.
    Defined operationally in Abstract/§4; authors note it is only a lower bound when models both memorize and reason.

reviewed 2026-07-30 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?." pith.science (2026). https://pith.science/paper/MFLLYBSX

@misc{pith2026260723440,
  author       = {Pith},
  title        = {Pith review of: Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MFLLYBSX}},
  note         = {Machine review of arXiv:2607.23440}
}
Share X Bluesky LinkedIn Reddit HN
abstract

In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($\Delta_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $\Delta_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $\Delta_{acc}$ of 23.6\%, while English-centric models tested have a mean $\Delta_{acc}$ of 5.1\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs' creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.

Figures

Figures reproduced from arXiv: 2607.23440 by Chongtian Shao, Hai Hu, Kejia Zhang, Siyuan Song, Tianjian Zhu, Xiaojing Zhao.

Figure 1
Figure 1. Figure 1: Overview of three experiments in this study. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Results for experiment 1: accuracy of 0-shot multiple-choice question for xiehouyu for LLMs and humans. Different shapes or colors refer to different types of xiehouyu. The lines represent human accuracy. Qwen2.5-1.5B and Qwen2.5-3B, and has no not￾icable influence on other models. For few-shot prompting (orange bars), increasing from 0-shot to 3-shot increases the performance for Qwen2.5 models for at mos… view at source ↗
Figure 3
Figure 3. Figure 3: Quality distributions across models and xiehouyu types and familiarity. els across xiehouyu types and familiarity are illus￾trated in [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Effects of different conditions on the overall accuracy in the multiple choice questions in Experiment 1. [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 3 canonical work pages

  1. [1]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding , author=. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  2. [2]

    Are Multilingual

    Liu, Chen and Koto, Fajri and Baldwin, Timothy and Gurevych, Iryna , booktitle=. Are Multilingual

  3. [3]

    Potential Idiomatic Expression

    Adewumi, Tosin and Vadoodi, Roshanak and Tripathy, Aparajita and Nikolaido, Konstantina and Liwicki, Foteini and Liwicki, Marcus , booktitle=. Potential Idiomatic Expression

  4. [4]

    Journal of Pragmatics , volume=

    Understanding and classifying two-part allegorical sayings: Metonymy, metaphor, and cultural constraints , author=. Journal of Pragmatics , volume=. 2008 , publisher=

  5. [5]

    Idioms , pages=

    Specialization and reinterpretation in idioms , author=. Idioms , pages=. 2014 , publisher=

  6. [6]

    2010 , school =

    Qu, Wei , title =. 2010 , school =

  7. [7]

    Processing of

    Ma, Lijun and Ma, Yunxiao and He, Xiaoqing and Liu, Haitao and Zhang, Jingyu , journal=. Processing of

  8. [8]

    2015 , publisher=

    Shu, Dingfang , journal=. 2015 , publisher=

  9. [9]

    Journal of Neurolinguistics , volume=

    Electrophysiological insights into the processing of figurative two-part allegorical sayings , author=. Journal of Neurolinguistics , volume=. 2013 , publisher=

  10. [10]

    2016 , publisher=

    Zhang, Grace , booktitle=. 2016 , publisher=

  11. [11]

    2006 , publisher=

    Glossary of semantics and pragmatics , author=. 2006 , publisher=

  12. [12]

    Metaphor and metonymy in comparison and contrast , volume=

    The interaction of metaphor and metonymy in composite expressions , author=. Metaphor and metonymy in comparison and contrast , volume=

  13. [13]

    Understanding

    Hessel, Jack and Marasovi. Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  14. [14]

    1997 , publisher=

    Understanding figurative and literal language: The graded salience hypothesis , author=. 1997 , publisher=

  15. [15]

    Journal of pragmatics , volume=

    On the priority of salient meanings: Studies of literal and figurative language , author=. Journal of pragmatics , volume=. 1999 , publisher=

  16. [16]

    Syntax and semantics , volume=

    Logic and conversation , author=. Syntax and semantics , volume=

  17. [17]

    Cognitive science , volume=

    Literal meaning and psychological theory , author=. Cognitive science , volume=. 1984 , publisher=

  18. [18]

    Journal of pragmatics , volume=

    A new look at literal meaning in understanding what is said and implicated , author=. Journal of pragmatics , volume=. 2002 , publisher=

  19. [19]

    Translation of

    Liu, Chiung-wen and Zhang, Grace Qiao , journal=. Translation of. 2006 , publisher=

  20. [20]

    Erkenntnis , pages=

    Literal meaning , author=. Erkenntnis , pages=. 1978 , publisher=

  21. [21]

    CLUECorpus2020: A Large-scale

    Liang Xu and Xuanwei Zhang and Qianqian Dong , journal=. CLUECorpus2020: A Large-scale. 2020 , volume=

  22. [22]

    2011 , publisher =

    Wen, Duanzheng , title =. 2011 , publisher =

  23. [23]

    A good pun is its own reword

    "A good pun is its own reword": Can Large Language Models Understand Puns? , author=. arXiv preprint arXiv:2404.13599 , year=

  24. [24]

    Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    A neural approach to pun generation , author=. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  25. [25]

    arXiv preprint arXiv:1904.06828 , year=

    Pun generation with surprise , author=. arXiv preprint arXiv:1904.06828 , year=

  26. [26]

    Findings of the Association for Computational Linguistics: NAACL 2022 , pages=

    ID10M: Idiom identification in 10 languages , author=. Findings of the Association for Computational Linguistics: NAACL 2022 , pages=

  27. [27]

    ChID: A large-scale

    Zheng, Chujie and Huang, Minlie and Sun, Aixin , journal=. ChID: A large-scale

  28. [28]

    Proceedings of the 28th International Conference on Computational Linguistics , pages=

    An analysis of language models for metaphor recognition , author=. Proceedings of the 28th International Conference on Computational Linguistics , pages=

  29. [29]

    Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Does GPT-3 grasp metaphors? identifying metaphor mappings with generative language models , author=. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  30. [30]

    arXiv preprint arXiv:2209.08141 , year=

    Psychologically-informed chain-of-thought prompts for metaphor understanding in large language models , author=. arXiv preprint arXiv:2209.08141 , year=

  31. [31]

    Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

    Large dataset and language model fun-tuning for humor recognition , author=. Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

  32. [32]

    Jentzsch, Sophie and Kersting, Kristian , journal=

  33. [33]

    understanding

    Do androids laugh at electric sheep? humor" understanding" benchmarks from the new yorker caption contest , author=. arXiv preprint arXiv:2209.06293 , year=

  34. [34]

    The roles of familiarity and context in processing

    Wang, Xiaolu and Wang, Yizhen and Tian, Wanning and Zheng, Wei and Chen, Xiaoli , journal=. The roles of familiarity and context in processing. 2021 , publisher=

  35. [35]

    Do Large Language Models Understand Conversational Implicature--A case study with a

    Yue, Shisen and Song, Siyuan and Cheng, Xinyuan and Hu, Hai , journal=. Do Large Language Models Understand Conversational Implicature--A case study with a

  36. [36]

    arXiv preprint arXiv:2009.03300 , year=

    Measuring massive multitask language understanding , author=. arXiv preprint arXiv:2009.03300 , year=

  37. [37]

    Li, Haonan and Zhang, Yixuan and Koto, Fajri and Yang, Yifei and Zhao, Hai and Gong, Yeyun and Duan, Nan and Baldwin, Timothy , journal=

  38. [38]

    arXiv preprint arXiv:2406.01574 , year=

    Mmlu-pro: A more robust and challenging multi-task language understanding benchmark , author=. arXiv preprint arXiv:2406.01574 , year=

  39. [39]

    arXiv preprint arXiv:2212.06801 , year=

    A fine-grained comparison of pragmatic language understanding in humans and language models , author=. arXiv preprint arXiv:2212.06801 , year=

  40. [40]

    Introducing Qwen1.5 , url =

    Qwen Team , month =. Introducing Qwen1.5 , url =

  41. [41]

    arXiv preprint arXiv:2309.10305 , url=

    Baichuan 2: Open Large-scale Language Models , author=. arXiv preprint arXiv:2309.10305 , url=

  42. [42]

    2024 , eprint=

    ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools , author=. 2024 , eprint=

  43. [43]

    2024 , eprint=

    InternLM2 Technical Report , author=. 2024 , eprint=

  44. [44]

    2024 , note =

    OpenAI , title =. 2024 , note =

  45. [45]

    arXiv preprint arXiv:2407.10671 , year=

    Qwen2 Technical Report , author=. arXiv preprint arXiv:2407.10671 , year=

  46. [46]

    Zhongguo xiehouyu da cidian [A Comprehensive Dictionary of

    Wen, Duanzheng , year =. Zhongguo xiehouyu da cidian [A Comprehensive Dictionary of

  47. [47]

    doi:10.5281/zenodo.3402023 , version =

    Bright Xu , title =. doi:10.5281/zenodo.3402023 , version =

  48. [48]

    Advances in neural information processing systems , volume=

    Chain-of-thought prompting elicits reasoning in large language models , author=. Advances in neural information processing systems , volume=

  49. [49]

    arXiv preprint arXiv:2412.15115 , year =

    Qwen2.5 Technical Report , author =. arXiv preprint arXiv:2412.15115 , year =

  50. [50]

    QwQ-32B: Embracing the Power of Reinforcement Learning , url =

    Qwen Team , month =. QwQ-32B: Embracing the Power of Reinforcement Learning , url =

  51. [51]

    DeepSeek-R1: Incentivizing Reasoning Capability in

    DeepSeek-AI , year=. DeepSeek-R1: Incentivizing Reasoning Capability in. 2501.12948 , archivePrefix=

  52. [52]

    2024 , url =

    Llama 3 Model Card , author=. 2024 , url =

  53. [53]

    arXiv preprint arXiv:2507.20534 , year=

    Kimi k2: Open agentic intelligence , author=. arXiv preprint arXiv:2507.20534 , year=

  54. [54]

    A Cross-Cultural and Multilingual Evaluation of Large Language Models on

    Xingwei Wang and Zihan Sun and Siyu Wang and Boxing Li and Linli Li , year=. A Cross-Cultural and Multilingual Evaluation of Large Language Models on. 2405.12345 , archivePrefix=

  55. [55]

    Prat and others , title =

    Chantel S. Prat and others , title =. UW News , year =

  56. [56]

    How Does The Human Mind Actually Solve Puzzles? , year =

  57. [57]

    Can Language Models Make Fun? A Case Study in C hinese Comical Crosstalk

    Li, Jianquan and Wu, XiangBo and Liu, Xiaokang and Xie, Qianqian and Tiwari, Prayag and Wang, Benyou. Can Language Models Make Fun? A Case Study in C hinese Comical Crosstalk. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.419

  58. [58]

    Holyoak , title =

    Nicholas Ichien and Dušan Stamenković and Keith J. Holyoak , title =. Metaphor and Symbol , volume =. 2024 , publisher =. doi:10.1080/10926488.2024.2380348 , URL =

  59. [59]

    2025 , eprint=

    Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models , author=. 2025 , eprint=

  60. [60]

    and Koller, Alexander

    Bender, Emily M. and Koller, Alexander. Climbing towards NLU : On Meaning, Form, and Understanding in the Age of Data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.463

  61. [61]

    Mirzadeh, Seyed Iman and Alizadeh, Keivan and Shahrokhi, Hooman and Tuzel, Oncel and Bengio, Samy and Farajtabar, Mehrdad , booktitle=

  62. [62]

    arXiv preprint arXiv:2507.10532 , year=

    Reasoning or memorization? unreliable results of reinforcement learning due to data contamination , author=. arXiv preprint arXiv:2507.10532 , year=

  63. [63]

    arXiv preprint arXiv:2601.02671 , year=

    Extracting books from production language models , author=. arXiv preprint arXiv:2601.02671 , year=

  64. [64]

    Advances in Neural Information Processing Systems , volume=

    Emergent and predictable memorization in large language models , author=. Advances in Neural Information Processing Systems , volume=

  65. [65]

    P hono T hink: Improving Large Language Models' Reasoning on C hinese Phonological Ambiguities

    Ma, Jianfei and Feng, Zhaoxin and Chersoni, Emmanuele and Song, Huacheng and Zhang, Ziqi. P hono T hink: Improving Large Language Models' Reasoning on C hinese Phonological Ambiguities. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.961

  66. [66]

    arXiv preprint arXiv:2412.21187 , year=

    Do not think that much for 2+ 3=? on the overthinking of o1-like llms , author=. arXiv preprint arXiv:2412.21187 , year=

  67. [67]

    Advances in Neural Information Processing Systems , volume=

    Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting , author=. Advances in Neural Information Processing Systems , volume=

  68. [68]

    arXiv preprint arXiv:2503.16419 , year=

    Stop overthinking: A survey on efficient reasoning for large language models , author=. arXiv preprint arXiv:2503.16419 , year=

  69. [69]

    arXiv preprint arXiv:2505.16122 , year=

    Plan and budget: Effective and efficient test-time scaling on large language model reasoning , author=. arXiv preprint arXiv:2505.16122 , year=

This paper was first reviewed by grok-4.5 on July 30, 2026.