REVIEW 4 major objections 11 minor 69 references
A large accuracy drop from rare to brand-new Chinese xiehouyu riddles marks memorization in frontier Chinese LLMs, even as some models beat humans on novel items and still lag at creating them.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 22:01 UTC pith:MFLLYBSX
load-bearing objection Solid contamination-aware Chinese riddle benchmark; the Δacc story is useful but only partly identified, and the authors mostly own that. the 4 major comments →
Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Using linguist-written novel xiehouyu to block contamination, the authors show that frontier Chinese models have a large mean accuracy drop (about 23.6 points) from low-familiarity existing homophonic items to novel ones, versus near-zero drop for native speakers and about 5 points for English-centric models—evidence they treat as memorization from larger Chinese training data—while Gemini 3.1 Pro still scores 92.6% on novel items (above human MCQ accuracy) and LLM-created xiehouyu receive worse reasonableness and funniness ratings than human creations.
What carries the argument
Δacc (accuracy on low-familiarity existing homophonic xiehouyu minus accuracy on novel ones): a lower-bound memorization index, justified because humans show almost no gap and are therefore argued to use the same reasoning process on both sets.
Load-bearing premise
The score gap between rare dictionary riddles and linguist-written new ones mainly reflects training-set memorization, not leftover difficulty, style, or distribution shift between those two sets—and the Chinese-versus-English contrast is driven mainly by Chinese data volume even though the model groups also differ in openness and scale.
What would settle it
Build a new matched novel set with the same human difficulty profile, or measure how often the low-familiarity dictionary items actually appear in public Chinese pretraining crawls: if the Low–New gap for Chinese models vanishes under tighter difficulty matching, or fails to track documented Chinese-data exposure, the memorization reading of Δacc weakens.
If this is right
- Benchmarks built from culturally circulated material need paired novel items of matched difficulty, or they will mix retrieval with reasoning.
- Claims of strong LLM reasoning on Chinese figurative language should be re-checked on uncontaminated, expert-written items.
- Homophonic Chinese wordplay remains a harder test than polysemy or direct proverb types for current models.
- Even models that match or beat humans on novel multiple-choice xiehouyu still underperform human linguists at creating reasonable, funny ones.
- Heavier Chinese pretraining can inflate scores on rare existing items without equal gains on truly new riddles.
Where Pith is reading between the lines
- Phonology-aware training or explicit pronunciation modules may be needed before models close the remaining gap on strict homophonic links.
- The same Low-versus-New template could be ported to other language-specific games (puns, two-part allegories, dialect wordplay) to audit memorization outside Chinese.
- Creation tasks may stay a stricter creativity filter than multiple-choice understanding even after contamination is controlled.
- Open reporting of Chinese pretraining mix would let future work test whether Δacc scales with documented Chinese token volume.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces X-Riddles, a benchmark of 1,143 Chinese xiehouyu: 900 dictionary-sampled items balanced across homophonic/polysemous/direct types and high/low familiarity, plus 243 novel homophonic items written by linguistics students and validated by the authors. Three experiments probe LLMs: MCQ riddle–answer matching (Exp 1), free-form explanation (Exp 2), and xiehouyu creation (Exp 3). The central index is Δacc = Acc(low-familiarity) − Acc(novel) on homophonic items: native speakers show Δ=2.9 (71.5→68.6), frontier Chinese models average 23.6 (e.g., DeepSeek-V3.2 95.8→64.9), and English-centric models 5.1. The authors interpret the Chinese-model gap as memorization from larger Chinese training data, while noting Gemini 3.1 Pro reaches 92.6% on novel items (above the human 68.6). Exp 2 reports degraded explanation quality and hallucinated homophones on novel items; Exp 3 finds DS-R1/o1 creations rated less reasonable and funny than human ones.
Significance. If the memorization reading holds, this is a valuable contribution to the retrieval-vs-reasoning debate and to evaluation methodology for culturally specific language. Named strengths: (i) expert-authored novel items designed to sidestep contamination, with an explicit validation protocol; (ii) a 100+ native-speaker MCQ baseline providing an external anchor for item difficulty (human Δ≈2.9); (iii) a matched-difficulty existing-vs-novel pairing that is a reusable decontamination template; (iv) broad model coverage with condition ablations (context, shots, CoT, literal vs figurative target); (v) a token-effort analysis yielding a falsifiable dissociation (Chinese models spend +40.6% tokens from Low to New while losing 23.6 points); (vi) full prompt transparency in appendices. The paper also reports a result against its own headline (Gemini 3.1 Pro, Δ=−0.2) and concedes the openness/scale confound in §4.3.4, which increases credibility. The creation experiment, if its evaluation is sound, is a rare controlled comparison of expert vs LLM linguistic creativity.
major comments (4)
- [§3, §4.3.3, Table 4] §3 and §4.3.3 (Table 4): 256 of 312 candidate novel items—hence the large majority of the retained 243—were seed-based, reusing the exact homophone pair of an existing dictionary xiehouyu (e.g., 舅/旧 from 外甥打灯笼——照旧). Exp 2 itself shows homophone identification is precisely where models fail; a model that memorized the seed gets the hardest link of the reasoning chain for free, so only the riddle→literal mapping is genuinely novel for ~4/5 of the 'uncontaminated' set. The abstract's 'to avoid data contamination' is therefore too strong, and Δacc's magnitude is not cleanly interpretable. A cheap, decisive check exists: report per-model accuracy and Δacc separately for the 56 free-created items vs the seed-based ones. If the Chinese/English contrast replicates on the free-created subset, the central claim is much better supported; either way, foreground the construction detail in the abstrac
- [§4.1, Table 4] All MCQ distractors are randomly sampled dictionary answers. Thus in the Low condition the correct option is an attested answer among attested answers, while in the New condition the correct option is a novel string among attested distractors. A model that prefers familiar/attested strings—plausible precisely for models trained on more Chinese text—would be helped on Low and hurt on New, producing a Δacc that reflects option-familiarity bias rather than, or in addition to, memorization of the items. The human Δ=2.9 controls intrinsic item difficulty but not model-specific weighting of this cue, since humans and models need not use surface familiarity the same way. A free-form answer-generation condition (no options) or a novel-distractor control on a subsample would discriminate the accounts; at minimum, this alternative should be analyzed and discussed.
- [§4.3.3–4.3.4, Table 4] The group contrast in Table 4 decomposes into two effects: Chinese models are higher on Low (90.6 vs 82.7) and lower on New (67.0 vs 77.5). Only the first component bears on memorization; the second is a reasoning gap (addressed separately via the overthinking discussion, Table 5). Because Δacc = Low − New sums both, the headline 'memorization' framing attributes the full ~18.5-point group difference in Δ to one mechanism, when roughly half of it is the New-side deficit that memorization does not explain. Please present the (Low, New) plane directly, state how much of the group Δ difference comes from each side, and temper the abstract sentence ('likely trained with much larger Chinese data, thus memorizing more') accordingly—particularly given the conceded openness/scale confound.
- [§6.1, Table 10] The abstract-level claim that LLM creations are less reasonable and funny than humans' rests on three of the authors rating 60 items per model against 312 human items, and the text does not state whether ratings were blind to source. Since the authors also validated (and supervised the writing of) the human items, non-blind rating would be a serious bias. Please state the blinding and item-interleaving procedure, and ideally add independent raters. Relatedly, Exp 3 evaluates only o1 and DS-R1—dated relative to the Exp 1 frontier set (Gemini 3.1 Pro, GPT-5.2, etc.)—so the generalization should be scoped or a current model added.
minor comments (11)
- [Table 5, §4.3.3] Text says 'from High to Low (166.3% vs. 38.8%)' but the column is N/H (High→New); the caption's direction for L/H ('increase from Low to High') is also confusing relative to the column headers.
- [Tables 3–5, 7, 10] No confidence intervals or statistical tests are reported anywhere. With n≈243–288 per split, per-model Δacc CIs are roughly ±6–8 points, so individual rankings (e.g., Doubao 13.7 vs GLM 16.6) are not reliable even though the group contrast is. Please add at least bootstrap CIs for Table 4.
- [Table 3] Total is 1,142 here vs 1,143 stated in §3/§8. High-familiarity cells are tiny (n=12/25/61), so the human 99.0% and models' 100.0 on High rest on a handful of items; the familiarity cutoff ≥2.5 would benefit from a sensitivity check (e.g., tertile split).
- [Table 4] The † on Grok-4 is unexplained in this table (the footnote appears only under Table 5).
- [Table 8, §5.2] The pattern is mixed—Q-max scores higher on New (67.5) than Low (45.4)—and per-cell n is only 12–40, so the text claim that homophone identification is worse 'especially for novel xiehouyu' should be qualified.
- [§5.1] Inter-rater agreement for the 56 explanation raters is not reported. Model versions also differ across experiments: Exp 2 used DS-R1/Qwen-max/Kimi-k2 (July 2025) while Exp 1 used V3.2/Qwen3.5-Plus/Kimi-k2.5 (Feb–Mar 2026); please harmonize or justify, and report access dates for all models.
- [Table 10] Human N=312 is the pre-filtering count (243 survived validation); state explicitly that the comparison uses raw human output against raw model output, and whether human and model items were interleaved during rating.
- [Appendix E] Character-set Jaccard overlap will be systematically inflated for Chinese given the small active character inventory; τ=0.5 is unjustified. Consider character-bigram or embedding-based novelty, or at least a justification of the threshold.
- [§3, §8] Release of X-Riddles (with licensing terms for the dictionary-derived items) and the scoring scripts is not stated; for a benchmark paper this would substantially increase utility and reproducibility.
- [passim] Typos/wording: 'homophonoic' (§4.2), 'noticable' (§4.3.2), 'are able of complex reasoning' (§1), 'close-sourced' (§4.1), Table 8 header 'accurately in identifying'. §4.2 mentions 'DeepSeek-R1' although Table 4 reports DeepSeek-V3.2.
- [§4.3.3] As convergent evidence for the indirect Δacc index, consider a direct contamination probe—e.g., greedy completion of the answer given the riddle, or n-gram overlap with known corpora for open-weight models—even on a subsample.
Circularity Check
Empirical benchmark with external human anchors; Δacc is a proposed proxy, not a result forced by definition or self-citation.
full rationale
This paper does not present a first-principles derivation whose outputs reduce to its inputs. The central quantity Δacc is the observed accuracy gap between low-familiarity dictionary items and linguist-written novel items; humans provide an external near-zero baseline (2.9), and model groups are compared on the same held-out splits. Calling Δacc an “index for memorization” is an interpretive framing with a validity assumption (Low≈New difficulty for pure reasoners; gap ≈ retrieval), not algebraic self-definition or a fitted parameter renamed as a prediction. No load-bearing uniqueness theorem, ansatz, or self-citation chain forces the Chinese-vs-English contrast or the Gemini result on novel items. Creation ratings are independent human judgments. Concerns about MCQ option statistics, seed-reused homophone pairs, or openness/scale confounds are identification/correctness issues, not circularity. Score 0; steps empty.
Axiom & Free-Parameter Ledger
free parameters (3)
- High-familiarity cutoff (mean human rating ≥ 2.5 on 1–4 scale) =
≥2.5
- Novel-item validation rule (remove if ≥2 of 3 authors judge unreasonable) =
majority-of-3 author veto
- Creation novelty overlap threshold τ=0.5 (Jaccard on character sets) =
0.5
axioms (5)
- domain assumption Native-speaker near-equal accuracy on Low and New homophonic items implies those sets engage the same reasoning mechanism and are difficulty-comparable for interpreting model Δacc as memorization.
- domain assumption Expert-created xiehouyu that passed author filters did not exist in any pretraining corpus before the study.
- domain assumption Homophonic, polysemous, and direct categories plus literal/figurative answer distinction capture the main reasoning routes in xiehouyu.
- domain assumption Standard multiple-choice accuracy, Likert explanation quality, and author reasonableness/funniness ratings are valid proxies for understanding and creative quality.
- ad hoc to paper MCQ option sampling and prompt formats in Appendix C do not systematically favor memorized surface forms beyond the intended task.
invented entities (2)
-
X-Riddles benchmark (900 dictionary + 243 novel annotated xiehouyu)
no independent evidence
-
Δacc memorization index (Acc_low − Acc_new on homophonic items)
no independent evidence
Cite this review
Pith. "Pith review of Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?." pith.science (2026). https://pith.science/paper/MFLLYBSX
@misc{pith2026260723440,
author = {Pith},
title = {Pith review of: Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?},
year = {2026},
howpublished = {\url{https://pith.science/paper/MFLLYBSX}},
note = {Machine review of arXiv:2607.23440}
}
read the original abstract
In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($\Delta_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $\Delta_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $\Delta_{acc}$ of 23.6\%, while English-centric models tested have a mean $\Delta_{acc}$ of 5.1\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs' creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.
Figures
Reference graph
Works this paper leans on
-
[1]
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding , author=. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[2]
Are Multilingual
Liu, Chen and Koto, Fajri and Baldwin, Timothy and Gurevych, Iryna , booktitle=. Are Multilingual
-
[3]
Potential Idiomatic Expression
Adewumi, Tosin and Vadoodi, Roshanak and Tripathy, Aparajita and Nikolaido, Konstantina and Liwicki, Foteini and Liwicki, Marcus , booktitle=. Potential Idiomatic Expression
-
[4]
Journal of Pragmatics , volume=
Understanding and classifying two-part allegorical sayings: Metonymy, metaphor, and cultural constraints , author=. Journal of Pragmatics , volume=. 2008 , publisher=
2008
-
[5]
Idioms , pages=
Specialization and reinterpretation in idioms , author=. Idioms , pages=. 2014 , publisher=
2014
-
[6]
2010 , school =
Qu, Wei , title =. 2010 , school =
2010
-
[7]
Processing of
Ma, Lijun and Ma, Yunxiao and He, Xiaoqing and Liu, Haitao and Zhang, Jingyu , journal=. Processing of
-
[8]
2015 , publisher=
Shu, Dingfang , journal=. 2015 , publisher=
2015
-
[9]
Journal of Neurolinguistics , volume=
Electrophysiological insights into the processing of figurative two-part allegorical sayings , author=. Journal of Neurolinguistics , volume=. 2013 , publisher=
2013
-
[10]
2016 , publisher=
Zhang, Grace , booktitle=. 2016 , publisher=
2016
-
[11]
2006 , publisher=
Glossary of semantics and pragmatics , author=. 2006 , publisher=
2006
-
[12]
Metaphor and metonymy in comparison and contrast , volume=
The interaction of metaphor and metonymy in composite expressions , author=. Metaphor and metonymy in comparison and contrast , volume=
-
[13]
Understanding
Hessel, Jack and Marasovi. Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[14]
1997 , publisher=
Understanding figurative and literal language: The graded salience hypothesis , author=. 1997 , publisher=
1997
-
[15]
Journal of pragmatics , volume=
On the priority of salient meanings: Studies of literal and figurative language , author=. Journal of pragmatics , volume=. 1999 , publisher=
1999
-
[16]
Syntax and semantics , volume=
Logic and conversation , author=. Syntax and semantics , volume=
-
[17]
Cognitive science , volume=
Literal meaning and psychological theory , author=. Cognitive science , volume=. 1984 , publisher=
1984
-
[18]
Journal of pragmatics , volume=
A new look at literal meaning in understanding what is said and implicated , author=. Journal of pragmatics , volume=. 2002 , publisher=
2002
-
[19]
Translation of
Liu, Chiung-wen and Zhang, Grace Qiao , journal=. Translation of. 2006 , publisher=
2006
-
[20]
Erkenntnis , pages=
Literal meaning , author=. Erkenntnis , pages=. 1978 , publisher=
1978
-
[21]
CLUECorpus2020: A Large-scale
Liang Xu and Xuanwei Zhang and Qianqian Dong , journal=. CLUECorpus2020: A Large-scale. 2020 , volume=
2020
-
[22]
2011 , publisher =
Wen, Duanzheng , title =. 2011 , publisher =
2011
-
[23]
"A good pun is its own reword": Can Large Language Models Understand Puns? , author=. arXiv preprint arXiv:2404.13599 , year=
-
[24]
Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
A neural approach to pun generation , author=. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[25]
arXiv preprint arXiv:1904.06828 , year=
Pun generation with surprise , author=. arXiv preprint arXiv:1904.06828 , year=
Pith/arXiv arXiv 1904
-
[26]
Findings of the Association for Computational Linguistics: NAACL 2022 , pages=
ID10M: Idiom identification in 10 languages , author=. Findings of the Association for Computational Linguistics: NAACL 2022 , pages=
2022
-
[27]
ChID: A large-scale
Zheng, Chujie and Huang, Minlie and Sun, Aixin , journal=. ChID: A large-scale
-
[28]
Proceedings of the 28th International Conference on Computational Linguistics , pages=
An analysis of language models for metaphor recognition , author=. Proceedings of the 28th International Conference on Computational Linguistics , pages=
-
[29]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Does GPT-3 grasp metaphors? identifying metaphor mappings with generative language models , author=. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[30]
arXiv preprint arXiv:2209.08141 , year=
Psychologically-informed chain-of-thought prompts for metaphor understanding in large language models , author=. arXiv preprint arXiv:2209.08141 , year=
-
[31]
Proceedings of the 57th annual meeting of the association for computational linguistics , pages=
Large dataset and language model fun-tuning for humor recognition , author=. Proceedings of the 57th annual meeting of the association for computational linguistics , pages=
-
[32]
Jentzsch, Sophie and Kersting, Kristian , journal=
-
[33]
Do androids laugh at electric sheep? humor" understanding" benchmarks from the new yorker caption contest , author=. arXiv preprint arXiv:2209.06293 , year=
-
[34]
The roles of familiarity and context in processing
Wang, Xiaolu and Wang, Yizhen and Tian, Wanning and Zheng, Wei and Chen, Xiaoli , journal=. The roles of familiarity and context in processing. 2021 , publisher=
2021
-
[35]
Do Large Language Models Understand Conversational Implicature--A case study with a
Yue, Shisen and Song, Siyuan and Cheng, Xinyuan and Hu, Hai , journal=. Do Large Language Models Understand Conversational Implicature--A case study with a
-
[36]
arXiv preprint arXiv:2009.03300 , year=
Measuring massive multitask language understanding , author=. arXiv preprint arXiv:2009.03300 , year=
Pith/arXiv arXiv 2009
-
[37]
Li, Haonan and Zhang, Yixuan and Koto, Fajri and Yang, Yifei and Zhao, Hai and Gong, Yeyun and Duan, Nan and Baldwin, Timothy , journal=
-
[38]
arXiv preprint arXiv:2406.01574 , year=
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark , author=. arXiv preprint arXiv:2406.01574 , year=
-
[39]
arXiv preprint arXiv:2212.06801 , year=
A fine-grained comparison of pragmatic language understanding in humans and language models , author=. arXiv preprint arXiv:2212.06801 , year=
-
[40]
Introducing Qwen1.5 , url =
Qwen Team , month =. Introducing Qwen1.5 , url =
-
[41]
arXiv preprint arXiv:2309.10305 , url=
Baichuan 2: Open Large-scale Language Models , author=. arXiv preprint arXiv:2309.10305 , url=
-
[42]
2024 , eprint=
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools , author=. 2024 , eprint=
2024
-
[43]
2024 , eprint=
InternLM2 Technical Report , author=. 2024 , eprint=
2024
-
[44]
2024 , note =
OpenAI , title =. 2024 , note =
2024
-
[45]
arXiv preprint arXiv:2407.10671 , year=
Qwen2 Technical Report , author=. arXiv preprint arXiv:2407.10671 , year=
-
[46]
Zhongguo xiehouyu da cidian [A Comprehensive Dictionary of
Wen, Duanzheng , year =. Zhongguo xiehouyu da cidian [A Comprehensive Dictionary of
-
[47]
doi:10.5281/zenodo.3402023 , version =
Bright Xu , title =. doi:10.5281/zenodo.3402023 , version =
-
[48]
Advances in neural information processing systems , volume=
Chain-of-thought prompting elicits reasoning in large language models , author=. Advances in neural information processing systems , volume=
-
[49]
arXiv preprint arXiv:2412.15115 , year =
Qwen2.5 Technical Report , author =. arXiv preprint arXiv:2412.15115 , year =
-
[50]
QwQ-32B: Embracing the Power of Reinforcement Learning , url =
Qwen Team , month =. QwQ-32B: Embracing the Power of Reinforcement Learning , url =
-
[51]
DeepSeek-R1: Incentivizing Reasoning Capability in
DeepSeek-AI , year=. DeepSeek-R1: Incentivizing Reasoning Capability in. 2501.12948 , archivePrefix=
-
[52]
2024 , url =
Llama 3 Model Card , author=. 2024 , url =
2024
-
[53]
arXiv preprint arXiv:2507.20534 , year=
Kimi k2: Open agentic intelligence , author=. arXiv preprint arXiv:2507.20534 , year=
-
[54]
A Cross-Cultural and Multilingual Evaluation of Large Language Models on
Xingwei Wang and Zihan Sun and Siyu Wang and Boxing Li and Linli Li , year=. A Cross-Cultural and Multilingual Evaluation of Large Language Models on. 2405.12345 , archivePrefix=
-
[55]
Prat and others , title =
Chantel S. Prat and others , title =. UW News , year =
-
[56]
How Does The Human Mind Actually Solve Puzzles? , year =
-
[57]
Can Language Models Make Fun? A Case Study in C hinese Comical Crosstalk
Li, Jianquan and Wu, XiangBo and Liu, Xiaokang and Xie, Qianqian and Tiwari, Prayag and Wang, Benyou. Can Language Models Make Fun? A Case Study in C hinese Comical Crosstalk. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. doi:10.18653/v1/2023.acl-long.419
-
[58]
Nicholas Ichien and Dušan Stamenković and Keith J. Holyoak , title =. Metaphor and Symbol , volume =. 2024 , publisher =. doi:10.1080/10926488.2024.2380348 , URL =
arXiv 2024
-
[59]
2025 , eprint=
Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models , author=. 2025 , eprint=
2025
-
[60]
Bender, Emily M. and Koller, Alexander. Climbing towards NLU : On Meaning, Form, and Understanding in the Age of Data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.463
-
[61]
Mirzadeh, Seyed Iman and Alizadeh, Keivan and Shahrokhi, Hooman and Tuzel, Oncel and Bengio, Samy and Farajtabar, Mehrdad , booktitle=
-
[62]
arXiv preprint arXiv:2507.10532 , year=
Reasoning or memorization? unreliable results of reinforcement learning due to data contamination , author=. arXiv preprint arXiv:2507.10532 , year=
-
[63]
arXiv preprint arXiv:2601.02671 , year=
Extracting books from production language models , author=. arXiv preprint arXiv:2601.02671 , year=
-
[64]
Advances in Neural Information Processing Systems , volume=
Emergent and predictable memorization in large language models , author=. Advances in Neural Information Processing Systems , volume=
-
[65]
P hono T hink: Improving Large Language Models' Reasoning on C hinese Phonological Ambiguities
Ma, Jianfei and Feng, Zhaoxin and Chersoni, Emmanuele and Song, Huacheng and Zhang, Ziqi. P hono T hink: Improving Large Language Models' Reasoning on C hinese Phonological Ambiguities. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.961
-
[66]
arXiv preprint arXiv:2412.21187 , year=
Do not think that much for 2+ 3=? on the overthinking of o1-like llms , author=. arXiv preprint arXiv:2412.21187 , year=
-
[67]
Advances in Neural Information Processing Systems , volume=
Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting , author=. Advances in Neural Information Processing Systems , volume=
-
[68]
arXiv preprint arXiv:2503.16419 , year=
Stop overthinking: A survey on efficient reasoning for large language models , author=. arXiv preprint arXiv:2503.16419 , year=
-
[69]
arXiv preprint arXiv:2505.16122 , year=
Plan and budget: Effective and efficient test-time scaling on large language model reasoning , author=. arXiv preprint arXiv:2505.16122 , year=
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.