{"id":"22e8a4c1-3a33-493c-a49a-2c7e69aff445","arxiv_id":"2607.23440","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Δacc between low-frequency and novel xiehouyu is ~23.6% for Chinese frontier LLMs vs ~5.1% for English-centric models and ~2.9% for humans, while LLM-created xiehouyu rate below human creations.","lead":"Chinese-origin frontier LLMs score far higher on rare existing Chinese xiehouyu riddles than on brand-new ones written by linguists, while humans and English-centric models do not. The gap is used as a memorization index and suggests contamination-aware eval is needed before claiming linguistic reasoning.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The central \"memorization\" attribution rests entirely on an indirect MCQ-based Δacc index that is never validated against any direct measure of retrieval; MCQ option support and seed-based construction of \"novel\" items leave the retrieval-vs-reasoning interpretation under-identified.","rationale":"The reader correctly identified the load-bearing premise — that Δacc indexes memorization rather than difficulty/style/distribution shift, and that the Chinese-vs-English contrast is confounded — and appropriately weighted the human baseline and the paper's own §4.3.4 caveats. My scrutiny confirms rather than overturns this: the human Δ=2.9 baseline is a genuine and fairly strong control for intrinsic item difficulty, the seed-reuse issue cuts partly in the paper's favor (it biases Δ downward for seed items, consistent with the \"lower bound\" framing), and Gemini 3.1 Pro's 92.6% on New items shows the novel set is solvable by genuine reasoning, so the benchmark itself is not broken. What remains unsecured is the causal attribution: no direct retrieval evidence is offered, the MCQ option-construction asymmetry between Low and New conditions is unexamined, and group differences in openness/scale/capability are conceded but not controlled. These are exactly the conditions the reader attached. The paper is honest about its limits, the empirical pattern (within-model Low→New drops of 13–39 points for Chinese frontier models vs. ≤10 for English-centric ones) is documented and interesting, and the proposed verification (cloze exact-match probe plus seed/free-style Δ split) is cheap to run on the open-weight models. If the probe confirms verbatim retrieval on Low items, the memorization claim is substantially strengthened; if not, the attribution needs revision. Either way, the CONDITIONAL verdict with HIGH confidence stands; no adjustment is warranted.","tokens_in":21256,"tokens_out":2946,"duration_ms":96496,"concrete_test":"Run a free-completion probe on the open-weight models (Qwen2.5 family, DeepSeek-V3.2): present each homophonic riddle in a natural cloze prompt (\"riddle ——\") with no options, greedy decoding, and score exact-string match of the answer-intended. If Low-item exact-match approaches MCQ accuracy (≈90%+) while New-item exact-match collapses far below MCQ accuracy, verbatim retrieval is directly confirmed for Low items. Secondarily, recompute Δacc separately for seed-based vs. the ~56 free-style New items: if Δ is much smaller for seed-based items, the New set is contaminated at the homophone-pair level and the 23.6% figure is not a clean memorization estimate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline inference — frontier Chinese models' ~23.6% Low→New accuracy drop indexes memorization of rare xiehouyu from larger Chinese training data — depends on Δacc being a valid memorization probe. Two soft spots bear most of the weight. (1) MCQ format confound: on Low items a model can succeed either by retrieving the stored riddle–answer pair or by eliminating distractors (randomly sampled answers of other xiehouyu). For existing items, the correct answer is a \"real-sounding\" dictionary answer and distractors are also dictionary answers, whereas for New items the correct answer is itself novel — so option-set statistics differ systematically between Low and New conditions. The human Δ of 2.9 controls for intrinsic item difficulty but not for a model-specific option-elimination asymmetry, since humans and models needn't weight surface cues the same way. (2) Contamination leakage in the \"uncontaminated\" set: 256 of 312 candidate new items (i.e., the large majority of the final 243) were seed-based, reusing the exact homophone pair (e.g., 舅/旧) from an existing dictionary xiehouyu. A model that memorized the seed's pun gets the critical phonological link for free; only the riddle mapping is new. This means Δacc is contaminated at the level of the very mechanism being measured, and its magnitude is not cleanly interpretable as either a lower bound or an estimate of memorization. The authors partially concede the openness/scale confound (§4.3.4) but never directly probe retrieval. This matches the reader's weakest-assumption concern; the human baseline narrows but does not close the gap.","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper introduces X-Riddles, a benchmark of 1,143 Chinese xiehouyu: 900 dictionary-sampled items balanced across homophonic/polysemous/direct types and high/low familiarity, plus 243 novel homophonic items written by linguistics students and validated by the authors. Three experiments probe LLMs: MCQ riddle–answer matching (Exp 1), free-form explanation (Exp 2), and xiehouyu creation (Exp 3). The central index is Δacc = Acc(low-familiarity) − Acc(novel) on homophonic items: native speakers show Δ=2.9 (71.5→68.6), frontier Chinese models average 23.6 (e.g., DeepSeek-V3.2 95.8→64.9), and English-centric models 5.1. The authors interpret the Chinese-model gap as memorization from larger Chinese training data, while noting Gemini 3.1 Pro reaches 92.6% on novel items (above the human 68.6). Exp 2 reports degraded explanation quality and hallucinated homophones on novel items; Exp 3 finds DS-R1/o1 creations rated less reasonable and funny than human ones.","tokens_in":21706,"tokens_out":10868,"duration_ms":79500,"significance":"If the memorization reading holds, this is a valuable contribution to the retrieval-vs-reasoning debate and to evaluation methodology for culturally specific language. Named strengths: (i) expert-authored novel items designed to sidestep contamination, with an explicit validation protocol; (ii) a 100+ native-speaker MCQ baseline providing an external anchor for item difficulty (human Δ≈2.9); (iii) a matched-difficulty existing-vs-novel pairing that is a reusable decontamination template; (iv) broad model coverage with condition ablations (context, shots, CoT, literal vs figurative target); (v) a token-effort analysis yielding a falsifiable dissociation (Chinese models spend +40.6% tokens from Low to New while losing 23.6 points); (vi) full prompt transparency in appendices. The paper also reports a result against its own headline (Gemini 3.1 Pro, Δ=−0.2) and concedes the openness/scale confound in §4.3.4, which increases credibility. The creation experiment, if its evaluation is sound, is a rare controlled comparison of expert vs LLM linguistic creativity.","major_comments":[{"comment":"§3 and §4.3.3 (Table 4): 256 of 312 candidate novel items—hence the large majority of the retained 243—were seed-based, reusing the exact homophone pair of an existing dictionary xiehouyu (e.g., 舅/旧 from 外甥打灯笼——照旧). Exp 2 itself shows homophone identification is precisely where models fail; a model that memorized the seed gets the hardest link of the reasoning chain for free, so only the riddle→literal mapping is genuinely novel for ~4/5 of the 'uncontaminated' set. The abstract's 'to avoid data contamination' is therefore too strong, and Δacc's magnitude is not cleanly interpretable. A cheap, decisive check exists: report per-model accuracy and Δacc separately for the 56 free-created items vs the seed-based ones. If the Chinese/English contrast replicates on the free-created subset, the central claim is much better supported; either way, foreground the construction detail in the abstrac","section":"§3, §4.3.3, Table 4"},{"comment":"All MCQ distractors are randomly sampled dictionary answers. Thus in the Low condition the correct option is an attested answer among attested answers, while in the New condition the correct option is a novel string among attested distractors. A model that prefers familiar/attested strings—plausible precisely for models trained on more Chinese text—would be helped on Low and hurt on New, producing a Δacc that reflects option-familiarity bias rather than, or in addition to, memorization of the items. The human Δ=2.9 controls intrinsic item difficulty but not model-specific weighting of this cue, since humans and models need not use surface familiarity the same way. A free-form answer-generation condition (no options) or a novel-distractor control on a subsample would discriminate the accounts; at minimum, this alternative should be analyzed and discussed.","section":"§4.1, Table 4"},{"comment":"The group contrast in Table 4 decomposes into two effects: Chinese models are higher on Low (90.6 vs 82.7) and lower on New (67.0 vs 77.5). Only the first component bears on memorization; the second is a reasoning gap (addressed separately via the overthinking discussion, Table 5). Because Δacc = Low − New sums both, the headline 'memorization' framing attributes the full ~18.5-point group difference in Δ to one mechanism, when roughly half of it is the New-side deficit that memorization does not explain. Please present the (Low, New) plane directly, state how much of the group Δ difference comes from each side, and temper the abstract sentence ('likely trained with much larger Chinese data, thus memorizing more') accordingly—particularly given the conceded openness/scale confound.","section":"§4.3.3–4.3.4, Table 4"},{"comment":"The abstract-level claim that LLM creations are less reasonable and funny than humans' rests on three of the authors rating 60 items per model against 312 human items, and the text does not state whether ratings were blind to source. Since the authors also validated (and supervised the writing of) the human items, non-blind rating would be a serious bias. Please state the blinding and item-interleaving procedure, and ideally add independent raters. Relatedly, Exp 3 evaluates only o1 and DS-R1—dated relative to the Exp 1 frontier set (Gemini 3.1 Pro, GPT-5.2, etc.)—so the generalization should be scoped or a current model added.","section":"§6.1, Table 10"}],"minor_comments":[{"comment":"Text says 'from High to Low (166.3% vs. 38.8%)' but the column is N/H (High→New); the caption's direction for L/H ('increase from Low to High') is also confusing relative to the column headers.","section":"Table 5, §4.3.3"},{"comment":"No confidence intervals or statistical tests are reported anywhere. With n≈243–288 per split, per-model Δacc CIs are roughly ±6–8 points, so individual rankings (e.g., Doubao 13.7 vs GLM 16.6) are not reliable even though the group contrast is. Please add at least bootstrap CIs for Table 4.","section":"Tables 3–5, 7, 10"},{"comment":"Total is 1,142 here vs 1,143 stated in §3/§8. High-familiarity cells are tiny (n=12/25/61), so the human 99.0% and models' 100.0 on High rest on a handful of items; the familiarity cutoff ≥2.5 would benefit from a sensitivity check (e.g., tertile split).","section":"Table 3"},{"comment":"The † on Grok-4 is unexplained in this table (the footnote appears only under Table 5).","section":"Table 4"},{"comment":"The pattern is mixed—Q-max scores higher on New (67.5) than Low (45.4)—and per-cell n is only 12–40, so the text claim that homophone identification is worse 'especially for novel xiehouyu' should be qualified.","section":"Table 8, §5.2"},{"comment":"Inter-rater agreement for the 56 explanation raters is not reported. Model versions also differ across experiments: Exp 2 used DS-R1/Qwen-max/Kimi-k2 (July 2025) while Exp 1 used V3.2/Qwen3.5-Plus/Kimi-k2.5 (Feb–Mar 2026); please harmonize or justify, and report access dates for all models.","section":"§5.1"},{"comment":"Human N=312 is the pre-filtering count (243 survived validation); state explicitly that the comparison uses raw human output against raw model output, and whether human and model items were interleaved during rating.","section":"Table 10"},{"comment":"Character-set Jaccard overlap will be systematically inflated for Chinese given the small active character inventory; τ=0.5 is unjustified. Consider character-bigram or embedding-based novelty, or at least a justification of the threshold.","section":"Appendix E"},{"comment":"Release of X-Riddles (with licensing terms for the dictionary-derived items) and the scoring scripts is not stated; for a benchmark paper this would substantially increase utility and reproducibility.","section":"§3, §8"},{"comment":"Typos/wording: 'homophonoic' (§4.2), 'noticable' (§4.3.2), 'are able of complex reasoning' (§1), 'close-sourced' (§4.1), Table 8 header 'accurately in identifying'. §4.2 mentions 'DeepSeek-R1' although Table 4 reports DeepSeek-V3.2.","section":"passim"},{"comment":"As convergent evidence for the indirect Δacc index, consider a direct contamination probe—e.g., greedy completion of the answer given the riddle, or n-gram overlap with known corpora for open-weight models—even on a subsample.","section":"§4.3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its own construction pipeline—the seed-based design is described openly in §3—so I have no concealment concerns; the issue is over-claiming in the abstract relative to what the design supports. Both requested checks (seed vs free-created breakdown; free-form or matched-distractor control) are feasible within the paper's existing scope, and the first requires only reanalysis. Fit with the venue is good. The author-rated creation experiment (Exp 3) would benefit from independent raters before the creativity claim appears in the abstract."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is the resource and the design, not a clean causal proof of memorization. They built X-Riddles (900 dictionary items + 243 linguist-written novel homophonic xiehouyu), got real human MCQ and familiarity baselines, and ran match / explain / create. The human near-zero Low→New gap (71.5→68.6) versus large gaps for several Chinese-origin frontier models (mean Δacc 23.6) is a clean empirical pattern, and Gemini 3.1 Pro clearing novel items at 92.6% while staying near-zero Δ is genuinely interesting. Creation ratings also land where you’d expect: models trail humans on reasonableness and funniness.\n\nWhat is new is not “contamination matters” in the abstract—that literature already exists—but a culturally specific, expert-authored uncontaminated slice plus an operational Low-minus-New index with human anchor and token-effort splits. Condition ablations on Qwen, explanation ratings, and the honesty in §4.3.4 about openness/scale confounds are all to their credit. Citations look normal for this area.\n\nSoft spots, in proportion: Δacc is an indirect probe, not a retrieval assay. Most “novel” items reuse seed homophone pairs, so phonological links are not fully new; MCQ distractors also differ in flavor between dictionary and novel answers. That weakens the strong “more Chinese data ⇒ memorization” headline without erasing the pattern. Novel items are only the hard homophonic slice, and creation is only two models with author ratings. None of that sinks the paper; it caps how far the causal story can travel.\n\nThis is for people who care about LLM eval under cultural data leakage, Chinese figurative language, or how to build uncontaminated language benchmarks. I’d bring it to reading group, cite the benchmark and the Δacc template, and send it to referees. Worth engaging; revise the interpretation, keep the data.","headline":"Solid contamination-aware Chinese riddle benchmark; the Δacc story is useful but only partly identified, and the authors mostly own that.","tokens_in":22788,"tokens_out":497,"would_cite":true,"duration_ms":20061,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A large accuracy drop from rare to brand-new Chinese xiehouyu riddles marks memorization in frontier Chinese LLMs, even as some models beat humans on novel items and still lag at creating them.","keywords":["xiehouyu","LLM reasoning","memorization","data contamination","Chinese wordplay","homophonic puns","creative generation","multiple-choice evaluation"],"falsifier":"Build a new matched novel set with the same human difficulty profile, or measure how often the low-familiarity dictionary items actually appear in public Chinese pretraining crawls: if the Low–New gap for Chinese models vanishes under tighter difficulty matching, or fails to track documented Chinese-data exposure, the memorization reading of Δacc weakens.","tokens_in":22553,"feed_emoji":"🧩","tokens_out":1013,"duration_ms":32147,"temperature":0.7,"pith_summary":"This paper asks whether large language models truly reason about non-literal Chinese wordplay or mostly retrieve answers seen in training. The authors build X-Riddles, mixing dictionary xiehouyu with hundreds of new homophonic riddles written by linguists so the novel set cannot have leaked into pretraining. On multiple-choice matching, native speakers score almost the same on low-familiarity existing items and on novel ones, which the authors take as shared reasoning. Frontier Chinese models instead lose on average about 24 percentage points from low-familiarity to novel items, while English-centric models lose only about 5; the authors read that split as heavier Chinese-data memorization. At the same time, the best model reaches 92.6% on novel items—above human accuracy—yet model-written xiehouyu are rated less reasonable and less funny than human ones. The work argues that reasoning claims need contamination-aware tests, and that creative language play in this Chinese form still favors human experts.","feed_headline":"Chinese LLMs drop ~24 points on riddles they never saw","feed_subtitle":"Novel xiehouyu separate memorization from reasoning; creation still trails human linguists.","key_machinery":"Δacc (accuracy on low-familiarity existing homophonic xiehouyu minus accuracy on novel ones): a lower-bound memorization index, justified because humans show almost no gap and are therefore argued to use the same reasoning process on both sets.","core_discovery":"Using linguist-written novel xiehouyu to block contamination, the authors show that frontier Chinese models have a large mean accuracy drop (about 23.6 points) from low-familiarity existing homophonic items to novel ones, versus near-zero drop for native speakers and about 5 points for English-centric models—evidence they treat as memorization from larger Chinese training data—while Gemini 3.1 Pro still scores 92.6% on novel items (above human MCQ accuracy) and LLM-created xiehouyu receive worse reasonableness and funniness ratings than human creations.","pith_inferences":["Phonology-aware training or explicit pronunciation modules may be needed before models close the remaining gap on strict homophonic links.","The same Low-versus-New template could be ported to other language-specific games (puns, two-part allegories, dialect wordplay) to audit memorization outside Chinese.","Creation tasks may stay a stricter creativity filter than multiple-choice understanding even after contamination is controlled.","Open reporting of Chinese pretraining mix would let future work test whether Δacc scales with documented Chinese token volume."],"forward_implications":["Benchmarks built from culturally circulated material need paired novel items of matched difficulty, or they will mix retrieval with reasoning.","Claims of strong LLM reasoning on Chinese figurative language should be re-checked on uncontaminated, expert-written items.","Homophonic Chinese wordplay remains a harder test than polysemy or direct proverb types for current models.","Even models that match or beat humans on novel multiple-choice xiehouyu still underperform human linguists at creating reasonable, funny ones.","Heavier Chinese pretraining can inflate scores on rare existing items without equal gains on truly new riddles."],"fun_headline_variants":["Chinese LLMs drop 23.6 points on novel xiehouyu riddles","Novel xiehouyu expose memorization in frontier Chinese models","Gemini hits 92.6% on unseen xiehouyu; creation still lags humans","Δacc shows Chinese models memorize low-frequency xiehouyu","LLMs trail linguists at creating new Chinese xiehouyu"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The score gap between rare dictionary riddles and linguist-written new ones mainly reflects training-set memorization, not leftover difficulty, style, or distribution shift between those two sets—and the Chinese-versus-English contrast is driven mainly by Chinese data volume even though the model groups also differ in openness and scale.","fun_headline_variants_meta":{"raw":{"variants":["Chinese LLMs drop 23.6 points on novel xiehouyu riddles","Novel xiehouyu expose memorization in frontier Chinese models","Gemini hits 92.6% on unseen xiehouyu; creation still lags humans","Δacc shows Chinese models memorize low-frequency xiehouyu","LLMs trail linguists at creating new Chinese xiehouyu"]},"model":"grok-4.5","effort":"low","cost_usd":0.003266,"raw_usage":{"total_tokens":1172,"prompt_tokens":887,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":32664000,"prompt_tokens_details":{"text_tokens":887,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":203,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":887,"tokens_out":82,"duration_ms":4555,"temperature":1.0,"reasoning_tokens":203,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T22:01:02.549278+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Build a new matched novel set with the same human difficulty profile, or measure how often the low-familiarity dictionary items actually appear in public Chinese pretraining crawls: if the Low–New gap for Chinese models vanishes under tighter difficulty matching, or fails to track documented Chinese-data exposure, the memorization reading of Δacc weakens.","supporting_citations":[],"review_version":1}