{"id":"cee8ee50-4264-4ffc-a1a7-2cc0587953cf","arxiv_id":"2608.11741","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new expert-audited dataset and benchmark for ancient Chinese character exegesis shows that multimodal LLMs improve substantially when fine-tuned on domain-specific VQA data.","lead":"This paper introduces ACCE, a four-level vision-language question-answering task for ancient Chinese character exegesis, along with a 500K-pair training dataset and a fully expert-verified 8K-pair benchmark. It reports that fine-tuning multimodal LLMs on the dataset substantially improves performance across all four levels.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"JieZi-Bench's early-script reference answers cannot come from the four cited dictionaries (Shuowen predates OBI discovery), so benchmark separation from training knowledge is unestablished and the fine-tuning gains may be partly circular.","rationale":"The reader's weakest assumption focused on the self-report of exhaustive expert verification and on dictionary consistency. I agree those are important, but the sharper structural problem is that the four cited benchmark dictionaries cannot, in principle, ground the early-script portion of JieZi-Bench. Because JieZi-Bench deliberately includes OBI, Bronze, and Warring States images, and because none of the four lexicographic sources contains systematic diachronic or component-level analysis for those script stages, the reference answers for those items must have been produced from modern scholarly knowledge. That knowledge is not 'held separate' from JieZi-Dataset in any verifiable sense; the training source, Hanzi Yuanliu Dazidian, is itself a modern etymological dictionary covering the same multi-script evolution. This makes the benchmark's independence claim unestablished rather than merely unverified. The proposed audit would settle it directly: if the early-script reference answers are traceable only to modern scholarship, then the fine-tuning improvements on those subsets may be inflated by distributional overlap with training, and the central claim of a reliable, independent benchmark needs to be revised. If the audit instead shows that the four dictionaries, plus expert synthesis, suffice and that the answers do not overlap with Hanzi Yuanliu Dazidian, the original conditional acceptance would be justified. I therefore keep the reader's conditional verdict but sharpen the condition: release the provenance mapping between each JieZi-Bench answer and its source dictionary or expert reconstruction, especially for early-script items.","tokens_in":27827,"tokens_out":7401,"duration_ms":80724,"concrete_test":"Run a source-traceability audit on a random sample of at least 100 JieZi-Bench images with script stage OBI, Bronze, or Warring States: two independent paleography experts attempt to reconstruct each L2-L4 reference answer using only the four cited dictionaries (Kangxi, Shuowen, Shuowen Zhu, Revised Mandarin). Record the fraction of answers that cannot be reconstructed without Hanzi Yuanliu Dazidian or equivalent modern scholarship. If that fraction is substantially above zero for early-script stages, the benchmark's source separation is not established for those stages; the authors should either re-source those answers or report them as expert-authored modern reconstructions with provenance.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing assumption is the independence of JieZi-Bench from JieZi-Dataset. Section 4.1 says benchmark reference answers are curated from Kangxi, Shuowen Jiezi, Shuowen Jiezi Zhu, and Revised Mandarin Chinese Dictionary, 'held separate from the training data.' But JieZi-Bench spans all six script stages (Sec 4.3, Fig. 3), with Bronze and Seal a significant proportion and OBI present. Shuowen Jiezi (ca. 100 CE) cannot contain oracle-bone analysis — OBI were not excavated and recognized until 1899 — and Kangxi/Revised Mandarin are not systematic sources for Bronze or Warring States component functions or diachronic trajectories. For every OBI, Bronze, and Warring States item, the L2-L4 reference answers must therefore have been written from modern paleographic knowledge, not taken from the four cited dictionaries. That modern knowledge is the same scholarly tradition, and largely the same etymological dictionary, used to build JieZi-Dataset. So the claimed 'held separate' guarantee is not established for early-script items, and the benchmark is not demonstrably independent of the training data. This is not merely a self-report issue; the paper's own Fig. 6 example (射) says Shuowen 'completely misinterpreted' the etymology, showing the benchmark cannot treat Shuowen as authoritative. If references were silently filled from the training-adjacent source, fine-tuning gains on JieZi-Bench could reflect memorization of the same modern etymological analyses rather than transferable exegesis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces Ancient Chinese Character Exegesis (ACCE), a four-level vision-language question answering task that spans basic character identification, glyph-form analysis, meaning exegesis, and diachronic evolution. The authors construct two resources: JieZi-Dataset, a roughly 500K QA-pair training set with about 130K glyph images across six script stages, built from the modern etymological dictionary Hanzi Yuanliu Dazidian through OCR, LLM-based structurization, and template-guided VQA generation with stage-wise expert spot-checking; and JieZi-Bench, an approximately 8K QA-pair benchmark on 1,024 images, claimed to be fully expert-verified and curated from four lexicographic sources held separate from the training data. The paper benchmarks a range of commercial and open MLLMs and reports that fine-tuning Qwen3.5-2B/4B/9B on JieZi-Dataset improves performance on all ACCE levels, with additional controls for unseen characters, unseen glyphs, in-context learning, retrieval-augmented generation, data scaling, and question paraphrasing.","tokens_in":28111,"tokens_out":6883,"duration_ms":73597,"significance":"If the benchmark-independence issue identified below is resolved, this is a substantial contribution to computational paleography and multimodal understanding. The paper formalizes a scholarly workflow that prior datasets reduce to recognition, provides the first multi-script, multi-level VQA training resource at this scale, and evaluates with several carefully designed controls: unseen-character and unseen-glyph splits, comparison against ICL and RAG baselines, a paraphrase-robustness test, and human-expert validation of both the LLM-as-a-judge protocol and BERTScore. The public release of code and data and the explicit reporting of verification error rates are strengths. The main empirical claim, that domain-specific fine-tuning improves ACCE performance, is plausible and mostly supported by the reported experiments; the weakest load-bearing point is the provenance and independence of the JieZi-Bench reference answers for early script stages.","major_comments":[{"comment":"The provenance description for JieZi-Bench is internally inconsistent. Section 4.1 states that JieZi-Bench is sourced solely from Kangxi Dictionary, Shuowen Jiezi, Shuowen Jiezi Zhu, and Revised Mandarin Chinese Dictionary, while Section 4.3 and Fig. 3 state that the benchmark spans all six script stages, including Oracle Bone, Bronze, and Warring States. Shuowen Jiezi (ca. 100 CE) cannot contain oracle-bone forms, which were not archaeologically recognized until 1899, and none of the four cited works is a systematic source for Bronze or Warring States component functions or diachronic trajectories. The manuscript must state, per script stage, where the benchmark glyph images and the L2-L4 reference answers actually come from. As written, the claim that the benchmark was held separate from the training data is not established for a substantial subset of the benchmark items.","section":"Section 4.1 and Section 4.3"},{"comment":"Because JieZi-Dataset is built from Hanzi Yuanliu Dazidian, a modern etymological dictionary that synthesizes contemporary paleographic scholarship, and because early-script reference answers in JieZi-Bench must, per the previous comment, be expert-written from the same modern scholarship, the unseen-character and unseen-glyph splits do not by themselves rule out memorization of etymological analyses rather than transferable exegetical reasoning. The central fine-tuning conclusion would be materially strengthened by (i) reporting the provenance of each benchmark answer as taken from a pre-modern lexicographic source versus written by experts from modern paleographic literature, (ii) measuring textual or semantic overlap between benchmark reference answers and Hanzi Yuanliu Dazidian entries, and (iii) re-running the headline fine-tuning comparison on the subset of benchmark items whose reference answers can be traced to sources independent of the training dictionary.","section":"Section 5.3 and Section 6"},{"comment":"The claim that JieZi-Bench is entirely expert-curated is not backed by reproducibility statistics for the verification process. The paper reports error rates found during correction but does not state how many experts participated, what their paleographic qualifications were, how often experts disagreed, or how disagreements were resolved. Given that the paper itself acknowledges that scholarly consensus varies for early scripts and that even Shuowen Jiezi can completely misinterpret an etymology (Fig. 6), reporting inter-annotator agreement on a sample of the benchmark would materially support the benchmark-reliability claim.","section":"Section 4.2 and Section 4.3"}],"minor_comments":[{"comment":"The empirical token-length threshold of 200 used to select benchmark entries should be justified, and the sensitivity of the benchmark composition to this threshold should be reported.","section":"Section 4.1"},{"comment":"Several generalization cells have very small sample sizes (e.g., Bronze UC n=26, Bronze UG n=87); confidence intervals or exact tests should be reported before drawing conclusions from differences between unseen-character and unseen-glyph conditions.","section":"Table 3"},{"comment":"The phrase substantially improves performance across all four levels is broadly supported, but the L3 ORIM gains for the 2B and 9B models are modest (+3.4 and +2.0 in Table 2); the claim could be nuanced to avoid overstating the effect on original-meaning tasks.","section":"Abstract and Section 5.3"},{"comment":"The near-zero baseline scores for Qwen3.5-4B are unusual and the copying-bias explanation is plausible, but the paper should clarify whether this anomalous behavior was observed consistently across all inference settings and whether the other model families exhibit any similar instability.","section":"Section 5.3 and Table 2"}],"recommendation":"major_revision","confidential_remarks":"The key gate is the Section 4.1 provenance issue. If the authors can provide per-script-stage provenance for JieZi-Bench reference answers and show that early-script items are not drawn from the same modern etymological tradition as JieZi-Dataset, the paper would be a strong fit for the journal. If the provenance cannot be documented, the headline fine-tuning claim is weakened enough that the benchmark's independence guarantee should be substantially revised. I would also suggest that the dataset release include provenance annotations per benchmark item, not only the QA pairs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first serious attempt to formalize the full exegesis workflow for ancient Chinese characters as a four-level VQA task, and it comes with an impressively large training set (500K QA pairs) and a substantial benchmark (8K QA pairs). The experiments are more careful than the typical dataset paper: they include paraphrase robustness tests, comparisons to ICL/RAG, scaling curves, and human validation of BERTScore and LLM-judge metrics. Credit where due: the task decomposition into L1-L4 with ten subtasks is sensible and fills a real gap. The dataset construction pipeline, with template-constrained generation and expert-in-the-loop spot checks, is a reasonable approach to keep hallucination down at 500K scale.\n\nThe main soft spot is exactly what the stress-test note points at: the benchmark's independence claim. Section 4.1 says reference answers are curated from Kangxi, Shuowen Jiezi, Shuowen Jiezi Zhu, and Revised Mandarin Chinese Dictionary, held separate from training data. But JieZi-Bench includes Oracle Bone and Bronze items. Those dictionaries cannot provide systematic analyses for scripts that were not yet excavated (OBI) or only partially covered (Bronze). The paper's own example in Fig. 6 (射) says Shuowen completely misinterpreted the etymology, so at least for that item the reference had to be corrected using modern paleographic knowledge. That modern knowledge is the same scholarly tradition, and likely the same etymological dictionary (Hanzi Yuanliu Dazidian), that was used to build JieZi-Dataset. So the \"held separate\" guarantee is not established for early-script items, and the fine-tuning gains on L2-L4 could partly reflect memorization of the same modern etymological analyses rather than transferable exegesis. This is not an accusation of fabrication; it's a missing transparency requirement. The paper should report, per script stage, exactly which source was used for each benchmark answer, and ideally run a leakage diagnostic (e.g., train on JieZi-Dataset without the modern dictionary entries for benchmark characters, or show that the model performs better on benchmark items whose references are genuinely cross-checkable across the four classical dictionaries).\n\nThe rest of the evaluation looks solid. The unseen-character and unseen-glyph splits are a good idea, and the fact that structural metrics don't collapse on unseen glyphs is encouraging. The paraphrase robustness results are reasonable. The paper is honest about the spot-checking of the training set, and the failure analysis of Qwen3.5-4B is refreshingly detailed.\n\nBottom line: the resource is valuable and the experiments are mostly well done, but the benchmark's independence is a load-bearing claim that needs fixing before I would trust the fine-tuning numbers as evidence of general exegesis. The authors should be asked to clarify sources per script type and to provide leakage diagnostics. That said, this paper deserves a serious referee; it's a solid contribution to computational paleography, and the gaps are addressable.\n\nRecommendation: send to peer review, but with a strong request for the source-separation documentation.","headline":"Strong resource paper whose benchmark independence claim for early script stages does not hold up as stated; fine-tuning gains may be partly circular.","tokens_in":28656,"tokens_out":3263,"would_cite":true,"duration_ms":32859,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that ancient Chinese character exegesis can be modeled as a four-level vision-language question-answering task, and that the JieZi-Dataset and JieZi-Bench resources make it the first such task to have a large-scale…","keywords":["Ancient Chinese Character Exegesis","vision-language question answering","paleographic dataset","multimodal large language models","diachronic evolution","glyph-form analysis","benchmark","expert audit"],"falsifier":"Independently re-derive answers for a random sample of JieZi-Bench items: give three paleography experts who did not build the benchmark the same glyph images and the same four dictionaries, without showing the released answers, and measure agreement with the released references; low agreement (e.g., below roughly 90%) would directly undermine the benchmark reliability claim.","tokens_in":27621,"feed_emoji":"📜","tokens_out":8398,"duration_ms":81358,"temperature":0.7,"pith_summary":"The paper's central claim is that the scholarly interpretation of an ancient Chinese glyph—identifying it, analyzing its components, explaining its original meaning, and tracing how it changed across script periods—can be formalized as a four-level vision-language question-answering task (ACCE) and backed by the first large-scale resources for that task. The authors build JieZi-Dataset, with over 500K QA pairs and 130K glyph images spanning six script stages, using an expert-in-the-loop pipeline in which an LLM generates questions from expert-designed templates plus dictionary source text, and experts spot-check each stage. They build JieZi-Bench, roughly 8K QA pairs whose reference answers come from separate authoritative dictionaries and every one of which was manually verified by experts. Benchmarking current multimodal models, they report that the models do reasonably on basic identification and script classification but struggle on form analysis, meaning, and diachronic evolution, and that fine-tuning on JieZi-Dataset improves performance on all four levels. A careful reader would care because this turns a labor-intensive humanities skill into a measurable machine-learning task with a public dataset and a reproducible evaluation.","feed_headline":"Ancient glyph exegesis becomes a 500K-question AI benchmark","feed_subtitle":"A new dataset lifts multimodal models on all four levels of ancient Chinese character analysis.","key_machinery":"The load-bearing object is the ACCE task decomposition itself: four progressive levels (L1 basic information, L2 glyph form, L3 meaning, L4 diachronic evolution) spanning ten subtasks (CHAR, SCRC, STRC, COMR, COMF, COMI, FORC, ORIM, COME, EVOI). The mechanism that carries the argument is the expert-in-the-loop generation pipeline: expert-designed QA templates plus dictionary source-text references constrain LLM generation so answers are grounded in verified metadata rather than parametric memory; stage-wise human verification (spot-checks for the training set, exhaustive checks for the benchmark) prevents error propagation. The named central entities are JieZi-Dataset, the 500K-pair training resource, and JieZi-Bench, the 8K-pair evaluation resource, and the separation of benchmark sources from training sources is what makes the fine-tuning gains interpretable.","core_discovery":"The central discovery is that the full exegesis workflow—not just recognition—can be stated as a structured VQA task with four progressive levels: basic information (character identity and script type), glyph form (structure, components, their functions and interpretations, formation category), glyph meaning (original meaning), and diachronic evolution (component evolution and holistic explanation). To support it, the paper contributes two complementary resources: a large training set whose generation is constrained by expert templates and verified dictionary text to suppress hallucination, and a smaller benchmark whose reference answers are curated from four authoritative lexicographic works and exhaustively checked by experts. On that benchmark, the paper finds that general multimodal models score moderately on script classification and formation classification but drop sharply on component-level, semantic, and evolutionary reasoning; after fine-tuning on JieZi-Dataset, a small model surpasses much larger general models on structural parsing, and the largest fine-tuned model achieves the strongest results across nearly all subtasks. The paper concludes that domain-specific, expert-audited training data is the decisive bottleneck for computational paleography, not model scale.","pith_inferences":["Beyond the paper's claims, the four-level decomposition looks portable to other undeciphered or under-resourced scripts: identification, form analysis, meaning, and diachronic change are general questions for cuneiform, Egyptian, and Maya writing, though those fields would need equivalent authoritative dictionaries to reproduce the pipeline.","The finding that form-based reasoning survives recognition failure suggests a two-stage design—glyph-agnostic structural analysis feeding a separate identification stage—might be more effective for rare glyphs than end-to-end generation, a hypothesis the paper's data could be used to test.","The validated automated metrics (BERTScore and an LLM judge with human-correlation around 0.6 to 0.7) could serve as cheaper evaluation tools for open-ended humanities VQA beyond ancient Chinese, saving human annotation effort in benchmark construction.","The paper's paraphrase test covers only two rewritten questions per instance; a stronger paraphrase attack with more diverse rewrites would clarify how much of the fine-tuning gain is genuine glyph-grounded exegesis versus template pattern matching."],"forward_implications":["JieZi-Bench gives the field a common yardstick: any future model can be scored on the same four levels and ten subtasks, with explicit splits for unseen characters, unseen glyphs, and unseen components.","Fine-tuning on JieZi-Dataset improves every subtask, so teams without access to huge general-purpose corpora can build competitive paleography models from this single public resource.","Because structural parsing holds up even when character identification fails on unseen glyphs, the data appears to teach transferable knowledge of visual form rather than rote character-to-label mapping; this points toward architectures that separate form analysis from identification.","Current models' lowest scores concentrate in component interpretation, original meaning, and evolution, which singles out the exact tasks where future data collection and modeling effort should focus."],"supporting_citations":[{"why":"Primary source of JieZi-Dataset: modern etymological dictionary covering glyph form, meaning, and diachronic evolution for 13K+ characters.","marker":"[21]"},{"why":"External dataset whose character and script labels are aligned with dictionary entries to add glyph diversity.","marker":"[52]"},{"why":"External image dataset providing additional historic-document glyphs for characters covered by the dictionary.","marker":"[66]"},{"why":"OCR pipeline used to digitize the scanned dictionary text into image-text pairs.","marker":"[17]"},{"why":"LLM prompt-based extraction that converts raw dictionary prose into structured metadata records.","marker":"[18]"},{"why":"Object detector trained to localize glyph images and classify script types from scanned pages.","marker":"[27]"},{"why":"Open-source model family used as the base for fine-tuning experiments; the improvements on JieZi-Bench establish the dataset's training value.","marker":"[41]"},{"why":"Closed-source model benchmarked on JieZi-Bench; its low character-recognition score supports the claim that general models lack paleographic knowledge.","marker":"[38]"}],"fun_headline_variants":["500K expert-checked questions push AI on ancient Chinese glyphs","Ancient Chinese exegesis gets a 500K-strong expert-vetted AI benchmark","Expert-audited 500K QA pairs benchmark ancient Chinese character exegesis","New AI benchmark packs 500K expert-vetted questions for ancient Chinese paleography"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that every JieZi-Bench QA pair was truly checked and revised by human experts and that the four source dictionaries agree on the entries that were kept, so the reference answers are trustworthy ground truth.","fun_headline_variants_meta":{"raw":{"variants":["500K expert-checked questions push AI on ancient Chinese glyphs","Ancient Chinese exegesis gets a 500K-strong expert-vetted AI benchmark","Expert-audited 500K QA pairs benchmark ancient Chinese character exegesis","New AI benchmark packs 500K expert-vetted questions for ancient Chinese paleography"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002674,"raw_usage":{"total_tokens":10260,"prompt_tokens":1045,"completion_tokens":9215,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":661,"completion_tokens_details":{"reasoning_tokens":9130}},"tokens_in":661,"tokens_out":9215,"duration_ms":65226,"temperature":1.0,"reasoning_tokens":9130,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:28:24.082092+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-derive answers for a random sample of JieZi-Bench items: give three paleography experts who did not build the benchmark the same glyph images and the same four dictionaries, without showing the released answers, and measure agreement with the released references; low agreement (e.g., below roughly 90%) would directly undermine the benchmark reliability claim.","supporting_citations":[{"cited_title":"2023.Hanzi Yuanliu Dazidian [Dictionary of Chinese Character Etymology]","cited_arxiv_id":null,"evidence_quote":"Primary source of JieZi-Dataset: modern etymological dictionary covering glyph form, meaning, and diachronic evolution for 13K+ characters."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"External dataset whose character and script labels are aligned with dictionary entries to add glyph diversity."},{"cited_title":"the human back/spine","cited_arxiv_id":null,"evidence_quote":"External image dataset providing additional historic-document glyphs for characters covered by the dictionary."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"LLM prompt-based extraction that converts raw dictionary prose into structured metadata records."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closed-source model benchmarked on JieZi-Bench; its low character-recognition score supports the claim that general models lack paleographic knowledge."}],"review_version":1}