{"id":"1aa9d0b1-9709-4939-9abd-6ce7599c30c6","arxiv_id":"2411.11354","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey of oracle character recognition, covering datasets, methods, challenges, and future directions.","lead":"This paper surveys the field of automated recognition of ancient Chinese oracle bone characters, covering datasets, methods, and open challenges. It is a reference guide for computer vision researchers and archaeologists working on this niche but historically important script.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's SOTA table is internally inconsistent about OBC306's best method, and the cited MAAN paper appears to be about Dongba characters; a reference and accuracy audit is needed.","rationale":"The reader's weakest assumption identifies the same core issue: the survey's value rests on the accuracy of its reported dataset statistics and SOTA accuracies, and the OBC306 attribution is internally inconsistent. My stress-test confirms this and sharpens it: reference [34] appears to describe Dongba character recognition, not oracle character recognition, which suggests a citation error rather than a mere wording slip. This makes the concern more load-bearing because it calls into question the reliability of the methodology survey's attributions, not just a single table entry. However, the issue is localized and correctable, and the survey's structure and coverage remain valuable. The reader already issued CONDITIONAL, and my analysis does not change that verdict; it strengthens the justification for the condition by adding a specific reference-domain check. I mark agreement as 'partial' because I agree with the reader's identified inconsistency but go further by pointing out the likely domain mismatch in reference [34].","tokens_in":20750,"tokens_out":1743,"duration_ms":17967,"concrete_test":"Retrieve the full texts of [32] (Diff-Oracle) and [34] (MAAN) and check: (a) whether [34] reports any experiment on OBC306 or any oracle-character dataset, and (b) what average and total accuracies each paper reports for OBC306. If [34] contains no OBC306 result, then the Section 4.1 sentence should be removed and all SOTA attributions in Table 1 should be re-checked against primary sources. If [34] does report a higher OBC306 accuracy than [32], then Table 1 and Section 4.4 need correction. In either case, verify the domain and dataset of [34] to confirm whether it is an oracle-character or Dongba-character method.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central contribution is a synthesis, so its utility depends on accurate secondary data. The paper is internally inconsistent about the state of the art on OBC306: Section 4.1 states that MAAN [34] achieves the highest accuracy on OBC306 [1], while Section 4.4 and Table 1 credit Diff-Oracle [32] with the best OBC306 results (average accuracy 88.07, total accuracy 94.12). Both statements cannot be true. More seriously, reference [34] is titled \"Multiple attentional aggregation network for handwritten Dongba character recognition,\" which describes a different script (Dongba pictographs), not oracle characters. If this citation is wrong, then the claim that a CNN-based method is SOTA on OBC306 is unsupported, and the methodology narrative in Section 4.1 misattributes a result to the oracle-character literature. Because the paper's stated purpose is to provide a reliable map of datasets, methods, and state-of-the-art results, an unresolved contradiction in the central benchmark undermines the core contribution. The issue is correctable, so conditional acceptance is appropriate pending verification of every SOTA entry in Table 1 against its primary source.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript surveys oracle character recognition (OrCR), organizing the field into three challenges (writing variability, data scarcity, and low image quality), a catalogue of roughly 20 datasets and online resources, a taxonomy of methodologies (traditional, deep, and hybrid, plus challenge-specific approaches), a discussion of seven related tasks, and future research directions. It presents two summary tables (datasets with reported accuracies, and representative methods) and claims to be the first systematic and structured survey of OrCR.","tokens_in":21025,"tokens_out":3198,"duration_ms":31475,"significance":"If the reported secondary data are reliable, this survey has clear value as an entry point for newcomers and a structured reference for practitioners: it assembles dataset statistics, state-of-the-art accuracies, and a method taxonomy in one place, and it identifies open problems such as open-set recognition and robust handling of label noise. The paper makes no new derivations or predictions, but its contribution as a synthesis is legitimate. However, the central deliverable is a map of other papers' results, so the survey's utility depends directly on the mutual consistency and correct attribution of those results; the inconsistencies described below affect the flagship dataset and a named state-of-the-art method and therefore need to be resolved before the survey can be relied upon.","major_comments":[{"comment":"The survey is internally inconsistent about the state of the art on OBC306: Section 4.1 states that MAAN [34] achieves the highest accuracy on OBC306 [1], while Table 1 credits Diff-Oracle [32] with the best OBC306 results (average accuracy 88.07, total accuracy 94.12) and Section 4.4 repeats that Diff-Oracle demonstrates the optimal accuracy on OBC306. Both attributions cannot be true, and since OBC306 is described as the widely used large-scale benchmark, the contradiction undermines the reliability of the benchmark table. The authors must verify the primary sources and correct the conflicting statements.","section":"Sections 4.1, 4.4, and Table 1"},{"comment":"Reference [34] is titled \"Multiple attentional aggregation network for handwritten Dongba character recognition\" (Expert Systems with Applications 213 (2023) 118865), which concerns a different script (Dongba pictographs), not oracle characters. The discussion in Section 4.1 that MAAN achieves the highest accuracy on OBC306, the description of MAAN as introducing a hybrid attentional mapping unit and spatial attentional aggregation unit for OrCR, and the listing of MAAN as a representative deep-learning method in Table 2 are therefore unsupported unless the authors intended a different paper. The same incorrect citation also appears in Section 3.2 in the claim that OBC306 is widely used [5, 33, 34]. Please re-check the reference and remove or replace it with a genuine OrCR method.","section":"Reference [34], Section 4.1, and Table 2"},{"comment":"The sentence \"Diff-Oracle [32] demonstrates the optimal accuracy on Oracle-241 [33] and OBC306 [1]\" contains a citation error: Oracle-241 is introduced in Section 3.2 as reference [6], not as reference [33] (AGTGAN). This appears to be a typo, but it is precisely the kind of attribution error that matters in a survey, and the reported accuracy for Oracle-241 in Table 1 (90.47 average, 91.11 total, credited to [32]) should also be double-checked against the primary source.","section":"Section 4.4, noise simulation discussion"}],"minor_comments":[{"comment":"The line \"Preprint submitted to Nuclear Physics B\" appears to be a template artifact left over from a different submission workflow; it should be removed or replaced with the intended journal or venue information.","section":"Title page"},{"comment":"The discussion states that \"methods like Orc-Bert and FFD rely heavily on extracting radicals from character images,\" but the earlier description of FFD (B-spline free-form deformation on stroke vectors) does not mention radical extraction. Please reconcile this statement with the method descriptions.","section":"Section 4.3, few/zero-shot discussion"},{"comment":"The OracleRC row has no entry under the #Samples column, although the text in Section 3.2 reports 2,005 character classes. Please provide the total sample count or explicitly state that it is not reported in the source.","section":"Table 1, OracleRC row"},{"comment":"The figure caption contains the stray word \"Top\" at the end of the first sentence; this appears to be a layout artifact and should be removed.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The survey's main contribution is the reliability of its secondary data, so the reference-level inconsistencies in the OBC306 SOTA attribution and the apparent mis-citation of a Dongba-character paper as an OrCR method are not merely cosmetic. I recommend that the editor request a full audit of Table 1's accuracy entries and of the method-to-benchmark attributions against their primary sources before the paper is considered further."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is the first systematic survey of oracle character recognition, and it earns its keep as an entry point. It consolidates roughly twenty datasets into a single table, organizes methods around the three challenges it identifies (writing variability, data scarcity, low image quality), covers seven related tasks, and lists online resources. The structure is clean and the challenge analysis is sensible. For someone walking into OrCR, this map is the product.\n\nThe soft spots sit in the secondary data, which is where a survey's value lives, and they are real. Section 4.1 says MAAN [34] reaches the highest accuracy on OBC306; Table 1 and Section 4.4 credit Diff-Oracle [32] with the best OBC306 numbers (88.07 average, 94.12 total). Both cannot be right. Worse, reference [34] is titled 'Multiple attentional aggregation network for handwritten dongba character recognition' — Dongba is a different script. The CNN-SOTA claim on OBC306 therefore hangs on a misattributed citation. There is a smaller slip in Section 4.4 where 'Oracle-241 [33]' should presumably cite [6]. These are correctable, but they are exactly the errors a survey must not carry, because readers will treat Table 1 as ground truth.\n\nOne thing to weigh: the method the table credits with SOTA on OBC306 and Oracle-241 is the authors' own Diff-Oracle, while the competing claim is the one that is miscited. I do not read that as deliberate, but it raises the bar for checking every SOTA entry against its primary source. The 'submitted to Nuclear Physics B' header is a leftover template artifact, cosmetic only.\n\nWho this is for: newcomers to OrCR and digital-humanities researchers who want one consolidated picture of datasets, benchmarks, and method families. It delivers that, assuming the SOTA table gets audited. This deserves a serious referee; the errors are fixable and the consolidation is worth having in the literature. My recommendation: accept with major revision, with a mandatory entry-by-entry verification of Table 1 against primary sources, fixing the MAAN/Dongba attribution, and cleaning up the citation slips.","headline":"First systematic OrCR survey with a genuinely useful dataset and method consolidation, but the SOTA table contradicts itself on OBC306 and cites a Dongba-character paper as oracle SOTA — needs a fact-check pass before the synthesis can be trusted.","tokens_in":21474,"tokens_out":5770,"would_cite":true,"duration_ms":50787,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's central claim is that oracle character recognition can be organized into a structured landscape of three challenges, about twenty datasets, and a method taxonomy, and that it offers the first systematic survey of that landscape.","keywords":["oracle bone script","oracle character recognition","handprinted characters","scanned character images","writing variability","data scarcity","deep learning","zero-shot recognition"],"falsifier":"A reader could check the original OBC306 papers and verify the two cited methods under the survey's stated accuracy definitions; the inconsistency between MAAN and Diff-Oracle would immediately show which attribution is wrong. More broadly, reproducing Table 1's reported accuracies from the cited papers would settle whether the survey's performance map is reliable.","tokens_in":20593,"feed_emoji":"🦴","tokens_out":6296,"duration_ms":60290,"temperature":0.7,"pith_summary":"This survey tries to establish a systematic map of oracle character recognition (OrCR), the automated reading of roughly 3,500-year-old Chinese inscriptions on turtle shells and animal bones. It organizes the field into three core challenges—writing variability, data scarcity, and low image quality—and reviews about twenty datasets, the methods built around each challenge, and seven neighboring tasks such as decipherment, detection, and bone rejoining. If the map is accurate, a new researcher can locate the standard benchmarks, the currently reported accuracies, and the open problems without reading dozens of primary papers. The paper also argues that no standard benchmark exists and that none of the three challenges is fully solved.","feed_headline":"Survey maps oracle character recognition: 20 datasets, 3 challenges","feed_subtitle":"A single guide now connects benchmarks, methods, and open problems for researchers new to these ancient scripts.","key_machinery":"The organizing mechanism is the challenge-driven taxonomy: every dataset is categorized as handprinted or scanned, every method is filed under general recognition, writing variability, data scarcity, or low image quality, and Tables 1 and 2 consolidate the numbers. Within that structure, the load-bearing analytical objects are the two evaluation metrics, total accuracy and average accuracy, because they determine whether a method's reported success on majority classes is masking failure on rare characters.","core_discovery":"Oracle character recognition has produced roughly twenty datasets and dozens of methods, but no prior work had organized the field into a coherent structure. The paper supplies that structure: three intrinsic challenges (writing variability, data scarcity, low image quality), a dataset-by-dataset summary with best reported accuracies, a method taxonomy tied to each challenge, and seven related tasks from decipherment to oracle bone rejoining. Its strongest factual claims are that OBC306 is the de facto standard scanned benchmark, that no unified benchmark exists because datasets lack shared class labels or modern-Chinese mappings, and that reported accuracies on scanned-image datasets remain clearly below those on handprinted datasets. It also argues that radical decomposition enables zero-shot recognition of unseen classes and that generative augmentation, particularly diffusion-based synthesis, is the leading response to both data scarcity and image noise.","pith_inferences":["If the conflict between MAAN and Diff-Oracle over the best OBC306 accuracy is representative, the survey's accuracy column should be treated as a starting index rather than a verified leaderboard, and a reader should check original papers before using any number as a baseline.","The survey's own analysis implies that a unified dataset built from radical-level annotations could bridge datasets that currently share no character classes, enabling transfer learning across the whole field.","A natural test of the paper's synthesis would be to train with diffusion-based noise simulation plus unsupervised domain adaptation and compare against either technique alone on OBC306; the survey's structure predicts the combination should win.","As generative methods improve, the boundary between OrCR and oracle character decipherment may blur, because generated stylized glyphs could be used to test decipherment hypotheses before physical evidence is found."],"forward_implications":["A newcomer can use the survey's dataset table to pick a starting benchmark: OBC306 for scale and scanned realism, HUST-OBS for clean handprinted data, and HWOBC for balanced classes.","Because the field lacks a common class label set, results reported on different oracle datasets are not directly comparable, so a shared benchmark or cross-dataset mapping is the next infrastructure step.","Reporting average accuracy alongside total accuracy will remain necessary for imbalanced oracle datasets, since the gap between the two exposes minority-class failures.","Radical-based zero-shot reasoning provides a concrete route to recognizing unseen oracle characters, and combining it with generative augmentation is a natural next step.","Training with both denoising and noise simulation, rather than either alone, is a promising path toward better recognition of real scanned oracle images."],"supporting_citations":[{"why":"Defines OBC306, the large scanned benchmark whose statistics and best accuracy anchor the survey's dataset table.","marker":"[1]"},{"why":"Defines Oracle-AYNU and the deep-metric-learning baseline used to illustrate handprinted data and imbalanced distributions.","marker":"[2]"},{"why":"Provides HUST-OBS, the open handprinted dataset recommended by the survey for avoiding image-quality issues.","marker":"[13]"},{"why":"Provides Oracle-50K/Oracle-FS and the Orc-Bert few-shot augmentor, supporting the data-scarcity discussion.","marker":"[17]"},{"why":"Provides HWOBC, the balanced handprinted benchmark whose reported accuracy illustrates the upper end of OrCR performance.","marker":"[18]"},{"why":"Supplies Diff-Oracle, the diffusion-based generator credited with the best OBC306 and Oracle-241 accuracies in the low-image-quality discussion.","marker":"[32]"},{"why":"Supplies MAAN, the attentional network the survey's general-methods section credits with the highest OBC306 accuracy, creating the attribution conflict.","marker":"[34]"}],"fun_headline_variants":["Survey maps 20 oracle character datasets and 3 core challenges","First systematic survey of oracle character recognition","Oracle character recognition: 20 datasets, 3 challenges in one survey","The state of oracle character recognition: benchmarks and open problems","A guide to oracle bone script recognition: 20 datasets and beyond"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness depends on the reported dataset statistics and performance numbers being faithful to the cited papers, and this assumption is already strained by an internal conflict: Section 4.1 says MAAN holds the best OBC306 accuracy, while Table 1 and Section 4.4 credit Diff-Oracle with that result.","fun_headline_variants_meta":{"raw":{"variants":["Survey maps 20 oracle character datasets and 3 core challenges","First systematic survey of oracle character recognition","Oracle character recognition: 20 datasets, 3 challenges in one survey","The state of oracle character recognition: benchmarks and open problems","A guide to oracle bone script recognition: 20 datasets and beyond"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000288,"raw_usage":{"total_tokens":1678,"prompt_tokens":921,"completion_tokens":757,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":674}},"tokens_in":537,"tokens_out":757,"duration_ms":7185,"temperature":1.0,"reasoning_tokens":674,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:36:29.519278+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could check the original OBC306 papers and verify the two cited methods under the survey's stated accuracy definitions; the inconsistency between MAAN and Diff-Oracle would immediately show which attribution is wrong. More broadly, reproducing Table 1's reported accuracies from the cited papers would settle whether the survey's performance map is reliable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies MAAN, the attentional network the survey's general-methods section credits with the highest OBC306 accuracy, creating the attribution conflict."}],"review_version":1}