Pith. sign in

REVIEW 7 cited by

An open dataset for the evolution of oracle bone characters: EVOBC

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12467 v2 pith:ZLAZR2CY submitted 2024-01-23 cs.AI

classification cs.AI
keywords charactersboneoracleevolutionscriptdatasetancientdeciphering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The earliest extant Chinese characters originate from oracle bone inscriptions, which are closely related to other East Asian languages. These inscriptions hold immense value for anthropology and archaeology. However, deciphering oracle bone script remains a formidable challenge, with only approximately 1,600 of the over 4,500 extant characters elucidated to date. Further scholarly investigation is required to comprehensively understand this ancient writing system. Artificial Intelligence technology is a promising avenue for deciphering oracle bone characters, particularly concerning their evolution. However, one of the challenges is the lack of datasets mapping the evolution of these characters over time. In this study, we systematically collected ancient characters from authoritative texts and websites spanning six historical stages: Oracle Bone Characters - OBC (15th century B.C.), Bronze Inscriptions - BI (13th to 221 B.C.), Seal Script - SS (11th to 8th centuries B.C.), Spring and Autumn period Characters - SAC (770 to 476 B.C.), Warring States period Characters - WSC (475 B.C. to 221 B.C.), and Clerical Script - CS (221 B.C. to 220 A.D.). Subsequently, we constructed an extensive dataset, namely EVolution Oracle Bone Characters (EVOBC), consisting of 229,170 images representing 13,714 distinct character categories. We conducted validation and simulated deciphering on the constructed dataset, and the results demonstrate its high efficacy in aiding the study of oracle bone script. This openly accessible dataset aims to digitalize ancient Chinese scripts across multiple eras, facilitating the decipherment of oracle bone script by examining the evolution of glyph forms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Chronicles-OCR is the first benchmark with 2,800 images across the complete evolutionary trajectory of Chinese characters, defining four tasks to evaluate VLLMs' cross-temporal visual perception.

  2. Decoding Ancient Oracle Bone Script via Generative Dictionary Retrieval

    cs.IR 2026-04 unverdicted novelty 7.0 of 10

    Generative dictionary retrieval decodes unseen Oracle Bone Script characters at 54.3% Top-10 accuracy by synthesizing plausible variants guided by character evolution principles.

  3. HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

    cs.CV 2026-07 conditional novelty 6.5 of 10

    HCSU supplies the first large decoupled Tie/Bei calligraphy dataset with expert aesthetic labels and finds SOTA LVLMs remain knowledgeable yet unperceptive on fine-grained style.

  4. AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning

    cs.HC 2026-07 conditional novelty 6.0 of 10

    AlphaOracle integrates morphological, contextual, and philological evidence into an interpretable pipeline that reportedly assists experts in deciphering oracle bone script, reading 'Lao' as a toponym or clan name.

  5. Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    The paper introduces the S-OBI benchmark for sentence-level oracle bone inscription understanding and reports that current MLLMs remain dependent on character-level recognition due to propagating visual errors.

  6. OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    OracleAnalyser applies post-training and a new Stable Focal Preference Optimization algorithm to a 3B MLLM for oracle bone script analysis, releasing datasets and a benchmark where the small model outperforms larger ones.

  7. Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    MSLA is a new attention mechanism that models multi-scale and cross-layer interactions to achieve more accurate OBI recognition than prior attention methods.

Pith tools