REVIEW 8 cited by
An open dataset for the evolution of oracle bone characters: EVOBC
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The earliest extant Chinese characters originate from oracle bone inscriptions, which are closely related to other East Asian languages. These inscriptions hold immense value for anthropology and archaeology. However, deciphering oracle bone script remains a formidable challenge, with only approximately 1,600 of the over 4,500 extant characters elucidated to date. Further scholarly investigation is required to comprehensively understand this ancient writing system. Artificial Intelligence technology is a promising avenue for deciphering oracle bone characters, particularly concerning their evolution. However, one of the challenges is the lack of datasets mapping the evolution of these characters over time. In this study, we systematically collected ancient characters from authoritative texts and websites spanning six historical stages: Oracle Bone Characters - OBC (15th century B.C.), Bronze Inscriptions - BI (13th to 221 B.C.), Seal Script - SS (11th to 8th centuries B.C.), Spring and Autumn period Characters - SAC (770 to 476 B.C.), Warring States period Characters - WSC (475 B.C. to 221 B.C.), and Clerical Script - CS (221 B.C. to 220 A.D.). Subsequently, we constructed an extensive dataset, namely EVolution Oracle Bone Characters (EVOBC), consisting of 229,170 images representing 13,714 distinct character categories. We conducted validation and simulated deciphering on the constructed dataset, and the results demonstrate its high efficacy in aiding the study of oracle bone script. This openly accessible dataset aims to digitalize ancient Chinese scripts across multiple eras, facilitating the decipherment of oracle bone script by examining the evolution of glyph forms.
Forward citations
Cited by 8 Pith papers
-
JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis
A new expert-audited dataset and benchmark for ancient Chinese character exegesis shows that multimodal LLMs improve substantially when fine-tuned on domain-specific VQA data.
-
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
HCSU supplies the first large decoupled Tie/Bei calligraphy dataset with expert aesthetic labels and finds SOTA LVLMs remain knowledgeable yet unperceptive on fine-grained style.
-
AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning
AlphaOracle integrates morphological, contextual, and philological evidence into an interpretable pipeline that reportedly assists experts in deciphering oracle bone script, reading 'Lao' as a toponym or clan name.
-
MCCD: A Multi-Attribute Chinese Calligraphy Character Dataset Annotated with Script Styles, Dynasties, and Calligraphers
A 329,715-image Chinese calligraphy dataset with character, style (10), dynasty (15), and calligrapher (142) labels, plus single- and multi-task recognition benchmarks.
-
OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?
OBI-Bench evaluates 23 large multimodal models on five oracle bone inscription tasks and finds they lag on fine-grained perception but approach untrained-human level in deciphering.
-
OBIFormer: A Fast Attentive Denoising Framework for Oracle Bone Inscriptions
OBIFormer reports state-of-the-art PSNR and SSIM on oracle bone inscription denoising benchmarks while using fewer parameters than prior transformer-based methods.
-
OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion
OracleSage, a LLaVA-based framework with hierarchical visual features and graph-based semantic reasoning, reaches 20.2% top-1 accuracy on the new OracleSem dataset, far below standard classifiers at around 90%.
-
A comprehensive survey of oracle character recognition: challenges, benchmarks, and beyond
A structured survey of oracle character recognition, covering datasets, methods, challenges, and future directions.
Discussion (0). Continue with ORCID to comment.