Pith. sign in

REVIEW 2 cited by

OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.17837 v1 pith:NFLNJMRZ submitted 2024-11-26 cs.CV

classification cs.CV
keywords semanticoraclesageunderstandingvisualbonecross-modalframeworkgraph-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition due to their complex pictographic structures and divergence from modern Chinese characters. We introduce OracleSage, a novel cross-modal framework that integrates hierarchical visual understanding with graph-based semantic reasoning. Specifically, we propose (1) a Hierarchical Visual-Semantic Understanding module that enables multi-granularity feature extraction through progressive fine-tuning of LLaVA's visual backbone, (2) a Graph-based Semantic Reasoning Framework that captures relationships between visual components and semantic concepts through dynamic message passing, and (3) OracleSem, a semantically enriched OBS dataset with comprehensive pictographic and semantic annotations. Experimental results demonstrate that OracleSage significantly outperforms state-of-the-art vision-language models. This research establishes a new paradigm for ancient text interpretation while providing valuable technical support for archaeological studies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters

    cs.CV 2025-09 conditional novelty 6.0 of 10

    PictOBI-20k, a new 15k-question benchmark, shows top large multimodal models reach only 53.7% accuracy at matching oracle bone pictographs to object photos, with vision encoders often outperforming the full models.

  2. Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A VLM-based framework with retrieval-augmented generation parses tombstone photos into structured semantic graphs, reaching 89.5 Smatch F1 versus 36.1 for the prior OCR pipeline.

Pith tools