Pith. sign in

REVIEW 2 cited by

OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.01175 v2 pith:XE3K4UZ6 submitted 2024-12-02 cs.CV cs.AI

classification cs.CVcs.AI
keywords lmmsobi-benchoracletasksboneancientcharactersdeciphering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce OBI-Bench, a holistic benchmark crafted to systematically evaluate large multi-modal models (LMMs) on whole-process oracle bone inscriptions (OBI) processing tasks demanding expert-level domain knowledge and deliberate cognition. OBI-Bench includes 5,523 meticulously collected diverse-sourced images, covering five key domain problems: recognition, rejoining, classification, retrieval, and deciphering. These images span centuries of archaeological findings and years of research by front-line scholars, comprising multi-stage font appearances from excavation to synthesis, such as original oracle bone, inked rubbings, oracle bone fragments, cropped single characters, and handprinted characters. Unlike existing benchmarks, OBI-Bench focuses on advanced visual perception and reasoning with OBI-specific knowledge, challenging LMMs to perform tasks akin to those faced by experts. The evaluation of 6 proprietary LMMs as well as 17 open-source LMMs highlights the substantial challenges and demands posed by OBI-Bench. Even the latest versions of GPT-4o, Gemini 1.5 Pro, and Qwen-VL-Max are still far from public-level humans in some fine-grained perception tasks. However, they perform at a level comparable to untrained humans in deciphering tasks, indicating remarkable capabilities in offering new interpretative perspectives and generating creative guesses. We hope OBI-Bench can facilitate the community to develop domain-specific multi-modal foundation models towards ancient language research and delve deeper to discover and enhance these untapped potentials of LMMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Just Noticeable Difference for Large Multimodal Models

    cs.CV 2025-07 conditional novelty 7.0 of 10

    Large multimodal models have measurable just-noticeable-difference thresholds, and most lag far behind humans in seeing small image changes.

  2. Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A VLM-based framework with retrieval-augmented generation parses tombstone photos into structured semantic graphs, reaching 89.5 Smatch F1 versus 36.1 for the prior OCR pipeline.

Pith tools