Pith. sign in

REVIEW 7 cited by

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2509.09731 v1 pith:J2ZPN4CE submitted 2025-09-10 cs.CL

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning

classification cs.CL
keywords chineseancientdocumentsvlmsancientdocknowledgedocumentlinguistic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while current Vision-Language Models (VLMs) struggle with their visual and linguistic complexity. Existing document benchmarks focus on English printed texts or simplified Chinese, leaving a gap for evaluating VLMs on ancient Chinese documents. To address this, we present AncientDoc, the first benchmark for Chinese ancient documents, designed to assess VLMs from OCR to knowledge reasoning. AncientDoc includes five tasks (page-level OCR, vernacular translation, reasoning-based QA, knowledge-based QA, linguistic variant QA) and covers 14 document types, over 100 books, and about 3,000 pages. Based on AncientDoc, we evaluate mainstream VLMs using multiple metrics, supplemented by a human-aligned large language model for scoring.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

    cs.CV 2026-07 conditional novelty 6.5

    HCSU supplies the first large decoupled Tie/Bei calligraphy dataset with expert aesthetic labels and finds SOTA LVLMs remain knowledgeable yet unperceptive on fine-grained style.

  2. Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis

    cs.CL 2026-06 unverdicted novelty 6.0

    New Sinhala OCR dataset from 1981-2019 legislative acts enables LightOnOCR-2-1B to reach 1.05% CER, beating Surya-OCR, Tesseract, and Google Document AI.

  3. Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting

    cs.CV 2026-05 unverdicted novelty 6.0

    SAME-Net adds a differentiable soft attention mask embedding module to achieve rectification-free end-to-end scene text spotting with 84.02% H-mean on Total-Text.

  4. Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

    cs.AI 2026-03 unverdicted novelty 6.0

    MLLMs exhibit a consistent recognition-reasoning inversion on discrete visual symbols across domains, underperforming on elementary perception while appearing competent on higher-level reasoning via linguistic compensation.

  5. Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction

    cs.CV 2026-04 unverdicted novelty 5.0

    A multilevel perceptual CRF model using Swin Transformer, HPF fusion, HA adapters, and dynamic scaling attention achieves state-of-the-art monocular depth estimation on NYU Depth v2, KITTI, and MatterPort3D with reduc...

  6. LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data

    cs.CV 2026-01 reject novelty 4.0

    A lightweight RGB-D cross-attention network is proposed for rail defect detection, but the SOTA accuracy and generalization claims are internally inconsistent and the implementation is not public.

  7. Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method

    cs.CV 2026-01 reject novelty 3.0

    Fine-tuning Stable Diffusion with DreamBooth-style knowledge and hypernetwork-guided crack control maps can synthesize substation meter defect images that boost a YOLOv8 defect detector's mAP when added to the training set.