Pith. sign in

REVIEW 2 major objections 33 references

Vision-language models know calligraphy terms yet still fail to ground style in real brushwork.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 21:20 UTC pith:TAUOSU4L

load-bearing objection Solid first resource that cleanly separates Tie/Bei and ships hierarchical expert text; the “knowledgeable but unperceptive” claim is useful but partly rests on unvalidated aesthetic labels. the 2 major comments →

arxiv 2607.04147 v1 pith:TAUOSU4L submitted 2026-07-05 cs.CV cs.AI

HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

classification cs.CV cs.AI
keywords calligraphy stylesfine-grained perceptionlarge vision-language modelshistorical manuscriptsstyle discriminationaesthetic reasoningcultural heritagemodal mixture
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that fine-grained historical calligraphy style understanding is blocked less by model size than by bad data: prior archives mix ink manuscripts with stone rubbings and only attach flat labels. HCSU supplies 39,307 curated character images from 49 major calligraphers across ten dynasties, deliberately separated into ink (Tie), rubbing (Bei), and raw (Wild) domains, plus hierarchical expert aesthetic descriptions of ink, stroke, and structure. Two protocols follow: an 8-way style discrimination task and a short, expert-style description task scored by semantic similarity and terminology/richness judges. Leading multimodal models beat chance and can recite professional vocabulary, but they stay sensitive to script, text, and source shortcuts and often invent aesthetic claims without tying them to visible brushwork. The benchmark is offered as a public test of whether future systems can do expert-level visual reasoning needed for cultural heritage preservation.

Core claim

State-of-the-art large vision-language models reach non-trivial accuracy on controlled 8-way calligraphy style discrimination and can emit terminology-heavy descriptions, yet they remain sensitive to script-level, textual, and medium-specific cues and frequently fail to ground aesthetic judgments in fine-grained brushwork evidence. The paper calls this the knowledgeable-but-unperceptive gap, sharpened by the finding that high-contrast stone rubbings are often easier for current models than authentic ink manuscripts that preserve continuous texture.

What carries the argument

HCSU: a media-separated calligraphy benchmark (Tie ink, Bei rubbings, Wild raw) of 39,307 images with hierarchical expert aesthetic descriptions that support both 8-way style discrimination and interpretable aesthetic-reasoning evaluation.

Load-bearing premise

The hierarchical expert aesthetic descriptions and the fixed 8-way, 2-shot, 90-images-per-class protocol are taken as fair, visually grounded measures of genuine style perception rather than protocol or annotation artifacts.

What would settle it

Train or prompt a model using only character identity and coarse script or dynasty labels, with no hierarchical brushwork descriptions and no media-separated ink supervision; if it matches or exceeds the reported proprietary accuracies on both Tie and Bei and matches expert descriptions under the same judge metrics, the claimed perception gap is not distinctive to HCSU-style evaluation.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Cultural-heritage vision benchmarks must keep ink manuscripts and stone rubbings as separate evaluation domains instead of mixing them.
  • Style tasks should score grounding in stroke, ink, and structure attributes, not only end-to-end calligrapher labels.
  • Vision encoders need to capture continuous ink texture and pressure cues, not only high-contrast geometric outlines.
  • Expert natural-language aesthetic critiques become a first-class measurable target alongside classification accuracy.
  • Scaling open multimodal models alone will not close the gap without domain-aligned visual pretraining and reasoning objectives.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same media-mixing and flat-label pattern likely weakens other historical-art recognition settings that treat seals, illuminations, or manuscript hands as pure OCR problems.
  • The Bei-over-Tie advantage is a portable diagnostic: any fine-grained art model that prefers high-contrast silhouettes over continuous texture is still geometry-biased.
  • Closing the gap may require explicit stroke- and ink-aware training signals rather than more general instruction data alone.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The paper introduces HCSU, a 39,307-image benchmark for fine-grained historical Chinese calligraphy style understanding spanning 49 calligraphers, 10 dynasties, and five scripts. It separates processed ink manuscripts (Tie), stone rubbings (Bei), and unconstrained Wild images to address modal mixture, and supplies hierarchical expert aesthetic annotations (ink/stroke/structure/charm/layout) beyond flattened labels. Two protocols are defined: 8-way style discrimination with reference images and descriptions, and 2-shot aesthetic description generation scored by BERTScore and LLM-as-a-Judge (Terminology/Richness). Evaluations of proprietary and open LVLMs report non-trivial but modest discrimination accuracy (~19–38% vs 12.5% chance), domain sensitivities (often Bei>Tie; Wild>processed Origin), and a mismatch between discrimination and generation, supporting a “knowledgeable but unperceptive” claim.

Significance. If the construction and protocols hold, HCSU is a useful community resource: it is the first large calligraphy benchmark that systematically decouples Tie from Bei, preserves ink-color appearance via a documented pipeline, and pairs images with multi-tier expert critiques for both discrimination and interpretable reasoning. Public release (with documented access policy for Wild) and multi-model baselines make the contribution actionable for cultural-heritage CV and LVLM fine-grained evaluation. The dual-task design and ablations (preprocessing; description source) are strengths even when they complicate a simple accuracy narrative.

major comments (2)
  1. §3.4 and §4.1 treat hierarchical expert descriptions (normalized mainly from Shu Lin Zao Jian) as visually grounded ground truth for both 8-way discrimination and LLM-as-a-Judge Term./Rich. scoring, but the manuscript reports only three-stage QC (catalogue check, terminology normalization, manual removal of generic/inconsistent text) with no inter-annotator agreement, no blind re-annotation of held-out crops, and no human study linking description phrases to local brushwork evidence. Table 5 already shows Expert text is not a universal accuracy booster (GPT-generated can match or beat Expert on Tie), which is consistent with art-historical priors that may be only weakly locked to the specific 256×256 instance. Without quantitative annotation fidelity, the central “fail to ground aesthetic judgments in fine-grained brushwork evidence” claim (§5–6, Abstract) is only partially supported and
  2. §3.2 calibrates K=8, 2-shot, and 90 images/class using a single probe (Qwen3-VL-235B) and reports no error bars, confidence intervals, or significance tests on Tables 2–5. Discrimination accuracies sit in a narrow band above chance; without multi-seed sampling of distractors/candidates or statistical tests, claims of model ranking, Bei–Tie gaps, and Wild>Origin robustness (§5.3–5.4) remain fragile. At minimum, report variability over candidate sets and clarify that protocol hyperparameters were probe-chosen rather than independently fixed.

Circularity Check

0 steps flagged

No significant circularity: empirical dataset/benchmark paper whose claims are measured outcomes, not derivations that reduce to fitted inputs.

full rationale

HCSU is a dataset-and-benchmark paper, not a first-principles derivation. Its load-bearing claims (non-trivial but limited LVLM accuracy on 8-way style discrimination; terminology/richness gaps on description generation; Bei–Tie and Wild–Origin sensitivities; the “knowledgeable but unperceptive” characterization) are empirical measurements of external models on a fixed protocol, not quantities obtained by fitting parameters and then re-predicting the same quantities. Protocol choices (K=8, 2-shot, 90 images per class) were calibrated with a probe model (Qwen3-VL-235B) for difficulty and efficiency (§3.2, Fig. 4); that is ordinary benchmark design, not a fitted input renamed as a prediction of the main results. Expert aesthetic descriptions are external annotations (normalized from Shu Lin Zao Jian plus three-stage QC) used as evaluation targets and candidate text; the paper itself reports that they are not universal accuracy boosters (Table 5) and does not treat model outputs as defining the labels. There is no self-definitional loop, no uniqueness theorem imported from the authors, no ansatz smuggled via self-citation, and no renaming of a known result as a new derivation. The paper is self-contained against its own external-model evaluations; circularity score is therefore 0.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

As a dataset-and-benchmark paper the load-bearing commitments are curation choices, processing assumptions, and evaluation-protocol decisions rather than free physical constants or new particles. The free parameters are the discrete design knobs fixed by probe experiments; the axioms are standard CV practices plus domain claims about what counts as style; the invented entities are the dataset itself and its annotation schema.

free parameters (4)
  • K (number of candidates in discrimination) = 8
    Chosen as 8 after probe-model sweep K=2..10 on accuracy vs random baseline; directly sets task difficulty and reported accuracies.
  • few-shot count for description generation = 2
    Fixed at 2 after 1/2/3/4/8-shot calibration on Terminology, Richness, BERTScore; largest gain observed from 1 to 2.
  • images per category in final benchmark = 90
    Set to 90 after scale-vs-accuracy curves on three backbones; controls evaluation set size.
  • output resolution after geometric normalization = 256×256
    All processed images forced to 256×256 via aspect-preserving pad + Lanczos; affects what visual detail models receive.
axioms (4)
  • domain assumption Expert hierarchical descriptions (ink_style, stroke_style, character structure, charm, layout) derived from Shu Lin Zao Jian and normalized by annotators constitute valid ground-truth aesthetic targets.
    Invoked throughout §3.4 and used as both discrimination candidates and generation references; quality-control stages are described but inter-annotator agreement numbers are not reported.
  • domain assumption Canonicalization (polarity inversion for Bei, Otsu, CCA denoising, white-canvas isolation) removes non-style confounders while preserving the style cues needed for fair evaluation.
    Core of the processing pipeline in §3.3; the paper itself notes that normalization can discard high-resolution texture and therefore also releases a Wild domain.
  • ad hoc to paper Selecting only 49 historically prominent calligraphers with ≥500 images yields a representative and sufficiently difficult style-discrimination problem.
    Curation filter in §3.2 that reduces 310 authors / 14 dynasties to the final 49 / 10; acknowledged as a coverage limitation.
  • standard math Standard multimodal evaluation practices (temperature 0.1, persona prompt as calligraphy expert, BERTScore + LLM-as-a-Judge with DeepSeek-chat) produce reliable comparative rankings.
    §4 experimental methodology; common in LVLM papers but inherits judge-model biases.
invented entities (2)
  • HCSU dataset (Tie / Bei / Wild domains + hierarchical expert annotations) independent evidence
    purpose: Provide the first large-scale, modal-decoupled resource for fine-grained calligraphy style discrimination and interpretable aesthetic reasoning.
    Central contribution; independent existence is the public Hugging Face release itself.
  • Three-tier style attribute schema (ink_style, stroke_style, character structure + overall charm/layout) no independent evidence
    purpose: Decompose abstract authorial signature into concrete perceptual attributes that support both classification and natural-language critique.
    Introduced in §3.4 and Fig. 2; grounded in traditional art-historical terminology but newly formalized as machine-readable fields.

pith-pipeline@v1.1.0-grok45 · 20189 in / 3415 out tokens · 31841 ms · 2026-07-11T21:20:28.866881+00:00 · methodology

0 comments
read the original abstract

Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels. To bridge this gap, we introduce HCSU, the first comprehensive dataset tailored for fine-grained Historical Calligraphy Style Understanding. HCSU comprises 39,307 meticulously curated character images from 49 historically prominent calligraphers across 10 dynasties, systematically decoupling authentic ink manuscripts (Tie) from stone rubbings (Bei) to resolve the long-standing modal mixture problem. Moving beyond conventional flattened labels, HCSU provides hierarchical expert-written aesthetic descriptions, enabling two rigorous evaluation protocols: fine-grained style discrimination and interpretable aesthetic reasoning. Extensive evaluations reveal a persistent gap between calligraphy-related knowledge and visually grounded style perception: state-of-the-art LVLMs show non-trivial performance but remain sensitive to script-level, textual, and source-specific cues, and often struggle to ground aesthetic judgments in fine-grained brushwork evidence. Ultimately, the HCSU benchmark exposes fundamental limitations in current multimodal architectures, aiming to inspire the evolution of expert-level visual reasoning for cultural heritage preservation. The dataset is available at https://huggingface.co/datasets/Tongji209/HCSU.

Figures

Figures reproduced from arXiv: 2607.04147 by Chen Ye, Yan Liu, Yinsheng Yao.

Figure 1
Figure 1. Figure 1: The challenge of disentangling style from content. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the HCSU Dataset. HCSU spans 10 historical dynasties and five major calligraphic scripts, forming a vast matrix of artistic styles. Beyond images, it provides expert-level annotations that decompose calligraphic expression into three technical dimensions—ink style, stroke style, and character structure—supplemented by overall charm and layout. including the CASIA-HWDB series [13,14] and large d… view at source ↗
Figure 3
Figure 3. Figure 3: Comprehensive statistics of the HCSU dataset. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Data-driven benchmark calibration. (a) Accuracy vs. images per class for three backbones. (b) Probe performance on the K-way task (K = 8 optimal). (c) Effect of shots on description quality; 2-shot offers the best efficiency–performance balance. 3.3 Data Processing To improve structural consistency across heterogeneous calligraphy artifacts— ink manuscripts (Tie) and stone rubbings (Bei)—we design a unifie… view at source ↗
Figure 5
Figure 5. Figure 5: Unified dataset construction pipeline. Top: ink manuscript (Tie). Bottom: stone rubbing (Bei). The Bei sample is polarity-inverted during canonicalization to match the dark-on-light format before geometric normalization. ground, we isolate the authentic stroke colors and textures onto a uniform white canvas, successfully eliminating complex background interference without sacri￾ficing stylistic fidelity. 3… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 11 linked inside Pith

  1. [1]

    Anthropic: Introducing claude 4.5 sonnet (2025),https://www.anthropic.com/, accessed: 2026-06-28

  2. [2]

    arXiv preprint arXiv:2308.12966 (2023)

    Bai, J., et al.: Qwen-vl: A versatile vision-language model for understanding, lo- calization, text reading, and beyond. arXiv preprint arXiv:2308.12966 (2023)

  3. [3]

    arXiv preprint arXiv:2502.13923 (2025)

    Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., et al.: Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923 (2025)

  4. [4]

    arXiv preprint arXiv:2511.21631 (2025)

    Bai, S., et al.: Qwen3-vl technical report. arXiv preprint arXiv:2511.21631 (2025)

  5. [5]

    Chen,K.,etal.:Cogvlm:Visualexpertforpretrainedlanguagemodels.In:NeurIPS (2024)

  6. [6]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., et al.: Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 24185–24198 (2024)

  7. [7]

    In: Findings of the Association for Computational Linguistics: ACL 2024 (2024)

    Geigle,G.,etal.:Africanoreuropeanswallow?benchmarkinglargevision-language models for fine-grained object classification. In: Findings of the Association for Computational Linguistics: ACL 2024 (2024)

  8. [8]

    arXiv preprint arXiv:2507.06261 (2025)

    Google DeepMind: Gemini 2.5: Pushing the frontier with advanced reasoning, mul- timodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025)

  9. [9]

    arXiv preprint arXiv:2401.12467 (2024)

    Guan, H., Wan, J., Yang, H., Zhu, X., Jin, L., et al.: An open dataset for the evolution of oracle bone characters: Evobc. arXiv preprint arXiv:2401.12467 (2024)

  10. [10]

    5-vl technical report

    Guo, D., Wu, F., Zhu, F., Leng, F., Shi, G., Chen, H., Fan, H., Wang, J., Jiang, J., Wang, J., et al.: Seed1. 5-vl technical report. arXiv preprint arXiv:2505.07062 (2025)

  11. [11]

    In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

    Kim, J., Ji, H.: Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. pp. 6187–6207 (2024)

  12. [12]

    In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Li, Y., Xu, Z., Chen, S., Huang, H., Li, Y., Ma, S., Jiang, Y., Li, Z., Zhou, Q., Zheng, H.T., Shen, Y.: Towards real-world writing assistance: A chinese character checking benchmark with faked and misspelled characters. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 8656–8668 (2024)

  13. [13]

    In: 2009 10th International Conference on Docu- ment Analysis and Recognition

    Liu, C.L., Yin, F., Wang, D.H., Wang, Q.F.: Casia-olhwdb1: A database of online handwritten chinese characters. In: 2009 10th International Conference on Docu- ment Analysis and Recognition. pp. 1206–1210. IEEE (2009)

  14. [14]

    Liu, C.L., Yin, F., Wang, Q.F., Wang, D.H.: Online and offline handwritten chinese characterrecognition:Benchmarkingonnewdatabases.PatternRecognition46(1), 155–162 (2013)

  15. [15]

    Advances in neural information processing systems36, 34892–34916 (2023)

    Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural information processing systems36, 34892–34916 (2023)

  16. [16]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Luo, Y., Tang, J., Huang, C., Hao, F., Lian, Z.: Callireader: contextualizing chinese calligraphy via an embedding-aligned vision-language model. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 23030–23040 (2025)

  17. [17]

    Wenwu chubanshe, Beijing (1984), original work published 1935

    Ma, Z.: Shu lin zao jian: shu lin ji shi (in Chinese). Wenwu chubanshe, Beijing (1984), original work published 1935

  18. [18]

    arXiv preprint arXiv:2410.21276 (2024)

    OpenAI: Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)

  19. [19]

    OpenAI Blog (2025),https://openai.com/index/ introducing-gpt-5-2/, accessed: 2026-06-28 HCSU 17

    OpenAI: Introducing gpt-5.2. OpenAI Blog (2025),https://openai.com/index/ introducing-gpt-5-2/, accessed: 2026-06-28 HCSU 17

  20. [20]

    arXiv preprint arXiv:2512.10384 (2025)

    Pang, C., Yu, H., Chen, Z., Lu, L., Lou, X.: Towards fine-grained recognition with large visual language models: Benchmark and optimization strategies. arXiv preprint arXiv:2512.10384 (2025)

  21. [21]

    Scientific Data12(1), 169 (2025)

    Shi, Y., Peng, D., Zhang, Y., Cao, J., Jin, L.: A large-scale dataset for chinese historical document recognition and analysis. Scientific Data12(1), 169 (2025)

  22. [22]

    Wah, C., Branson, S., Welinder, P., Perona, P., Belongie, S.: The Caltech-UCSD Birds-200-2011 Dataset. Tech. Rep. CNS-TR-2011-001, California Institute of Technology (2011)

  23. [23]

    Wang, Q.F., Yin, F., Liu, C.L.: Handwritten chinese text recognition by integrating multiplecontexts.IEEETransactionsonPatternAnalysisandMachineIntelligence 34(8), 1469–1481 (2012)

  24. [24]

    IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

    Xu, P., et al.: Lvlm-ehub: A comprehensive evaluation benchmark for large vision- language models. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  25. [25]

    In: 2013 12th international conference on document analysis and recognition

    Yin, F., Wang, Q.F., Zhang, X.Y., Liu, C.L.: Icdar 2013 chinese handwriting recog- nition competition. In: 2013 12th international conference on document analysis and recognition. pp. 1464–1470. IEEE (2013)

  26. [26]

    Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Wang, G., Li, H., Zhu, J., Chen, J., et al.: Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652 (2024)

  27. [27]

    In: ICCVW (2025)

    Ypsilantis, et al.: Infusing fine-grained visual knowledge to vision-language models. In: ICCVW (2025)

  28. [28]

    arXiv preprint arXiv:2509.09731 (2025)

    Yu, H., Wu, Y., Shi, F., Liao, L., Lu, J., et al.: Benchmarking vision-language models on chinese ancient documents: From ocr to knowledge reasoning. arXiv preprint arXiv:2509.09731 (2025)

  29. [29]

    arXiv preprint arXiv:2504.14988 (2025)

    Yu, H.T., Peng, Y., Belongie, S., Wei, X.S.: Benchmarking large vision-language models on fine-grained image tasks: A comprehensive evaluation. arXiv preprint arXiv:2504.14988 (2025)

  30. [30]

    Zhang, W., Ma, H., Liu, L., Lu, Y., Suen, C.Y.: Callinet: a triplet network for chinesecalligraphystyleclassification.InternationalJournalonDocumentAnalysis and Recognition (IJDAR) (2025)

  31. [31]

    arXiv preprint arXiv:2405.19363 (2024)

    Zhang, Y., et al.: Vision LLMs are bad at hierarchical visual understanding, and LLMs are the bottleneck. arXiv preprint arXiv:2405.19363 (2024)

  32. [32]

    Pattern Recognition p

    Zhang, Y., Shi, Y., Zhang, P., Zhao, Y., Yang, Z., Jin, L.: Megahan97k: A large- scale dataset for mega-category chinese character recognition with over 97k cate- gories. Pattern Recognition p. 111757 (2025)

  33. [33]

    In: International Conference on Document Analysis and Recognition

    Zhao, Y., Zhang, Y., Jin, L.: Mccd: A multi-attribute chinese calligraphy character dataset annotated with script styles, dynasties, and calligraphers. In: International Conference on Document Analysis and Recognition. pp. 538–555. Springer (2025)