REVIEW 2 major objections 33 references
Vision-language models know calligraphy terms yet still fail to ground style in real brushwork.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 21:20 UTC pith:TAUOSU4L
load-bearing objection Solid first resource that cleanly separates Tie/Bei and ships hierarchical expert text; the “knowledgeable but unperceptive” claim is useful but partly rests on unvalidated aesthetic labels. the 2 major comments →
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
State-of-the-art large vision-language models reach non-trivial accuracy on controlled 8-way calligraphy style discrimination and can emit terminology-heavy descriptions, yet they remain sensitive to script-level, textual, and medium-specific cues and frequently fail to ground aesthetic judgments in fine-grained brushwork evidence. The paper calls this the knowledgeable-but-unperceptive gap, sharpened by the finding that high-contrast stone rubbings are often easier for current models than authentic ink manuscripts that preserve continuous texture.
What carries the argument
HCSU: a media-separated calligraphy benchmark (Tie ink, Bei rubbings, Wild raw) of 39,307 images with hierarchical expert aesthetic descriptions that support both 8-way style discrimination and interpretable aesthetic-reasoning evaluation.
Load-bearing premise
The hierarchical expert aesthetic descriptions and the fixed 8-way, 2-shot, 90-images-per-class protocol are taken as fair, visually grounded measures of genuine style perception rather than protocol or annotation artifacts.
What would settle it
Train or prompt a model using only character identity and coarse script or dynasty labels, with no hierarchical brushwork descriptions and no media-separated ink supervision; if it matches or exceeds the reported proprietary accuracies on both Tie and Bei and matches expert descriptions under the same judge metrics, the claimed perception gap is not distinctive to HCSU-style evaluation.
If this is right
- Cultural-heritage vision benchmarks must keep ink manuscripts and stone rubbings as separate evaluation domains instead of mixing them.
- Style tasks should score grounding in stroke, ink, and structure attributes, not only end-to-end calligrapher labels.
- Vision encoders need to capture continuous ink texture and pressure cues, not only high-contrast geometric outlines.
- Expert natural-language aesthetic critiques become a first-class measurable target alongside classification accuracy.
- Scaling open multimodal models alone will not close the gap without domain-aligned visual pretraining and reasoning objectives.
Where Pith is reading between the lines
- The same media-mixing and flat-label pattern likely weakens other historical-art recognition settings that treat seals, illuminations, or manuscript hands as pure OCR problems.
- The Bei-over-Tie advantage is a portable diagnostic: any fine-grained art model that prefers high-contrast silhouettes over continuous texture is still geometry-biased.
- Closing the gap may require explicit stroke- and ink-aware training signals rather than more general instruction data alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HCSU, a 39,307-image benchmark for fine-grained historical Chinese calligraphy style understanding spanning 49 calligraphers, 10 dynasties, and five scripts. It separates processed ink manuscripts (Tie), stone rubbings (Bei), and unconstrained Wild images to address modal mixture, and supplies hierarchical expert aesthetic annotations (ink/stroke/structure/charm/layout) beyond flattened labels. Two protocols are defined: 8-way style discrimination with reference images and descriptions, and 2-shot aesthetic description generation scored by BERTScore and LLM-as-a-Judge (Terminology/Richness). Evaluations of proprietary and open LVLMs report non-trivial but modest discrimination accuracy (~19–38% vs 12.5% chance), domain sensitivities (often Bei>Tie; Wild>processed Origin), and a mismatch between discrimination and generation, supporting a “knowledgeable but unperceptive” claim.
Significance. If the construction and protocols hold, HCSU is a useful community resource: it is the first large calligraphy benchmark that systematically decouples Tie from Bei, preserves ink-color appearance via a documented pipeline, and pairs images with multi-tier expert critiques for both discrimination and interpretable reasoning. Public release (with documented access policy for Wild) and multi-model baselines make the contribution actionable for cultural-heritage CV and LVLM fine-grained evaluation. The dual-task design and ablations (preprocessing; description source) are strengths even when they complicate a simple accuracy narrative.
major comments (2)
- §3.4 and §4.1 treat hierarchical expert descriptions (normalized mainly from Shu Lin Zao Jian) as visually grounded ground truth for both 8-way discrimination and LLM-as-a-Judge Term./Rich. scoring, but the manuscript reports only three-stage QC (catalogue check, terminology normalization, manual removal of generic/inconsistent text) with no inter-annotator agreement, no blind re-annotation of held-out crops, and no human study linking description phrases to local brushwork evidence. Table 5 already shows Expert text is not a universal accuracy booster (GPT-generated can match or beat Expert on Tie), which is consistent with art-historical priors that may be only weakly locked to the specific 256×256 instance. Without quantitative annotation fidelity, the central “fail to ground aesthetic judgments in fine-grained brushwork evidence” claim (§5–6, Abstract) is only partially supported and
- §3.2 calibrates K=8, 2-shot, and 90 images/class using a single probe (Qwen3-VL-235B) and reports no error bars, confidence intervals, or significance tests on Tables 2–5. Discrimination accuracies sit in a narrow band above chance; without multi-seed sampling of distractors/candidates or statistical tests, claims of model ranking, Bei–Tie gaps, and Wild>Origin robustness (§5.3–5.4) remain fragile. At minimum, report variability over candidate sets and clarify that protocol hyperparameters were probe-chosen rather than independently fixed.
Circularity Check
No significant circularity: empirical dataset/benchmark paper whose claims are measured outcomes, not derivations that reduce to fitted inputs.
full rationale
HCSU is a dataset-and-benchmark paper, not a first-principles derivation. Its load-bearing claims (non-trivial but limited LVLM accuracy on 8-way style discrimination; terminology/richness gaps on description generation; Bei–Tie and Wild–Origin sensitivities; the “knowledgeable but unperceptive” characterization) are empirical measurements of external models on a fixed protocol, not quantities obtained by fitting parameters and then re-predicting the same quantities. Protocol choices (K=8, 2-shot, 90 images per class) were calibrated with a probe model (Qwen3-VL-235B) for difficulty and efficiency (§3.2, Fig. 4); that is ordinary benchmark design, not a fitted input renamed as a prediction of the main results. Expert aesthetic descriptions are external annotations (normalized from Shu Lin Zao Jian plus three-stage QC) used as evaluation targets and candidate text; the paper itself reports that they are not universal accuracy boosters (Table 5) and does not treat model outputs as defining the labels. There is no self-definitional loop, no uniqueness theorem imported from the authors, no ansatz smuggled via self-citation, and no renaming of a known result as a new derivation. The paper is self-contained against its own external-model evaluations; circularity score is therefore 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- K (number of candidates in discrimination) =
8
- few-shot count for description generation =
2
- images per category in final benchmark =
90
- output resolution after geometric normalization =
256×256
axioms (4)
- domain assumption Expert hierarchical descriptions (ink_style, stroke_style, character structure, charm, layout) derived from Shu Lin Zao Jian and normalized by annotators constitute valid ground-truth aesthetic targets.
- domain assumption Canonicalization (polarity inversion for Bei, Otsu, CCA denoising, white-canvas isolation) removes non-style confounders while preserving the style cues needed for fair evaluation.
- ad hoc to paper Selecting only 49 historically prominent calligraphers with ≥500 images yields a representative and sufficiently difficult style-discrimination problem.
- standard math Standard multimodal evaluation practices (temperature 0.1, persona prompt as calligraphy expert, BERTScore + LLM-as-a-Judge with DeepSeek-chat) produce reliable comparative rankings.
invented entities (2)
-
HCSU dataset (Tie / Bei / Wild domains + hierarchical expert annotations)
independent evidence
-
Three-tier style attribute schema (ink_style, stroke_style, character structure + overall charm/layout)
no independent evidence
read the original abstract
Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels. To bridge this gap, we introduce HCSU, the first comprehensive dataset tailored for fine-grained Historical Calligraphy Style Understanding. HCSU comprises 39,307 meticulously curated character images from 49 historically prominent calligraphers across 10 dynasties, systematically decoupling authentic ink manuscripts (Tie) from stone rubbings (Bei) to resolve the long-standing modal mixture problem. Moving beyond conventional flattened labels, HCSU provides hierarchical expert-written aesthetic descriptions, enabling two rigorous evaluation protocols: fine-grained style discrimination and interpretable aesthetic reasoning. Extensive evaluations reveal a persistent gap between calligraphy-related knowledge and visually grounded style perception: state-of-the-art LVLMs show non-trivial performance but remain sensitive to script-level, textual, and source-specific cues, and often struggle to ground aesthetic judgments in fine-grained brushwork evidence. Ultimately, the HCSU benchmark exposes fundamental limitations in current multimodal architectures, aiming to inspire the evolution of expert-level visual reasoning for cultural heritage preservation. The dataset is available at https://huggingface.co/datasets/Tongji209/HCSU.
Figures
Reference graph
Works this paper leans on
-
[1]
Anthropic: Introducing claude 4.5 sonnet (2025),https://www.anthropic.com/, accessed: 2026-06-28
2025
-
[2]
arXiv preprint arXiv:2308.12966 (2023)
Bai, J., et al.: Qwen-vl: A versatile vision-language model for understanding, lo- calization, text reading, and beyond. arXiv preprint arXiv:2308.12966 (2023)
Pith/arXiv arXiv 2023
-
[3]
arXiv preprint arXiv:2502.13923 (2025)
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., et al.: Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923 (2025)
Pith/arXiv arXiv 2025
-
[4]
arXiv preprint arXiv:2511.21631 (2025)
Bai, S., et al.: Qwen3-vl technical report. arXiv preprint arXiv:2511.21631 (2025)
Pith/arXiv arXiv 2025
-
[5]
Chen,K.,etal.:Cogvlm:Visualexpertforpretrainedlanguagemodels.In:NeurIPS (2024)
2024
-
[6]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., et al.: Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 24185–24198 (2024)
2024
-
[7]
In: Findings of the Association for Computational Linguistics: ACL 2024 (2024)
Geigle,G.,etal.:Africanoreuropeanswallow?benchmarkinglargevision-language models for fine-grained object classification. In: Findings of the Association for Computational Linguistics: ACL 2024 (2024)
2024
-
[8]
arXiv preprint arXiv:2507.06261 (2025)
Google DeepMind: Gemini 2.5: Pushing the frontier with advanced reasoning, mul- timodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025)
Pith/arXiv arXiv 2025
-
[9]
arXiv preprint arXiv:2401.12467 (2024)
Guan, H., Wan, J., Yang, H., Zhu, X., Jin, L., et al.: An open dataset for the evolution of oracle bone characters: Evobc. arXiv preprint arXiv:2401.12467 (2024)
Pith/arXiv arXiv 2024
-
[10]
Guo, D., Wu, F., Zhu, F., Leng, F., Shi, G., Chen, H., Fan, H., Wang, J., Jiang, J., Wang, J., et al.: Seed1. 5-vl technical report. arXiv preprint arXiv:2505.07062 (2025)
Pith/arXiv arXiv 2025
-
[11]
In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Kim, J., Ji, H.: Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. pp. 6187–6207 (2024)
2024
-
[12]
In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Li, Y., Xu, Z., Chen, S., Huang, H., Li, Y., Ma, S., Jiang, Y., Li, Z., Zhou, Q., Zheng, H.T., Shen, Y.: Towards real-world writing assistance: A chinese character checking benchmark with faked and misspelled characters. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 8656–8668 (2024)
2024
-
[13]
In: 2009 10th International Conference on Docu- ment Analysis and Recognition
Liu, C.L., Yin, F., Wang, D.H., Wang, Q.F.: Casia-olhwdb1: A database of online handwritten chinese characters. In: 2009 10th International Conference on Docu- ment Analysis and Recognition. pp. 1206–1210. IEEE (2009)
2009
-
[14]
Liu, C.L., Yin, F., Wang, Q.F., Wang, D.H.: Online and offline handwritten chinese characterrecognition:Benchmarkingonnewdatabases.PatternRecognition46(1), 155–162 (2013)
2013
-
[15]
Advances in neural information processing systems36, 34892–34916 (2023)
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural information processing systems36, 34892–34916 (2023)
2023
-
[16]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Luo, Y., Tang, J., Huang, C., Hao, F., Lian, Z.: Callireader: contextualizing chinese calligraphy via an embedding-aligned vision-language model. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 23030–23040 (2025)
2025
-
[17]
Wenwu chubanshe, Beijing (1984), original work published 1935
Ma, Z.: Shu lin zao jian: shu lin ji shi (in Chinese). Wenwu chubanshe, Beijing (1984), original work published 1935
1984
-
[18]
arXiv preprint arXiv:2410.21276 (2024)
OpenAI: Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)
Pith/arXiv arXiv 2024
-
[19]
OpenAI Blog (2025),https://openai.com/index/ introducing-gpt-5-2/, accessed: 2026-06-28 HCSU 17
OpenAI: Introducing gpt-5.2. OpenAI Blog (2025),https://openai.com/index/ introducing-gpt-5-2/, accessed: 2026-06-28 HCSU 17
2025
-
[20]
arXiv preprint arXiv:2512.10384 (2025)
Pang, C., Yu, H., Chen, Z., Lu, L., Lou, X.: Towards fine-grained recognition with large visual language models: Benchmark and optimization strategies. arXiv preprint arXiv:2512.10384 (2025)
arXiv 2025
-
[21]
Scientific Data12(1), 169 (2025)
Shi, Y., Peng, D., Zhang, Y., Cao, J., Jin, L.: A large-scale dataset for chinese historical document recognition and analysis. Scientific Data12(1), 169 (2025)
2025
-
[22]
Wah, C., Branson, S., Welinder, P., Perona, P., Belongie, S.: The Caltech-UCSD Birds-200-2011 Dataset. Tech. Rep. CNS-TR-2011-001, California Institute of Technology (2011)
2011
-
[23]
Wang, Q.F., Yin, F., Liu, C.L.: Handwritten chinese text recognition by integrating multiplecontexts.IEEETransactionsonPatternAnalysisandMachineIntelligence 34(8), 1469–1481 (2012)
2012
-
[24]
IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
Xu, P., et al.: Lvlm-ehub: A comprehensive evaluation benchmark for large vision- language models. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[25]
In: 2013 12th international conference on document analysis and recognition
Yin, F., Wang, Q.F., Zhang, X.Y., Liu, C.L.: Icdar 2013 chinese handwriting recog- nition competition. In: 2013 12th international conference on document analysis and recognition. pp. 1464–1470. IEEE (2013)
2013
-
[26]
Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Wang, G., Li, H., Zhu, J., Chen, J., et al.: Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652 (2024)
Pith/arXiv arXiv 2024
-
[27]
In: ICCVW (2025)
Ypsilantis, et al.: Infusing fine-grained visual knowledge to vision-language models. In: ICCVW (2025)
2025
-
[28]
arXiv preprint arXiv:2509.09731 (2025)
Yu, H., Wu, Y., Shi, F., Liao, L., Lu, J., et al.: Benchmarking vision-language models on chinese ancient documents: From ocr to knowledge reasoning. arXiv preprint arXiv:2509.09731 (2025)
Pith/arXiv arXiv 2025
-
[29]
arXiv preprint arXiv:2504.14988 (2025)
Yu, H.T., Peng, Y., Belongie, S., Wei, X.S.: Benchmarking large vision-language models on fine-grained image tasks: A comprehensive evaluation. arXiv preprint arXiv:2504.14988 (2025)
Pith/arXiv arXiv 2025
-
[30]
Zhang, W., Ma, H., Liu, L., Lu, Y., Suen, C.Y.: Callinet: a triplet network for chinesecalligraphystyleclassification.InternationalJournalonDocumentAnalysis and Recognition (IJDAR) (2025)
2025
-
[31]
arXiv preprint arXiv:2405.19363 (2024)
Zhang, Y., et al.: Vision LLMs are bad at hierarchical visual understanding, and LLMs are the bottleneck. arXiv preprint arXiv:2405.19363 (2024)
Pith/arXiv arXiv 2024
-
[32]
Pattern Recognition p
Zhang, Y., Shi, Y., Zhang, P., Zhao, Y., Yang, Z., Jin, L.: Megahan97k: A large- scale dataset for mega-category chinese character recognition with over 97k cate- gories. Pattern Recognition p. 111757 (2025)
2025
-
[33]
In: International Conference on Document Analysis and Recognition
Zhao, Y., Zhang, Y., Jin, L.: Mccd: A multi-attribute chinese calligraphy character dataset annotated with script styles, dynasties, and calligraphers. In: International Conference on Document Analysis and Recognition. pp. 538–555. Springer (2025)
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.