REVIEW 3 major objections 4 minor 36 references
Signs of the Past, Patterns of the Present: On the Automatic Classification of Old Babylonian Cuneiform Signs
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A standard image classifier trained on depth renderings of Old Babylonian clay tablets identifies cuneiform signs with 87.1% top-1 accuracy after fine-tuning, and training across three proveniences lifts out-of-distribution accuracy to…
desk verdict First Old Babylonian cuneiform sign classification benchmark with genuinely useful imaging and transfer experiments, but the headline accuracy is crop-level and likely inflated by a per-crop split that leaks tablet-level cues. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a ResNet50 image classifier trained on square crops of individual signs cut from 2D+ visualizations, where the crucial input type is the SketchB visualization: a rotation-invariant sketch derived from the surface normal map. The argument is carried by controlled training-set ablations over visualizations (twelve render types), proveniences (single versus combined), and a fine-tuning stage that removes augmentations and lowers the learning rate to match the test distribution. A secondary mechanism is the TSNE projection of the 2048-dimensional penultimate-layer features, used to show that clusters correspond to script variants documented in sign lists.
What would settle it
Take a random set of crops from the test data, have two trained Assyriologists independently assign sign classes without seeing the corpus labels, and compare each human's agreement with the model's accuracy against the same reference labels. If human-human agreement is no higher than human-model agreement, the $87.1\%$ figure is partly an artifact of ambiguous labels rather than of visual recognition.
Extended reading notes
Core claim
The authors claim that a standard convolutional network can classify Old Babylonian cuneiform signs cut from dome-captured visualizations, provided the rendering exposes depth: non-photorealistic sketches and normal maps beat any single lighting angle by 4–5 percentage points. The model reaches $87.1\% \pm 0.3$ top-1 on Nippur after fine-tuning, and the base model ranges $77.2\%$ to $85.6\%$ across Dūr-Abiešuḫ, Nippur, and Sippar. Training on all three proveniences raises accuracy on the held-out Marad site to $93.4\%$ of in-distribution performance, whereas single-provenience training drops to $51.4\%$–$79.0\%$ of in-distribution levels. Qualitative TSNE analysis shows that the model's feature space mirrors palaeographic variants that Assyriologists recognize, including the known ANSZE/GIRI3 overlap. The paper concludes that this is a viable assistive tool for reading tablets, with top-5 performance strong enough for a suggestion tool, while top-1 remains limited without context.
Load-bearing premise
The reported accuracies assume that each cropped sign has exactly one correct class, namely the contextual reading assigned by a human annotator in the Cuneur software; if those readings are inconsistent or the Unicode-based sign list splits or merges signs differently from actual palaeography, every accuracy number inherits that bias.
Editorial extensions
If this is right
- A top-5 accuracy of $96.5\%$ means a human-in-the-loop tool can offer five likely signs per crop, which is already useful for reading unfamiliar proveniences.
- Including multiple proveniences in training is the single most effective step for generalizing to unseen sites, raising out-of-distribution accuracy to $93.4\%$ of in-distribution performance.
- Fine-tuning the base model on the target provenience's own data improves results for every city tested, making specialization a cheap final step after broad training.
- Acquisition standards should prioritize depth-revealing visualizations (normal maps, sketches) and variety of provenience over raw image quantity.
- Signs on the curved left and right edges of tablets are the hardest to classify under directional light, so even illumination or depth-based renderings are preferable for those zones.
Reading between the lines
- The authors do not test legacy flatbed scans or photographs; a direct extension would be to measure how much accuracy drops on such images and then apply domain adaptation from SketchB renderings to legacy data.
- The embedding space that TSNE visualizes could be turned into a quantitative palaeographic instrument, for example dating undated tablets by nearest-neighbor positions among dated exemplars—something the paper only gestures at.
- The out-of-distribution claim rests on a single unseen site, Marad; a stronger stress test would hold out an entire region or a chronologically distinct corpus, which could lower the $93.4\%$ figure.
- Because labels are contextual readings, the accuracy ceiling may be set by inter-annotator agreement; measuring that agreement would place the reported scores in perspective.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper trains ResNet50 models to classify individual cuneiform signs cropped from 2D+ visualizations of Old Babylonian tablets from three proveniences (Nippur, Dūr-Abiešuḫ, Sippar), evaluates a held-out Marad set for out-of-distribution performance, compares twelve visualization types, studies the effect of fine-tuning, and uses TSNE plots for qualitative palaeographic analysis. The headline results are a top-1 accuracy of 87.1% ± 0.3 and top-5 accuracy of 96.5% ± 0.1 on the Nippur test set after fine-tuning, with base-model accuracies of 77.2–85.6% top-1 across the three in-distribution proveniences and an OOD performance of 93.4% of the in-distribution performance when training on all three.
Significance. If the evaluation is trustworthy, this is a useful first benchmark for automated Old Babylonian sign classification. The paper's strengths include repeated runs with reported standard deviations, a systematic comparison of lighting angles and depth visualizations, an external OOD test set (Marad) completely withheld from training, and a publicly available code repository and suggestion tool. The results also speak directly to data acquisition standards for cuneiform tablets, which is a practical contribution to digital Assyriology.
major comments (3)
- [§3.4] The 80/20 split is performed per sign category at the level of cropped signs, not at the level of tablets or tablet sides. Because crops are square bounding boxes that deliberately include partial neighboring signs (§3.3), crops from the same tablet side can appear in both training and test partitions, sharing lighting, surface texture, scribal hand, and often overlapping context. A ResNet50 can exploit such tablet-specific cues, so the reported accuracies, especially the fine-tuned Nippur row in Table 1 (87.1% ± 0.3), may be optimistic. The standard deviations only reflect random re-splitting of crops, not the choice of which tablets or tablet sides are held out. Please report results with a tablet-level or tablet-side-level split (for example, leave-one-tablet-out or grouped split), or at minimum quantify the fraction of test crops that share a tablet side with training crops.
- [§3.2.2, §3.1.1] The class labels are contextual readings assigned by human annotators using the Nuolenna sign list. As the paper acknowledges, many signs have multiple readings and visually similar signs from different classes can overlap, as the ANSZE/GIRI3 confusion in §4.3.2 shows. This means the reported accuracy conflates visual classification with the annotators' context-based reading decisions. The paper should report inter-annotator agreement on a subset of signs, or at least explicitly discuss how label noise and class overlap affect the interpretation of the accuracy numbers. Without this, the absolute accuracies are difficult to interpret as purely palaeographic classification performance.
- [§4.3.3, Table 1] The fine-tuning procedure is not fully specified. The text states that the base model is 'fine-tuned using either all data or only the data of one specific provenience,' but it does not explicitly state that the held-out 20% test crops are excluded from the fine-tuning set. If the test partition is included in fine-tuning, the results in Table 1 would be invalid; if not, this should be stated clearly. Please clarify the exact data split used for each fine-tuning experiment, including whether the test set is the same 20% crop-level partition described in §3.4.
minor comments (4)
- [§2] In the related work section, 'Cobanaglu' should be 'Cobanoglu' to match reference [7].
- [§4.3.1] The phrase 'do to their combined state' should be 'due to their combined state.'
- [References] Reference [8] contains a typo in the title: 'Old Babylonian Peiod' should be 'Old Babylonian Period.'
- [§3.2.2] The data availability statement says 'The data and scripts published with this paper can be found on our Zenodo/Github page (link),' but the link is a placeholder. Please provide the actual DOI or URL for the dataset and scripts.
Circularity Check
No circularity: the accuracy and transfer claims are measured on held-out test crops and a fully withheld Marad set, not derived from fitted parameters or from self-citations.
full rationale
The paper's quantitative claims are supervised-learning evaluations. The label space is defined externally by the Nuolenna Unicode sign list (Section 3.2.2), and the per-class 80/20 split (Section 3.4) creates held-out crops; the Marad set is kept outside training at all times (Section 4.3.3). The reported top-1/top-5 numbers are therefore empirical measurements on examples not used for the reported fit, not quantities forced by construction. The self-citations (e.g., Hameeuw et al. [11] for visualization naming, Rattenborg et al. [22] for corpus spread, Willems et al. [32] for light-dome acquisition) are provenance and methodology references, not load-bearing arguments that the results are correct by definition; no uniqueness theorem or ansatz is imported from the authors' prior work. TSNE-based variant analysis is explicitly post hoc and the authors caution about TSNE artifacts (Section 3.4). The main nontrivial risk is the crop-level rather than tablet-level split (Section 3.4), which could let tablet-specific surface/lighting cues appear in both train and test and inflate in-distribution numbers; this is an evaluation-design concern, not circularity. Likewise, the placeholder '(link)' for the promised Zenodo/Github resources and the deferral of annotation-data evaluation to 'Smidt forthcoming' are completeness issues, not circular steps.
Assumptions & free parameters
assumptions (3)
- domain assumption The Nuolenna Unicode sign list provides a consistent and correct mapping from visual signs to class labels.
- domain assumption Random rotation by 0-360 degrees is a valid augmentation because sign orientation carries no class information in the training distribution.
- domain assumption Each cropped sign is an independent sample despite multiple crops coming from the same tablet or same annotator.
Cite this review
Pith. "Pith review of Signs of the Past, Patterns of the Present: On the Automatic Classification of Old Babylonian Cuneiform Signs." pith.science (2026). https://pith.science/paper/NRMMENJZ
@misc{pith2026250713959,
author = {Pith},
title = {Pith review of: Signs of the Past, Patterns of the Present: On the Automatic Classification of Old Babylonian Cuneiform Signs},
year = {2026},
howpublished = {\url{https://pith.science/paper/NRMMENJZ}},
note = {Machine review of arXiv:2507.13959}
}
read the original abstract
The work in this paper describes the training and evaluation of machine learning (ML) techniques for the classification of cuneiform signs. There is a lot of variability in cuneiform signs, depending on where they come from, for what and by whom they were written, but also how they were digitized. This variability makes it unlikely that an ML model trained on one dataset will perform successfully on another dataset. This contribution studies how such differences impact that performance. Based on our results and insights, we aim to influence future data acquisition standards and provide a solid foundation for future cuneiform sign classification tasks. The ML model has been trained and tested on handwritten Old Babylonian (c. 2000-1600 B.C.E.) documentary texts inscribed on clay tablets originating from three Mesopotamian cities (Nippur, D\=ur-Abie\v{s}uh and Sippar). The presented and analysed model is ResNet50, which achieves a top-1 score of 87.1% and a top-5 score of 96.5% for signs with at least 20 instances. As these automatic classification results are the first on Old Babylonian texts, there are currently no comparable results.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Sean E. Anderson and Marc Levoy. 2002. Unwrapping and visualizing cuneiform tablets. IEEE Computer Graphics and Applications 22, 6 (2002), 82–88
work page 2002
-
[2]
Marine Béranger. 2023. Dur-Abi-ešuh and the Abandonment of Nippur During the Late Old Babylonian Period: A Historical Survey. Journal of Cuneiform Studies 75 (2023), 27–47. doi:10.1086/725217
-
[3]
Bartosz Bogacz and Hubert Mara. 2022. Digital Assyriology—Advances in Visual Cuneiform Analysis. Journal on Computing and Cultural Heritage 15, 2, Article 38 (may 2022), 22 pages. doi:10.1145/3491239
-
[4]
Kai Brandenbusch, Eugen Rusakov, and Gernot A. Fink. 2021. Context aware generation of cuneiform signs. In Document Analysis and Recognition - International Conference on Document Analysis and Recognition , Josep Lladós, Daniel Lopresti, and Seiichi Uchida (Eds.). Springer International Publishing, Cham, Switzerland, 65–79
work page 2021
-
[5]
Dominique Charpin. 2010. Writing, Law, and Kingship in Old Babylonian Mesopotamia . University of Chicago Press, Chicago and London
work page 2010
-
[6]
Dominique Charpin, Dietz Otto Edzard, Marten Stol, Pascal Attinger, Walther Sallaberger, and Markus Wäfler. 2004. Mesopotamien - Die altbabylonische Zeit. Orbis Biblicus et Orientalis, Vol. 160. Academic Press and Vandenhoeck & Ruprecht, Fribourg and Göttingen
work page 2004
-
[7]
Yunus Cobanoglu, Luis Sáenz, Ilya Khait, and Enrique Jiménez. 2024. Sign detection for cuneiform tablets. it-Information Technology 66, 1 (2024), 28–38
work page 2024
-
[8]
Rients de Boer. 2013. Marad in the Early Old Babylonian Peiod: Its Kings, Chronology, and Isin’s Influence. Journal of Cuneiform Studies 65 (2013), 73–90. https://www.jstor.org/stable/10.5615/jcunestud.65.2013.0073
arXiv 2013
Show all 36 references
-
[9]
Tobias Dencker, Pablo Klinkisch, Stefan M Maul, and Björn Ommer. 2020. Deep learning of cuneiform sign detection with weak supervision using transliteration alignment. Plos one 15, 12 (2020), e0243039
2020
-
[10]
Ilya Gershevitch. 1979. The Alloglottography of Old Persian. Transactions of the Philological Society 77 (1979), 114–190
1979
-
[11]
Hendrik Hameeuw, Katrien De Graef, Gustav Ryberg Smidt, Anne Goddeeris, Timo Homburg, and Krishna Kumar Thirukokaranam Chandrasekar. 2024. Preparing multi-layered visualisations of Old Babylonian cuneiform tablets for a machine learning OCR training model towards automated sig...
2024 doi
-
[12]
Hendrik Hameeuw and Geert Willems. 2011. New visualization techniques for cuneiform texts and sealings. Akkadica 132, 2 (2011), 163–178
2011
-
[13]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE Computer Society, Los Alamitos, CA, USA, 770–778
2016
-
[14]
Timo Homburg, Robert Zwick, Hubert Mara, and Kai-Christian Bruhn. 2022. Annotated 3D-Models of Cuneiform Tablets. Journal of Open Archaeology Data (JOAD) 10, 92 (2022), 4 pages. doi:10.5334/joad.92
2022 doi
-
[15]
John Huehnergard. 2011. A Grammar of Akkadian (3rd ed.). Eisenbrauns, Winona Lake, Indiana
2011
-
[16]
Tommi Jauhiainen, Heidi Jauhiainen, Tero Alstola, and Krister Lindén. 2019. Language and Dialect Identification of Cuneiform Texts. In Proceedings of the Sixth Workshop on NLP for Similar Languages, Varieties and Dialects , Marcos Zampieri, Preslav Nakov, Shervin Malmasi, Niko...
2019 doi
-
[17]
René Labat. 1976. Manuel D’Épigraphie Akkadienne (5th ed.). Librairie Orientaliste Paul Geuthner, Paris
1976
-
[18]
Karel Van Lerberghe and Gabriella Voet. 2009. A Late Old Babylonian Temple Archive from Dūr-Abiešuḫ . Cornell University Studies in Assyriology and Sumerology, Vol. 8. CDL Press, Bethesda, Maryland
2009
-
[19]
Piotr Michalowski. 2007. The Lives of the Sumerian Language (2nd print ed.). Oriental Institute Seminars, Vol. 2. The Oriental Institute of the University of Chicago, Chicago, Illinois, Chapter 10, 163–188
2007
-
[20]
Rachel Mikulinsky, Morris Alper, Shai Gordin, Enrique Jiménez, Yoram Cohen, and Hadar Averbuch-Elor. 2025. ProtoSnap: Prototype Alignment for Cuneiform Signs. arXiv: 2502.00129 [cs.CV] https://arxiv.org/abs/2502.00129
2025 arXiv
-
[21]
Joseph Nockels, Paul Gooding, and Melissa Terras. 2024. The implications of handwritten text recognition for accessing the past at scale. Journal of Documentation 80, 7 (2024), 148–167
2024
-
[22]
Rune Rattenborg, Gustav Ryberg Smidt, Carolin Johansson, Nils Melin-Kronsell, and Seraina Nett. 2023. The Archaeological Distribution of the Cuneiform Corpus. A Provisional Quantitative and Geospatial Survey. Altorientalische Forschungen 50, 2 (2023), 178–205. doi:doi:10.1515/...
2023 doi
-
[23]
Christopher Rest, Denis Fisseler, Frank Weichert, Turna Somel, and Gerfrid GW Müller. 2022. Illumination-based augmentation for cuneiform deep neural sign classification. Journal on Computing and Cultural Heritage (JOCCH) 15, 3 (2022), 1–20
2022
-
[24]
Christian Reul, Dennis Christ, Alexander Hartelt, Nico Balbach, Maximilian Wehner, Uwe Springmann, Christoph Wick, Christine Grundig, Andreas Büttner, and Frank Puppe. 2019. OCR4all—An open-source tool providing a (semi-) automatic OCR workflow for historical printings. Applie...
2019
-
[25]
Eugen Rusakov, Turna Somel, Gernot A Fink, and Gerfrid GW Müller. 2020. Towards query-by-eXpression retrieval of cuneiform signs. In 2020 17th International Conference on Frontiers in Handwriting Recognition (ICFHR) . IEEE, Dortmund, 43–48. Sign of the Past, Patterns of the Pr...
2020
-
[26]
Walther Sallaberger. 2004. Das Ende des Sumerischen: Tod und Nachleben einer altmesopotamischen Sprache . Münchner Forschungen zur historischen Sprachwissenschaft, Vol. 2. Hempen Verlag, Bremen, 108–140
2004
-
[27]
Thea Sommerschield, Yannis Assael, John Pavlopoulos, Vanessa Stefanak, Andrew Senior, Chris Dyer, John Bodel, Jonathan Prag, Ion Androutsopoulos, and Nando De Freitas. 2023. Machine learning for ancient languages: A survey. Computational Linguistics 49, 3 (2023), 703–747
2023
-
[28]
Ernst Stötzner, Timo Homburg, and Hubert Mara. 2023. CNN based Cuneiform Sign Detection Learned from Annotated 3D Renderings and Mapped Photographs with Illumination Augmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . IEEE Computer Societ...
2023
-
[29]
Michael P. Streck. 2010. Großes Fach Altorientalistik: Der Umfang des keilschriftlichen Textkorpus. Mitteilungen der Deutschen Orient-Gesellschaft zu Berlin 142 (2010), 35–58
2010
-
[30]
Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008), 2579–2605
2008
-
[31]
Martin Wattenberg, Fernanda Viégas, and Ian Johnson. 2016. How to use t-SNE effectively. Distill 1, 10 (2016), e2
2016
-
[32]
Geert Willems, Frank Verbiest, Wim Moreau, Hendrik Hameeuw, Karel Van Lerberghe, and Luc Van Gool. 2005. Easy and cost-effective cuneiform digitizing. In The 6th International Symposium on Virtual Reality, Archaeology and Cultural Heritage (V AST 2005). Eurographics Associatio...
2005
-
[33]
Williams, Grace Su, Sandra R
Edward C. Williams, Grace Su, Sandra R. Schloen, Miller C. Prosser, Susanne Paulus, and Sanjay Krishnan. 2025. DeepScribe: Localization and Classification of Elamite Cuneiform Signs Via Deep Learning. J. Comput. Cult. Herit. (mar 2025), 31 pages. doi:10.1145/3716850 Just Accepted
2025 doi
-
[34]
Christopher Woods. 2007. Bilingualism, Scribal Learning, and the Death of Sumerian (2nd print ed.). Oriental Institute Seminars, Vol. 2. The Oriental Institute of the University of Chicago, Chicago, Illinois, Chapter 6, 95–124
2007
-
[35]
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE Computer Society, Los Alamitos, CA, USA, 1492–1500
2017
-
[36]
Vasiliy Yugay, Kartik Paliwal, Yunus Cobanoglu, Luis Sáenz, Ekaterine Gogokhia, Shai Gordin, and Enrique Jiménez. 2024. Stylistic classification of cuneiform signs using convolutional neural networks. IT-Information Technology 66, 1 (2024), 15–27. A Cune-AI-form Tool As part o...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.