REVIEW 3 major objections 6 minor 1 cited by
Advancing Offline Handwritten Text Recognition: A Systematic Review of Data Augmentation and Generation Techniques
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This systematic review of offline handwritten text recognition claims that GAN-based data generation has become the dominant augmentation strategy and maps the field's datasets, metrics, and remaining gaps.
desk verdict A useful qualitative map of HTR augmentation methods, but the PRISMA study corpus is inconsistent and undisclosed, so the headline statistics cannot be checked. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the PRISMA systematic-review protocol: a four-phase pipeline that moves from 1,302 identified records, through 848 unique records after duplicate removal, to 91 title/abstract-screened papers, and finally to 55 papers that pass a full-text review and an 11-item binary quality checklist (EC8 requires a score of at least 7). This protocol is what converts an unstructured literature into countable claims such as the 60.9% GAN share and the 19.4% IAM usage share; the taxonomy diagram and the proportion figures carry the survey's conclusions.
What would settle it
A reader can test the central claim by counting the entries in the paper's own reference list (43 numbered items) against the stated final corpus (55, with Figure 2 also showing 51) and by verifying the reported PRISMA counts at each phase; a mismatch would mean the quantitative findings are not grounded in the claimed corpus.
Extended reading notes
Core claim
The central discovery the paper advances is a structured evidence map of offline handwritten data augmentation and generation: after four PRISMA phases (identification, screening, full-text review, and quality filtering with an 11-item checklist), it retains 55 studies and organizes them into a taxonomy of traditional/model-based methods, deep learning approaches (GANs, autoencoders, sequence models, transformers, diffusion models), and augmentation techniques (geometric transforms, noise injection, StackMix, GAN-based augmentation). The paper claims that GAN-based methods account for over 60% of surveyed approaches, that the IAM Handwriting Database is the most frequently used dataset at roughly 20% of studies, and that the main challenges are style variability, dataset bias, low-resource data scarcity, and GAN training instability. It concludes that GANs are becoming the central tool for HTR data synthesis and that future work should integrate generation more tightly with recognition models, especially for low-resource scripts.
Load-bearing premise
The load-bearing premise is that the PRISMA screening was carried out exactly as described, so the final corpus really is the claimed set of studies; if the study count is wrong, the percentages and the evidence map built on them no longer hold.
Editorial extensions
If this is right
- GAN-based augmentation should be treated as the current default baseline for offline HTR data synthesis; new methods will need to beat or match it on realistic handwriting measures.
- Low-resource scripts are the clearest test bed: transfer learning, multilingual GANs, and StackMix-style augmentation are the concrete techniques the survey points to for languages with little annotated data.
- Evaluation needs a shared protocol combining quantitative metrics (CER, WER, FID) with human assessment, because no single metric currently captures style authenticity and diversity.
- Datasets such as IAM and RIMES will remain reference benchmarks, while historical and multilingual datasets (Bentham, HKR, CASIA, IFN/ENIT) fill the diversity gap for future work.
- Diffusion models and transformer-based architectures are the emerging alternatives most likely to complement or replace GANs in the next generation of handwriting generators.
Reading between the lines
- The review's findings imply that a controlled benchmark study, training one recognizer on the same script while varying only the augmentation family (geometric, GAN, diffusion), would be the natural next experiment to test which approach actually generalizes best; the survey itself does not run this experiment.
- Because the final-corpus count is reported inconsistently (55 in one place versus 51 in another, with 43 references listed), the quantitative shares should be treated as provisional until the corpus is reconciled; re-running the proportions on a verified list would check their stability.
- A testable extension suggested by the taxonomy is to combine cross-lingual transfer with diacritic-aware augmentation for scripts like Arabic and Urdu, which the included studies address separately but never jointly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a PRISMA-based systematic review of data augmentation and generation techniques for offline handwritten text recognition (HTR). The authors report screening 1,302 records from five databases, retaining 55 studies after duplicate removal, title/abstract screening, full-text review, and an 11-item quality-assessment filter. The review organizes the retained work into traditional/model-based methods, deep learning approaches (GANs, diffusion, transformers, RNNs), and augmentation techniques; summarizes commonly used datasets and evaluation metrics; and answers six research questions on evolution, effectiveness, challenges, low-resource languages, datasets, and quality assessment. The paper's main contribution is an evidence map of the field with quantitative claims about method prevalence (GANs >60%) and dataset usage (IAM ~20%).
Significance. If the reported corpus were fully specified and the counts consistent, the survey would provide a useful evidence map for researchers working on HTR data augmentation and generation. The paper assembles a broad taxonomy (Figure 4), a timeline of methodological evolution (Figure 3), a dataset table (Table 2), and a structured discussion of six research questions. The authors are also to be credited for following a PRISMA-style flow and for defining explicit inclusion/exclusion and quality criteria. However, the significance of the survey depends entirely on the reproducibility of the study-selection process; as submitted, the primary study list is not enumerable and the flow counts contradict each other, so the quantitative findings are not verifiable.
major comments (3)
- [Section 2.2.2; Figures 1 and 2(b)] The number of retained primary studies is internally inconsistent. Section 2.2.2 states that the full-text review narrowed the corpus to 55 papers, and Figure 1 reports a final corpus of 55 with a database breakdown of ACM 12, IEEE 13, ScienceDirect 5, Springer 7, arXiv 18. Figure 2(b), however, reports 51 'relevant studies' after exclusion, with a breakdown of 17, 5, 7, 10, and 12 for arXiv, ScienceDirect, Springer, ACM, and IEEE. Because all RQ answers and frequency figures are aggregates over this corpus, the review's evidence base is not uniquely defined and the reported counts cannot both be correct.
- [References; Section 5, Figures 7 and 8] The 55 retained studies are never enumerated. The reference list contains only 43 numbered entries, and no appendix, supplementary file, or table lists the selected studies with their database origins and QC scores. Consequently, the denominator is unknown and the claim that GAN-based methods account for over 60% of approaches (Section 5, Figure 7) and that IAM appears in nearly 20% of studies (Figure 8) cannot be independently recomputed. For a systematic review, this is a load-bearing omission rather than a formatting issue.
- [Figure 1; Section 2.1, Phase 4; Section 2.2.2] The role of the quality threshold EC8 is contradictory. Figure 1 states that 36 of the 91 full-text papers were excluded with 'EC8 = quality < 7/11,' and that the remaining 55 all met the threshold. Section 2.2.2 similarly says all 55 satisfied EC8. However, Phase 4 of Section 2.1 states that 'no studies were excluded, as all met at least seven of the predefined Quality Criteria.' If no studies were excluded by EC8, the 91-to-55 reduction must come entirely from the other exclusion criteria; if 36 were excluded by EC8, the Phase 4 statement is false. The PRISMA flow accounting needs to be reconciled.
minor comments (6)
- [Table 2] Table 2 labels the RIMES Database as English, but Section 3.2 correctly describes it as a French handwriting dataset; the table should be corrected.
- [Section 2.1, Phase 3] The phrase 'the eight standard threshold was established' is incomplete or garbled; it should refer to the 7/11 quality threshold defined by EC8.
- [Throughout] There are numerous typographical and formatting errors, including 'T able 1' and 'T able 2' in the table captions, 'revivers' for 'reviewers' in Phase 2, and missing spaces in phrases such as 'combinestwodifferenthandwritingsamplestocreatesyntheticdatapoints' in Section 4.2.
- [Figure 6] Figure 6 presents sample outputs from several named models but does not cite the original sources of the images; for a survey, permissions and citations for reproduced figures should be provided.
- [Section 2.2.1 and Figure 2] The text says the database distribution is illustrated in Figure 2, but Figure 2(a) shows the counts after duplicate removal rather than the initial 1,302 records; the caption should clarify that the 848 unique records are displayed.
- [Section 3.2 and Table 2] The IAM database is described in the text as containing 63,000 words, while Table 2 reports 115,000 words; these numbers should be reconciled with the official IAM statistics.
Circularity Check
No circular derivation found; the review's frequency claims rest on an inconsistently reported corpus, which is a reproducibility defect rather than a circularity defect.
full rationale
This paper is a PRISMA-style literature review; it does not derive quantitative results from first principles, fit parameters to data, or import a load-bearing uniqueness theorem from the authors' own prior work. The reference list contains no self-citations by the present authors, so no self-citation chain is involved. The survey's headline statistics, such as GANs accounting for over 60% of methods and IAM appearing in nearly 20% of studies, are descriptive aggregates over the selected corpus; even if that corpus is internally inconsistent (Section 2.2 reports 55 retained papers, Figure 2(b) shows 51, and the reference list contains only 43 entries), the statistics are not equivalent to their inputs by construction. A survey is expected to characterize its selected studies; doing so is not circular. The PRISMA inclusion criteria are selection filters, not fitted parameters, and the paper does not rename a known result as a new derivation. The identified corpus discrepancy is a genuine evidence-quality concern: the 55-paper corpus cannot be reconstructed from the manuscript, so all frequency figures currently lack a verifiable denominator. That concern belongs under correctness and reproducibility, not circularity. Accordingly, no circular step can be quoted, and the honest finding is a circularity score of 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The PRISMA protocol, designed for medical systematic reviews, is applicable to a computer vision literature survey and was implemented correctly.
- ad hoc to paper The author-defined quality criteria QC1 through QC11 are valid and sufficient to judge study quality, and the 7/11 threshold is meaningful.
- domain assumption The two reviewers applied exclusion criteria EC1 through EC8 consistently and correctly.
Cite this review
Pith. "Pith review of Advancing Offline Handwritten Text Recognition: A Systematic Review of Data Augmentation and Generation Techniques." pith.science (2026). https://pith.science/paper/HM54NAAI
@misc{pith2026250706275,
author = {Pith},
title = {Pith review of: Advancing Offline Handwritten Text Recognition: A Systematic Review of Data Augmentation and Generation Techniques},
year = {2026},
howpublished = {\url{https://pith.science/paper/HM54NAAI}},
note = {Machine review of arXiv:2507.06275}
}
read the original abstract
Offline Handwritten Text Recognition (HTR) systems play a crucial role in applications such as historical document digitization, automatic form processing, and biometric authentication. However, their performance is often hindered by the limited availability of annotated training data, particularly for low-resource languages and complex scripts. This paper presents a comprehensive survey of offline handwritten data augmentation and generation techniques designed to improve the accuracy and robustness of HTR systems. We systematically examine traditional augmentation methods alongside recent advances in deep learning, including Generative Adversarial Networks (GANs), diffusion models, and transformer-based approaches. Furthermore, we explore the challenges associated with generating diverse and realistic handwriting samples, particularly in preserving script authenticity and addressing data scarcity. This survey follows the PRISMA methodology, ensuring a structured and rigorous selection process. Our analysis began with 1,302 primary studies, which were filtered down to 848 after removing duplicates, drawing from key academic sources such as IEEE Digital Library, Springer Link, Science Direct, and ACM Digital Library. By evaluating existing datasets, assessment metrics, and state-of-the-art methodologies, this survey identifies key research gaps and proposes future directions to advance the field of handwritten text generation across diverse linguistic and stylistic landscapes.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study
A style classifier routing Manchu pages to checkpoints from a fine-tuning history matches the per-domain best CERs (0.30/1.57/4.83%), but the specialists were selected on the same frozen test sets.
Reference graph
Works this paper leans on
-
[1]
A. Boteanu, D. Cheng, and S. Kadioğlu. Read-write-learn: Self-learning for handwriting recognition. In Proceedings of the ACM Symposium on Document Engineering 2023, pages 1–4, Limerick, Ireland, Aug
work page 2023
- [2]
- [3]
-
[4]
D. Moher. Preferred reporting items for systematic reviews and meta-analyses: The prisma statement. Annals of Internal Medicine, 2009
work page 2009
-
[5]
B. Kitchenham, O. Pearl Brereton, D. Budgen, M. Turner, J. Bailey, and S. Linkman. Systematic literature reviews in software engineering – a systematic literature review.Inf. Softw. Technol., 51(1):7–15, Jan 2009
work page 2009
-
[6]
B. Kitchenham et al. Systematic literature reviews in software engineering – a tertiary study.Inf. Softw. Technol., 52(8):792–805, Aug 2010
work page 2010
-
[7]
M. N. Abdi and M. Khemakhem. A model-based approach to offline text-independent arabic writer identification and verification.Pattern Recognit., 48(5):1890–1903, May 2015
work page 1903
-
[8]
W. Li, Y. Song, and C. Zhou. Computationally evaluating and synthesizing chinese calligraphy.Neuro- computing, 135:299–305, Jul 2014
work page 2014
Show all 43 references
-
[9]
Jiang, Z
Y. Jiang, Z. Lian, Y. Tang, and J. Xiao. Dcfont: an end-to-end deep chinese font generation system. In SIGGRAPH Asia 2017 Technical Briefs, pages 1–4, Bangkok, Thailand, Nov 2017. ACM
2017
-
[10]
A. U. Dey and G. Harit. Generating synthetic handwriting using n-gram letter glyphs. InProceedings of the Tenth Indian Conference on Computer Vision, Graphics and Image Processing, pages 1–8, Guwahati Assam, India, Dec 2016. ACM
2016
-
[11]
D. G. Balreira and M. Walter. Handwriting synthesis from public fonts. In2017 30th SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), pages 246–253, Niteroi, Oct 2017. IEEE
2017
-
[12]
M. A. Souibgui et al. One-shot compositional data generation for low resource handwritten text recognition. arXiv, Oct 2021. arXiv:2105.05300
2021 arXiv
-
[13]
M. T. Parvez and S. A. Alsuhibany. Segmentation-validation based handwritten arabic captcha generation. Comput. Secur., 95:101829, Aug 2020
2020
-
[14]
S. Zheng. Analysis of generating handwriting based on gan model with different structures. In2021 IEEE International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology (CEI), pages 654–658, Fuzhou, China, Sep 2021. IEEE
2021
-
[15]
Elanwar and M
R. Elanwar and M. Betke. Generative adversarial networks for handwriting image generation: a review. Vis. Comput., Jul 2024
2024
-
[16]
L. Kang, P. Riba, Y. Wang, M. Rusiñol, A. Fornés, and M. Villegas. Ganwriting: Content-conditioned generation of styled handwritten word images.arXiv, Jul 2020. arXiv:2003.02567
2020 arXiv
-
[17]
S. Yuan, R. Liu, M. Chen, B. Chen, Z. Qiu, and X. He. Se-gan: Skeleton enhanced gan-based model for 23 brush handwriting font generation.arXiv, Apr 2022. arXiv:2204.10484
2022 arXiv
-
[18]
Fogel, H
S. Fogel, H. Averbuch-Elor, S. Cohen, S. Mazor, and R. Litman. Scrabblegan: Semi-supervised varying length handwritten text generation.arXiv, Mar 2020. arXiv:2003.10557
2020 arXiv
-
[19]
C. Wang, Y. Tang, Z. Jiang, and W. Zhang. An end-end method for handwritten xibo font generation. In 2021 4th International Conference on Artificial Intelligence and Pattern Recognition, pages 662–668, Xiamen, China, Sep 2021. ACM
2021
-
[20]
Elaraby, S
N. Elaraby, S. Barakat, and A. Rezk. A conditional gan-based approach for enhancing transfer learning performance in few-shot hcr tasks.Sci. Rep., 12(1):16271, Sep 2022
2022
-
[21]
Patil, T
V. Patil, T. Ghosh, S. Abdelhak, and C.-H. S. Kuo. Natural handwriting style generation with author adaptation. In 2022 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pages 955–961, Prague, Czech Republic, Oct 2022. IEEE
2022
-
[22]
N. Riaz, H. Arbab, A. Maqsood, A. Ul-Hasan, and F. Shafait. Conv-transformer architecture for unconstrained off-line urdu handwriting recognition. 2022
2022
-
[23]
Kotani, S
A. Kotani, S. Tellex, and J. Tompkin. Generating handwriting via decoupled style descriptors.arXiv, Sep 2020. arXiv:2008.11354
2020 arXiv
-
[24]
A. F. De Sousa Neto, B. L. D. Bezerra, G. C. D. De Moura, and A. H. Toselli. Data augmentation for offline handwritten text recognition: A systematic literature review.SN Comput. Sci., 5(2):258, Feb 2024
2024
-
[25]
Shonenkov, D
A. Shonenkov, D. Karachev, M. Novopoltsev, M. Potanin, and D. Dimitrov. Stackmix and blot augmentations for handwritten text recognition.arXiv, Aug 2021. arXiv:2108.11667
2021 arXiv
-
[26]
Sareen, R
B. Sareen, R. Ahuja, and A. Singh. Cnn-based data augmentation for handwritten gurumukhi text recognition. Multimed. Tools Appl., 83(28):71035–71053, Feb 2024
2024
-
[27]
Zdenek and H
J. Zdenek and H. Nakayama. Jokergan: Memory-efficient model for handwritten text generation with text line awareness. InProceedings of the 29th ACM International Conference on Multimedia, pages 5655–5663, Virtual Event, China, Oct 2021. ACM
2021
-
[28]
Shao and C.-L
Y. Shao and C.-L. Liu. Teaching machines to write like humans using l-attributed grammar.Eng. Appl. Artif. Intell., 90:103489, Apr 2020
2020
-
[29]
I. B. Mustapha, S. Hasan, H. Nabus, and S. M. Shamsuddin. Conditional deep convolutional generative adversarial networks for isolated handwritten arabic character generation.Arab. J. Sci. Eng., 47(2):1309– 1320, Feb 2022
2022
-
[30]
Alonso, B
E. Alonso, B. Moysset, and R. Messina. Adversarial generation of handwritten text images conditioned on sequences. arXiv, Mar 2019. arXiv:1903.00277
2019 arXiv
-
[31]
C. Luo, Y. Zhu, L. Jin, Z. Li, and D. Peng. Slogan: Handwriting style synthesis for arbitrary-length and out-of-vocabulary text. arXiv, Feb 2022. arXiv:2202.11456
2022 arXiv
-
[32]
Elarian, I
Y. Elarian, I. Ahmad, S. Awaida, W. G. Al-Khatib, and A. Zidouri. An arabic handwriting synthesis system. Pattern Recognit., 48(3):849–861, Mar 2015
2015
-
[33]
Pippi, S
V. Pippi, S. Cascianelli, and R. Cucchiara. Handwritten text generation from visual archetypes.arXiv, Mar 2023. arXiv:2303.15269
2023 arXiv
-
[34]
L. Kang, P. Riba, M. Rusinol, A. Fornes, and M. Villegas. Distilling content from style for handwritten word recognition. In 2020 17th International Conference on Frontiers in Handwriting Recognition 24 (ICFHR), pages 139–144, Dortmund, Germany, Sep 2020. IEEE
2020
-
[35]
Shonenkov, D
A. Shonenkov, D. Karachev, M. Novopoltsev, M. Potanin, D. Dimitrov, and A. Chertok. Handwritten text generation and strikethrough characters augmentation.arXiv, Dec 2021. arXiv:2112.07395
2021 arXiv
-
[36]
Mattick, M
A. Mattick, M. Mayr, M. Seuret, A. Maier, and V. Christlein. Smartpatch: Improving handwritten word imitation with patch discriminators. volume 12821, pages 268–283, 2021
2021
-
[37]
L. Kang, P. Riba, M. Rusiñol, A. Fornés, and M. Villegas. Content and style aware generation of text-line images for handwriting recognition.IEEE Trans. Pattern Anal. Mach. Intell., 44(12):8846–8860, Dec 2022
2022
-
[38]
Davis, C
B. Davis, C. Tensmeyer, B. Price, C. Wigington, B. Morse, and R. Jain. Text and style conditioned gan for generation of offline handwriting lines.arXiv, Sep 2020. arXiv:2009.00678
2020 arXiv
-
[39]
X. Liu, G. Meng, S. Xiang, and C. Pan. Handwritten text generation via disentangled representations. IEEE Signal Process. Lett., 28:1838–1842, 2021
2021
-
[40]
Z. Lian, B. Zhao, and J. Xiao. Automatic generation of large-scale handwriting fonts via style learning. In SIGGRAPH ASIA 2016 Technical Briefs, pages 1–4, Macau, Nov 2016. ACM
2016
-
[41]
M.-K. N. Huu, S.-T. Ho, V.-T. Nguyen, and T. D. Ngo. Multilingual-gan: A multilingual gan-based approach for handwritten generation. In2021 International Conference on Multimedia Analysis and Pattern Recognition (MAPR), pages 1–6, Hanoi, Vietnam, Oct 2021. IEEE
2021
-
[42]
H. Ding, B. Luan, D. Gui, K. Chen, and Q. Huo. Improving handwritten ocr with training samples gen- erated by glyph conditional denoising diffusion probabilistic model.arXiv, May 2023. arXiv:2305.19543
2023 arXiv
-
[43]
T. S. F. Haines, O. Mac Aodha, and G. J. Brostow. My text in your handwriting.ACM Trans. Graph., 35(3):1–18, Jun 2016. 25
2016
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.