{"id":"842d46a7-a19d-4709-8994-7fb16305cbb0","arxiv_id":"2508.09936","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Comparing GAN, diffusion, and autoregressive handwriting generators shows the autoregressive model helps recognition most with very little real data, while the diffusion model wins once more real data is available.","lead":"This paper compares three AI handwriting generators to see which one best helps a handwriting recognition system learn from tiny sets of real historical text. It gives practical guidance on when to prioritize realistic handwriting style versus variety in generated training data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central crossover claim is contradicted by the paper's own tables: Emuru stays best on Washington at all fine-tuning sizes, and on Leopardi DiffPen only overtakes around 50% (~652 images), not at >130.","rationale":"The reader's weakest assumption focused on paradigm-level representativeness (VATr++, DiffPen, Emuru differing in training data and output granularity). That is a valid external-validity concern, but the more load-bearing problem is internal: the paper's own quantitative tables contradict the headline ranking and threshold. The strongest claim is specifically the empirical crossover: Emuru is best under ~130 real images, DiffPen is best above. Washington is a direct counterexample—Emuru remains best at 50%, 75%, and 100% of training data, i.e., at 263, 394, and 526 real images. Leopardi also only shows DiffPen overtaking at ~652 images, which is far above the stated threshold. Since the central contribution of the paper is this quantitative guideline, and it fails on one of three datasets and is materially wrong on another, the claim cannot be accepted as stated. The filtering analysis and generation-quality metrics remain useful, so a revised version that reports per-dataset findings without a universal crossover claim, and ideally with uncertainty estimates, could be salvageable. But the current paper's main conclusion is not supported by the evidence it presents.","tokens_in":17439,"tokens_out":9436,"duration_ms":91475,"concrete_test":"Recompute, for each fine-tuning fraction in Tables 2-4, the best CER per HTG model across all filtering strategies; plot the best-CER difference (DiffPen minus Emuru) against the absolute number of real training images. If the difference never becomes negative on Washington at any fraction beyond 130 images, the central claim fails. As a second check, rerun the 50% and 100% Washington conditions with three seeds and report mean +/- std to test whether Emuru's apparent advantage (0.7 and 0.4 CER points) is significant.","verdict_should_be":"REJECT","load_bearing_attack":"Section 4.4 states that with fewer than 130 real images Emuru is best and that beyond this threshold DiffPen wins because diversity dominates; Section 5 repeats this as a general guideline. The reported tables do not support this consistently. In Washington (Table 3), taking the minimum CER over all filtering rows per model, Emuru beats DiffPen at every fine-tuning size: 4.4 vs 5.1 at 50% (263 images), 4.2 vs 4.6 at 75%, and 3.5 vs 3.9 at 100% (526 images). Thus Emuru wins even far above the claimed 130-image threshold. In Leopardi (Table 2), Emuru's best is 10.0 at 25% (326 images) vs DiffPen's 10.1, and DiffPen only becomes best at 50% (652 images), not at the stated >130 threshold. Saint Gall is the only dataset with a crossover near 25% (~117 images). The '130 images' cutoff and the style-vs-diversity explanation therefore appear to be artifacts of particular filtering choices and dataset-specific behavior, not a robust empirical law. No error bars or repeated seeds are reported, so even the small gaps (e.g., 3.5 vs 3.9) may be noise, but the absence of any Washington crossover is independent of that uncertainty.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the use of synthetic handwritten text generated by three styled HTG models (VATr++, DiffPen, Emuru) as pretraining data for low-resource handwritten text recognition. The authors propose a pipeline that generates author-specific synthetic line images, filters them by readability (TrOCR CER) or style fidelity (HWD percentiles), pretrains a CRNN-based HTR model (DefCRNN), and then fine-tunes on progressively smaller subsets of three historical datasets (Leopardi, Washington, Saint Gall). The central claims are that Emuru yields the most style-faithful synthetic data and the best zero-shot and very-low-resource fine-tuning performance; that there is a threshold around 130 real fine-tuning images below which style similarity is crucial and above which diversity (and therefore DiffPen) becomes dominant; and that HWD-based filtering gives no clear benefit while moderate CER-based filtering is beneficial. The paper also reports generation metrics (FID, KID, HWD, DeltaCER) for the three generators.","tokens_in":17791,"tokens_out":8668,"duration_ms":85285,"significance":"If the empirical ranking and the 130-image threshold were reliable, the paper would provide a practically useful benchmark and actionable guidance for HTG-based pretraining for HTR, spanning multiple languages and fine-tuning regimes. The negative result on HWD filtering is also a useful data point, and the large experimental matrix is a strength. However, the central quantitative threshold is not consistently supported by the paper's own tables, and the absence of repeated runs plus a model-design confound prevent strong paradigm-level conclusions. The paper is best viewed as a valuable comparative study of three specific off-the-shelf systems, with conclusions that need to be rephrased and supported more carefully.","major_comments":[{"comment":"The claimed 'fewer than 130 real images' crossover is contradicted by the reported results. In Washington, taking the minimum CER over all filtering rows for each model, Emuru beats DiffPen at every fine-tuning size: 4.4 vs 5.1 at 50% (263 real images), 4.2 vs 4.6 at 75% (394), and 3.5 vs 3.9 at 100% (526). Thus Emuru wins even far above the stated 130-image threshold. In Leopardi, DiffPen only becomes better at 50% (652 images), not above 130: at 25% (326 images) Emuru's best CER is 10.0 vs DiffPen's 10.1. In Saint Gall the crossover occurs at 50% (234 images), not at 25% (117). These numbers also undermine the Section 5 recommendation that 'more than 130' training images makes DiffPen preferable. The 130-image cutoff appears to be an artifact of particular filtering choices and dataset-specific behavior, not a general empirical law.","section":"Section 4.4, Tables 2–4"},{"comment":"All CER numbers are single runs with no error bars or significance testing. Several load-bearing comparisons are small: e.g., 10.0 vs 10.1 on Leopardi at 25%, 4.2 vs 4.6 on Washington at 75%, and 3.5 vs 3.9 at 100%. Without repeated seeds, one cannot determine whether these differences are meaningful. The paper should report means and standard deviations over at least three seeds, or explicitly state that the close margins are not statistically assessed. This is particularly important because the threshold-based guideline in Section 4.4 and Section 5 depends on exactly these small gaps.","section":"Section 4.4, Tables 2–4"},{"comment":"The paper attributes performance differences to generative paradigms, but the three models differ in several other respects. VATr++ operates at word level with 15 style-sample inputs; DiffPen operates at word level and is patched into lines, uses 5 style samples, and was trained with a different data mixture; Emuru is a line-level autoregressive model trained exclusively on synthetic data. The observed advantage of Emuru in low-resource settings could stem from output granularity, training-data composition, or conditioning protocol rather than autoregression per se. Since the abstract and conclusions use paradigm labels ('adversarial', 'diffusion', 'autoregressive'), these confounds are load-bearing. The paper should either rephrase the conclusions to refer to 'the three compared systems' or provide additional controlled evidence that isolates the generative paradigm.","section":"Section 3.2 and Section 5"}],"minor_comments":[{"comment":"The sentence claiming that Emuru-generated images 'exhibit the lowest readability according to the TrOCR model' is unclear and appears inconsistent with the stated DeltaCER values. If Emuru's images are harder for TrOCR despite having low absolute DeltaCER, the text should explain this explicitly; otherwise it is likely a typo for 'highest readability' or 'lowest CER gap.'","section":"Section 4.4"},{"comment":"The 130-image threshold is presented using approximate examples (10% of Leopardi corresponds to 130 images; 25% of Saint Gall corresponds to 125). Given its importance, the paper should specify exactly which fine-tuning fractions and datasets the threshold applies to, and how it was derived from the tables.","section":"Section 4.4"},{"comment":"The conclusion that HWD-based filtering gives no clear benefit should be qualified by the fact that HWD thresholds are percentile-based per model, not absolute style-fidelity cutoffs. Comparing the 'HWD25%' of one model with that of another does not compare equal absolute fidelity levels. An analysis with common absolute HWD thresholds, or a statement of this limitation, would strengthen the filtering section.","section":"Section 3.3 / Section 5"},{"comment":"The caption reads 'the respected target dataset'; this should be 'the respective target dataset.'","section":"Figure 3"},{"comment":"Reference [50] has a truncated title: 'Alfie: Democratising RGBA Image Generation with No.' The full title should be restored.","section":"References"},{"comment":"The paper would benefit from a reproducibility section or link to code/checkpoints. The tables report the exact numbers, but no scripts or pretrained models are provided; for a benchmark-style study, releasing the pipeline would substantially increase the paper's value.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The experimental matrix is extensive and the paper is likely to be a useful reference for practitioners, but the headline quantitative guideline (the 130-image crossover) is not supported by the paper's own tables. The authors should reanalyze the data, rephrase the conclusions, and ideally add repeated runs. I also note that the first author is a co-creator of Emuru and of the HWD metric used for generation evaluation and filtering. This is not by itself a defect, but the paper should include a disclosure statement and, ideally, corroborate the HWD-based conclusions with an independent style similarity measure or external evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a useful benchmark paper, but the headline crossover rule — Emuru under 130 images, DiffPen beyond — doesn't hold up against the paper's own tables. On Washington, Emuru beats DiffPen at every fine-tuning size, even at 100% (526 lines) where Emuru gets 3.5 CER vs DiffPen's 3.9. On Leopardi, DiffPen only takes the lead at 50% (~652 lines), not at the stated 130. Saint Gall is the only dataset with a crossover near 25% (~117 lines). So the 130-image threshold and the style-vs-diversity explanation look like a reading of selected filtering rows rather than a robust empirical law. That's a load-bearing claim, because Sections 4.4 and 5 present it as the main takeaway.\n\nWhat the paper does well: it fills a real gap. A systematic comparison of GAN, diffusion, and autoregressive HTG models for low-resource HTR fine-tuning was missing. The filtering analysis is the most valuable part — the negative result that HWD-based style filtering gives no clear benefit is credible and useful, and the recommendation to use a relaxed CER threshold is reasonable. The authors also report generation metrics and direct-transfer results, and they make the full tables available in the paper, which is exactly why the inconsistency is visible.\n\nSoft spots beyond the threshold claim: no error bars or repeated seeds, so differences under 1 CER are treated as meaningful without any variance estimate. The three models are treated as representatives of their paradigms, but they differ in training data, output granularity (word vs line), and number of style references (15 vs 5 vs 1). I don't think the paper can attribute the performance differences to 'autoregressive vs diffusion' generically. There is also a mild self-evaluation concern — the authors created Emuru and HWD — but they use external metrics for filtering and the ranking on generation scores is consistent, so I don't see that as disqualifying.\n\nBottom line: this deserves a serious referee and could be a solid contribution after major revision. The empirical material is there; the interpretation needs to be reined in. If the authors reframe the conclusions to describe dataset-specific behavior rather than a universal threshold, and add at least a few seeds or a discussion of variance, the paper would be a reasonable guide for practitioners.\n\nRecommendation: send to peer review, but expect heavy revision.","headline":"Useful benchmark, but the headline crossover claim (Emuru <130 images, DiffPen beyond) doesn't survive the paper's own tables.","tokens_in":18283,"tokens_out":2492,"would_cite":true,"duration_ms":25395,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Style-faithful synthetic handwriting beats diverse synthetic data when fewer than 130 real lines are available for fine-tuning.","keywords":["Handwritten Text Recognition","Handwritten Text Generation","low-resource learning","synthetic pretraining data","style fidelity","historical manuscripts","diffusion models","autoregressive generation"],"falsifier":"Re-run the comparison with an autoregressive and a diffusion generator trained on the same synthetic corpus, both generating whole lines, and measure HTR CER after fine-tuning on fewer than 130 real images; if the diffusion model performs as well as or better than the autoregressive one, the paradigm-level conclusion fails. Alternatively, take DiffPen and filter its outputs to match Emuru's HWD distribution; if its low-data advantage does not disappear, style similarity is not the operative factor.","tokens_in":17367,"feed_emoji":"✍️","tokens_out":7272,"duration_ms":68035,"temperature":0.7,"pith_summary":"Handwritten Text Recognition models struggle on small historical collections whose handwriting differs from typical training data. This paper asks which style of synthetic handwriting generator—adversarial, diffusion, or autoregressive—produces the most useful pretraining images for such low-resource settings. It finds that when fewer than about 130 real transcribed lines are available, the autoregressive model Emuru, whose synthetic images most closely match the target handwriting style, gives the best recognition results after fine-tuning. With more real data, the more stylistically varied images from the diffusion model DiffPen become the better pretraining source, because diversity outweighs style similarity. The practical upshot is a quantitative guideline: match style for very small collections, prioritize variety once enough real lines exist.","feed_headline":"Style beats variety below 130 real lines","feed_subtitle":"Which synthetic handwriting data to pretrain on flips at 130 real training lines.","key_machinery":"The central object is an evaluation pipeline rather than a single mathematical identity: each of three off-the-shelf styled HTG models—VATr++ (GAN), DiffPen (diffusion), Emuru (autoregressive)—generates line images conditioned on a few real style samples; these images pretrain the DefCRNN recognizer (a deformable-convolution CNN-LSTM line recognizer) before fine-tuning on fractions of real data. The load-bearing comparisons are the style-fidelity/readability scatter (HWD vs CER) and the fine-tuning curves across data fractions, which isolate the crossover at roughly 130 images.","core_discovery":"Running a pretrain-then-fine-tune pipeline with the DefCRNN recognizer on three historical single-author datasets (Leopardi, Washington, Saint Gall), the paper reports a consistent crossover: Emuru, the autoregressive generator, yields the lowest CER when fine-tuning uses fewer than 130 real images, and DiffPen, the diffusion generator, yields the lowest CER beyond that threshold. The paper interprets this as style similarity driving performance in extreme low-data regimes, while synthetic-data diversity drives it once fine-tuning data is more plentiful. It also finds that filtering synthetic samples by handwriting distance (HWD) has no clear benefit, while readability-based filtering (CER)","pith_inferences":["An extension the paper leaves implicit: the paradigm-level ranking may be confounded by differences in training-data composition and output granularity among the three models, so the conclusion should be re-tested with generators matched on those axes.","The roughly 130-image crossover is likely tied to the recognizer and datasets used; a practical calibration curve could be built for a given archive by sweeping fine-tuning size before committing to a generator.","Emuru's strong style fidelity despite being trained only on synthetic data hints that diverse synthetic corpora may be a general route to zero-shot style generalization, not just a property of autoregressive architectures."],"forward_implications":["For collections with fewer than about 130 labeled lines, use a style-faithful autoregressive generator such as Emuru to build pretraining data; matching the target handwriting matters more than variety.","For collections with more labeled lines, a diffusion-based generator such as DiffPen gives better HTR fine-tuning results by providing stylistic diversity.","Filtering synthetic pretraining data by handwriting-style distance (HWD) does not improve recognition and can be skipped.","Readability-based filtering helps only when the threshold is relaxed (CER around 0.30); strict thresholds remove too many samples and hurt or prevent convergence.","In zero-shot direct transfer without fine-tuning, Emuru-generated data yields the best recognition, indicating style fidelity is the key factor when no real target labels are used."],"supporting_citations":[{"why":"Emuru, the autoregressive HTG model whose style-faithful synthetic lines drive the low-data fine-tuning results.","marker":"[47]"},{"why":"DiffPen, the diffusion HTG model whose stylistic diversity yields better HTR results when more real fine-tuning lines are available.","marker":"[40]"},{"why":"VATr++, the GAN-based HTG model that serves as the adversarial-paradigm comparison point.","marker":"[64]"},{"why":"DefCRNN, the deformable-convolution CNN-LSTM recognizer used for all pretraining and fine-tuning experiments.","marker":"[9]"},{"why":"TrOCR, the off-the-shelf recognizer used to compute CER-based readability filtering of synthetic samples.","marker":"[34]"},{"why":"HWD, the handwriting-distance score used for style-fidelity filtering and for measuring style similarity of generated samples.","marker":"[46]"},{"why":"Earlier result establishing the relevance of style and language match when choosing pretraining data for single-writer fine-tuning, which motivates the synthetic-data protocol.","marker":"[45]"}],"fun_headline_variants":["Synthetic HTR data: style under 130 lines, variety above","130-line crossover decides best handwriting generator","For HTR fine-tuning, less than 130 lines? Choose Emuru","Diffusion wins HTR after 130 real lines; autoregressive before","Handwriting style vs diversity flips at 130 training lines"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The three HTG models are treated as representatives of their generative paradigms, even though they differ in training-data composition, output granularity, and conditioning, so the observed ranking could be caused by those differences rather than by the paradigm.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic HTR data: style under 130 lines, variety above","130-line crossover decides best handwriting generator","For HTR fine-tuning, less than 130 lines? Choose Emuru","Diffusion wins HTR after 130 real lines; autoregressive before","Handwriting style vs diversity flips at 130 training lines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000592,"raw_usage":{"total_tokens":2580,"prompt_tokens":681,"completion_tokens":1899,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":1811}},"tokens_in":425,"tokens_out":1899,"duration_ms":12639,"temperature":1.0,"reasoning_tokens":1811,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:41:48.850485+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison with an autoregressive and a diffusion generator trained on the same synthetic corpus, both generating whole lines, and measure HTR CER after fine-tuning on fewer than 130 real images; if the diffusion model performs as well as or better than the autoregressive one, the paradigm-level conclusion fails. Alternatively, take DiffPen and filter its outputs to match Emuru's HWD distribution; if its low-data advantage does not disappear, style similarity is not the operative factor.","supporting_citations":[{"cited_title":"Zero-Shot Styled Text Image Generation, but Make It Autoregressive","cited_arxiv_id":null,"evidence_quote":"Emuru, the autoregressive HTG model whose style-faithful synthetic lines drive the low-data fine-tuning results."},{"cited_title":"DiffusionPen: Towards Controlling the Style of Handwritten Text Generation.ECCV, 2024","cited_arxiv_id":null,"evidence_quote":"DiffPen, the diffusion HTG model whose stylistic diversity yields better HTR results when more real fine-tuning lines are available."},{"cited_title":"V ATr++: Choose Your Words Wisely for Handwritten Text Genera- tion.IEEE Trans","cited_arxiv_id":null,"evidence_quote":"VATr++, the GAN-based HTG model that serves as the adversarial-paradigm comparison point."},{"cited_title":"Boosting Modern and Historical Handwrit- ten Text Recognition with Deformable Convolutions.IJDAR, pages 1–15, 2022","cited_arxiv_id":null,"evidence_quote":"DefCRNN, the deformable-convolution CNN-LSTM recognizer used for all pretraining and fine-tuning experiments."},{"cited_title":"TrOCR: Transformer-based optical character recognition with pre- trained models.AAAI, 2023","cited_arxiv_id":null,"evidence_quote":"TrOCR, the off-the-shelf recognizer used to compute CER-based readability filtering of synthetic samples."},{"cited_title":"HWD: A Novel Evaluation Score for Styled Handwritten Text Generation","cited_arxiv_id":null,"evidence_quote":"HWD, the handwriting-distance score used for style-fidelity filtering and for measuring style similarity of generated samples."},{"cited_title":"How to choose pretrained handwriting recognition models for single writer fine-tuning","cited_arxiv_id":null,"evidence_quote":"Earlier result establishing the relevance of style and language match when choosing pretraining data for single-writer fine-tuning, which motivates the synthetic-data protocol."}],"review_version":1}