{"id":"07115c34-60a1-4032-a6fa-421b2082eb1a","arxiv_id":"2502.08417","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of handwritten text recognition that categorizes methods by reading-order complexity and compares reported performance on the IAM database.","lead":"This paper surveys handwritten text recognition (HTR), organizing methods into word/line recognition and paragraph/document recognition. It compiles datasets, metrics, and reported error rates to give practitioners a roadmap for the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table III's IAM comparison mixes heterogeneous evaluation protocols, so the claim that beyond-line error rates are tightly clustered and that line-level adaptation is key is not established.","rationale":"The reader's conditional verdict is appropriate. This survey is not presenting a new empirical method or a theorem; its value is taxonomic and expository. The central empirical claim about beyond-line IAM performance is the most load-bearing part of the paper, and it rests on externally reported CER numbers that are not demonstrated to be comparable. The internal inconsistency in Section IV-C (VAN/OrigamiNet values swapped relative to Table III) and the missing DAN CER are concrete symptoms of that fragility. However, the concern is localized and does not invalidate the survey's taxonomy or its descriptive value. The correct response is to require the authors to audit and document the evaluation protocols behind each table row, and to correct the identified internal errors, while retaining the overall conditional acceptance. The reader identified precisely these weak points, so my stress test agrees with the weakest-assumption analysis.","tokens_in":32477,"tokens_out":4335,"duration_ms":47882,"concrete_test":"Perform a protocol audit of Table III: for every row (JLSAT, SAAR, OrigamiNet, SPAN, FPHR, TBPHTR, VAN, DAN, MSDocTr-Lite), open the original papers and record the exact IAM evaluation condition: which split (e.g., IAM test set A or B, paragraphs versus full pages), whether line-level segmentation or line bounding boxes were used at evaluation time, the character set, and the language model/lexicon. Re-tabulate the CER values under a single shared protocol, then recompute the range that supports the 'tightly clustered' statement. If excluding non-comparable protocols changes the ranking by more than 1 CER point, Section IV-C's conclusion that line-level adaptation is the key factor should be revised or explicitly scoped to paragraph-level IAM under matched settings.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is that the CER values in Tables II and III are mutually comparable. Section IV-C draws its central empirical conclusion—that beyond-line IAM error rates are tightly clustered and that the key factor is adapting a document into line-level structure—directly from these rows. However, the table is assembled from original papers with heterogeneous protocols: different IAM splits (line versus paragraph/full-page), different character sets and tokenizers, different language models/lexicons, and very different synthetic-pretraining budgets. Table II records only coarse columns for these factors, and Table III has no protocol column at all. The fragility is visible in Section IV-C itself: the prose assigns 4.7% to VAN and 4.6% to Origaminet, while Table III lists Origaminet 4.7 and VAN 4.6, and the DAN row is left without a CER even though DAN is described as the open-source state-of-the-art unconstrained model. These are not merely cosmetic issues: if the VAN and Origaminet numbers correspond to different evaluation modes, the 4.6–6.4 cluster is not a comparison of like with like. The paper's own Section V-A concedes that heterogeneous synthetic data and tokenization practices make comparisons unfair, yet the central claim in IV-C depends on those same heterogeneous numbers. The conclusion that line-level adaptation is the key factor therefore remains underdetermined by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey proposes a taxonomy of Handwritten Text Recognition (HTR) methods organized by recognition granularity, dividing the field into \"up to line-level\" (words and lines) and \"beyond line-level\" (paragraphs and documents). The beyond-line category is further subdivided into attention masking, line unfolding, and unconstrained approaches. The paper provides definitions of reading order complexity, a historical narrative from handcrafted to deep-learning systems, a review of datasets and evaluation metrics, and a comparative performance analysis over the IAM benchmark. The central empirical claim, stated in Section IV-C, is that beyond-line IAM error rates are tightly clustered and that the key factor is adapting the document into a line-level structure using traditional methods.","tokens_in":32684,"tokens_out":2850,"duration_ms":28837,"significance":"If the taxonomy and the performance comparison were sound, this would be a useful reference for the HTR community: the reading-order-based definitions of line/paragraph/document levels are clear and the historical narrative is well grounded in the cited literature. The survey also usefully highlights the proliferation of synthetic pre-training and the lack of standardized evaluation protocols. However, the empirical conclusion about beyond-line methods rests on a comparison that is currently not reliable, because the tables mix heterogeneous evaluation protocols and contain internal inconsistencies. The taxonomic contribution is defensible, but the benchmarking analysis needs repair before the survey can serve as a trustworthy summary of state-of-the-art results.","major_comments":[{"comment":"The prose states that \"the V AN and Origaminet, exhibit error rates of 4.7 and 4.6% respectively,\" but Table III lists Origaminet at 4.7% and VAN at 4.6%; the assignment is inverted, directly affecting the sentence that identifies the current state-of-the-art models.","section":"Section IV-C, Table III"},{"comment":"DAN is described as the open-source state-of-the-art unconstrained model, yet its CER is left as \"—\" in Table III, and the discussion states that DAN excludes IAM; this removes the strongest unconstrained comparison point and weakens the claim that \"performance remains tightly clustered across methods,\" since the DAN data point is absent from the displayed cluster.","section":"Section IV-C, Table III"},{"comment":"The central empirical conclusion that beyond-line IAM error rates are tightly clustered (4.6–6.4%) and that line-level adaptation is the key factor is drawn from Table III rows that compile CER values from papers using different IAM evaluation protocols (line-level vs. paragraph/full-page), different character sets and tokenizers, different language models/lexicons, and very different synthetic-pretraining budgets; neither Table II nor Table III includes a protocol column, and Section V-A itself concedes that heterogeneous synthetic data and tokenization practices make comparisons unfair. The cluster claim is therefore underdetermined by the evidence as presented; at minimum, the tables need explicit protocol columns and the prose must qualify the comparison.","section":"Section IV-C and Tables II/III"}],"minor_comments":[{"comment":"The caption reads \"CT)\" where it should read \"CTC\"; this typo should be corrected.","section":"Figure 9 caption"},{"comment":"The word \"sinthetic\" is misspelled and should be \"synthetic.\"","section":"Section IV-A2"},{"comment":"\"second autor\" should be \"second author.\"","section":"Acknowledgments"},{"comment":"The row for C-BGRU-GRU-Att. [150] has a blank CER; either provide the value from the source or explain why it is omitted.","section":"Table II"},{"comment":"The notation for the Vertical Attention Network alternates between \"V AN\" and \"VAN\"; please use one consistent form.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The survey's taxonomy and historical synthesis are solid, but the benchmark section is not yet reliable enough for the conclusions drawn. The two self-citations ([180], [194]) are used appropriately for metrics discussion and generalization outlook, and I do not see a circularity problem. The fit with the journal's scope is fine; a revised version that fixes the table/prose inconsistencies and adds protocol columns would be a strong contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a competent, well-structured survey that fills a real gap. The reading-order complexity framing—one direction for lines, two for paragraphs, three for documents—is a useful way to organize the HTR literature, and the masking/unfolding/unconstrained split for beyond-line methods is a fair re-description of what is out there. The definitions are clear, the historical narrative matches the cited work, and Section III plus the timeline figure are genuinely helpful for anyone entering the field. The consolidated IAM tables are useful even with their warts.\n\nThe soft spots are concentrated in Section IV-C and the benchmark tables. The prose and Table III swap the CER values for VAN and Origaminet (prose says VAN 4.7, Origaminet 4.6; the table says the reverse). That is a small fix but a concrete one. More importantly, the central empirical conclusion—that beyond-line IAM error rates are tightly clustered and that line-level adaptation is the key factor—rests on CER numbers drawn from papers with different evaluation protocols: different IAM splits, character sets, language models, and synthetic-data budgets. Table II has columns for LM and lexicon, but Table III has no protocol column at all. The authors concede in Section V-A that synthetic data and tokenization make comparisons unfair, so the claim in IV-C is weaker than the evidence supports. I would soften that wording or add a caveat about protocol comparability.\n\nOne thing the stress-test note gets wrong: the missing DAN row in Table III is not an omission. Section IV-C explicitly explains that DAN excludes IAM because the dataset only reaches paragraph-level difficulty while DAN targets more complex documents. So that part is fine.\n\nI would send this to review. The taxonomy and survey content are solid, and the core contribution does not depend on the cluster claim being airtight. For a student or researcher new to HTR, this is one of the better current entry points. I would cite it for the taxonomy and the consolidated benchmark overview.","headline":"A useful reading-order taxonomy for HTR surveys; the benchmark comparison is the weak joint and should be fixed before publication.","tokens_in":33241,"tokens_out":2059,"would_cite":true,"duration_ms":22419,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey organizes handwritten text recognition by reading-order complexity and reports that beyond-line methods cluster tightly on IAM, with line-adaptation approaches leading.","keywords":["handwritten text recognition","reading order","line-level transcription","document-level transcription","connectionist temporal classification","segmentation-free recognition","benchmarking","survey"],"falsifier":"Re-run the methods compared in Tables II and III under a single protocol: identical IAM splits, identical character sets, identical language-model settings, and identical synthetic pre-training. If an unconstrained model such as DAN or FPHR then matches or beats the masking and unfolding models on a document-level corpus, the survey's conclusion that line-level adaptation is the key factor would be overturned.","tokens_in":32252,"feed_emoji":"📝","tokens_out":7038,"duration_ms":66551,"temperature":0.7,"pith_summary":"This survey argues that the field of Handwritten Text Recognition is best organized by the number of reading-order directions a model must follow. It divides methods into those that transcribe up to one line (words and lines, a single reading direction) and those that go beyond the line (paragraphs and documents, where a second or third direction appears). Within that framework it compares published results on the IAM benchmark and reports that beyond-line error rates cluster tightly, with line-adaptation approaches (masking and unfolding) ahead of unconstrained document-level models. The takeaway is that for paragraph-level IAM, the decisive step is still converting the page into line-level units and then using established line transcription techniques. A reader should care because this gives a simple axis for comparing a fragmented literature and a concrete claim about where the current practical bottleneck is.","feed_headline":"Reading order, not image size, divides handwriting recognition","feed_subtitle":"New survey maps HTR methods by reading-order complexity and finds paragraph-level error rates cluster near 5% CER.","key_machinery":"The central object is the reading-order hierarchy (Definitions 1 through 3) together with the image-to-sequence collapse function $r(\\cdot)$, which maps a 2D feature map to a 1D sequence. Up-to-line methods rely on $r(\\cdot)$ collapsing vertical features into frames; beyond-line methods either learn a mask that selects lines before collapsing (attention masking), reshape the feature map to unfold lines into one long sequence (line unfolding), or bypass the collapse entirely with a Transformer decoder that learns an arbitrary reading order (unconstrained). The survey uses this machinery to classify every method and to explain why the vertical collapse is the critical technical barrier at paragraph and document level.","core_discovery":"The paper's central discovery is taxonomic: the meaningful complexity boundary in HTR is not image size but reading order. It defines line-level HTR as input with one reading direction, paragraph-level as two directions (line direction plus line-to-line direction), and document-level as an arbitrary third direction. It then maps the entire methodological literature onto this axis: up-to-line methods split into handcrafted pipelines (explicit or implicit segmentation) and end-to-end models (CTC, sequence-to-sequence, hybrid); beyond-line methods split into attention masking, line unfolding, and unconstrained approaches. On the empirical side, the survey collects IAM character error rates and finds that beyond-line systems are tightly clustered, with the best masking and unfolding systems (VAN and Origaminet) reaching about 4.6 to 4.7 percent CER, while the best unconstrained system (MSDocTr-Lite) reports 6.4 percent. It concludes that the key factor lies in adapting the document into a line-level structure for transcription using traditional methods.","pith_inferences":["Inference: the same reading-order axis could organize related tasks like layout analysis and document understanding, where reading order is usually treated as a separate post-processing module.","Inference: if the line-adaptation result generalizes, then synthetic pre-training for document-level HTR should be designed to produce line-adaptable representations, not only full-page images.","Inference: the taxonomy predicts that a document-level model that can switch reading orders without retraining would be a qualitative advance, since current unconstrained models learn one order from data."],"forward_implications":["If the taxonomy is right, future method papers should state which reading-order level they target, since the comparison axis is not image size but the number of reading directions.","If the performance conclusion is right, practitioners building paragraph-level HTR on IAM-like data should prefer masking or unfolding plus CTC over unconstrained document models, because they are simpler and match or beat them.","If the tight clustering is real, IAM no longer discriminates among beyond-line approaches; evaluations should shift to datasets like Rimes or Bozen that exercise complex layouts.","If the framework holds, benchmarking should standardize tokenization and synthetic-data reporting, because those choices currently change the meaning of the reported character error rate."],"supporting_citations":[{"why":"Introduces attention masking as the first segmentation-free paragraph-level HTR model (JLSAT), defining the masking category.","marker":"[25]"},{"why":"Presents Scan, Attend and Read (SAAR), the companion attention-masking model that the survey contrasts with JLSAT.","marker":"[67]"},{"why":"Proposes Origaminet, the first line-unfolding approach, which the survey reports as one of the two best beyond-line models on IAM.","marker":"[69]"},{"why":"Introduces SPAN, a size-independent unfolding model; the survey uses it to trace the unfolding category's limitations.","marker":"[70]"},{"why":"Proposes FPHR, the first unconstrained full-page model, establishing 2D positional encoding and synthetic pre-training.","marker":"[71]"},{"why":"Describes the Document Attention Network (DAN), the most popular unconstrained model, which the survey contrasts with line-adaptation methods.","marker":"[68]"},{"why":"Presents the Vertical Attention Network (VAN), the masking model that the survey lists among the best beyond-line systems.","marker":"[24]"},{"why":"Introduces MSDocTr-Lite, the best unconstrained system in the survey's beyond-line comparison.","marker":"[72]"},{"why":"Supplies the line-level conclusion that a CNN backbone plus Transformer encoder, CTC decoder, and explicit language model is the most effective line-level strategy.","marker":"[55]"},{"why":"Introduces the IAM database, the common benchmark on which the survey's comparative CER tables are based.","marker":"[157]"}],"fun_headline_variants":["Handwriting recognition's real hurdle is reading order, not size","Survey: reading order, not image size, splits HTR methods","Paragraph-level HTR error rates cluster near 5% CER","Reading order defines HTR's hard frontier, not image size"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's central performance conclusions assume that the CER values collected from different papers are directly comparable, even though the original works use different dataset splits, character sets, language models, and synthetic pre-training; the paper itself acknowledges this lack of a unified comparison framework.","fun_headline_variants_meta":{"raw":{"variants":["Handwriting recognition's real hurdle is reading order, not size","Survey: reading order, not image size, splits HTR methods","Paragraph-level HTR error rates cluster near 5% CER","Reading order defines HTR's hard frontier, not image size"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000592,"raw_usage":{"total_tokens":2775,"prompt_tokens":944,"completion_tokens":1831,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":1759}},"tokens_in":560,"tokens_out":1831,"duration_ms":14956,"temperature":1.0,"reasoning_tokens":1759,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T05:09:35.109249+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the methods compared in Tables II and III under a single protocol: identical IAM splits, identical character sets, identical language-model settings, and identical synthetic pre-training. If an unconstrained model such as DAN or FPHR then matches or beats the masking and unfolding models on a document-level corpus, the survey's conclusion that line-level adaptation is the key factor would be overturned.","supporting_citations":[{"cited_title":"The iam-database: an english sentence database for offline handwriting recognition,","cited_arxiv_id":null,"evidence_quote":"Introduces the IAM database, the common benchmark on which the survey's comparative CER tables are based."}],"review_version":1}