{"id":"62c96ae2-8c3b-49a5-9da1-689c10c726b9","arxiv_id":"2506.07960","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A deep-learning pipeline extracted over six million migration records from handwritten Finnish church books, producing an open dataset with substantial error rates.","lead":"Researchers built an automated system that read about 200,000 photographed pages of Finnish church migration records and turned them into a dataset of over six million movement entries. The dataset is openly available, but the reading system still makes many errors, so researchers must clean and check the data carefully before using it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest claim is not the HTR error rate but the unquantified mapping from 6.2 million raw rows to usable records; Elimäki shows the linkage step, not the CER, is the bottleneck, and its outcomes are not propagated to the dataset-level claim.","rationale":"The reader identified the HTR CER as the weakest assumption. I partially agree, but the more load-bearing gap is the missing end-to-end record-level accuracy and the unpropagated case-study outcome. A CER of 0.19 on 7-character cells is not by itself disqualifying, and the reader's own wording already flags the Elimäki linkage results; the paper is transparent and the pipeline clearly works at scale. However, the phrase 'suitable for research' in the abstract and Section 6 is stronger than what Section 5 demonstrates, because the only full-pipeline evaluation required manual LLM-output review and still lost 40% of rows. That is a correctness/scope risk, not an internal inconsistency, and the authors themselves list most of these limitations in future work. No fabrication or ad hominem concerns are warranted. The appropriate verdict remains CONDITIONAL: the dataset is a real and useful resource, but the paper should either quantify the cleaning effort needed to turn raw rows into research-ready records, or soften the 'suitable for research' claim to 'suitable after per-parish validation and cleaning.'","tokens_in":16305,"tokens_out":1567,"duration_ms":17742,"concrete_test":"Replicate the Elimäki pipeline output on a second parish with mixed handdrawn and preprinted books, and compute the fraction of rows that survive the same LLM-plus-manual standardization without any manual review, reporting per-field (date, name, parish) accuracy against a hand-checked 200-row gold sample. If the survival rate and per-field accuracy are comparable to Elimäki's 60% and if the year-extraction failure rate is below, say, 5% of pages, the dataset-level claim survives; if survival is substantially lower or year failures cluster, the paper should be revised to describe the released dataset as raw extractions requiring per-parish cleaning.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central feasibility claim invites the reader to infer that one can take the 6.2 million extracted rows and use them for demographic research. The paper reports component metrics (CER 0.19, table/row/column F1 in the 90s) but never an end-to-end precision for a complete record: a usable record requires correct row segmentation, correct column alignment, correct cell text for date/name/parish, and successful linking to a known parish. Those errors compound multiplicatively. The Elimäki case study is the only place where the full pipeline is evaluated, and there the linkage step reduces the data to 60% of initial records, with only 8% of raw parish names matching at edit distance 0 and 66% linkable after LLM cleaning plus manual review. That manual review is not part of the automated pipeline and was not applied to the other 467 parishes. In addition, the year field in Elimäki fails for 1914–1915, so temporal analyses silently lose years; the paper does not quantify how widely such page-level year failures propagate in the full dataset. Section 4 explicitly reports only 52% of test images have the correct number of tables and columns, and 91% row extraction excluding double-page splits. The paper therefore supports 'large-scale extraction produces a large noisy dataset' but not yet 'a structured dataset suitable for research' without per-parish cleaning, and the cost and quality of that cleaning are unmeasured.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a fully automated deep-learning pipeline for extracting structured migration records from roughly 200,000 images of Finnish church moving records (1800–1920). The pipeline combines de-skewing, table and line detection, cell-type classification, a National Archives HTR model for text recognition, and an LLM-based year-sequence correction. The authors evaluate each component on held-out annotated data and report strong component-level results (e.g., table detection F1 97.0, row detection F1 95.5, column detection F1 97.1), but also report that only 52% of test images have the correct number of tables and columns, that overall character error rate is 0.19, and that an Elimäki case study retains only 60% of extracted rows after cleaning and manual review. The paper releases the extracted dataset on Zenodo and the pipeline and annotations on GitHub, with the central claim that large-scale automated extraction of structured data from handwritten historical records is feasible.","tokens_in":16710,"tokens_out":3270,"duration_ms":40807,"significance":"If the central claim were fully supported, this would be an important contribution to historical demography and digital humanities: a multi-million-row migration dataset covering two centuries, with open code and data, would enable studies of internal migration, urbanization, and disease spread at an unprecedented scale. The authors are transparent about many limitations, and the release of the pipeline, annotated data, and dataset is a valuable scholarly resource. However, the significance is tempered because the dataset-level claim of being 'suitable for research' is only partially supported by the evidence presented: component accuracies are high, but end-to-end errors in table structure, text recognition, and place-name linking are substantial, and the one end-to-end case study required manual review and still discarded 40% of the initial rows. The work is best interpreted as demonstrating a scalable pipeline that produces a large, noisy extraction requiring per-parish cleaning, rather than a ready-to-use research dataset, unless additional dataset-level validation is provided.","major_comments":[{"comment":"The abstract states that applying the pipeline 'resulted in a structured dataset suitable for research,' and Section 6 concludes that the work 'demonstrates that large-scale, automated extraction of structured data from handwritten historical records is feasible.' These claims are not supported by the paper's own end-to-end numbers. Section 4 reports that in 48% of test images (92/192) the predicted number of tables or columns does not match the ground truth, and Section 5 shows that after cleaning and manual review only 60% of Elimäki rows remain usable. No end-to-end measure of full-record correctness (correct row segmentation, correct column alignment, correct cell text, correct year, and successful place linkage) is reported. The paper should either soften the claims to describe a large-scale pipeline producing a noisy dataset that requires cleaning, or provide dataset-level quality metrics that directly quantify the fraction of complete, research-usable records.","section":"Abstract and Section 6"},{"comment":"The component evaluations in Tables 5–9 are performed on separate test sets and are informative, but they do not quantify the compounding of errors across stages. For example, a record is usable only if the row and column structure is correct, the HTR output for the date/name/parish is accurate, and the year is correctly assigned. Section 4 explicitly notes 48% image-level table/column errors and only approximately 91% row extraction excluding double-page splits. The paper never reports the accuracy of a complete extracted record (e.g., precision and recall of rows whose date, name, parish, and year are all correct). Without such an end-to-end metric, the 'suitable for research' claim rests on an unmeasured multiplicative error chain. I ask the authors to provide a record-level evaluation or to explicitly limit the claim to specific downstream tasks that tolerate character-level noise.","section":"Section 3.7 and Section 4"},{"comment":"The case study reveals a silent temporal-data loss that is not quantified for the full dataset: 'there are no records of arrivals to Elimäki for the years 1914 and 1915. This is due to the year identification failing in few pages, and the records from these years are merged with records from 1916 and 1917.' Since year is a critical variable for demographic analysis, the paper should either quantify how many pages across the full dataset suffer from this year-merge problem or otherwise demonstrate that such failures are rare outside Elimäki. As written, a user of the released dataset cannot know which years or pages are affected, which is a load-bearing uncertainty for the dataset's research usability.","section":"Section 5.1"},{"comment":"Table 11 shows that only 8% of extracted parish names match a known parish at edit distance 0, and 23% at edit distance ≤1. The subsequent LLM-based cleaning plus manual review raises linkage to 66% in the Elimäki case study, but that manual review is not part of the automated pipeline and was not applied to the other 467 parishes. The paper's claim that the pipeline is 'automated extraction' is technically accurate, but the path from raw extraction to a research-ready dataset is not automated and its cost, accuracy, and generalizability are unmeasured. The paper should state clearly that automatic place-name linking is a major unsolved stage for the rest of the dataset, or provide an automatic linking method with quantitative evaluation.","section":"Section 5, Table 11"}],"minor_comments":[{"comment":"The 91% row extraction figure excludes the 16 double-page split images; it would be helpful to also report row extraction including those images, or to explain how double-page splits are handled in the released CSV files.","section":"Section 4"},{"comment":"The text recognition evaluation excludes 342 lines containing question marks (unreadable to human annotators); the paper should state how many lines in the full dataset are likely to be affected by this unreadability, as it may affect the reliability of the extracted text.","section":"Section 3.7.4"},{"comment":"The year-detection evaluation in Table 10 reports precision and recall for year mentions on annotated pages, but does not report the downstream per-record year accuracy after LLM sequence correction; a per-record or per-page-level year accuracy would be more directly relevant to users.","section":"Section 3.6"},{"comment":"The Elimäki case study notes that 64% of rows followed the expected layout and an additional 34% were realigned using heuristics; the paper does not describe these heuristics in enough detail for a reader to reproduce or assess them, and a description or reference would improve reproducibility.","section":"Section 5"},{"comment":"The manuscript contains several minor typographical and formatting issues, such as the inconsistent spacing in author names and the non-standard use of 'A ¨ ıda' and 'a ∗†'; I recommend a careful copyedit.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is a potentially valuable dataset and pipeline paper, and the authors have been refreshingly transparent about error rates. However, the central claim of 'suitable for research' is not yet supported by end-to-end metrics, and the Elimäki case study — with its manual review, 40% row loss, and silent year gaps — suggests that the released 6.2-million-row dataset should not be presented as a clean research-ready resource without substantial caveats. I believe the paper can be made publishable by revising the claims to match the evidence, adding a record-level end-to-end evaluation (even on a small sample), and explicitly characterizing the cleaning needed for research use. The current manuscript is too optimistic on this load-bearing point, and the discrepancy between the abstract and the reported results should be resolved before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe thing to know about this paper is that it delivers a genuinely new, openly released dataset of 6.2 million Finnish migration records extracted from roughly 200,000 handwritten pages, and it does so with unusually candid error reporting. The pipeline itself is not novel—YOLO, Mask R-CNN, TrOCR, and an LLM for year correction—but the scale and the transparency are real contributions.\n\nWhat the paper does well: it evaluates each component on held-out data, reports the ugly numbers (CER 0.19, only 52% of test images with correct table/column counts), and then adds a case study that walks through the full pipeline on one parish. That case study is the most useful part. It shows that only 60% of the 18,809 initial rows survive cleaning, and that only 8% of raw parish names match known parishes at edit distance zero. The authors also release the annotated data and the code, so the work is reproducible.\n\nThe soft spot is the claim that the output is 'suitable for research.' That is only partially supported. The stress-test note is on target: the bottleneck is not handwriting recognition but the unquantified record-linkage step. The Elimäki case requires LLM cleaning plus manual review to get 66% of parish names linked; that manual review was not applied to the other 467 parishes. So the 6.2 million rows are raw extractions, not clean research records. The paper never provides an end-to-end precision estimate for the whole dataset, and the year failures (1914–15 in Elimäki) show that page-level errors can silently propagate.\n\nI don't think this is a fatal problem. The paper is honest about many of these issues, and the dataset will be a useful resource for historical demographers and digital humanities researchers. But the abstract and conclusion overstate readiness. The authors should either add dataset-level quality metrics (with error bars) or soften the claim to 'extraction is feasible; per-parish cleaning is required.' A serious referee would catch this and ask for the revised wording.\n\nWho is this for? Anyone working on large-scale HTR for tabular historical documents, and any historian interested in Finnish migration. It deserves a peer-review round with requested major revisions. I would not cite it in my own work (it's not my area), but domain researchers will.\n\nRecommendation: engage with it. Accept for peer review, but the revision should focus on the gap between raw rows and usable records.","headline":"A valuable new dataset and an unusually honest pipeline paper, but 'suitable for research' overreaches: the unmeasured record-linkage step, not the HTR model, is the bottleneck.","tokens_in":17082,"tokens_out":3144,"would_cite":false,"duration_ms":37079,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a fully automated deep-learning pipeline can convert roughly 200,000 handwritten Finnish church migration records into a structured dataset of over six million entries.","keywords":["handwritten text recognition","document layout analysis","historical migration records","Finnish church records","deep learning pipeline","demographic dataset","place name normalization"],"falsifier":"Take a random sample of released rows, say 1,000, manually transcribe the corresponding source-page images, and compare field by field after applying the paper's cleaning steps; if the agreement on date, migration direction, and standardized place name falls below the level needed for the intended demographic analysis, the claim that the dataset is suitable for research fails.","tokens_in":16124,"feed_emoji":"📜","tokens_out":7783,"duration_ms":81982,"temperature":0.7,"pith_summary":"This paper tries to establish that handwritten archival tables can be converted into structured, research-ready data without per-page manual transcription. The authors build a deep learning pipeline that de-skews scanned pages, detects table structure, classifies cells, recognizes handwritten text, and infers the year of each record, then apply it to roughly 200,000 images of Finnish church moving records from 1800 to 1920. The result is an open dataset of over six million migration entries. A case study of one parish shows how the extracted data can quantify arrivals and departures and map migration destinations, while also revealing how much cleaning and place-name normalization the raw output requires.","feed_headline":"Six million migration records pulled from 200,000 handwritten pages","feed_subtitle":"A deep-learning pipeline turns handwritten church records into over six million structured migration entries.","key_machinery":"The pipeline itself is the central mechanism. De-skewing is posed as a six-keypoint pose detection task solved with YOLO in two stages, so each page of an opening is straightened independently. Table and cell detection use YOLO followed by a density-based clustering step to reconstruct missing cells; Mask R-CNN performs pixel-level line detection to define rows and columns; a cell-type classifier skips empty and repetition cells; and a transformer-based handwriting recognition model fine-tuned for historical Finnish and Swedish transcribes each cell. Year extraction uses a YOLO detector plus the same handwriting recognition model, with a language model correcting the recognized years by enforcing a coherent sequence across pages. Because these stages are modular, each can be improved or replaced independently.","core_discovery":"The central claim is that large-scale automated extraction of structured data from handwritten historical records is feasible. The pipeline reaches component-level accuracies that support this: table detection F1 of 97.0, row detection F1 of 95.5, column detection F1 of 97.1, and cell-level handwriting recognition with a character error rate of 0.19. Applied to the full archive, it produced roughly 6.2 million rows from about 200,000 images in four days of parallel computing. The authors show the output can support demographic analysis, but the Elimäki case study makes clear that turning raw rows into usable records requires substantial post-processing: after duplication removal, field filtering, and place-name standardization, 60% of that parish's 18,809 extracted rows were usable, and only 8% of extracted parish-name spellings matched known parishes exactly before normalization.","pith_inferences":["The 'suitable for research' claim should be read as conditionally true: at the current character error rate, any study that relies on raw text fields will inherit noise, and the usable fraction of rows will vary by parish and layout.","Because the case-study cleaning used manual review and language-model normalization, a fully automated national-scale version of that cleaning step does not yet exist; the released dataset may therefore be most reliable for aggregate counts and less reliable for individual-level linkage.","The year-sequence correction via a language model suggests a reusable trick for other historical series where a field is monotonic or sequential; the same idea could correct misrecognized dates or page numbers in other archives.","Since free-text and half-table records, about 12% of the images, were excluded from the pipeline, migration events recorded outside tabular layouts are missing; a complete picture of Finnish internal migration would need a complementary method for those pages."],"forward_implications":["Historians can study internal migration volumes, flows, and destination patterns across Finland for 1800–1920 without manual transcription of the source books.","Linked with digitized birth, death, and disease records, the dataset enables studies of how migration shaped the spread of infectious diseases in pre-industrial Finland.","The modular pipeline can be adapted to other handwritten tabular archives, such as censuses, tax registers, or military rolls, with retraining on their layouts.","The Elimäki case demonstrates a post-processing recipe — duplication removal, heuristic column realignment, language-model-based place-name normalization, and manual review — that can turn noisy extractions into a clean parish-level migration dataset.","The open release of both the pipeline and the dataset lets other researchers reproduce the extraction and build their own cleaning layers."],"supporting_citations":[{"why":"Supplies the transformer-based handwriting recognition model that transcribes every cell, the core text extraction component.","marker":"Kansallisarkisto (2024)"},{"why":"Provides the digitized archive of about 200,000 images that defines the input dataset and its scope.","marker":"Finland's Family History Association (FFHA), 2025"},{"why":"Introduces the TrOCR architecture on which the handwriting recognition model is based.","marker":"Li et al. (2023)"},{"why":"Introduces Mask R-CNN, used for pixel-level line detection that defines table rows and columns.","marker":"He et al. (2017)"},{"why":"Establishes the YOLO-based table detection approach adapted for detecting tables and cells in the pipeline.","marker":"Huang et al. (2019)"},{"why":"Describes the annotation platform used to create the manual annotations that train and evaluate the pipeline.","marker":"Colutto et al. (2019)"}],"fun_headline_variants":["AI turns 200k church pages into 6M migration records","Handwritten Finnish church records mined for 6M migrations","Deep learning extracts 6M migrations from 200k images","From 200,000 handwritten pages to 6 million data rows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole enterprise depends on the assumption that the handwriting recognition output, with roughly one character in five wrong, is accurate enough that cleaning and normalization can produce a dataset whose remaining errors do not distort the demographic conclusions drawn from it.","fun_headline_variants_meta":{"raw":{"variants":["AI turns 200k church pages into 6M migration records","Handwritten Finnish church records mined for 6M migrations","Deep learning extracts 6M migrations from 200k images","From 200,000 handwritten pages to 6 million data rows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1277,"prompt_tokens":884,"completion_tokens":393,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":321}},"tokens_in":500,"tokens_out":393,"duration_ms":5207,"temperature":1.0,"reasoning_tokens":321,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:20:27.197145+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of released rows, say 1,000, manually transcribe the corresponding source-page images, and compare field by field after applying the paper's cleaning steps; if the agreement on date, migration direction, and standardized place name falls below the level needed for the intended demographic analysis, the claim that the dataset is suitable for research fails.","supporting_citations":[{"cited_title":"APACrefauthors \\ 2025","cited_arxiv_id":null,"evidence_quote":"Provides the digitized archive of about 200,000 images that defines the input dataset and its scope."},{"cited_title":", Yan, Q","cited_arxiv_id":null,"evidence_quote":"Establishes the YOLO-based table detection approach adapted for detecting tables and cells in the pipeline."}],"review_version":1}