{"id":"1f0bce69-b5f2-4af9-be21-baa4e755adea","arxiv_id":"2507.19092","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The best OCR pipeline for Slovene folklore depends on document type: olmOCR for clean typewritten pages, Tesseract plus LLM post-processing for complex layouts and degraded newspapers.","lead":"The paper compares two OCR pipelines on Slovene folklore texts: a single-stage vision-language model (olmOCR) and a two-stage Tesseract with LLM cleanup. It suggests the best choice depends on document layout and scan quality, and warns that LLM cleanup can silently modernize dialect words.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manual review without sample sizes, ground truth, or inter-rater checks cannot support the per-corpus pipeline rankings in Table 1; the central claim needs quantitative validation.","rationale":"I read the paper as an applied comparison whose central claim is necessarily comparative. The claim is plausible, and the use of real Slovene corpora and locally deployed olmOCR is a strength. However, the entire empirical basis is the authors' visual judgment. The absence of quantitative metrics is not a stylistic omission; it is the load-bearing support for every recommendation. The concrete test would replace anecdote with measurement and directly tests whether the rankings survive. No ad hominem is intended; the concern is about evidence, not integrity. The reader's weakest_assumption is essentially the same, so agreement is 'agree'. Since the paper remains conditionally acceptable pending such validation, no change to the reader's CONDITIONAL verdict is needed.","tokens_in":9108,"tokens_out":4030,"duration_ms":43187,"concrete_test":"Create ground-truth transcriptions for a random sample of at least 20 pages per corpus. Run both pipelines on identical images; for Ciciban also run Tesseract+LLM+LayoutParser. Compute character and word error rates, and a dialect-preservation score: the fraction of pre-identified dialectal/archaic tokens preserved after LLM post-processing. Report confidence intervals and have a second annotator independently score a subset with agreement statistics. If olmOCR ties or beats Tesseract+LLM on Kmetijske/Ciciban once errors are measured, or if CIs overlap, Table 1's rankings and the Section 5 conclusion need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 5: 'no single OCR pipeline is universally optimal' plus the per-corpus preferences in Table 1) depends on the assumption that the authors' manual visual review in Section 3.2 reliably ranks accuracy, layout preservation, and linguistic authenticity. That assumption is unsupported: no page counts, error metrics, ground truth, prompts, model versions for the LLM calls, or inter-rater agreement are reported. The only evidence offered is three selected examples (Figures 5-7), which are illustrative rather than systematic. Because the evaluators knew which pipeline produced each output, the rankings are vulnerable to confirmation bias. The reporting also contains a small but telling inconsistency: Section 4.2.1 says ChatGPT changed 'žlahnega' to 'glavnega', while Figure 5 displays 'glahnega' in the highlighted output. Finally, the paper's own examples document semantic drift (e.g., 'storjice' → 'zgodbe') in the pipeline it nevertheless recommends for Kmetijske, so the notion of 'preferred' is never defined against the stated authenticity goal. This makes the central comparative conclusion underdetermined by the published evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares two OCR pipelines for digitizing Slovene folkloristic and historical materials: a single-stage olmOCR pipeline and a two-stage Tesseract-plus-LLM pipeline, with optional LayoutParser for complex layouts. The evaluation covers three corpora: a 19th-century newspaper (Kmetijske in rokodelske novice), a mid-20th-century children's magazine (Ciciban), and typewritten fairy tales from the Institute of Folklore Studies. Based on qualitative manual review of OCR outputs, the paper concludes that no single pipeline is universally optimal and recommends Tesseract+LLM for the newspaper and children's magazine, and olmOCR for clean typewritten fairy tales, while cautioning about LLM-induced semantic drift and normalization of dialectal terms.","tokens_in":9471,"tokens_out":3595,"duration_ms":36153,"significance":"If the conclusions were supported by rigorous evidence, the paper would provide practically useful guidance for digital-heritage practitioners working with historical Slavic-language materials, and its emphasis on the trade-off between readability and linguistic authenticity is a worthwhile contribution. The authors also deserve credit for testing on authentic, publicly available corpora rather than synthetic data, and for explicitly discussing privacy and local deployment advantages of olmOCR. However, the central comparative claims currently rest entirely on manual review with no quantitative metrics, no sample sizes, no gold-standard comparison, and no inter-annotator reliability check, so the specific recommendations in Table 1 are not yet established by the published evidence.","major_comments":[{"comment":"The per-corpus rankings in Table 1 and the central conclusion in Section 5 rest on the manual visual review described in Section 3.2, but the paper reports no number of pages evaluated per corpus, no character error rates or word error rates, no comparison against a ground-truth transcription, and no inter-annotator agreement statistic. Because the ranking claims are load-bearing, this lack of quantitative support is a major issue. For the revision, the authors should report the size of each evaluation sample, compute standard error metrics against a manually verified transcription for at least a random subset of each corpus, and have at least two independent reviewers score outputs on the stated criteria, with agreement reported.","section":"§3.2 (Output Evaluation), §4.3, Table 1"},{"comment":"Section 4.2.1 states that ChatGPT changed 'žlahnega' to 'glavnega', but the highlighted enhanced output in Figure 5 contains 'glahnega' rather than 'glavnega'. This is a concrete inconsistency in the key example used to illustrate semantic drift. The text or the figure must be corrected, and all highlighted tokens in Figure 5 should be verified against the actual model outputs before the example is used as evidence.","section":"§4.2.1, Figure 5"},{"comment":"The paper acknowledges that the Tesseract+LLM pipeline introduced 'storjice → zgodbe' normalization and 'semantic drift' on Kmetijske in rokodelske novice, yet Table 1 still recommends this pipeline for that corpus. Since Section 3.2 lists 'language integrity' as an evaluation criterion, the trade-off between readability and authenticity is never made explicit. The authors should define a scoring rubric that includes authenticity preservation and show that the recommended pipeline wins on that rubric; otherwise the recommendation is underdetermined by the stated criteria.","section":"§4.1.2, §4.2.1, Table 1"},{"comment":"The LLM post-processing stage is not described reproducibly: the paper says 'ChatGPT-4 (GPT-4.o version model)' and mentions 'other LLMs (such as Gemini and Claude)' in Section 4.1.2, but it never provides the exact model identifiers, the prompts used, decoding parameters, or any results separated by model. This makes the experiments impossible to reproduce and leaves unclear whether the findings are specific to one model configuration or general across LLMs. The revision should include the full prompts, model versions and access dates, and per-model results.","section":"§3.2 (Pipeline B), §4.1.2"}],"minor_comments":[{"comment":"The phrase 'the use olmOCR' is missing 'of'; it should read 'the use of olmOCR'.","section":"§4.3, bullet list"},{"comment":"The phrase 'collected by theInštitut za narodopisje' is missing a space after 'the'.","section":"§3.1.3"},{"comment":"The caption says 'Top: original scan. Middle: raw OCR output using Tesseract. Bottom: ChatGPT-4 enhanced output', but the visible figure appears to contain two text paragraphs rather than three clearly separated panels; the layout should be clarified, and the yellow highlights may not be visible in grayscale print.","section":"Figure 5 caption"},{"comment":"The cost estimate 'under $190 per million pages' lacks a price date and a specification of the hardware or API assumptions; please add a reference or a calculation footnote.","section":"§3.2 (Pipeline A)"},{"comment":"The corpus name is spelled 'Kmetijske in Rokodelske novice' in Table 1 but 'Kmetijske in rokodelske novice' elsewhere; please standardize the capitalization.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is currently a qualitative work-in-progress report rather than a fully supported comparative study. The central question is within scope for a digital-heritage venue, and the authors' practical orientation is valuable. If the authors can add quantitative validation, sample sizes, and inter-annotator agreement, the paper could become acceptable. In its present form, the evidence basis is too thin for the strength of the conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a readable, honest applied comparison of olmOCR against Tesseract+LLM on three real Slovene folklore collections. What is genuinely new is the corpus work: nobody had reported these pipeline behaviors on Kmetijske in rokodelske novice, Ciciban, and the typewritten fairy-tale collection. The paper also does a good job of situating itself, and the warning about LLM post-processing normalizing dialectal words (storjice to zgodbe) is a real caution for heritage practitioners.\n\nThe soft spot is the one the stress-test names, and it is in proportion: the central rankings in Table 1 rest on unquantified manual review. No page counts, error rates, sample sizes, prompts, or model versions for the LLM calls are reported. Figures 5-7 are selected examples, and the evaluators knew which pipeline produced each output, so confirmation bias is a live concern. There is also a small but telling inconsistency: Section 4.2.1 says ChatGPT changed žlahnega to glavnega, but Figure 5 shows glahnega in the highlighted output. That makes the illustrative example harder to trust.\n\nThe deeper issue is that the pipeline recommended for Kmetijske (Tesseract+LLM) is the one whose own example shows semantic drift. So 'preferred' is doing a lot of work: preferred for readability, not for linguistic fidelity. The general conclusion that no single pipeline is universally optimal is almost certainly true, but it is not the claim that needs support; the specific per-corpus preferences are underdetermined by the published evidence.\n\nWho gets value: digital heritage practitioners who want a pilot-level orientation before choosing OCR workflows. Methodologically oriented readers will be frustrated by the missing metrics. With revision—reporting what was actually reviewed, adding quantitative evaluation or at least a clear limitation statement, and fixing the example mismatch—this could be a solid practice paper.\n\nMy recommendation: send it to peer review with the expectation of heavy revision. The question matters and the authors did real work on real archives, so referee time is not wasted. But it should not be accepted as-is, and the authors should be pushed to either add error metrics or scale the conclusions back.","headline":"A useful pilot comparison of OCR pipelines on real Slovene folklore corpora, but the per-corpus recommendations outrun the evidence because the evaluation is entirely qualitative.","tokens_in":9816,"tokens_out":3361,"would_cite":false,"duration_ms":38983,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Comparing two OCR approaches on three Slovene folkloristic corpora, the paper concludes that no single pipeline is universally optimal: olmOCR suits clean typewritten fairy tales, while Tesseract plus LLM post-processing suits historical…","keywords":["optical character recognition","OCR","large language models","post-OCR correction","historical documents","Slovene language","folkloristic texts","layout analysis"],"falsifier":"A quantitative re-run of the same three corpora against manually verified ground truth, measuring character error rate and the rate at which dialectal words are changed, would settle the claim: if Tesseract plus LLM post-processing matches or beats olmOCR on typewritten fairy tales, or if olmOCR beats Tesseract plus LLM post-processing on Ciciban, the document-sensitive ranking would fail.","tokens_in":8930,"feed_emoji":"📜","tokens_out":6156,"duration_ms":56874,"temperature":0.7,"pith_summary":"This paper compares two ways of turning scanned Slovene folkloristic documents into digital text: a single-stage tool called olmOCR and a two-stage pipeline that combines Tesseract OCR with large-language-model post-processing. It argues that neither approach is best everywhere; the winning choice depends on scan quality, layout complexity, and linguistic register. For clean typewritten fairy tales, olmOCR transcribes more faithfully, while for a 19th-century newspaper and a children's magazine, Tesseract followed by an LLM produces more readable results, with layout parsing recommended for the magazine. The paper warns that LLM post-processing can quietly replace dialectal or archaic words with modern equivalents, trading linguistic authenticity for readability. The practical upshot is a document-sensitive selection strategy for folklore digitization.","feed_headline":"No single OCR pipeline is best for Slovene folklore texts","feed_subtitle":"For clean typewritten stories olmOCR wins; newspapers and magazines need Tesseract plus LLM post-processing.","key_machinery":"The comparison runs on two pipelines. olmOCR is a locally deployed vision-language model that converts scans directly into plain text; Tesseract is a conventional OCR engine whose raw output is passed to a large language model for error correction and paragraph reconstruction, with LayoutParser optionally used first to detect and segment text regions. The decisive mechanism is the interaction between these pipelines and document type: clean uniform pages favor direct recognition, while noisy or complex pages benefit from LLM repair at the cost of normalizing dialectal or archaic language.","core_discovery":"Across three Slovene corpora of different eras and layouts, the paper finds that no OCR pipeline is universally optimal. olmOCR outperforms Tesseract plus LLM post-processing on mid-20th-century typewritten fairy tales, preserving sentence structure and diacritics with fewer hallucinations. Tesseract plus LLM post-processing wins on a 19th-century newspaper and on a children's magazine, where the LLM restores readable, coherent text but occasionally modernizes historical terms. The determining variables are scan quality, layout complexity, and linguistic register, and the paper recommends a document-sensitive pipeline choice rather than a single tool.","pith_inferences":["A practical extension of the paper's comparison is that archives should store both the raw OCR output and the LLM-enhanced version, so linguistic authenticity is never destroyed by normalizing post-processing.","The observed trade-off suggests a testable hypothesis for neighbouring languages: dialectal or archaic spellings will be the first casualties of LLM post-correction in any historical corpus, not just Slovene.","The authors' qualitative judgments could be turned into a quantitative benchmark by measuring character error rate and dialect-term preservation on the same corpora, allowing archives to set explicit thresholds for when to prefer fidelity over readability."],"forward_implications":["Digitization projects on Slovene folklore should not commit to a single OCR tool; the choice should follow the document's scan quality, layout, and linguistic register.","On clean, uniformly typewritten pages, a single-stage vision-language OCR such as olmOCR preserves dialectal forms with fewer hallucinations and can be run locally for privacy.","On complex layouts such as children's magazines with mixed poetry, prose, and images, Tesseract plus an LLM post-processing step, with layout parsing when needed, produces the most readable and coherent text.","LLM post-processing should be used with human validation or constrained prompts, because it can replace dialectal or archaic words with modern equivalents.","Scan quality is the dominant limiting factor: conventional preprocessing such as grayscale conversion, binarization, and dilation does not reliably salvage low-quality scans."],"supporting_citations":[{"why":"Describes olmOCR, the single-stage vision-language pipeline that wins on typewritten fairy tales.","marker":"[PBD*25]"},{"why":"Surveys LLM-based post-OCR correction for historical documents and frames the trade-offs motivating the two-stage pipeline.","marker":"[KLKG25]"},{"why":"Documents layout and typography challenges in historical documents, grounding the need for layout-aware processing.","marker":"[FKG25]"},{"why":"Reports that LLM post-correction can harm accuracy through hallucinated outputs, the caution underlying the authenticity warnings.","marker":"[BER*24]"},{"why":"Shows GPT-based correction can reduce character error rates, supporting the promise of the LLM stage.","marker":"[Bou24]"},{"why":"Treats historical-to-modern normalization as a translation task, illustrating the normalization risk the paper observes.","marker":"[Ehr24]"},{"why":"Supplies the historical newspaper corpus used in the Kmetijske comparisons.","marker":"[kme02]"},{"why":"Supplies the children's magazine corpus used in the Ciciban comparisons.","marker":"[Cic45]"}],"fun_headline_variants":["No universal OCR winner for Slovene folklore texts","OCR choice hinges on era, layout, and register","olmOCR best for fairy tales, Tesseract+LLM for press","No one-size-fits-all OCR for folklore archives"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire ranking rests on the authors' manual, informal reading of a small set of outputs, with no reported error rates, sample sizes, or second reviewers, so the conclusions stand only if those subjective judgments are an accurate measure of transcription quality.","fun_headline_variants_meta":{"raw":{"variants":["No universal OCR winner for Slovene folklore texts","OCR choice hinges on era, layout, and register","olmOCR best for fairy tales, Tesseract+LLM for press","No one-size-fits-all OCR for folklore archives"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000476,"raw_usage":{"total_tokens":2294,"prompt_tokens":809,"completion_tokens":1485,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":1418}},"tokens_in":425,"tokens_out":1485,"duration_ms":11505,"temperature":1.0,"reasoning_tokens":1418,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:00:01.230020+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A quantitative re-run of the same three corpora against manually verified ground truth, measuring character error rate and the rate at which dialectal words are changed, would settle the claim: if Tesseract plus LLM post-processing matches or beats olmOCR on typewritten fairy tales, or if olmOCR beats Tesseract plus LLM post-processing on Ciciban, the document-sensitive ranking would fail.","supporting_citations":[],"review_version":2}