{"id":"e296462e-a43b-4d01-9914-043fe011846d","arxiv_id":"2509.03510","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A 2,128-writer Persian handwriting database with coded family relationships is presented, along with preliminary tests using image features to find similar handwriting among relatives.","lead":"This paper introduces a new Persian handwriting database collected from 2,128 members of 210 families, with coded links between relatives stored in the metadata. It is meant as a resource for testing whether handwriting style runs in families.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Family-relationship labels are unverified; if wrong, every heritability analysis built on the database is invalid.","rationale":"This is a database paper, so the central claim is that the resource is new and usable for studying handwriting heritability. That claim is load-bearing on the accuracy of the family-relationship metadata. The reader's weakest assumption correctly identifies this: the paper provides no verification that the declared pedigrees match true genetic relatedness. The coding scheme is well-specified, but the reliability of self-reported kinship in a large extended-family collection is an empirical risk. If labels are wrong, the database's core novelty collapses; no amount of better feature extraction or more sophisticated similarity analysis can fix mislabeled ground truth. I considered the alternative concern that Section 5's similarity-detection experiment lacks a baseline and statistical support. That is real, but it is less load-bearing: the experiments are explicitly framed as initial demonstrations, and a weak experiment does not invalidate the database itself. In contrast, label errors directly undermine the database's primary purpose. The proposed re-contact test is a concrete way to estimate the label error rate, and the complementary inter-family baseline would also settle whether the stated 'similarities' claim has empirical support. Since the reader already requested conditional acceptance, my analysis does not shift the verdict.","tokens_in":14196,"tokens_out":4470,"duration_ms":50779,"concrete_test":"Independently re-verify the family metadata for a random subset of at least 30 of the 210 families. Re-contact each center and, without showing the original forms, ask them to redraw the family tree and list each member's relationship to the center. Compare this re-collected genealogy to the database's encoded relationship codes. Count disagreements that change kinship type (blood vs. non-blood, first-degree vs. second-degree). If any such mismatch is found, the label error rate is non-zero and the heritability metadata is unreliable; if no mismatches appear, the concern is substantially mitigated. A complementary check: using the same DGF/Euclidean protocol, compare intra-family distances to distances between matched unrelated writers; the 'similarities detected' claim requires intra-family distances to be significantly smaller.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The database's unique value is the family-relationship metadata: each writer is linked to the family center via codes such as '0_2.1_4.1' that are embedded in image filenames and GT files. The paper relies on 210 volunteer centers to self-report which relatives share 'common genetic roots' and what the exact relationship is. Section 2 says non-genetic relatives were excluded, but no procedure verifies genetic relatedness—no pedigree documentation, no DNA testing, no independent audit of the family trees. The coding schema in Table 1 is internally consistent, but internal consistency does not imply biological correctness. In heritability research, mislabeling (a step-parent coded as father, an adopted sibling as sister, or recall errors in a large extended family) creates systematic errors that can inflate or erase familial resemblance. Every downstream analysis—family correlations, heritability estimates, forensic comparison—depends on these labels being true. The initial experiments in Section 5 cannot catch such errors because they compare only within declared families and lack a non-family control baseline. Thus the most load-bearing assumption is not the DGF feature or the Euclidean distance, but the accuracy of the relationship codes themselves.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a Persian offline handwriting database built around recorded family relationships. The resource contains handwritten digits, letters, geometric shapes, and free paragraphs from 2,128 writers belonging to 210 families, with relationship metadata encoded in filenames and ground-truth files. A family-relationship coding scheme, a form-extraction pipeline, and the database structure are described. Initial experiments compare within-family handwriting using DGF features and Euclidean distance on text paragraphs, and the most/least similar pairs are shown for four families. The authors claim this is the first comprehensive database enabling research on heritability and family effects on handwriting.","tokens_in":14505,"tokens_out":6039,"duration_ms":61151,"significance":"If the resource is as described, it is a genuinely useful and potentially unique contribution: a freely available multi-modal handwriting corpus with family-relationship metadata, which could open new directions in writer identification, forensic document examination, and studies of handwriting heritability. The relationship coding scheme is well organized, and the data collection effort is substantial. However, the experimental demonstration of familial similarity is currently anecdotal, and the validity of the family-relationship labels is a central risk that the manuscript does not address. These issues can be remedied, but they require substantive revision.","major_comments":[{"comment":"The abstract's claim that 'similarities among their features and writing styles are detected' is supported only by four selected visual examples. No aggregate distance distributions, statistical tests, confidence intervals, or comparisons against unrelated writer pairs are provided. Without a baseline, showing the most similar family member is uninformative, because a nearest neighbor exists for any set of pairs. The authors should either remove the claim or replace it with quantitative evidence, e.g., within-family versus between-family distance distributions and a permutation test.","section":"Section 5, Figs. 13-16"},{"comment":"The experimental protocol contains post-hoc choices: the 100-zone split was selected 'after conducting several experiments,' and Euclidean distance was chosen because it was 'more consistent with human eyes' verification.' These choices were therefore made after seeing the results, making the demonstration potentially overfitted and not a confirmatory test. A valid protocol should preselect the feature and distance measure, or validate the choices on a held-out subset of families.","section":"Sections 5.1 and 5.3"},{"comment":"The database's unique value rests on the accuracy of family-relationship labels, yet these labels are self-reported by 210 volunteer centers and are not independently verified. The paper states that members without 'common genetic roots' were excluded, but no pedigree documentation, DNA validation, or audit protocol is described. Systematic mislabeling (e.g., step-parent recorded as parent, adopted sibling as sibling) would invalidate downstream heritability analyses. The authors should provide a validation protocol, acknowledge this limitation, and explain the expected impact of label error.","section":"Sections 2 and 4.2"},{"comment":"The reported counts are internally inconsistent with the multiple-collection scheme. With 210 centers filling two forms at T1, T2, and T3, and the remaining 1,918 writers filling two forms once, the total is 5,096 forms, not 4,256 (=2,128×2). Similarly, Table 3 reports 2,128 digits per digit class, but centers wrote digits at three time points, so the per-digit count should be 2,548 unless T2/T3 samples are excluded. The typo '=2218*2' compounds the problem. All database statistics should be reconciled and rechecked.","section":"Sections 2, 3.3, 4.1.1, Table 3"}],"minor_comments":[{"comment":"The feature is called 'DGF' but the text repeatedly says 'DFG operator.' Please correct the typo.","section":"Section 5.2"},{"comment":"The table appears to pair Persian letter names with incorrect glyphs: 'Che' is shown with آ, 'Yeh' with ش, 'Shin' with ق, 'Ghaf' with ه, 'He' with ی, and 'Alef' with چ. If this is only a table-layout error, it should be fixed; if it reflects the actual labeled data, it is a serious labeling error.","section":"Table 4"},{"comment":"'Grand-Truths' in the section title is a typo for 'Ground-Truths.'","section":"Section 4.2"},{"comment":"'compressive publicly available database' should be 'comprehensive.'","section":"Section 7"},{"comment":"The code for the mother is written as '0_1.1.' with a trailing dot; the encoding description should be made consistent.","section":"Section 2.1"},{"comment":"The footnote says only 'A sample version of this database can be downloaded.' Please clarify whether the full database is freely available, and how researchers can obtain it.","section":"Footnote 1"}],"recommendation":"major_revision","confidential_remarks":"The database itself has clear potential value, and the family-relationship coding schema is a strong contribution. The main risks are the unverified relationship labels and the anecdotal experimental evidence. If the authors can provide a quantitative similarity analysis with appropriate baselines and a clear statement of the kinship-label uncertainty, the manuscript could become acceptable. The numeric inconsistencies in the database counts and the Table 4 glyph pairing should be corrected before a revised submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This one is unusual: the contribution is the database, not the experiments. The family-relationship coding schema (e.g., '0_2.1_4.1' for a paternal uncle relative to the center) is genuinely new as far as I can tell from their comparison table—no prior handwriting database records family links. The scale is real: 210 families, 2,128 writers, digits, letters, shapes, and paragraphs, with three image modalities and rich GT. For OCR and writer-identification researchers working on Persian, this could be a useful resource, and the multi-month re-sampling of centers (T1–T3) is a nice touch.\n\nWhat the paper does well: the collection protocol is described in enough detail to be replicable; the counts in Tables 3–5 are consistent; the coding scheme is elegant and appears to reconstruct the family tree; and the comparison table is honest about prior work. I believe the novelty claim.\n\nThe soft spots are real but not fatal, in proportion.\n\nFirst, the demonstration experiment is anecdotal. They compute Euclidean distances between center and family members, then show four examples of 'most similar' and 'most dissimilar' rather than reporting distributions or comparing against unrelated writers. The choice of 100 zones and Euclidean distance was made after seeing results, and the phrase 'more consistent with human eyes verification' is post-hoc. This does not invalidate the database, but the abstract's phrase 'similarities ... are detected' overstates what the experiment shows. It's more like: using a simple feature, some pairings look similar to a human.\n\nSecond, the relationship labels are self-reported. The centers were trained, and non-genetic relatives were excluded, but there is no DNA testing or pedigree audit. For studies of family resemblance (shared environment plus genes), this may be acceptable; for heritability claims, it's a serious caveat that should be stated explicitly and ideally validated on a subset.\n\nThird, typo-level: Section 4.1 says 4,256 = 2218*2, but likely should be 2,128*2. Minor.\n\nThe paper is for the handwriting and forensic communities, especially those working in Persian and other scripts who want to ask questions about familial influence on writing. It deserves a serious referee, because the resource itself fills a gap. The referee should ask for a more rigorous validation of the relationship labels and a proper experiment with baseline comparisons before accepting broad claims about inheritance. I'd recommend sending to peer review.","headline":"A genuinely new family-relationship-labeled handwriting database, but the similarity evidence is anecdotal and the relationship labels rely on untested self-report.","tokens_in":14910,"tokens_out":2186,"would_cite":false,"duration_ms":22242,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new handwriting database records family trees for 2,128 Persian writers, enabling tests of whether handwriting is inherited.","keywords":["offline handwriting recognition","Persian handwritten database","heritability of handwriting","family relationship coding","writer identification","directional gradient feature","handwriting similarity"],"falsifier":"Take a random subset of writer pairs labeled as close relatives in the database and test their actual genetic relatedness with documented family records or DNA markers; if a substantial fraction of high-similarity pairs are not biologically related, or if unrelated control pairs matched for education, age, and region show the same feature distances as declared relatives, the paper's central premise—that the database measures family/genetic effects—fails.","tokens_in":14138,"feed_emoji":"✍️","tokens_out":6624,"duration_ms":66195,"temperature":0.7,"pith_summary":"The paper's contribution is a resource: the first handwriting database designed so that family relationship is part of the ground truth for every sample. It gathers digits, alphabet letters, geometric shapes, and free copies of the same Persian paragraph from 2,128 members of 210 families, and records each writer's relation to a chosen family 'center' with a compact code. The authors also run a feasibility experiment on the free texts using directional gradient features and Euclidean distance; in every family they can find a relative whose handwriting resembles the center's, which they read as evidence that family-level similarity is detectable in the data. If the database is sound, it matters because applications such as writer identification, signature verification, and forensic document examination currently treat handwriting similarities as individuality signals, without a way to control for family resemblance or to study it.","feed_headline":"Handwriting database ties 2,128 Persian writers to family trees","feed_subtitle":"Digits, letters, shapes, and texts carry coded family relationships so researchers can test handwriting heritability.","key_machinery":"The load-bearing mechanism is the center-based family-relationship coding schema. Each family is drawn as a tree hung from one designated 'center'; first-degree relations get single codes (0 center; 1 mother; 2 father; 3 sister; 4 brother; 5 daughter; 6 son; 7 wife; 8 husband), and more distant relatives are encoded by chaining these with underscores and dot-multiplicity markers—for example '0_2.1_4.1' names the center's father's first brother. This code is baked into every extracted image file name, so any digit, letter, shape, or text can be queried by kinship. Around this, the paper wraps a form design with corner markers that allow skew correction and field extraction, and a preliminary","core_discovery":"The central claim is that a handwriting database can be, and has been, built in which family relationship is a first-class ground-truth attribute for every sample. For each of 210 families, one 'center' person recruited relatives, and all of them completed two specially designed forms; members with no common genetic roots with the center were excluded. The database contains 21,280 digits, 68,096 alphabet letters, 17,024 geometric shapes, and 2,128 free-text paragraphs, each stored in true-color, grayscale, and binary formats. Every extracted image is named with a code that encodes the writer's relationship to the center, such as '0_2.1_4.1' for the center's father's first brother. Using dire","pith_inferences":["The strongest untested version of the paper's own goal would compare declared relatives with matched unrelated writers from the same region, school, and age; without such a control, the observed family resemblance could be caused by shared environment and handwriting instruction rather than shared genes.","The relationship codes are self-reported and explicitly exclude non-genetic relatives, but the paper provides no pedigree or DNA verification; before this dataset is used to claim heritability, a validation subset checking a sample of the coded links would be needed.","A direct heritability analysis using the coding schema's degree of kinship (siblings vs. cousins vs. grandparent–grandchild) could turn the database from a demonstration of family resemblance into an estimate of how resemblance scales with genetic distance.","Because the same paragraph was copied by all writers, controlling content makes the feature comparison clean, but it also conflates copying skill with intrinsic handwriting; an online tablet version, which the authors mention as future work, could add kinematic features that distinguish them."],"forward_implications":["Writer-identification and verification systems can now be evaluated on the harder case of distinguishing relatives, using the relationship labels to measure false-match rates among family members.","Forensic document examiners gain a public benchmark for asking whether a disputed sample could have been produced by a different member of the same family.","The relationship-coding protocol is script-independent, so equivalent family-aware handwriting databases can be built for Arabic, Latin, or other scripts and compared.","Since centers wrote the same forms three times over consecutive months, the database also enables separating within-writer variability from between-relative variability."],"supporting_citations":[{"why":"Supplies the DGF feature extraction method and the prior Persian handwriting database that this work extends with family relationships.","marker":"[40]"},{"why":"A widely used large Farsi digit database serving as a benchmark comparison for this database's digit collection.","marker":"[38]"},{"why":"A standard English off-line text database used in the comparison table to show existing corpora lack family metadata.","marker":"[30]"},{"why":"A major Arabic offline handwritten text database used in the comparison table, again lacking family relationships.","marker":"[35]"},{"why":"Establishes handwriting individuality via feature analysis, motivating the within-family similarity experiment.","marker":"[23]"},{"why":"Cites behavioral-genetics evidence for genetic and environmental influences on writing, motivating heritability questions.","marker":"[11]"}],"fun_headline_variants":["Persian handwriting database maps genetic ties across 210 families","210 families' handwriting reveals heritability clues","New database links handwriting styles to family trees","Genetic handwriting? 210-family Persian database aims to find out","Handwriting heritability: 2,128 samples from 210 families released"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The database's entire purpose rests on the assumption that the self-reported family-tree codes match actual biological relatedness; the paper excludes non-genetic relatives from collection but does not verify pedigrees or DNA, so if those labels are wrong, every heritability conclusion built on the data collapses.","fun_headline_variants_meta":{"raw":{"variants":["Persian handwriting database maps genetic ties across 210 families","210 families' handwriting reveals heritability clues","New database links handwriting styles to family trees","Genetic handwriting? 210-family Persian database aims to find out","Handwriting heritability: 2,128 samples from 210 families released"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000534,"raw_usage":{"total_tokens":2385,"prompt_tokens":702,"completion_tokens":1683,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":1603}},"tokens_in":446,"tokens_out":1683,"duration_ms":12147,"temperature":1.0,"reasoning_tokens":1603,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:50:19.718731+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random subset of writer pairs labeled as close relatives in the database and test their actual genetic relatedness with documented family records or DNA markers; if a substantial fraction of high-similarity pairs are not biologically related, or if unrelated control pairs matched for education, age, and region show the same feature distances as declared relatives, the paper's central premise—that the database measures family/genetic effects—fails.","supporting_citations":[{"cited_title":"Solimanpour, J","cited_arxiv_id":null,"evidence_quote":"Supplies the DGF feature extraction method and the prior Persian handwriting database that this work extends with family relationships."},{"cited_title":"Ziaratban, K","cited_arxiv_id":null,"evidence_quote":"A widely used large Farsi digit database serving as a benchmark comparison for this database's digit collection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A major Arabic offline handwritten text database used in the comparison table, again lacking family relationships."},{"cited_title":"Srihari, S","cited_arxiv_id":null,"evidence_quote":"Establishes handwriting individuality via feature analysis, motivating the within-family similarity experiment."},{"cited_title":"Olson, J","cited_arxiv_id":null,"evidence_quote":"Cites behavioral-genetics evidence for genetic and environmental influences on writing, motivating heritability questions."}],"review_version":1}