{"id":"69c3d9a7-0d90-4a0b-8b88-34df76b1a0cd","arxiv_id":"1908.02477","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A character-level neural translation model reconstructs Latin proto-words from Romance cognates with 64 percent exact accuracy and learns sound-change rules consistent with historical linguistics.","lead":"Researchers trained a neural network to reverse-engineer Latin words from their modern Romance descendant words in French, Spanish, Italian, Portuguese, and Romanian. The model reconstructs the ancestor word exactly about 64 percent of the time, and its errors follow the same sound-change patterns historians of language already know.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline numbers rest on an unreleased test set and mostly unverified Wiktionary/eSpeak labels; a systematic label-error audit is needed before the 0.65 edit-distance claim can be taken at face value.","rationale":"The reader's weakest assumption—dataset correctness—is also the most load-bearing condition for the central claim. The paper's own Appendix A.1 limits manual verification to 300 entries and 200 transcriptions, and footnote 6 discloses that the original Ciobanu-Dinu entries are not public. Since the headline numbers are measured against those labels, systematic label noise could shift edit distances, alter the error taxonomy in Table 2, and change the synthetic-rule conclusions. The footnote 7 comparison (0.881 before cleaning vs. 0.612 after cleaning) demonstrates that preprocessing decisions of this kind can move the metric substantially, so the concern is concrete rather than hypothetical. I do not see an internal inconsistency or a reason to reject the paper: the direct comparison on the original dataset, the sample-based quality checks, the released additions, and the synthetic rule test all provide real support. The right outcome remains a conditional acceptance: the paper should be accepted with the requirement that the exact test set be made available or independently audited, and that variance over runs be reported. This matches the reader's CONDITIONAL verdict, so no verdict change is needed.","tokens_in":13924,"tokens_out":9219,"duration_ms":111002,"concrete_test":"Ask the authors to release the exact test-set entries (or a version with the Ciobanu-Dinu entries under a data agreement) and have two annotators independently verify every test entry: (i) each daughter form is a true inherited cognate of the Latin etymon, (ii) the Latin gold form is correctly declined or conjugated, and (iii) the eSpeak IPA transcription matches the cited sources. Re-run the trained model on the verified subset and also retrain with 5 random seeds on the full cleaned data. If the verified-subset average edit distance differs from 0.65 by more than ±0.1, or if the seed-to-seed spread exceeds 0.05, the headline comparison is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numbers in Section 6 (average edit distance 0.65, 64.1% exact match) are computed on a test split of 1,055 entries drawn from a dataset whose original 3,218 Ciobanu-Dinu entries are not publicly available (footnote 6) and whose Wiktionary/eSpeak portion was only manually sampled: 300 entries for error counting and 200 IPA transcriptions (Appendix A.1). This is load-bearing because every downstream claim—the comparison to Ciobanu and Dinu (2018), the error taxonomy in Table 2, and the synthetic-rule evaluation—uses these labels as ground truth. If the unverified majority contains systematic errors (e.g., Wiktionary descendants that are learned borrowings rather than inherited cognates, incorrect accusative/infinitive normalization, or eSpeak vowel-length artifacts in the phonetic labels), then the reported edit distances and the 'phonologically meaningful' error patterns could be artifacts of the noise rather than properties of the model. The paper's own footnote 7 shows the direct comparison is sensitive to data cleaning (0.881 vs. 0.612 on the original dataset), confirming that label and preprocessing decisions of exactly this kind move the headline metric by more than 0.2 edit-distance units. The absence of any seed or run variance reporting further means we cannot tell whether 0.65 is a stable estimate or a favorable run.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses automated proto-word reconstruction: given cognate forms in five Romance daughter languages, a character-level neural encoder-decoder predicts the Latin proto-word. The authors contribute a new dataset of 8,799 comparative entries, combining a cleaned version of the Ciobanu and Dinu (2014b) dataset with Wiktionary-derived additions, in both orthographic and eSpeak-transcribed IPA forms. On a held-out test set, the orthographic model achieves an average edit distance of 0.65, normalized edit distance of 0.064, and 64.1% exact reconstructions, which the authors compare favorably to the 1.07 average edit distance reported by Ciobanu and Dinu (2018). The paper also provides an error taxonomy, a synthetic test of 33 documented phonological-change rules, a hierarchical clustering of learned phoneme embeddings, and an attention analysis across languages.","tokens_in":14201,"tokens_out":7255,"duration_ms":77788,"significance":"If the results hold, this is a valuable contribution to computational historical linguistics. The new dataset is substantially larger than existing resources and the cleaning protocol demonstrably improves data quality (43 errors in a 300-entry sample of the original data versus 4 in the cleaned data). The synthetic rule evaluation, in which the model correctly handles 22 of 33 documented sound changes, and the learned-embedding analysis provide evidence that the model internalizes phonologically meaningful regularities rather than merely memorizing surface correspondences. The paper is also refreshingly explicit about the limits of the supervised setting relative to the historical linguist's task. However, the strength of these conclusions is tempered by evaluation-protocol issues: the headline comparison to prior work is cross-dataset, the test set is not fully reproducible from the public release, and the ground-truth labels rest on relatively small manual verification samples.","major_comments":[{"comment":"The abstract and Section 6 claim that neural sequence models outperform conventional methods 'applied to this task so far,' but the only quantitative basis is a cross-dataset comparison to Ciobanu and Dinu (2018), who evaluated on a different, unreleased test set. No baseline is run on the new 1,055-entry test set, so the reader cannot tell whether the 0.65 edit distance reflects model quality or dataset ease. Footnote 7 mitigates the concern by reporting a smaller model on the original dataset (0.881 vs. 1.07), but that is still a comparison across implementations and hyperparameters. Add same-data baselines to Table 1, such as the CRF of Ciobanu and Dinu (2018) re-run on the new dataset or a simple alignment/majority baseline, to support the superiority claim on a common benchmark.","section":"Section 6, Table 1"},{"comment":"The test split of 1,055 entries is not fully reproducible: the release excludes entries that appeared in the original Ciobanu and Dinu (2014b) dataset, and the random split is not seeded or otherwise described. Moreover, all reported metrics come from a single training run, with no variance estimate. Because the headline numbers (0.65 average edit distance, 64.1% exact match) are the paper's central quantitative claim, the authors should report results over multiple random seeds (mean and standard deviation), state the seed used for the split, and either release the full test set (including cleaned original entries if legally possible) or release the model's test-set predictions so the exact numbers can be independently verified.","section":"Section 4 and footnote 6"},{"comment":"The quality of the ground-truth labels is established by manual checks of only 300 entries for error counting and 200 IPA transcriptions out of 8,799 entries and over 41,000 distinct words. The paper's own footnote 7 shows that preprocessing decisions of the kind in question change average edit distance by more than 0.2 units (0.881 vs. 0.612 on the original dataset), so label noise is not a negligible factor. A structured audit of a larger sample, stratified by source (Ciobanu-Dinu vs. Wiktionary) and by language, is needed to rule out systematic errors in the unverified majority, particularly for the Wiktionary additions and the eSpeak vowel-length and quality transcriptions.","section":"Appendix A.1"}],"minor_comments":[{"comment":"The table header contains typos: 'Ortographic' should be 'Orthographic' and 'datsaets' should be 'datasets'.","section":"Table 1"},{"comment":"In the paragraph on 'Other vowel changes,' 'This lef to reconstruction errors' should read 'This led to reconstruction errors'; also, 'A separated case is that of Greek words' should be 'A separate case.'","section":"Section 6.1"},{"comment":"In the discussion of vowel-length datasets, 'lenghts' should be 'lengths' (two occurrences).","section":"Section 6"},{"comment":"The running example uses 'lactem' as the Latin word for 'milk'; classical Latin is 'lac' (genitive 'lactis'). If the dataset normalizes to a later or analogical accusative form, this should be stated explicitly for readers unfamiliar with the normalization.","section":"Section 3"},{"comment":"Training details such as number of epochs, learning rate, optimizer, batch size, and early-stopping criterion are not reported; these details are needed for reproducibility.","section":"Section 5.1"},{"comment":"The figure caption and text should define precisely how the counts are normalized ('with respect to time step, letter frequency and language frequency in the corpus') and state whether the reported values are percentages or ratio scores.","section":"Section 6.5 / Figure 4"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and interesting paper that is likely to be a useful reference for computational historical linguistics. The main gap is that the central quantitative claims are not yet backed by a same-data baseline and a fully reproducible evaluation: the test set is partially unreleased, a single run is reported, and the label audit is based on small samples. These issues are addressable within the scope of a revision (e.g., adding baselines, reporting multiple seeds, releasing predictions, and expanding the manual audit), so I do not see grounds for rejection. I would encourage the editor to request those additions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid NAACL paper and the best current baseline for Romance proto-word reconstruction. The value is in the dataset and the analysis, not the architecture.\n\nWhat's actually new: they clean and extend the Ciobanu-Dinu cognate dataset from 3,218 to 8,799 sets, fix the accusative/infinitive normalization problem, add IPA via eSpeak, and run a standard attention encoder-decoder. The model does well: 64.1% exact reconstruction on the orthographic test set, average edit distance 0.65. The analysis is the real contribution: the error taxonomy is linguistically sensible, the synthetic rule test (22/33) gives an independent check that the model captures documented sound changes, and the embedding clusters look phonologically meaningful. They are also honest about the main confound: the comparison to Ciobanu and Dinu (2018) mixes a new dataset with a new model, and their own footnote reports a smaller model on the original data gets 0.881 while the cleaned data gets 0.612. Credit for stating that directly.\n\nSoft spots, in proportion. First, no variance numbers: single split, no seeds, so we do not know if 0.65 is stable. That is minor since the effect sizes appear large, but it is easy to fix. Second, the dataset is only partly public: their additions are on GitHub, but the original Ciobanu-Dinu entries are not, which makes full reproduction harder than it should be. Third, label quality: the test labels come from Wiktionary and eSpeak with manual verification of only 300 cognate sets and 200 IPA transcriptions. The reported 43 to 4 error reduction in the 300-entry sample is reassuring, but a systematic error in the unverified majority could shift edit distances by a few hundredths. The stress-test note about this is fair but not fatal; the synthetic rule experiment is independent of that label noise and supports the main claims.\n\nWho this is for: computational historical linguists and anyone building benchmarks for low-resource sequence-to-sequence. I would send this to review; it deserves referee time and I would likely accept after minor revisions asking for variance estimates and fuller data release.","headline":"A solid, honest empirical paper that sets a new benchmark for Romance proto-word reconstruction; the headline numbers are plausible but rest on a partially released dataset and no variance reporting.","tokens_in":14752,"tokens_out":2117,"would_cite":true,"duration_ms":23083,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A character-level neural encoder-decoder trained on Romance cognates reconstructs unseen Latin proto-words, achieving 64.1% exact reconstruction and average edit distance 0.65 on the orthographic test set.","keywords":["proto-language reconstruction","comparative method","Romance languages","character-level neural networks","historical linguistics","sound change","cognates","attention mechanism"],"falsifier":"Randomly sample test entries from the parts of the dataset that were not manually verified, have expert linguists check the cognate sets against etymological dictionaries and correct the IPA transcriptions by hand, then recompute the reported edit distances; if the corrections move the average orthographic edit distance from 0.65 by more than a small amount or push exact reconstruction below the earlier baseline, the central claim fails.","tokens_in":13704,"feed_emoji":"🏛️","tokens_out":7741,"duration_ms":79093,"temperature":0.7,"pith_summary":"The paper asks whether the comparative method of historical linguistics—reconstructing an ancestral proto-word from its surviving cognates—can be automated by a character-level neural network. It introduces a dataset of 8,799 Latin-to-Romance comparative entries, with cognates in French, Italian, Spanish, Portuguese and Romanian, in both written and IPA-transcribed forms. On unseen test cognates, the network reconstructs the exact Latin word 64.1% of the time in the orthographic setting, with average edit distance 0.65, improving on earlier reported results; phonetic-input performance is lower mainly because Latin vowel-length distinctions are hard to recover. Error analysis shows the model's mistakes cluster around well-documented sound changes such as high-mid vowel alternations, segment deletion, cluster simplification and morphological regularization, and a synthetic rule test confirms it internalizes many systematic phonological shifts. If this stands, automatic reconstruction can support historical linguists by scaling up proto-lexicon recovery and by highlighting which sound changes are genuinely opaque from daughter languages alone.","feed_headline":"Neural net rebuilds Latin proto-words with 64% exact accuracy","feed_subtitle":"A character-level sequence model beats earlier methods and learns documented sound changes along the way.","key_machinery":"The model is a character-level encoder-decoder with attention, built from GRU networks of 150 cells. The encoder processes the cognate set as a sequence of characters from all daughter languages, each character represented as a shared character embedding combined with a language-embedding vector, so that the same written letter can mean different sounds in different languages. The decoder generates the Latin proto-word character by character, using dot-product attention to pick out the most relevant cognate characters at each step, and an MLP with 200 hidden units produces the next character. The evaluation machinery is edit distance between the predicted and gold Latin form, reported as exact-match rate, average edit distance, and normalized edit distance.","core_discovery":"The central claim is that a supervised character-level encoder-decoder with attention can perform proto-word reconstruction well enough to beat earlier computational approaches, and that the model's internal representations and errors correspond to real historical phonology. Trained on sets of cognates tagged with their languages, the model reads the daughter forms and emits the Latin ancestor one character at a time. On the orthographic test set it reaches 64.1% exact reconstruction, 84.0% within one edit operation, with average edit distance 0.65 and normalized distance 0.064; the phonetic variant reaches 50.0% exact and normalized distance 0.100, with the gap largely explained by Latin tense-lax vowel contrasts that daughter languages neutralize. The paper further claims that roughly 80% of orthographic and 75% of phonetic errors fall into linguistically named categories, that the network correctly predicts 22 of 33 documented phonological-change rules on a synthetic test, and that its learned phoneme embeddings cluster into a phonologically meaningful hierarchy (vowels vs. consonants, voiced vs. voiceless pairs, allophones) without explicit supervision.","pith_inferences":["If the pattern generalizes, the same architecture could be tested on other attested proto-languages, such as Germanic or Slavic families, to see whether the 64% exact-reconstruction rate is typical or specific to the conservative orthography and close relatedness of Romance languages.","Because the model's success depends on having access to the full cognate set, a natural next step is to couple it with automatic cognate detection and measure how reconstruction accuracy degrades as cognate sets become noisier or incomplete.","The attention analysis suggests a concrete, testable extension: if a daughter language is most attended at a position, removing its cognate should hurt reconstruction at that position more than removing a low-attention language's cognate; this prediction could be checked with ablation experiments.","If model errors reliably flag opaque changes, the same approach could be used to triage large dictionaries for human etymological review, making the comparative method faster without replacing the linguist."],"forward_implications":["For well-attested language families, large-scale reconstruction of proto-lexicons becomes a semi-automatic process: linguists supply or verify cognate sets, and the model proposes ancestor forms that can be checked.","The systematic error categories make the model a diagnostic instrument: where reconstruction fails, the failure identifies sound changes that daughter languages have made opaque, indicating where extra evidence such as inscriptions or other branches is needed.","The phonologically structured embeddings suggest that training on reconstruction tasks is a way to learn phoneme taxonomies without labeled data, which could transfer to under-resourced languages.","The public 8,799-entry dataset gives the field a common benchmark, so future reconstruction methods can be compared on the same cognate sets and metrics."],"supporting_citations":[{"why":"Supplies the initial 3,218 complete cognate sets that the new dataset extends and cleans.","marker":"(Ciobanu and Dinu, 2014b)"},{"why":"Reports the CRF-based baseline (average edit distance 1.07, normalized 0.13, 50% exact) that the paper compares against.","marker":"(Ciobanu and Dinu, 2018)"},{"why":"Provides the attention-based encoder-decoder architecture on which the reconstruction model is built.","marker":"(Bahdanau et al., 2015)"},{"why":"Supplies the GRU encoder-decoder formulation used for the character-level model.","marker":"(Cho et al., 2014)"},{"why":"Earlier probabilistic model of sound change on phylogenetic trees, the main alternative approach the paper situates itself against.","marker":"(Bouchard-Côté et al., 2013)"},{"why":"Sound-change charts used to identify problems in the existing dataset and to interpret the model's error categories.","marker":"(Boyd-Bowman, 1980)"},{"why":"Reference for Latin vowel length and tense-lax contrast, used to explain the phonetic-task performance gap.","marker":"(Allen and Allen, 1989)"},{"why":"Etymological dictionary used for manual verification of Wiktionary-derived cognates.","marker":"(DIEZ and Donkin, 1864)"}],"fun_headline_variants":["Neural net rebuilds Latin proto-words, learns sound shifts","AI reconstructs extinct Latin words, beating traditional methods","64% exact: neural model revives Latin proto-words","Deep learning recovers Latin ancestors and documented phonology","Proto-Latin rebuilt by neural nets, with realized sound rules"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results stand on the dataset's correctness: the Wiktionary-derived cognate sets were manually checked only in part and the automatic IPA transcriptions were verified on only 200 words, so systematic errors in the unverified entries would distort every reported reconstruction score and every conclusion about learned sound changes.","fun_headline_variants_meta":{"raw":{"variants":["Neural net rebuilds Latin proto-words, learns sound shifts","AI reconstructs extinct Latin words, beating traditional methods","64% exact: neural model revives Latin proto-words","Deep learning recovers Latin ancestors and documented phonology","Proto-Latin rebuilt by neural nets, with realized sound rules"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000589,"raw_usage":{"total_tokens":2738,"prompt_tokens":894,"completion_tokens":1844,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":1761}},"tokens_in":510,"tokens_out":1844,"duration_ms":14088,"temperature":1.0,"reasoning_tokens":1761,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:42:25.740588+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomly sample test entries from the parts of the dataset that were not manually verified, have expert linguists check the cognate sets against etymological dictionaries and correct the IPA transcriptions by hand, then recompute the reported edit distances; if the corrections move the average orthographic edit distance from 0.65 by more than a small amount or push exact reconstruction below the earlier baseline, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports the CRF-based baseline (average edit distance 1.07, normalized 0.13, 50% exact) that the paper compares against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier probabilistic model of sound change on phylogenetic trees, the main alternative approach the paper situates itself against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sound-change charts used to identify problems in the existing dataset and to interpret the model's error categories."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reference for Latin vowel length and tense-lax contrast, used to explain the phonetic-task performance gap."}],"review_version":1}