{"id":"1eff788e-963a-4a85-9cfb-bd3ef37e8ed9","arxiv_id":"2608.05850","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MameLoshnLM, trained by continuing pretraining Llama 3.1 8B on a curated Yiddish corpus, outperforms similar-scale open models on a new multi-task Yiddish benchmark and better retains Yiddish-specific loshn-koydesh vocabulary and morphology than general multilingual models.","lead":"The authors release MameLoshnLM, an 8B Yiddish language model, along with a curated Yiddish pretraining corpus (Oytser) and a multi-task benchmark (Kashes). A generalist might read it because it shows how noisy web data can degrade a low-resource language model and how targeted data curation can recover native lexical and morphological patterns.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark claim depends on uncontaminated Kashes tasks, but only Kashes-mt source documents are excluded from Oytser; overlap of WikiANN and YiTB test texts with Oytser is not audited.","rationale":"The reader's weakest assumption is that Kashes is a valid and uncontaminated measure of Yiddish competence, and that is also the most load-bearing premise I can identify. The paper's strongest evidence for genuine competence is its benchmark superiority and the linguistic probes; both routes pass through tasks whose training-data overlap is unverified. The Kashes-mt exclusion is explicit, but the paper is silent about WikiANN, which is definitionally derived from the Yiddish Wikipedia that Oytser includes, and about YiTB, whose corpus provenance is not audited against Oytser. I considered the single-seed few-shot issue as an alternative, but the 5-shot average margin (62.6 vs. 56.8 for Llama and 57.0 for Gemma) is large enough that contamination is the more decisive threat to the central claim. I also noted that raw pretraining text does not contain annotation labels, so overlap is not automatically fatal for labeled tasks; that is precisely why the magnitude of overlap must be measured rather than assumed. The reader's CONDITIONAL verdict is appropriate: the benchmark numbers should not be taken at face value until the overlap audit is reported. My read therefore leaves the verdict unchanged.","tokens_in":27569,"tokens_out":11313,"duration_ms":122187,"concrete_test":"Run a document- and sentence-level overlap audit between Oytser and every Kashes test set, with special attention to WikiANN's 300 sentences and YiTB's 1,079 sentences. Use normalized 13-character shingles after script normalization and de-diacritization, and report per-dataset exact and near-duplicate rates. Then re-run the Table 4 evaluation for all models on the non-overlapping subset. If overlap exceeds roughly 1% of a test set, or if removing overlapping items changes any reported MameLoshnLM score by more than 1–2 points, the benchmark-level claim should be conditioned on this decontamination; if no overlap is found, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that MameLoshnLM outperforms comparable open models on Kashes and that the gain reflects genuine Yiddish competence. The only explicitly decontaminated evaluation task is Kashes-mt: Section 4.1 and Figure 1 state that every source document was excluded from Oytser. No equivalent audit is reported for WikiANN, which Table 3 describes as tagged entities from the 2018 Yiddish Wikipedia dump, even though Oytser contains the Yiddish Wikipedia as a training source (Table 2, 15.3K documents). Similarly, the YiTB treebank tasks (POS tagging, dependency parsing, lemmatization, transliteration) are not checked against Oytser's Yiddish Book Center corpus or Wikipedia split. If WikiANN or YiTB test sentences appear in pretraining, the model can exploit surface familiarity with the test distribution rather than generalizable linguistic competence, which would inflate the Yiddish-centered columns of Table 4 and the lemmatization-based morphological probes in Section 8. The central claim might survive a decontamination re-run, but the paper currently does not provide the overlap analysis needed to know whether the reported numbers are clean.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MameLoshnLM, an 8B-parameter Yiddish language model obtained by continued pretraining of Llama 3.1 8B on Oytser, a newly assembled Yiddish corpus that combines web-native sources with digitized Yiddish books. The paper also introduces Kashes, a multi-task benchmark spanning translation, linguistic analysis, named entity recognition, and machine-translated language-understanding tasks. The central reported result is that MameLoshnLM outperforms five open baselines of similar scale on Kashes on average, with the strongest gains on Yiddish-centered tasks, and that the gains reflect genuine Yiddish competence as shown by lexical (loshn-koydesh) and morphological probes. A secondary contribution is an audit of mC4's Yiddish split, which the paper claims contains substantial machine-translated and misidentified-Hebrew content.","tokens_in":27674,"tokens_out":7145,"duration_ms":66520,"significance":"If the benchmark numbers are clean, this is a substantial contribution to low-resource NLP: it provides the first open Yiddish LLM, a high-quality pretraining corpus, a multi-task Yiddish benchmark, and a replicable continued-pretraining recipe. The paper is unusually transparent: the model, corpus, and benchmark artifacts are released; the mC4 audit is documented in detail; the self-built Kashes-mt benchmark is explicitly decontaminated by excluding all source documents from the pretraining corpus; and the morphological probes include a well-designed within-task control (regular vs. suppletive auxiliary). The mC4 audit and the loshn-koydesh analysis are valuable independent findings that go beyond the benchmark numbers.","major_comments":[{"comment":"The benchmark-level claim of genuine Yiddish competence depends on uncontaminated test sets, but the paper only reports decontamination for Kashes-mt (§4.1: \"we exclude every source document from the Oytser pretraining corpus\"). No overlap audit is reported for WikiANN or the YiTB treebank tasks, even though Table 3 describes WikiANN as \"tagged entities from the 2018 Yiddish Wikipedia dump\" and Table 2 lists the Yiddish Wikipedia as a training source in Oytser (15,300 documents). The YiTB tasks (POS, dependency parsing, lemmatization, transliteration) could similarly overlap with the Wikipedia or YBC portions of Oytser. If any test sentences or documents appear in pretraining, the reported scores in Table 4 (e.g., WikiANN F1 59.7, POS 88.6, lemmatization 31.9) and the morphological probes in Section 8 (Tables 5 and 12) could be inflated by surface memorization rather than reflecting generalizable competence. This is load-bearing for the paper's central claim that continued pretraining yields genuine linguistic gains rather than benchmark overfitting. The authors should perform and report an explicit overlap analysis (exact and near-duplicate document and sentence matching, e.g., 8-gram overlap) between Oytser and each Kashes test set, and re-report the affected numbers with any overlapping instances removed, or otherwise demonstrate that overlap is negligible.","section":"Section 4, Table 3; Section 3.2, Table 2"}],"minor_comments":[{"comment":"Several row-level gaps between MameLoshnLM and the closest baseline are smaller than one point (e.g., dependency parsing LAS 40.6 vs. 40.3 for Qwen3; POS 88.6 vs. 87.6 for Gemma-2). Reporting confidence intervals, standard deviations across evaluation runs, or a significance test for these comparisons would strengthen the claim of consistent superiority.","section":"Section 7.1, Table 4"},{"comment":"The mC4 audit's machine-translation classification relies on URL fingerprints and a locale-sibling count, but the threshold for flagging a domain as machine-translated is not precisely stated beyond the qualitative description \"fewer than one\" language edition for native domains. A precise cutoff and the distribution of the sibling-count metric would improve reproducibility.","section":"Section 3.1, Appendix D"},{"comment":"The LK content-word rate uses the number of content tokens as the denominator, but compound LK phrases are counted as a single item in the numerator. The counting convention for the denominator should be stated explicitly so that it is clear whether compound components are also excluded from the denominator.","section":"Section 8, Appendix E.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is strong in its resource contributions and transparency, and the Kashes-mt decontamination step shows good practice. The main gating concern is the lack of an overlap audit for WikiANN and YiTB relative to Oytser; this is directly fixable by the authors and should be requested before acceptance. The mC4 audit, while somewhat subjective, is carefully documented and is a valuable contribution on its own."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine resource paper, not a hype job. First open Yiddish LM, a curated 5.3B-token corpus, a nine-task benchmark, and a clean diagnostic showing general multilingual models underproduce loshn-koydesh vocabulary and miss Yiddish-specific morphology. The mC4 audit is careful, the Kashes-mt construction is transparent, and the statistical testing is appropriate. The claim that continued pretraining on authentic Yiddish recovers native-like lexical and morphological competence is well supported by the probes.\n\nThe soft spots are real but concentrated. The biggest is contamination oversight: WikiANN is built from the 2018 Yiddish Wikipedia dump, and Oytser contains the Yiddish Wikipedia as a training source. YiTB treebank tasks are not audited against the Yiddish Book Center corpus or the Wikipedia split either. Only Kashes-mt source documents are explicitly excluded. The paper should either report an overlap analysis or re-run those tasks with contamination control. This is a genuine gap, but it does not break the central translation and LK findings, which use Kashes-mt and external lexicons.\n\nSecond, the paper claims English mixing preserves base capabilities but reports no English or general benchmark. One number on MMLU or a similar general benchmark would settle it. Third, few-shot results come from a single seed and are unstable at 1- and 3-shot; the headline 5-shot table is the load-bearing one, and the spread (e.g., dependency parsing LAS from 14.8 at 1-shot to 40.6 at 5-shot) suggests variance that deserves reporting.\n\nThat said, this deserves a serious referee. The resources are useful, the analyses are honest, Appendix D admits subjectivity in the mC4 audit, and the central finding—curated native text beats noisy web data for low-resource adaptation—is important beyond Yiddish. The decontamination issue is addressable in revision, and the paper provides the training corpus and benchmark publicly, so the verification work is feasible.\n\nRecommendation: send to peer review; expect revision on contamination audit and an English capability check.","headline":"Solid first Yiddish LM/corpus/benchmark worth refereeing; the main fix is a contamination audit for WikiANN and YiTB, which the paper currently lacks.","tokens_in":28329,"tokens_out":1505,"would_cite":true,"duration_ms":14961,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Continued pretraining on curated Yiddish text beats multilingual models on Yiddish tasks.","keywords":["Yiddish language model","continued pretraining","low-resource NLP","evaluation benchmark","data contamination","loshn-koydesh lexicon","morphological competence","mC4 audit"],"falsifier":"Check for verbatim or near-verbatim overlap between the Oytser pretraining corpus and the Kashes test sets beyond Kashes-mt: if WikiANN test sentences (drawn from the 2018 Yiddish Wikipedia) or YiTB treebank sentences appear in the corresponding Oytser sources, the reported NER and linguistic-analysis scores are inflated. A retraining experiment with those overlapping documents removed would show whether the gains survive.","tokens_in":27255,"feed_emoji":"📖","tokens_out":6194,"duration_ms":49691,"temperature":0.7,"pith_summary":"This paper tries to establish that a dedicated Yiddish language model, built by continuing the pretraining of Llama 3.1 8B on a newly curated Yiddish corpus called Oytser, outperforms open multilingual models of similar size on a new multi-task Yiddish benchmark called Kashes. Because the benchmark gains concentrate on Yiddish-centered tasks, and because the model produces loshn-koydesh (Hebrew/Aramaic-origin) vocabulary and Yiddish-specific inflection at rates closer to native text, the authors argue the gains reflect genuine linguistic competence rather than generic scale. The paper also reports an audit of the mC4 corpus's Yiddish split showing that less than half of its documents are genuine Yiddish, with the rest machine-translated or misidentified Hebrew. If the claims hold, Yiddish gets its first open language model, a reusable pretraining corpus, a multi-task evaluation resource, and a replicable recipe for languages that are historically rich but digitally underrepresented.","feed_headline":"Curated Yiddish text beats multilingual models on Yiddish tasks","feed_subtitle":"A new corpus and benchmark show the gains come from native vocabulary and morphology, not scale.","key_machinery":"The load-bearing mechanism is the combination of Oytser, a corpus that mixes contemporary web-native Yiddish with more than 12,000 OCRed books from the Yiddish Book Center, and continued pretraining of Llama 3.1 8B on a Yiddish-dominant mixture (72% Yiddish words, 28% English). The evaluation and analysis are carried by Kashes, a nine-task benchmark spanning translation, linguistic analysis, NER, and understanding, together with two probes: a loshn-koydesh lexicon for measuring Hebrew/Aramaic-origin vocabulary in translation output, and lemmatization on the UD Yiddish treebank broken down by morphological category, including a within-task control comparing the suppletive auxiliary 'zayn' with the regular 'hobn'.","core_discovery":"The central claim is that targeted resource construction plus continued pretraining can substantially narrow the gap between what general multilingual models achieve in a low-resource language and what a language-specific model achieves. Specifically, MameLoshnLM, trained for one epoch on roughly 5.3 billion Yiddish tokens with a small English auxiliary mixture, reaches an average score of 62.6 on the Kashes benchmark, ahead of Llama 3.1 8B (56.8) and Gemma-2 9B (57.0), with its largest leads on English-to-Yiddish translation, part-of-speech tagging, dependency parsing, transliteration, and two of three named-entity-recognition tasks. The authors further claim that the advantage is qualitative as well as quantitative: the model recovers the loshn-koydesh lexical layer that general models deplete, and handles Yiddish-specific morphology such as ge- participles, Hebrew-origin plurals, and the suppletive auxiliary 'zayn' much better than the base model, which tends to copy surface forms unchanged.","pith_inferences":["A contamination audit beyond Kashes-mt would directly test the benchmark-level claim: WikiANN is built from the 2018 Yiddish Wikipedia and the YiTB treebank is a curated literary resource, and the paper does not report whether their test sentences also appear in Oytser; if they do, the reported NER and linguistic-analysis gains would be partly attributable to memorization.","The same methodology—an mC4-style audit, a curated corpus, continued pretraining, and lexicon-plus-morphology probes—could be ported to other languages with a classical or liturgical lexical layer, such as Ladino or Judeo-Arabic, to test whether the loshn-koydesh depletion pattern generalizes.","The per-word LK recall gap suggests a concrete failure mode for evaluation: COMET scores correlate only weakly with LK recall, so translation quality metrics may miss systematic loss of heritage vocabulary; other low-resource benchmarks may need lexicon-based metrics to catch this."],"forward_implications":["Yiddish gains an open 8B model, a high-quality pretraining corpus, and a multi-task benchmark, enabling downstream work in Yiddish NLP and digital humanities.","The mC4 audit establishes that labeled language splits in web-scale corpora can be dominated by machine-translated spam and script-confused text; similar audits for other low-resource languages would likely be prudent before training on them.","Keeping a small English share in the pretraining mixture preserves base-model capabilities better than reallocating that share to historically related languages such as German and Hebrew.","The linguistic probes show that a language-specific model produces native-like lexical and morphological patterns at much higher rates, suggesting that curated native data can recover competence that noisy web data erodes."],"supporting_citations":[{"why":"Supplies mC4, the multilingual corpus whose Yiddish split is audited and found to be mostly machine-translated or misidentified Hebrew.","marker":"Xue et al. (2021)"},{"why":"Provides the YiTB treebank used for POS tagging, dependency parsing, lemmatization, and transliteration tasks, and for the morphological competence probes.","marker":"Andrews (2025)"},{"why":"Supplies WikiANN NER, built from the 2018 Yiddish Wikipedia dump, used as a Kashes task.","marker":"Rahimi et al. (2019)"},{"why":"Supplies EHRI-NER, the Holocaust-domain named-entity recognition dataset in Kashes.","marker":"Dermentzi & Scheithauer (2024)"},{"why":"Supplies newNLP NER, annotations over historical Yiddish newspapers, used as a Kashes task.","marker":"Berkovitch & Rusinek (2021)"},{"why":"Supplies the Aya machine-translated PIQA, WikiQA, and PAWS-Wiki tasks used for language understanding in Kashes.","marker":"Singh et al. (2024)"},{"why":"Supplies FLORES+, the English-derived Yiddish translation benchmark that MameLoshnLM and baselines are compared on.","marker":"NLLB Team et al. (2024)"},{"why":"Supplies the CC100 English data used in the continued-pretraining mixture to preserve base capabilities.","marker":"Conneau et al. (2020)"},{"why":"Defines Llama 3.1 8B, the base model that MameLoshnLM is produced from by continued pretraining.","marker":"Grattafiori et al. (2024)"},{"why":"Supplies SentAlign, the alignment tool used to build the natively authored Kashes-mt translation benchmark.","marker":"Steingrímsson et al. (2023)"}],"fun_headline_variants":["Curated Yiddish corpus lifts model past multilingual rivals","Yiddish-focused LM beats multilingual models on native tasks","MameLoshnLM: native corpus yields gains over generic multilingual models","First open 8B Yiddish model outperforms multilingual baselines","Yiddish model's edge: native lexicon and morphology, not scale"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Kashes benchmark measures genuine Yiddish competence rather than memorization of test material: the paper excludes Kashes-mt source documents from Oytser but does not report excluding or auditing the Yiddish Wikipedia source against WikiANN, or the YBC books against the YiTB treebank test sets.","fun_headline_variants_meta":{"raw":{"variants":["Curated Yiddish corpus lifts model past multilingual rivals","Yiddish-focused LM beats multilingual models on native tasks","MameLoshnLM: native corpus yields gains over generic multilingual models","First open 8B Yiddish model outperforms multilingual baselines","Yiddish model's edge: native lexicon and morphology, not scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000215,"raw_usage":{"total_tokens":1446,"prompt_tokens":982,"completion_tokens":464,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":378}},"tokens_in":598,"tokens_out":464,"duration_ms":4871,"temperature":1.0,"reasoning_tokens":378,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T22:21:59.943400+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check for verbatim or near-verbatim overlap between the Oytser pretraining corpus and the Kashes test sets beyond Kashes-mt: if WikiANN test sentences (drawn from the 2018 Yiddish Wikipedia) or YiTB treebank sentences appear in the corresponding Oytser sources, the reported NER and linguistic-analysis scores are inflated. A retraining experiment with those overlapping documents removed would show whether the gains survive.","supporting_citations":[{"cited_title":"YiTB : the yiddish tree bank, 2025","cited_arxiv_id":null,"evidence_quote":"Provides the YiTB treebank used for POS tagging, dependency parsing, lemmatization, and transliteration tasks, and for the morphological competence probes."},{"cited_title":"Repurposing holocaust-related digital scholarly editions to develop multilingual domain-specific named entity recognition tools","cited_arxiv_id":null,"evidence_quote":"Supplies EHRI-NER, the Holocaust-domain named-entity recognition dataset in Kashes."}],"review_version":1}