{"id":"16a94ebc-1c1d-4b2b-9c08-06969804229f","arxiv_id":"2501.03952","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Gemma 2 open-weight models nearly match commercial AI on Baltic-language tasks, but all tested open-weight multilingual models still make frequent lexical errors in generated text.","lead":"This paper measures how well open-weight AI models handle Lithuanian, Latvian, and Estonian for machine translation, multiple-choice reading, and free-form writing, and finds Gemma 2 rivals commercial cloud models while smaller variants lag. It matters because Baltic governments and companies need privacy-preserving AI that can run locally instead of sending data to foreign-hosted services.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'errors in at least 1 in 20 words for all open-weight multilingual LLMs' is not established: Table 4 gives Gemma 2 27B 4.08% on Lithuanian, and all rates come from 10 questions per language with no confidence intervals or inter-annotator agreement.","rationale":"After reading in good faith, the paper's practical contribution is the comparison of open-weight models against commercial systems on standard benchmarks. That portion is internally consistent: FLORES-200 devtest (1,012 sentences) and Belebele (900 questions) are established benchmarks, and the ranking Gemma 2 27B above the other open-weight families is consistent across Tables 1 and 2. I have no serious objection to that ranking or to the quantization-robustness observation for Gemma 2. The novel, attention-getting claim is the '1 in 20 words' lexical-hallucination floor for all open-weight multilingual LLMs. This is the claim most likely to be cited, and it is the one least supported by evidence. Table 4 is a manual pilot: ten questions per language, no per-question distribution, no inter-annotator reliability, no confidence intervals, and no released prompts or outputs. The raw numbers even contradict the universal floor as stated, since Gemma 2 27B scores 4.08% on Lithuanian. The reader identified the same fragility (small sample, no IAA, no CI) and correctly chose CONDITIONAL. My read does not change that verdict: the MT/MCQA findings remain useful, but the abstract's headline generalization needs to be narrowed or re-verified. I therefore keep the verdict unchanged and propose a concrete replication check that would settle whether the floor is real.","tokens_in":14196,"tokens_out":6486,"duration_ms":59965,"concrete_test":"Re-run the free-form evaluation with at least 50 balanced questions per language across the same model set, have two independent native-speaker annotators label every output with the same error schema, and compute Cohen's kappa plus Wilson 95% confidence intervals for each word-error rate. Then test the abstract's universal floor: does the lower bound of every evaluated open-weight model's interval remain at or above 5%? Separately, tabulate invented-word-only rates; if Gemma 2 27B's Lithuanian rate stays near 0.4% or any model's interval crosses 5%, the 'at least 1 in 20 words' claim must be narrowed and the 'lexical hallucination' wording should be revised to a broader text-error framing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most striking quantitative claim is the lexical-hallucination floor: open-weight multilingual LLMs have 'errors in at least 1 in 20 words' (abstract and conclusions). This rests entirely on Table 4, produced by two native-speaker linguists on ten free-form questions per language. The paper itself labels the parallel factual-accuracy results a pilot 'lacking statistical significance due to the small sample size.' No inter-annotator agreement, confidence intervals, or per-question error distributions are reported. Denominators are uneven (Gemma 2 27B: LT 1273 words/135 sentences vs LV 1171/97; Llama 3.1 8B: LT 724/68 vs LV 362/27), so a few long or hard answers can shift percentages by more than a point. More importantly, Table 4 contradicts the universal floor as stated: Gemma 2 27B has 4.08% errors in Lithuanian, below the 5% '1 in 20' threshold. Table 4 also includes only three open-weight multilingual models (Llama 3.1 8B/70B and Gemma 2 27B), so generalizing to 'all open-weight multilingual LLMs' is an extrapolation. Finally, the label 'lexical hallucinations' is broader than the evidence: annotators counted grammatical, inflectional, and syntactic errors as well as invented words, while the separately reported invented-word rates are much lower (e.g., 0.39/100 for Gemma 2 27B in Lithuanian, 1.45/100 in Latvian). The MT and MCQA results, based on FLORES-200 devtest and Belebele with 1,012 and 900 items, are more trustworthy and do support the Gemma-2 ranking, so the unsupported universal threshold is the weakest load-bearing element of the paper's central surprise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic evaluation of locally deployable open-weight LLMs (Llama 3/3.1/3.2, Gemma 2, Phi 3, NeMo) for Lithuanian, Latvian, and Estonian, with Czech and English as comparison languages. Using FLORES-200 devtest (1,012 sentences) for machine translation with COMET scores, Belebele (900 questions) for multiple-choice question answering, and a small human evaluation of free-form answers by two native-speaker linguists per language, it compares the open models against GPT-3.5 Turbo, GPT-4o, and DeepL. The main findings are that Gemma 2 27B approaches commercial systems in MT and MCQA, that quantization degrades Gemma 2 less than Llama models, and that open-weight multilingual models remain prone to frequent word-level errors in generated text, summarized in the abstract as 'errors in at least 1 in 20 words for all open-weight multilingual LLMs.'","tokens_in":14519,"tokens_out":4509,"duration_ms":39024,"significance":"The MT and MCQA results are valuable and generally credible: they use standard benchmarks, compare a coherent set of model families and precisions, and consistently identify Gemma 2 as the strongest open-weight family for these languages. The quantization comparison is a useful practical contribution, and the inclusion of Lt-Llama 2 fine-tuned models provides a concrete baseline for language specialization. The paper does not present a new method, and its main novelty is the empirical coverage of three under-resourced languages. The headline lexical-hallucination claim, however, is not supported with the same rigor as the benchmark results, because it rests on a small pilot annotation study with no confidence intervals or inter-annotator agreement. The practical implications for sovereign AI deployments are real if the findings hold, but the paper should state them with appropriate uncertainty.","major_comments":[{"comment":"The universal claim that open-weight multilingual LLMs produce lexical hallucinations with 'errors in at least 1 in 20 words' is not established by Table 4. Gemma 2 27B has a 4.08% error rate on Lithuanian, below the 5% threshold, and all rates derive from only ten free-form questions per language. The paper itself labels the parallel factual-accuracy results a pilot lacking statistical significance; the word-error rates come from the same small sample, yet no confidence intervals, per-question error distributions, or inter-annotator agreement are reported. The denominators also vary substantially across rows (e.g., 300 vs. 1,273 words for Lt-Llama 2 vs. Gemma 2 in Lithuanian), so the reported percentages are sensitive to a few long answers. I recommend reporting confidence intervals or at least per-question ranges and restricting the conclusion to the models actually evaluated.","section":"Abstract; §3, Table 4"},{"comment":"The term 'lexical hallucinations' overstates what was measured. The evaluation counted grammatically incorrect words, incorrect inflections, invented words, and words in syntactically incorrect structures, but the separately reported invented-word rates are much lower than the total error rates (e.g., Gemma 2 27B: 0.39/100 invented vs. 4.08/100 total in Lithuanian; 1.45 vs. 5.98 in Latvian). The abstract and conclusions should either use a more neutral term such as 'word-level linguistic errors' or restrict the hallucination claim to the invented-word counts.","section":"§2 (Text Quality) and §4"},{"comment":"The generalization from the human evaluation to the Baltic states as a whole is too broad. Table 4 contains no Estonian rows, so the claim about all three languages is unsupported for Estonian, and only three open-weight multilingual models (Llama 3.1 8B/70B and Gemma 2 27B) appear in the table. The conclusion that 'most multilingual models are still surprisingly prone to lexical hallucinations' should be qualified to the evaluated languages and model families, or additional Estonian annotation and more model families should be added.","section":"§2 and §3 (Table 4)"}],"minor_comments":[{"comment":"In the sentence 'As a result, Gamma 2 models show little performance degradation,' 'Gamma 2' should be 'Gemma 2'.","section":"§3"},{"comment":"The header contains 'OpenaAI GPT 3.5-Turbo'; correct it to 'OpenAI GPT-3.5 Turbo'.","section":"Table 1"},{"comment":"The claim that Gemma 2 shows a 'statistically insignificant drop' in MT performance is not backed by a reported significance test; please provide the test or phrase it as an observation about effect size.","section":"§3"},{"comment":"The precision labels '4bit', '8bit', and '16bit' are used inconsistently; consider using '4-bit', '8-bit', and '16-bit' throughout.","section":"§2"},{"comment":"The affiliation contains 'Lithu ania' with a spurious space; correct it.","section":"Author affiliations"}],"recommendation":"major_revision","confidential_remarks":"The paper is a competent empirical study and does not pose novelty or attribution concerns. The main risk is that the '1 in 20 words' formulation is likely to be quoted out of context, since it is the most striking claim in the abstract and conclusions; I would ask the authors to make the pilot nature of the human evaluation and the model/language coverage explicit in the abstract itself, not only in the body."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you care about how open-weight models actually behave on smaller European languages. The core empirical work is solid: they run four model families across MT and MCQA on Lithuanian, Latvian, and Estonian, using FLORES-200 devtest and Belebele, with COMET for MT. The tables are consistent and the main finding—Gemma 2 27B is the best open-weight family, close to GPT-3.5/4o, and robust to 4-bit quantization—is a genuinely new and useful data point for anyone thinking about sovereign AI deployments in the Baltics. The comparison across model sizes and quantization levels is well done.\n\nThe soft spot is exactly where the abstract makes its most striking claim. The statement that open-weight multilingual LLMs produce 'errors in at least 1 in 20 words' does not hold up against their own Table 4: Gemma 2 27B shows 4.08% errors on Lithuanian, which is below the 5% threshold. That table also comes from only ten free-form questions per language, judged by two native speakers, with no inter-annotator agreement, no confidence intervals, and uneven denominators across models. The paper itself admits the factual-accuracy results are a pilot lacking statistical significance, and the same human-eval data underlies the word-error rates. On top of that, the label 'lexical hallucinations' is too broad—annotators counted grammatical and syntactic errors, while the separately reported invented-word rates are much lower (e.g., 0.39 per 100 words for Gemma 2 27B in Lithuanian). So the '1 in 20 words' claim should be treated as preliminary, not a floor.\n\nThat said, the problem is confined to the human-eval generalization. The MT and MCQA results, based on standard benchmarks with hundreds of items, are trustworthy and support the ranking. I don't see any circularity or invented results; the citation pattern is appropriate, and the paper is honest about its limitations.\n\nWho is this for? Practitioners in the Baltic states choosing local models, and researchers working on low-resource multilingual evaluation. It's not a methodological breakthrough, but it's a competent benchmark that fills a gap. I'd send it to peer review, but the authors should be required to soften the abstract claim, report the human eval with proper caveats (or expand it), and ideally release the annotation data.","headline":"Solid benchmark of open-weight LLMs for Baltic languages, but the '1 in 20 words' floor is overstated and should be fixed before publication.","tokens_in":15127,"tokens_out":2393,"would_cite":true,"duration_ms":20723,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Open-weight LLMs can nearly match commercial AI for Baltic languages, but every open model still makes at least one lexical error per 20 words in generated text.","keywords":["open-weight LLMs","Baltic languages","machine translation","multiple-choice question answering","lexical hallucination","model quantization","low-resource languages","Gemma 2"],"falsifier":"Run the same ten prompts plus ninety new ones through each open-weight model at default settings, and have three independent native-speaker linguists mark errors; the 'at least 1 in 20 words' claim fails if any non-fine-tuned model's error rate falls below 5 percent with a confidence interval excluding 5 percent.","tokens_in":13966,"feed_emoji":"🌐","tokens_out":12771,"duration_ms":97633,"temperature":0.7,"pith_summary":"The paper asks whether open-weight large language models, which can be deployed on local hardware, are good enough for Lithuanian, Latvian, and Estonian to be used in privacy-sensitive public-sector applications. It benchmarks Llama 3, Gemma 2, Phi 3, and NeMo against commercial cloud services on machine translation, multiple-choice question answering, and free-form generation, using FLORES-200, Belebele, and a human error-counting protocol. The answer is yes for Gemma 2 27B at 4-bit precision, which matches GPT-3.5 Turbo and comes close to GPT-4o and DeepL on translation and reading comprehension. The paper also reports that every open-weight multilingual model it tested produces at least one lexical error per 20 words in free-form answers, while a language-specific fine-tuned model drops that rate to about one percent. A sympathetic reading is that local, sovereign AI for Baltic languages is within reach, but only if applications can tolerate or mitigate a steady trickle of invented and grammatically wrong words.","feed_headline":"Open-weight AI rivals cloud models for Baltic languages","feed_subtitle":"Yet human checks find at least one word error in every 20 in free-form answers.","key_machinery":"The load-bearing mechanism is the three-part evaluation protocol. Translation quality is scored with COMET on the FLORES-200 devtest set; comprehension is scored as accuracy on the Belebele multiple-choice benchmark; and generation quality is measured by having two native-speaker linguists count, per language, the number of words that are grammatically incorrect, wrongly inflected, invented, or syntactically misplaced in answers to ten open-ended prompts. The lexical hallucination claim comes from the last of these: the per-word error rate converts directly to the '1 in 20 words' figure. The other key object is the 4-bit quantization variant, which the paper uses to show that Gemma 2's architecture loses almost nothing when compressed, while Llama's does not.","core_discovery":"Stated on the paper's own terms, the discovery is that Gemma 2 27B, run at 4-bit quantization, performs close to the top commercially available models across all three Baltic languages: it reaches an average COMET score of 0.89 on FLORES-200 translation (versus 0.90 for GPT-4o and 0.91 for DeepL) and an average Belebele MCQA accuracy of 0.912 (versus 0.944 for GPT-4o and 0.797 for GPT-3.5 Turbo). At the same time, all evaluated open-weight multilingual models, including Gemma 2, show error rates of roughly 4 to 19 percent in free-form text generation in Lithuanian and Latvian as judged by native-speaker linguists, meaning at least one in every 20 words is grammatically incorrect, wrongly inflected, syntactically misplaced, or invented. The exception is a Lithuanian-specific fine-tune of Llama 2, whose error rate is about 1 percent with no invented words. The paper also demonstrates that quantization affects models differently: Gemma 2 suffers negligible drops from 16-bit to 4-bit, whereas Llama 3 loses measurable accuracy, with larger losses for the Baltic languages than for English or Czech.","pith_inferences":["If the 1-in-20 word error rate is roughly stable across domains, even 'good' open models need a hallucination-detection layer before use in government documents; the paper's pilot is too small to certify stability, but that is the natural reading of the reported rates.","The Czech–Baltic gap implies a data-volume threshold: once a language crosses a certain number of training tokens, open-model quality jumps; this predicts that future models trained on more Baltic web data will show disproportionate gains on exactly the tasks tested here.","Because the two annotators were linguists, the error counts may be stricter than what ordinary users notice; a user study measuring perceived acceptability could show whether the 1-in-20 bound overstates practical harm.","The quantization-robustness difference between Gemma and Llama suggests internal representation redundancy differs across model families, which could guide future model selection for low-resource multilingual deployment."],"forward_implications":["Gemma 2 27B at 4-bit is a practical local replacement for cloud APIs on Baltic-language machine translation and multiple-choice QA, at a fraction of the memory and no data-leakage risk.","Applications that generate free-form Baltic text should not go live with non-fine-tuned open models; the measured lexical error rate of 4 to 19 percent is too high for customer-facing or legal text.","Language-specific fine-tuning, as demonstrated by Lt-Llama 2, can reduce lexical errors to about 1 percent, so building high-quality Baltic-language corpora is the clear path to sovereign AI.","Quantization-friendly architectures like Gemma 2 make local deployment much cheaper; Llama's sensitivity to 4-bit means local Llama deployments should prefer 8-bit or full precision if quality matters.","The consistent gap between Czech and Baltic results suggests training-data volume, not linguistic structure, is the limiting factor for open models in smaller European languages."],"supporting_citations":[{"why":"Supplies the FLORES-101 benchmark that underlies the parallel sentences used for translation evaluation.","marker":"Goyal et al., 2021"},{"why":"Extends the benchmark to the FLORES-200 devtest set (1,012 sentences) used for all translation directions.","marker":"Costa jussà et al., 2022"},{"why":"Provides the Belebele multiple-choice dataset (900 questions per language) used for reading comprehension accuracy.","marker":"Bandarkar et al., 2024"},{"why":"Defines the COMET metric that scores all machine translation outputs.","marker":"Rei et al., 2020, 2022"},{"why":"Defines the Llama 3 and Llama 3.1 model variants evaluated across all tasks.","marker":"Dubey et al., 2024"},{"why":"Defines the Gemma and Gemma 2 model families that the paper finds to be the top open-weight performers.","marker":"Team et al., 2024; Mesnard et al., 2024"},{"why":"Defines the Phi-3 family that performs poorly on non-English languages and often fails the JSON output requirement.","marker":"Abdin et al., 2024"},{"why":"Provides the Lt-Llama 2 Lithuanian fine-tune that sets the low-error baseline in free-form text generation.","marker":"Nakvosas et al., 2024"},{"why":"Documents GPT-4o, the commercial cloud model whose performance the paper uses as comparison target.","marker":"OpenAI et al., 2024"}],"fun_headline_variants":["Open-weight AI rivals cloud models for Baltic languages, but errs in 1 in 20 words","Baltic languages: local open-weight models nearly match GPT-4o, but hallucinate","Gemma 2 leads open-weight models for Baltic languages, but invents 5% of words","For Baltic tongues, local AI comes close to GPT-4o but hallucinates 5% of words"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The binding assumption is that ten free-form questions per language, scored by two native-speaker linguists without reported agreement statistics, produce error rates representative of each model's general output in that language.","fun_headline_variants_meta":{"raw":{"variants":["Open-weight AI rivals cloud models for Baltic languages, but errs in 1 in 20 words","Baltic languages: local open-weight models nearly match GPT-4o, but hallucinate","Gemma 2 leads open-weight models for Baltic languages, but invents 5% of words","For Baltic tongues, local AI comes close to GPT-4o but hallucinates 5% of words"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001217,"raw_usage":{"total_tokens":5026,"prompt_tokens":983,"completion_tokens":4043,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":3941}},"tokens_in":599,"tokens_out":4043,"duration_ms":23714,"temperature":1.0,"reasoning_tokens":3941,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:42:55.056196+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same ten prompts plus ninety new ones through each open-weight model at default settings, and have three independent native-speaker linguists mark errors; the 'at least 1 in 20 words' claim fails if any non-fine-tuned model's error rate falls below 5 percent with a confidence interval excluding 5 percent.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Belebele multiple-choice dataset (900 questions per language) used for reading comprehension accuracy."},{"cited_title":"Open Llama2 Model for the Lithuanian Language","cited_arxiv_id":"2408.12963","evidence_quote":"Provides the Lt-Llama 2 Lithuanian fine-tune that sets the low-error baseline in free-form text generation."}],"review_version":1}