{"id":"9f2847c1-aeaf-44cc-ac2e-7c8d8dc2787a","arxiv_id":"2506.07617","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"The authors release a Hutsul-Ukrainian corpus and show LoRA-fine-tuned 7B models beat GPT-4o on automated and LLM-based metrics, but the evaluation is contaminated by overlapping training and test sources.","lead":"This paper introduces the first Hutsul-to-standard Ukrainian parallel corpus and dictionary, then fine-tunes small open-source language models to translate standard Ukrainian into the Hutsul dialect. It reports that these models beat GPT-4o on in-domain metrics, but the test set shares the same source novel as the training data, so the headline result is not trustworthy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Test-set contamination is decisive: train and test both draw on Дідо Иванчік (Secs. 3.1, 3.3, 3.4), so the reported BLEU/LLM-judge advantage over GPT-4o measures in-domain memorization rather than general Hutsul translation ability.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the test set is not independent of training and augmentation. I agree, and I find this is the single point on which the central claim depends. If the test set were representative of Hutsul generally, fine-tuned 7B models beating GPT-4o would be a meaningful result; because the test set comes from the same novel that dominates manual data, synthetic retrieval, and rule extraction, the result is equally consistent with the models having memorized the evaluation distribution. Secondary issues — GPT-4o serving as both a data generator and judge, no error bars, no human ratings, and the ambiguous \"zero-shot\" label for a baseline that received RAG context and dictionary entries — would each need attention in revision, but none is as consequential as the domain overlap. I therefore see no reason to change the reader's REJECT verdict; a focused source-disjoint evaluation could convert the rejection into a conditional acceptance if the claimed advantage survives.","tokens_in":10682,"tokens_out":6218,"duration_ms":74173,"concrete_test":"Re-run the full evaluation under a source-disjoint protocol. Remove every sentence pair originating from Дідо Иванчік from the manual training set, generate synthetic data only from Hutsul sources outside that novel, and evaluate on a newly collected Hutsul test set drawn from independent sources (e.g., folklore websites, ethnographic transcripts, other literary works). Recompute Table 2 for Mistral (manual+synthetic), Mistral (manual only), LLaMA variants, and GPT-4o. If the fine-tuned models no longer beat GPT-4o on BLEU, chrF++, TER, and dialect quality, the claimed advantage is an artifact of training/test overlap. As a first diagnostic, also report the fraction of the 1,900 test sentences having exact or near-identical matches in the training corpus.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is that the headline comparison is evaluated on data from the same text distribution used to build the models. Section 3.4 states: \"Test and validation sets contain only human-annotated sentence pairs from Дідо Иванчік.\" Section 3.3 identifies the same novel as \"the primary corpus for retrieval and the source of linguistic examples\" for synthetic generation, and Section 3.1 says \"a significant portion\" of the manual corpus is based on it. Thus the 1,900 test sentences are drawn from a book that is simultaneously the core of the manual training set, the retrieval base for RAG-augmented training data, and the source of the grammar rules used for generation. The fine-tuned models can score high BLEU/chrF++ and high \"dialect quality\" ratings simply by reproducing the novel's sentence inventory and style; GPT-4o receives no such training advantage. The abstract's claim that small fine-tuned models outperform zero-shot baselines \"across both automatic and LLM-evaluated metrics\" therefore exceeds what the experiment can establish. The paper's own Limitations section concedes that synthetic data lacks coverage of modern domains and that automatic metrics \"may overestimate linguistic validity,\" but the central comparative claim still presumes a representative test set. That presumption is contradicted by the split description.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the first parallel corpus for Hutsul dialect and standard Ukrainian, consisting of 9,852 manually aligned sentence pairs and a 7,320-entry dialect-to-standard dictionary, together with a retrieval-augmented generation (RAG) pipeline that expands the corpus with 52,142 synthetic pairs. Using LoRA, the authors fine-tune Mistral-7B-Instruct v0.3 and LLaMA-3.1 8B Instruct on either the manual corpus alone or the manual-plus-synthetic corpus, and compare them against GPT-4o prompted with RAG context and dictionary entries on a 1,900-sentence held-out test set, using BLEU, chrF++, TER, and a GPT-4o-based fluency/adequacy/dialect-quality judge. The paper's central claim is that these 7B–8B fine-tuned models outperform GPT-4o on standard-to-Hutsul translation across both automatic and LLM-based metrics, with the best configuration (Mistral on manual-plus-synthetic) reaching 74.35 BLEU and a 3.60 dialect-quality rating.","tokens_in":10997,"tokens_out":9666,"duration_ms":106181,"significance":"The resource contributions are genuinely valuable: the first parallel Hutsul–Ukrainian corpus, a substantial dictionary, a reproducible RAG-based synthetic generation pipeline with alignment-based filtering, and a public release of data, models, and code are all strengths that will benefit future work on this underexplored dialect. However, the headline comparative claim—that small open models outperform GPT-4o on Hutsul translation—is not established by the present evaluation design. Training data, synthetic generation, and the test set all draw on the same literary source, and the primary evaluation metric is an unvalidated GPT-4o judge; the reported margins are therefore consistent with in-distribution memorization and judge bias rather than general translation capability. The paper's durable value lies in its resources and pipeline, while the comparative results require a redesigned evaluation before they can support the stated conclusion.","major_comments":[{"comment":"The test set of 1,900 sentences is, by the paper's own description in §3.4, composed only of human-annotated sentence pairs from 'Dido Yvanchik', while §3.1 states that a significant portion of the manual training corpus is also based on that novel, and §3.3 (steps 1 and 2) describes the same novel as the primary retrieval corpus and the source of linguistic examples for synthetic generation. The fine-tuned models are therefore both trained and synthetically augmented on the very text distribution from which the test sentences are drawn, whereas the GPT-4o baseline has no such training exposure. Under this design, the Table 2 margins (74.35 vs 56.64 BLEU; 3.60 vs 3.22 dialect rating) can be explained by in-distribution memorization or style replication, so the abstract's claim that fine-tuned models 'outperform zero-shot baselines such as GPT-4o across both automatic and LLM-evaluated metrics' is not supported as a statement about general Hutsul translation ability.","section":"§3.4, §3.1, §3.3, Table 2"},{"comment":"Section 5.1 declares the GPT-4o-based fluency, adequacy, and dialect scores to be the paper's primary evaluation metrics, yet no human validation of this judge is provided, and the judge is the same model family (GPT-4o) that generated the synthetic training data in §3.3 and that serves as the comparison baseline in §5.2; the judgment prompt in §5.1 also supplies the reference translation to the judge. The reported dialect-quality differences in Table 2 (3.22–3.60 on a 1–5 scale) are small, and no variance, confidence intervals, or significance tests are reported for either the automatic or the LLM-based metrics. The Limitations section itself concedes that automatic metrics 'may overestimate linguistic validity' and that GPT-4o's preferences may align with standard Ukrainian; because the LLM scores carry the primary comparative claim, the unvalidated-judge concern is load-bearing rather than ancillary.","section":"§5.1, §3.3, §5.2, Table 2"},{"comment":"The qualitative example in §5.4 contradicts the ranking in Table 2: Mistral trained on manual data only receives Dialect=4 and Adequacy=5 in the example, while Mistral trained on manual-plus-synthetic receives Dialect=3, the reverse of the Table 2 ordering (3.35 vs 3.60 aggregate dialect scores), and the per-example BLEU scores (7.77–34.39) are far below the Table 2 averages (56.64–74.35). The example-level reversal of dialect scores is inconsistent with the aggregate ordering, and the large gap between example and aggregate BLEU shows that the reference-based metrics are highly unstable at the sentence level; this reinforces the need for variance reporting before any cross-model conclusion can be drawn.","section":"§5.4, Table 2"}],"minor_comments":[{"comment":"The abstract calls GPT-4o a 'zero-shot baseline' while the Introduction and §5.2 describe the comparison as few-shot prompting with RAG context and dictionary entries; these descriptions should be reconciled, since §5.2 shows the baseline is not zero-shot.","section":"Abstract, §5.2"},{"comment":"The filtering thresholds in §3.3 (SequenceMatcher similarity 0.45; U-src<0.1, U-tgt<0.1, X<0.2) are presented as empirical choices without sensitivity analysis; given that Table 1 shows the filtered synthetic data has very different alignment statistics from the original corpus, the robustness of downstream results to these thresholds should be assessed.","section":"§3.3"},{"comment":"Table 2 reports no confidence intervals, bootstrap estimates, or significance tests; for automatic metrics computed over 1,900 sentences this information is standard practice and would clarify whether the dialect-rating differences between systems are meaningful.","section":"Table 2, §5.1"},{"comment":"The dictionary of 7,320 pairs described in §3.2 is never used in the fine-tuning or evaluation description, so its role in the pipeline should be stated explicitly or removed from the contribution list.","section":"§3.2"},{"comment":"The references contain several formatting defects, including 'and 1 others' placeholders, 'Ziweietal.Liu.2023', and 'HugoTouvron,ThibautLavril,AlpYurtsever,and1others', which should be corrected.","section":"References"},{"comment":"Section 5.1 is internally inconsistent: it first calls the LLM-based adequacy and dialect scores the 'primary evaluation metrics' and then states that automatic metrics are 'used in a supporting role', which reverses the hierarchy described two sentences earlier.","section":"§5.1"},{"comment":"Typos and usage errors include 'To insure translation quality' (§5.1), 'This process have created some data alignment challenges' (§3.3), and 'enlarged' (§3.3); these should be cleaned up.","section":"§3.3, §5.1"}],"recommendation":"reject","confidential_remarks":"The reader's rejection is well founded: the test-set contamination is real, and the stress-test concern lands. The paper's transparency about the split and its limitations is to its credit, but disclosure does not make the comparative claim valid. The resource contributions (first parallel corpus for Hutsul, dictionary, open release, reproducible synthetic pipeline) are worthwhile and could support a resubmission reframed as a dataset/resource paper with an evaluation explicitly scoped to the in-domain setting. Any revised version would need a test set drawn from sources other than 'Dido Yvanchik', at least a small human evaluation of the LLM judge, and variance reporting for the metrics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the Hutsul–Ukrainian corpus, the dictionary, and the released fine-tuned models are the real contribution; the central claim that small fine-tuned models beat GPT-4o on Hutsul translation is not established, because the test set is drawn from the same novel used to build the training and synthetic data.\n\nWhat the paper does well: it is the first LLM adaptation for a Ukrainian dialect, and it ships concrete artifacts—9,852 human-aligned sentence pairs, a 7,320-entry dictionary, 52,142 synthetic pairs, code, and model weights. The automatic and LLM-judged numbers are internally consistent, and the qualitative examples show the fine-tuned models producing plausible Hutsul. The Limitations section is candid about synthetic-data domain gaps and about automatic metrics possibly overestimating linguistic validity. That honesty earns credit.\n\nThe soft spot is load-bearing. Section 3.4 says the test and validation sets contain only human-annotated pairs from “Дiдо Иванчiк,” while Sections 3.1 and 3.3 say that same novel is the core of the manual corpus, the source of the grammar rules used for synthetic generation, and the retrieval base for the RAG pipeline. So the fine-tuned models are evaluated on the distribution they were trained to mimic; their high BLEU, chrF++, and dialect-quality ratings reflect in-domain fit, not general Hutsul ability. GPT-4o, given no such training advantage, is not a fair comparator for generalization. The abstract’s “outperform across both automatic and LLM-evaluated metrics” therefore overstates what the experiment can show. The choice of GPT-4o as both data generator and judge is a secondary concern, and the absence of human evaluation and error bars makes the LLM-judge scores harder to interpret. None of this invalidates the dataset, but it does invalidate the comparative claim as stated.\n\nWho gets value: researchers working on low-resource dialect translation or dataset construction, and anyone teaching evaluation pitfalls. I would not desk-reject this; the resource deserves referee time. But the path to acceptance has to include redoing the evaluation on a test set held out from sources other than “Дiдо Иванчiк,” or at minimum reporting results partitioned by whether a test sentence’s source appears in training. Adding a small human evaluation and reporting variance would also help. The honest verdict: keep the resource, fix the evaluation, and the paper becomes a useful contribution rather than an overclaimed one.","headline":"Useful new Hutsul-Ukrainian resource, but the headline GPT-4o comparison is not established because the test set draws on the same novel that dominates training and synthetic data.","tokens_in":11514,"tokens_out":2452,"would_cite":true,"duration_ms":30311,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuned 7B–8B open models beat zero-shot GPT-4o on every reported metric for standard-to-Hutsul translation, with Mistral at 74.35 BLEU and 3.60 dialect quality versus GPT-4o's 56.64 and 3.22.","keywords":["Hutsul dialect","dialectal machine translation","low-resource NLP","retrieval-augmented generation","synthetic parallel data","LoRA fine-tuning","Ukrainian dialect","LLM-based evaluation"],"falsifier":"Translate a set of Hutsul texts written outside the novel that anchors the corpus — for example, blog posts, transcribed Carpathian speech, or folk tales from another region — with the best fine-tuned model and GPT-4o, and have native Hutsul speakers rate the outputs. If GPT-4o matches or exceeds the fine-tuned model on these out-of-corpus texts, the central claim of general superiority fails.","tokens_in":10483,"feed_emoji":"🗣️","tokens_out":7015,"duration_ms":65725,"temperature":0.7,"pith_summary":"The paper tries to establish that a small, openly available language model fine-tuned on a modest amount of Hutsul–Ukrainian parallel data can translate standard Ukrainian into the Hutsul dialect better than a large commercial model prompted zero-shot. To do this, it builds the first Hutsul–Ukrainian parallel corpus, a dialect dictionary, and a retrieval-augmented pipeline that manufactures 52,142 additional training pairs from grammar rules and retrieved examples. On a 1,900-sentence test set, the best Mistral model reaches a BLEU of 74.35 and a dialect-quality rating of 3.60, versus GPT-4o's 56.64 and 3.22. The authors argue this shows specialized fine-tuning is a practical route for low-resource dialect translation where large models alone fall short.","feed_headline":"Small fine-tuned open models beat GPT-4o at Hutsul translation","feed_subtitle":"Mistral-7B tuned on manual plus synthetic data reaches 74 BLEU and 3.60 dialect quality, versus GPT-4o's 56 and 3.22.","key_machinery":"The load-bearing mechanism is the advanced RAG synthetic-data pipeline: GPT-4o first extracts structured Hutsul grammar rules from the dialectal novel and auxiliary sources, then for each standard Ukrainian sentence from UberText retrieves the top-3 semantically similar Hutsul sentences from an indexed corpus, and is prompted to produce a dialect translation; a sequence-similarity and alignment-based filter (U-src, U-tgt, X) removes low-quality pairs. This mechanism expands the 9,852 manually aligned pairs into about 62,000 training pairs. The models are then adapted by LoRA fine-tuning, and evaluation combines BLEU, chrF++, and TER with GPT-4o ratings of fluency, adequacy, and dialectal quality.","core_discovery":"The central claim is that fine-tuned 7B–8B open-source models outperform GPT-4o on standard-to-Hutsul translation across every reported automatic and LLM-judged metric. The best configuration, Mistral-7B fine-tuned on the combined manual and synthetically augmented corpus, scores 74.35 BLEU, 81.89 chrF++, and a dialect quality rating of 3.60, while GPT-4o scores 56.64, 65.90, and 3.22. The paper also finds that synthetic data from the RAG pipeline substantially improves automatic metrics, and that even manual-only fine-tuning beats the commercial baseline, with adequacy staying near 4.7 for all models while dialectal quality is the most data-sensitive axis.","pith_inferences":["Because the test set comes from the same novel that anchors the training, grammar-rule extraction, and retrieval index, the reported margin over GPT-4o is best read as in-distribution performance; a cross-corpus test set would show how far the models generalize to other Hutsul writing.","The pipeline's recipe — extract rules, retrieve similar examples, generate synthetic pairs, filter by alignment — is a candidate template for other low-resource dialect pairs, and could be tested by applying it to Boyko or Lemko Ukrainian.","If GPT-4o's judgments are biased toward standard Ukrainian forms, then valid dialectal variants may be underrated; collecting native-speaker ratings would recalibrate the dialect-quality scores.","The alignment-based filtering thresholds (U-src < 0.1, U-tgt < 0.1, X < 0.2) could serve as a cheap generic filter for synthetic dialect data, though the thresholds would likely need retuning per dialect."],"forward_implications":["Fine-tuned local 7B models can beat a large commercial zero-shot model on a low-resource dialect, so usable dialect translation does not require API access or large-scale compute.","The RAG-based synthetic augmentation expands the training set to about 62,000 pairs and yields large jumps in BLEU and chrF++, from 62.36 to 74.35 BLEU for Mistral.","The released corpus, dictionary, and models give the Hutsul dialect its first computational resources, and the same recipe is claimed to be adaptable to other Ukrainian dialects.","Because synthetic data mainly boosts surface-level metrics, dialectal quality gains are more modest, suggesting that authentic retraining data remains important."],"supporting_citations":[{"why":"Supplies the LoRA fine-tuning method used to adapt the open-source models.","marker":"Hu et al. (2021)"},{"why":"Defines the BLEU metric used as one of the automatic translation quality measures.","marker":"Papineni et al. (2002)"},{"why":"Defines chrF++, the character-level metric used for morphologically rich dialect evaluation.","marker":"Popović (2015)"},{"why":"Defines TER, the edit-rate metric used to capture structural divergence.","marker":"Snover et al. (2006)"},{"why":"Provides the UberText corpus from which standard Ukrainian source sentences are sampled for synthetic generation.","marker":"Chaplynskyi (2023)"},{"why":"Supplies fast-align, used to compute the alignment metrics that filter low-quality synthetic pairs.","marker":"Dyer et al. (2013)"},{"why":"Provides the framework for using LLMs as evaluators of dialectal machine translation.","marker":"Aepli et al. (2023)"},{"why":"Supplies word-formation rules in Hutsul dialects that are incorporated into the grammar-rule prompt for synthetic data generation.","marker":"Greshchuk (2016)"}],"fun_headline_variants":["Mistral-7B beats GPT-4o on Hutsul translation","Small fine-tuned open models beat GPT-4o for Hutsul","7B LoRA model outperforms GPT-4o on Hutsul dialect","Open 7B model beats GPT-4o on rare Hutsul dialect"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that its 1,900-sentence test set, taken entirely from the same novel that supplied training, synthetic generation, and retrieval, fairly measures Hutsul translation ability; if that novel is not representative of the dialect, the reported advantage over GPT-4o may not generalize.","fun_headline_variants_meta":{"raw":{"variants":["Mistral-7B beats GPT-4o on Hutsul translation","Small fine-tuned open models beat GPT-4o for Hutsul","7B LoRA model outperforms GPT-4o on Hutsul dialect","Open 7B model beats GPT-4o on rare Hutsul dialect"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000621,"raw_usage":{"total_tokens":2878,"prompt_tokens":943,"completion_tokens":1935,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":1849}},"tokens_in":559,"tokens_out":1935,"duration_ms":16403,"temperature":1.0,"reasoning_tokens":1849,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:29:42.037112+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Translate a set of Hutsul texts written outside the novel that anchors the corpus — for example, blog posts, transcribed Carpathian speech, or folk tales from another region — with the best fine-tuned model and GPT-4o, and have native Hutsul speakers rate the outputs. If GPT-4o matches or exceeds the fine-tuned model on these out-of-corpus texts, the central claim of general superiority fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the BLEU metric used as one of the automatic translation quality measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines TER, the edit-rate metric used to capture structural divergence."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the UberText corpus from which standard Ukrainian source sentences are sampled for synthetic generation."},{"cited_title":"A Benchmark for Evaluating Machine Translation Metrics on Dialects Without Standard Orthography","cited_arxiv_id":"2311.16865","evidence_quote":"Provides the framework for using LLMs as evaluators of dialectal machine translation."},{"cited_title":"hutsul dialectal vocabulary in ukrainian belletristic language","cited_arxiv_id":null,"evidence_quote":"Supplies word-formation rules in Hutsul dialects that are incorporated into the grammar-rule prompt for synthetic data generation."}],"review_version":1}