{"id":"62f3bf7c-a234-4e17-bd66-5188ab64ff2a","arxiv_id":"2607.19101","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Machine-translated Russian data did not improve Finnish difficulty prediction on native Finnish test texts, contradicting the paper's headline claim.","lead":"Researchers tested whether machine-translated Russian texts could help train a Finnish AI to judge text difficulty. Their own tables show translated data helps only when the test texts are also translated; on real native Finnish texts, adding it slightly hurts accuracy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's claim of augmentation improving accuracy is contradicted by Table 5: native-only training beats native+MT on the native Finnish test (MSE 5.12 vs 7.57, R² 97.30 vs 96.01).","rationale":"The reader's strongest claim correctly identifies the contradiction between the abstract and Table 5. However, the reader's weakest assumption — that MT preserves CEFR difficulty — is a separate, secondary concern. The most load-bearing issue is the direct empirical contradiction: even if the translated labels were perfectly valid, the paper's own results show that adding MT data hurts performance on native Finnish text. This is a stronger and more immediate reason to reject than the label-validity concern. I partially agree with the reader because their strongest claim points to the same contradiction, but their weakest assumption points elsewhere. The concrete test I propose would decisively confirm whether the native-test degradation is real and significant; without such verification, the paper's central claim lacks support. The reader's verdict of REJECT remains appropriate, so no adjustment is needed.","tokens_in":8515,"tokens_out":3710,"duration_ms":34679,"concrete_test":"Retrain the Finnish BERT regression model under the exact conditions of Exp 1 (native-only) and Exp 5 (native+MT) with at least 5 random seeds, keeping the native Finnish test set fixed. Report mean and standard deviation of MSE and R² per condition, and run a paired bootstrap test on per-document squared errors. If native+MT does not significantly outperform native-only on the native test, the abstract's claim that augmentation improves accuracy is unsupported. This directly tests whether the reported degradation (MSE 5.12 → 7.57) is reproducible and statistically meaningful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — that augmenting scarce native Finnish data with MT-translated Russian data improves difficulty estimation — is directly contradicted by the paper's own experiments. In Table 5, Exp 1 (native-only training) achieves MSE 5.12 and R² 97.30 on the native Finnish test set, while Exp 5 (native+MT training) gives MSE 7.57 and R² 96.01 on the same test set. Section 4.2 acknowledges the native-only result as 'strong in-domain performance' but then spins the combined model as 'comparable accuracy' and concludes that MT augmentation 'substantially enhances low-resource Finnish performance.' The reported numbers show degradation, not improvement, on the evaluation that matters for the stated application (native input texts). The only configurations where augmentation yields large gains are those that test on MT-translated data (Exp 6 vs Exp 4), which is distribution matching rather than evidence for real-world low-resource deployment. Additionally, §3.2 admits that MT may not preserve CEFR difficulty and provides only manual spot-checking as validation, so the augmented labels themselves are unverified. Even if label preservation were true, the performance claim fails on the paper's own native-test comparison. Sections 1, 4, and 5 repeatedly overstate the results by omitting the native-only baseline when claiming improvement, making the abstract's unqualified claim unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using machine-translated (MT) Russian texts as synthetic augmented training data for Finnish text-difficulty assessment, where native annotated Finnish data are scarce. It trains BERT-based regression models over native Finnish, MT-translated Russian, and combined training sets, reporting MSE, MAE, and R² on native and MT test sets. The central claim, as stated in the abstract and Section 4, is that augmenting scarce native data with MT data significantly improves difficulty-estimation accuracy for Finnish. The paper also explores ablations with different proportions and CEFR-level subsets of MT data.","tokens_in":8853,"tokens_out":3129,"duration_ms":29982,"significance":"If the central claim were true, the approach would offer a practical way to bootstrap difficulty-assessment systems for low-resource languages by leveraging existing Russian annotations. The experimental matrix is fairly broad, including multiple train/test configurations and ablations, and the paper is candid about some of its assumptions (§3.2) and limitations (§6). However, the paper's own results contradict the headline: the best performance on native Finnish test data comes from native-only training, not from MT augmentation. The only large gains from augmentation occur when the test set is itself machine-translated, which is a distribution-matching artifact rather than evidence of generalization. The lack of quantitative validation of the central label-preservation assumption further weakens the empirical chain. The paper does not provide code, data, or machine-checked artifacts, so reproducibility is limited.","major_comments":[{"comment":"The abstract claims that augmenting scarce native data with machine-translated corpora significantly improves accuracy. Table 5 directly contradicts this on the native Finnish test set: Exp. 1 (FI native-only training) gives MSE=5.12 and R²=97.30, while Exp. 5 (FI native+MT training) gives MSE=7.57 and R²=96.01. On the evaluation that matters for the stated application (native input texts), adding MT data degrades performance. Section 4.2 acknowledges Exp. 1 as 'strong in-domain performance' but then describes the combined model as 'comparable accuracy' and concludes that translation-based augmentation is effective. The reported numbers do not support the abstract's unqualified claim.","section":"Abstract; §4.2; Table 5"},{"comment":"The result highlighted as a 'dramatic performance gain' (MSE=4.22, R²=94.26) is obtained by training on native+MT and testing on MT-translated texts generated by the same OpusMT pipeline used to create the augmented training data. This is a distribution-overlap effect: the test set shares the same translation artifacts and label-injection procedure as the training set. It provides no evidence for generalization to real-world native Finnish input, and therefore cannot support the paper's conclusion that MT augmentation improves low-resource Finnish difficulty assessment.","section":"§4.2; Table 5, Exp. 6"},{"comment":"The paper explicitly states that machine translation does not guarantee preservation of CEFR difficulty level across languages, and relies only on 'manual inspection of a sample by native language experts' to assert that the OpusMT model preserves difficulty. No sample size, inter-annotator agreement, or quantitative validation is reported. If CEFR levels are not preserved, every augmented label is potentially wrong, invalidating the training data in Experiments 1–6 and 9–16. This is a load-bearing assumption for both research questions and is not adequately supported.","section":"§3.2"},{"comment":"The ablation section interprets Exp. 10 (B1–B2 injection, R²=96.68) and Exp. 13 (40% MT, R²=96.20) as 'strongest improvement' and 'slightly outperforming' full augmentation. However, both are still below the native-only baseline of Exp. 1 (R²=97.30) on the same native test set. The omitted comparison against the native-only baseline makes the narrative in §4.3 misleading and further weakens the claim that MT data helps.","section":"§4.3; Table 6"}],"minor_comments":[{"comment":"The text states the native corpus contains 4544 texts, while Table 1 and Table 3 report a native total of 3460. Similarly, the text says the total dataset contains 12,865 texts, but Table 1 reports 10,917. Please reconcile these numbers.","section":"§3.1; Tables 1 and 3"},{"comment":"Typo: 'mean quared error' should be 'mean squared error'.","section":"Table 5 caption"},{"comment":"References such as 'Figure 1.4' and 'Figure 1.5' are unusual; standard notation would be 'Figure 1(d)' or 'Figure 1, Exp. 4'.","section":"§4.2; Figure 1"},{"comment":"The phrase 'while applying dataset balancing' is vague. The balancing procedure, its parameters, and its effect on the comparison should be described explicitly, since it is a possibly important confound.","section":"§4.2"},{"comment":"The table lists six linguistic features but the text notes they are not used in modeling. Consider clarifying their role (descriptive only) to avoid confusion.","section":"Table 4"}],"recommendation":"reject","confidential_remarks":"The central empirical claim of the paper is contradicted by its own Table 5: native-only training outperforms native+MT on the native Finnish test set. The only favorable results come from testing on MT-generated data, which is not a valid real-world scenario for the stated application. The label-preservation assumption is also essentially unvalidated. These are not fixable with a modest revision; the conclusions would need to be substantially redrawn. The paper also lacks code/data for reproducibility, and the novelty is incremental. I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take on arXiv:2607.19101. The headline claim — that machine-translated data significantly improves Finnish difficulty prediction — is contradicted by the paper's own Table 5. On the native Finnish test set, native-only training gives MSE 5.12 and R² 97.30; adding MT data gives MSE 7.57 and R² 96.01. The clear gains appear only when the test set is itself machine-translated, which is distribution matching, not evidence of real-world benefit.\n\nWhat is genuinely useful: the authors assemble a Finnish CEFR corpus from multiple public sources, translate Russian annotated corpora with OpusMT, and run a systematic set of ablations (label injection, proportional sampling). That is a solid experimental framework for an established idea — cross-lingual MT augmentation — applied to a new language pair. The data description is transparent about sources, and the limitations section does note the MT-preservation caveat.\n\nSoft spots are load-bearing. First, the central claim: Section 4.2 calls native-only 'strong in-domain performance' but then describes the combined model as 'comparable accuracy' and concludes augmentation 'substantially enhances.' Those are not supported by the reported numbers. Second, the MT test set in Exp 6 is generated by the same translation pipeline used for augmentation, so the gain is unsurprising. Third, the assumption that MT preserves CEFR difficulty is only manually spot-checked; the paper itself says it's not guaranteed. Fourth, dataset counts conflict: the text says 4,544 native texts, Table 1 sums to 3,460; the text says 12,865 total, Table 1 sums to 10,917; dev is 1,001 in one place and 1,011 in another. No code or data are released.\n\nWho is this for? Researchers working on low-resource readability assessment might read it as a cautionary example, but only if the abstract were rewritten. As written, it is not a reliable source.\n\nMy recommendation: desk-reject in current form. The internal contradiction alone justifies that, and the count discrepancies don't help. If the authors reframe it as a negative result and fix the reporting, it could become a useful workshop paper. I would not send it to a serious referee now.","headline":"The paper's own Table 5 contradicts its headline claim: native-only training beats native+MT on the native Finnish test set.","tokens_in":9317,"tokens_out":4720,"would_cite":false,"duration_ms":42516,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine-translated Russian corpora can supplement scarce Finnish CEFR training data, this paper argues.","keywords":["text difficulty assessment","CEFR","machine translation augmentation","low-resource NLP","Finnish","Russian","BERT regression","readability"],"falsifier":"If a quantitative comparison of source Russian CEFR labels with labels assigned to the Finnish translations by human experts (or by a strong Finnish difficulty model) showed a systematic shift in level for more than a small fraction of texts, the augmented labels would be unreliable and the reported gains would be artifacts. A direct check: take a held-out set of Russian texts with known CEFR levels, translate them, have expert raters label the Finnish versions, and measure the confusion matrix.","tokens_in":8363,"feed_emoji":"📚","tokens_out":4821,"duration_ms":49142,"temperature":0.7,"pith_summary":"The paper tackles the shortage of expert-labeled CEFR data in Finnish by translating Russian texts that already carry difficulty labels and adding them to the training set. It trains BERT-based regressors to predict a continuous difficulty score, and reports that combining native Finnish with translated Russian data improves robustness and accuracy compared with training on either alone. On the paper's own results, the clearest gain is on machine-translated test data; on native Finnish test text the combined model is close to, but slightly below, a native-only model. The intended contribution is a recipe for low-resource languages that lack expert annotations.","feed_headline":"Translated Russian data boosts Finnish difficulty models","feed_subtitle":"A BERT regressor trained on native plus machine-translated Finnish matches native-only accuracy and wins on translated tests.","key_machinery":"An augmentation pipeline that maps 11 CEFR sub-levels to a continuous 1-6 scale and uses an off-the-shelf Russian-to-Finnish MT system to create synthetic Finnish texts carrying Russian difficulty labels. The translated corpus triples the training data; a BERT regression head then learns to predict the score. The assumption that the MT system preserves difficulty level across translation is the mechanism that makes the labels usable.","core_discovery":"The paper claims that cross-lingual data augmentation via machine translation transfers CEFR difficulty labels from Russian to Finnish well enough to train a useful difficulty estimator. Using 3,460 native Finnish documents and 7,457 translated documents, a fine-tuned BERT regression model reaches R²=96.01% on native Finnish when trained on both, versus 97.30% for native-only training, and reaches R²=94.26% on machine-translated Finnish text, where native-only training collapses. The author's reading is that MT data supplements scarce native data and regularizes the model; the paper also reports that mid-level (B1-B2) translated texts help most, and that 40% proportional MT sampling is as ef","pith_inferences":["The paper's claim that augmentation 'significantly improves' accuracy is not supported by the native-test comparison in Table 5; the improvement appears only when the test set is itself machine-translated, so the practical benefit may depend on how much translated text the deployed system will see.","A stronger validation would test on Finnish texts translated by humans from Russian, or on native Finnish with known original Russian labels, to separate language-transfer effects from label-preservation artifacts.","The B1-B2 dominance suggests a label-preservation filter: training examples whose source-level or round-trip translation changes difficulty should be down-weighted or excluded.","One could use the same pipeline in reverse (Finnish to Russian) or on another language pair to check whether the result is specific to Russian-Finnish or general."],"forward_implications":["If true, low-resource languages can bootstrap difficulty models by translating existing high-resource annotated corpora rather than commissioning new expert annotations.","Translated data alone is insufficient for native-language prediction; some native data is still required, so MT augmentation is complementary, not a replacement.","Mid-level CEFR texts transfer best, so selective sampling by source difficulty should be part of any augmentation strategy.","A difficulty model trained this way could act as a critic in LLM simplification pipelines and personalized learning tools for Finnish.","Proportional augmentation around 40% translated data gives accuracy comparable to full augmentation, suggesting a saturation point beyond which extra MT data adds noise."],"fun_headline_variants":["Russian translations bolster Finnish text difficulty scoring","Translation as augmentation: Russian data aids Finnish CEFR models","Borrowed Russian labels train a Finnish difficulty regressor","Cross-lingual trick: Russian translations improve Finnish grading","MT-augmented training fills Finnish CEFR data gap"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire method rests on the assumption that the machine translation system preserves the CEFR difficulty level of a Russian text when rendering it in Finnish; the paper validates this only by manual inspection of a sample, not by any quantitative check.","fun_headline_variants_meta":{"raw":{"variants":["Russian translations bolster Finnish text difficulty scoring","Translation as augmentation: Russian data aids Finnish CEFR models","Borrowed Russian labels train a Finnish difficulty regressor","Cross-lingual trick: Russian translations improve Finnish grading","MT-augmented training fills Finnish CEFR data gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00087,"raw_usage":{"total_tokens":3572,"prompt_tokens":678,"completion_tokens":2894,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":422,"completion_tokens_details":{"reasoning_tokens":2825}},"tokens_in":422,"tokens_out":2894,"duration_ms":22923,"temperature":1.0,"reasoning_tokens":2825,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T13:25:51.310404+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a quantitative comparison of source Russian CEFR labels with labels assigned to the Finnish translations by human experts (or by a strong Finnish difficulty model) showed a systematic shift in level for more than a small fraction of texts, the augmented labels would be unreliable and the reported gains would be artifacts. A direct check: take a held-out set of Russian texts with known CEFR levels, translate them, have expert raters label the Finnish versions, and measure the confusion matrix.","supporting_citations":[],"review_version":1}