LoRA-fine-tuning a 12B open model on 15,000 GPT-o3 fable translations yields rubric scores close to GPT-o3 (4.83 vs 4.92) for English-to-Romanian literary translation at roughly one percent of the API cost.
BLEU is Not Suitable for the Evaluation of Text Simplification
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
BLEU is widely considered to be an informative metric for text-to-text generation, including Text Simplification (TS). TS includes both lexical and structural aspects. In this paper we show that BLEU is not suitable for the evaluation of sentence splitting, the major structural simplification operation. We manually compiled a sentence splitting gold standard corpus containing multiple structural paraphrases, and performed a correlation analysis with human judgments. We find low or no correlation between BLEU and the grammaticality and meaning preservation parameters where sentence splitting is involved. Moreover, BLEU often negatively correlates with simplicity, essentially penalizing simpler sentences.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
LoRA-fine-tuning a 12B open model on 15,000 GPT-o3 fable translations yields rubric scores close to GPT-o3 (4.83 vs 4.92) for English-to-Romanian literary translation at roughly one percent of the API cost.