A new Sinhala text simplification dataset with 3,000 human-written simplifications is released, and intermediate-task transfer learning on mT5/mBART beats prior zero-resource baselines.
A Constrained Sequence-to-Sequence Neural Model for Sentence Simplification
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Sentence simplification reduces semantic complexity to benefit people with language impairments. Previous simplification studies on the sentence level and word level have achieved promising results but also meet great challenges. For sentence-level studies, sentences after simplification are fluent but sometimes are not really simplified. For word-level studies, words are simplified but also have potential grammar errors due to different usages of words before and after simplification. In this paper, we propose a two-step simplification framework by combining both the word-level and the sentence-level simplifications, making use of their corresponding advantages. Based on the two-step framework, we implement a novel constrained neural generation model to simplify sentences given simplified words. The final results on Wikipedia and Simple Wikipedia aligned datasets indicate that our method yields better performance than various baselines.
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SiTSE: Sinhala Text Simplification Dataset and Evaluation
A new Sinhala text simplification dataset with 3,000 human-written simplifications is released, and intermediate-task transfer learning on mT5/mBART beats prior zero-resource baselines.