After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.
Detecting Spelling and Grammatical Anomalies in Russian Poetry Texts
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The quality of natural language texts in fine-tuning datasets plays a critical role in the performance of generative models, particularly in computational creativity tasks such as poem or song lyric generation. Fluency defects in generated poems significantly reduce their value. However, training texts are often sourced from internet-based platforms without stringent quality control, posing a challenge for data engineers to manage defect levels effectively. To address this issue, we propose the use of automated linguistic anomaly detection to identify and filter out low-quality texts from training datasets for creative models. In this paper, we present a comprehensive comparison of unsupervised and supervised text anomaly detection approaches, utilizing both synthetic and human-labeled datasets. We also introduce the RUPOR dataset, a collection of Russian-language human-labeled poems designed for cross-sentence grammatical error detection, and provide the full evaluation code. Our work aims to empower the community with tools and insights to improve the quality of training datasets for generative models in creative domains.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.