Pith. sign in

REVIEW 1 cited by

Detecting Spelling and Grammatical Anomalies in Russian Poetry Texts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.04507 v1 pith:CUW4KO2A submitted 2025-05-07 cs.CL

classification cs.CL
keywords datasetstextsdetectionmodelsqualitytraininganomalycreative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The quality of natural language texts in fine-tuning datasets plays a critical role in the performance of generative models, particularly in computational creativity tasks such as poem or song lyric generation. Fluency defects in generated poems significantly reduce their value. However, training texts are often sourced from internet-based platforms without stringent quality control, posing a challenge for data engineers to manage defect levels effectively. To address this issue, we propose the use of automated linguistic anomaly detection to identify and filter out low-quality texts from training datasets for creative models. In this paper, we present a comprehensive comparison of unsupervised and supervised text anomaly detection approaches, utilizing both synthetic and human-labeled datasets. We also introduce the RUPOR dataset, a collection of Russian-language human-labeled poems designed for cross-sentence grammatical error detection, and provide the full evaluation code. Our work aims to empower the community with tools and insights to improve the quality of training datasets for generative models in creative domains.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English

    cs.CL 2025-08 conditional novelty 6.0 of 10

    After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.

Pith tools