REVIEW 3 cited by
StyloMetrix: An Open-Source Multilingual Tool for Representing Stylometric Vectors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This work aims to provide an overview on the open-source multilanguage tool called StyloMetrix. It offers stylometric text representations that cover various aspects of grammar, syntax and lexicon. StyloMetrix covers four languages: Polish as the primary language, English, Ukrainian and Russian. The normalized output of each feature can become a fruitful course for machine learning models and a valuable addition to the embeddings layer for any deep learning algorithm. We strive to provide a concise, but exhaustive overview on the application of the StyloMetrix vectors as well as explain the sets of the developed linguistic features. The experiments have shown promising results in supervised content classification with simple algorithms as Random Forest Classifier, Voting Classifier, Logistic Regression and others. The deep learning assessments have unveiled the usefulness of the StyloMetrix vectors at enhancing an embedding layer extracted from Transformer architectures. The StyloMetrix has proven itself to be a formidable source for the machine learning and deep learning algorithms to execute different classification tasks.
Forward citations
Cited by 3 Pith papers
-
Stylometry recognizes human and LLM-generated texts in short samples
Stylometric features and tree-based classifiers separate human-written Wikipedia summaries from LLM-generated texts with high cross-validated accuracy on a new seven-class benchmark, though performance drops on other ...
-
Tuning for TraceTarnish: Techniques, Trends, and Testing Tangible Traits
TraceTarnish's adversarial stylometry leaves detectable changes in function-word and content-word frequencies and type-token ratio, but detection requires comparing pre- and post-transformation text.
-
StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features
A non-neural stylometric detector using LightGBM on spaCy-derived features reached a final mean score of 0.897 on the PAN 2025 task, below the 0.922 TF-IDF SVM baseline.
Discussion (0). Continue with ORCID to comment.