Pith. sign in

REVIEW 3 cited by

StyloMetrix: An Open-Source Multilingual Tool for Representing Stylometric Vectors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.12810 v1 pith:VUZ23QA3 submitted 2023-09-22 cs.CL

classification cs.CL
keywords stylometrixlearningdeepvectorsalgorithmsclassificationclassifierlayer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work aims to provide an overview on the open-source multilanguage tool called StyloMetrix. It offers stylometric text representations that cover various aspects of grammar, syntax and lexicon. StyloMetrix covers four languages: Polish as the primary language, English, Ukrainian and Russian. The normalized output of each feature can become a fruitful course for machine learning models and a valuable addition to the embeddings layer for any deep learning algorithm. We strive to provide a concise, but exhaustive overview on the application of the StyloMetrix vectors as well as explain the sets of the developed linguistic features. The experiments have shown promising results in supervised content classification with simple algorithms as Random Forest Classifier, Voting Classifier, Logistic Regression and others. The deep learning assessments have unveiled the usefulness of the StyloMetrix vectors at enhancing an embedding layer extracted from Transformer architectures. The StyloMetrix has proven itself to be a formidable source for the machine learning and deep learning algorithms to execute different classification tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stylometry recognizes human and LLM-generated texts in short samples

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Stylometric features and tree-based classifiers separate human-written Wikipedia summaries from LLM-generated texts with high cross-validated accuracy on a new seven-class benchmark, though performance drops on other ...

  2. Tuning for TraceTarnish: Techniques, Trends, and Testing Tangible Traits

    cs.CR 2025-12 unverdicted novelty 3.0 of 10

    TraceTarnish's adversarial stylometry leaves detectable changes in function-word and content-word frequencies and type-token ratio, but detection requires comparing pre- and post-transformation text.

  3. StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features

    cs.CL 2025-07 conditional novelty 3.0 of 10

    A non-neural stylometric detector using LightGBM on spaCy-derived features reached a final mean score of 0.897 on the PAN 2025 task, below the 0.922 TF-IDF SVM baseline.

Pith tools