REVIEW 2 major objections 7 minor 41 references
StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features
T0 review · 2 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A stylometric pipeline with gradient-boosted trees scores 0.897 on the PAN 2025 AI-detection task, below the TF-IDF baseline of 0.922.
desk verdict A competent and honest PAN 2025 system description that confirms cheap non-neural detectors still trail the TF-IDF baseline; the main gap is the absence of any training/test overlap analysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the feature set: normalized frequencies of n-grams of spaCy en_core_web_lg annotations—lemmas (uni- to trigrams, excluding named entities), part-of-speech tags (uni- to quadrigrams including punctuation), dependency-based bigrams (neighbourhood defined by distance in the dependency tree), and morphological unigrams where named entities are replaced by their types—each feature class capped at 1500 items, giving 4594 features total, optionally culled by document frequency. These are fed to LightGBM with DART boosting, bagging fraction 0.8 every 3 iterations, and capacity controlled by num_leaves, num_iterations, and max_depth, scaled upward for the 563,571-text training corpus. The mechanism is that the boosted trees learn to weight which linguistic-frequency patterns discriminate machine from human text.
What would settle it
Run the big-cv model on texts from the omitted domains (tweets or legal documents, where dependency parsing is less reliable) or on obfuscated variants that preserve annotation accuracy, and check whether the mean score collapses to the level of the ELOQUENT-obfuscated results; alternatively, retrain with the same features but gold-standard annotations on a held-out sample and see whether scores improve substantially.
Extended reading notes
Core claim
On the PAN 2025 'AI Detection Sensitivity' subtask, the paper's best model—dubbed big-cv, a LightGBM with DART boosting, 20 leaves, 1500 iterations, depth 12, and probability scores averaged over ten cross-validation folds—achieves an arithmetic-mean test score of 0.921 on the unobfuscated test set and 0.897 in the final macro-averaged evaluation across all datasets, including the obfuscated ELOQUENT contributions. The performance pattern across the small, medium, and big variants shows that larger model capacity consistently raised scores on validation and unobfuscated test data, while the obfuscated ELOQUENT set dropped scores, leading to the paper's two headline observations: capacity helps, obfuscation hurts. The authors therefore position the result as a trade-off between the smaller cost and greater explainability of boosted trees and the better generalization of neural-based systems.
Load-bearing premise
The pipeline assumes that spaCy's automatic annotations are accurate and consistent across all text domains, generators, and the test distribution, and that the pooled training corpora—which omit genres such as tweets and legal documents—are a faithful, unbiased sample of human and machine text.
Editorial extensions
If this is right
- Larger LightGBM capacity (more leaves, iterations, and depth) raises detection performance on validation and unobfuscated test sets.
- Obfuscation, as added in the ELOQUENT set, considerably reduces the model's mean evaluation score.
- Augmenting the training set with obfuscated samples should further improve robustness.
- Adding TF-IDF features or standardizing feature frequencies, following stylometric practice, could close the gap to the TF-IDF baseline.
- The main computational bottleneck is feature extraction on the large training set; classifier training, inference, and explanation are inexpensive.
Reading between the lines
- Because the features are linguistic-frequency counts, the approach is likely to transfer poorly to languages or domains where spaCy annotation quality drops, even if the classifier is retrained.
- The gap to the TF-IDF SVM suggests that the baseline's surface token frequencies capture a signal—possibly formatting artifacts and token-sequence regularities—that the grammar-based features discard.
- The reported vulnerability to obfuscation indicates that paraphrase attacks systematically distort the annotated n-gram distributions, so the most direct testable improvement is adversarial augmentation with paraphrased machine texts.
- The authors note that some genres were excluded from training, which implies an upper bound on generalization; a variant trained on all available genres would test whether the pooled-corpus strategy, rather than the feature set, drives the score.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript describes StyOch, a non-neural system submitted to Subtask 1 (binary AI detection) of the PAN 2025 Voight-Kampff task. The system trains LightGBM classifiers on a pooled corpus of 563,571 texts from public MGT benchmarks, using spaCy-derived frequency features: lemma n-grams, POS n-grams, dependency bigrams, and morphological/NER annotations. The authors report validation and test scores for four model variants, with the selected big-cv model achieving a final macro-averaged mean evaluation score of 0.897, below the TF-IDF SVM baseline (0.922) and the best system (0.989). The paper concludes that larger model capacity and cross-validation improve performance, while obfuscation substantially degrades it.
Significance. If the reported results are taken at face value, the paper is a useful system-description datapoint: a cheap, explainable, feature-based pipeline without neural components reaches competitive but not state-of-the-art performance on a standardized shared task. The paper is honest about its limitations, explicitly noting that it did not beat the TF-IDF baseline and that obfuscation hurt performance. Strengths include evaluation on held-out test data through TIRA, reporting of six metrics, transparent comparison of four model variants, and a large public training corpus. The main weaknesses are the absence of any deduplication or overlap analysis between the training pool and the test data, and the lack of statistical precision or a clearly stated model-selection rule.
major comments (2)
- [Section 4.1, Table 3] The central empirical claim is the final mean score of 0.897 in Table 4(c), but the paper never reports a deduplication or overlap analysis between the 563,571 training texts and the PAN 2025 test or ELOQUENT evaluation collections. Several training sources (MAGE, M4, HC3 Plus, MULTITuDE) draw human and machine texts from Wikipedia, news, reviews, and other open corpora that may overlap with the test distribution. Because the classifier uses frequency-based n-gram features, even near-duplicate texts could share enough feature mass to inflate the measured score. Please add an exact and near-duplicate overlap analysis (e.g., hash-based deduplication or n-gram overlap) or explicitly justify why the task design guarantees disjointness.
- [Section 4.2, Conclusion] Model selection and statistical uncertainty are not reported. The text says that both validation datasets 'could be used for classifier evaluation and selection,' but it does not state which set drove the choice of big-cv; note that big-cv-culled has a higher Validation 2 mean (0.951) than big-cv (0.933) while scoring lower on the test set (0.915 vs 0.921). Additionally, no confidence intervals, standard deviations, or significance tests accompany the capacity comparison. Please report the 10-fold CV variability and clarify the selection rule, otherwise the conclusion that 'larger capacity of boosted trees increased the detection performance' is not fully supported.
minor comments (7)
- [Table 1] The word-count entry for M4 is the unresolved placeholder 'XXX'; it should be filled in or replaced with a dash if unavailable.
- [Section 4.1] The sentence 'The final evaluation was also appended with the False Positive Rate (FNR) and False Negative Rate (FNR)' contains a duplicated acronym; it should read 'False Positive Rate (FPR) and False Negative Rate (FNR).'
- [Section 4.2] The conclusion sentence 'Although the our model have not reached the baseline TF-IDF scores' is grammatically incorrect; please revise.
- [Section 3.1] The sentence 'The total number of LLM labels available in that dataset was 348' is ambiguous; specify whether this refers to the pooled training corpus or to a particular dataset.
- [Section 3.2] There are typographical errors in the feature-description section, including 'we also testes so-calledculling' and the formatting of 'culling' as a threshold; please proofread.
- [Section 3.3] The phrase 'it is possible – and in fact in can be beneficial' contains a typo ('in can'); please correct to 'it can be.'
- [Author affiliation block] The email address field contains an unresolved template artifact ('/envel⌢pe-⌢penje...'); please replace it with a clean address or remove it.
Circularity Check
No significant circularity: the system is evaluated on held-out TIRA data, and self-citations are contextual rather than load-bearing.
full rationale
The paper makes no first-principles derivation claim; its central result is an empirically measured score on the PAN 2025 held-out test set (Table 4c, mean 0.897). The pipeline is transparent: spaCy features are extracted, LightGBM is trained on pooled public benchmarks, and scores are reported on Validation 1, Validation 2 (TIRA), and the final test set. Hyperparameters (num_leaves, num_iterations, max_depth) are chosen from pre-submission experiments on unrelated or smaller datasets and are explicitly not optimized on the test set (Section 3.3). The use of validation results to select the big-cv model is standard model selection, and the paper reports both validation and test outcomes, so the reported test score is not forced by construction. Self-citations such as [22] motivate the decision to expand training data and the choice of stylometric feature classes, but they do not constitute the evidence for the performance claim; the evidence is the independent TIRA evaluation. The skeptical concern about possible training/test overlap is a data-leakage or external-validity risk, not circular reasoning, and no circular step can be exhibited from the paper's own equations or citations. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (5)
- LGBM num_leaves =
10, 12, 20
- LGBM num_iterations =
100, 500, 1500
- LGBM max_depth =
8, 10, 12
- Feature culling minimum document frequency =
0.1 (culled model)
- Per-class feature cap =
1500 items per feature class
assumptions (3)
- domain assumption spaCy en_core_web_lg provides accurate tokenization, POS tagging, dependency parsing, morphology, and NER annotations for English text.
- domain assumption The labels in the pooled public datasets correctly identify human-written and machine-generated texts.
- domain assumption The TIRA validation and test sets are representative of the target distribution and are correctly labeled.
Cite this review
Pith. "Pith review of StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features." pith.science (2026). https://pith.science/paper/LLWA536D
@misc{pith2026250712064,
author = {Pith},
title = {Pith review of: StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features},
year = {2026},
howpublished = {\url{https://pith.science/paper/LLWA536D}},
note = {Machine review of arXiv:2507.12064}
}
read the original abstract
This submission to the binary AI detection task is based on a modular stylometric pipeline, where: public spaCy models are used for text preprocessing (including tokenisation, named entity recognition, dependency parsing, part-of-speech tagging, and morphology annotation) and extracting several thousand features (frequencies of n-grams of the above linguistic annotations); light-gradient boosting machines are used as the classifier. We collect a large corpus of more than 500 000 machine-generated texts for the classifier's training. We explore several parameter options to increase the classifier's capacity and take advantage of that training set. Our approach follows the non-neural, computationally inexpensive but explainable approach found effective previously.
Reference graph
Works this paper leans on
-
[1]
B. D. Lund, T. Wang, N. R. Mannuru, B. Nie, S. Shimray, Z. Wang, Chatgpt and a new academic reality: Artificial intelligence-written research papers and the ethics of the large language models in scholarly publishing, Journal of the Association for Information Science and Technology 74 (2023) 570–581
work page 2023
-
[2]
L. De Angelis, F. Baglivo, G. Arzilli, G. P. Privitera, P. Ferragina, A. E. Tozzi, C. Rizzo, Chatgpt and the rise of large language models: the new ai-driven infodemic threat in public health, Frontiers in public health 11 (2023) 1166120
work page 2023
-
[3]
J. Bevendorff, Y. Wang, J. Karlgren, M. Wiegmann, A. Tsivgun, J. Su, Z. Xie, M. Abassy, J. Mansurov, R. Xing, M. N. Ta, K. A. Elozeiri, T. Gu, R. V. Tomar, J. Geng, E. Artemova, A. Shelmanov, N. Habash, E. Stamatatos, I. Gurevych, P. Nakov, M. Potthast, B. Stein, Overview of the “Voight-Kampff” Generative AI Authorship Verification Task at PAN and ELOQUEN...
work page 2025
-
[4]
J. Bevendorff, D. Dementieva, M. Fröbe, B. Gipp, A. Greiner-Petter, J. Karlgren, M. Mayerl, P. Nakov, A. Panchenko, M. Potthast, A. Shelmanov, E. Stamatatos, B. Stein, Y. Wang, M. Wiegmann, E. Zangerle, Overview of PAN 2025: Voight-Kampff Generative AI Detection, Multilingual Text Detoxification, Multi-Author Writing Style Analysis, and Generative Plagiar...
work page 2025
-
[5]
E. Crothers, N. Japkowicz, H. L. Viktor, Machine-Generated Text: A Comprehensive Survey of Threat Models and Detection Methods, IEEE Access 11 (2023) 70977–71002. URL: https://doi.org/ 10.1109/ACCESS.2023.3294090. doi:10.1109/ACCESS.2023.3294090
-
[6]
J. Wu, S. Yang, R. Zhan, Y. Yuan, L. S. Chao, D. F. Wong, A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions, Computational Linguistics (2025) 1–64. URL: https://doi.org/10.1162/coli_a_00549. doi:10.1162/coli_a_00549
-
[7]
J. Bevendorff, M. Wiegmann, J. Karlgren, L. Dürlich, E. Gogoulou, A. Talman, E. Stamatatos, M. Potthast, B. Stein, Overview of the "Voight-Kampff" Generative AI Authorship Verification Task at PAN and ELOQUENT 2024, in: G. Faggioli, N. Ferro, P. Galuscáková, A. G. S. d. Herrera (Eds.), Working Notes of the Conference and Labs of the Evaluation Forum (CLEF...
work page 2024
-
[8]
M. Guo, Z. Han, H. Chen, J. Peng, A Machine-Generated Text Detection Model Based on Text Multi-Feature Fusion, in: G. Faggioli, N. Ferro, P. Galuscáková, A. G. S. d. Herrera (Eds.), Working Notes of the Conference and Labs of the Evaluation Forum (CLEF 2024), Grenoble, France, 9-12 September, 2024, volume 3740 of CEUR Workshop Proceedings, CEUR-WS.org, 20...
work page 2024
Show all 41 references
-
[9]
Miralles, A
P. Miralles, A. Martín, D. Camacho, Team aida at PAN: Ensembling Normalized Log Probabilities, in: G. Faggioli, N. Ferro, P. Galuščáková, A. G. S. Herrera (Eds.), Working Notes Papers of the CLEF 2024 Evaluation Labs, CEUR-WS.org, 2024, pp. 2807–2813. URL: http://ceur-ws.org/V...
2024
-
[10]
Yadagiri, D
A. Yadagiri, D. Kalita, A. Ranjan, A. K. Bostan, P. Toppo, P. Pakray, Team cnlp-nits-pp at PAN: Leveraging BERT for Accurate Authorship Verification: A Novel Approach to Textual Attribution, in: G. Faggioli, N. Ferro, P. Galuščáková, A. G. S. Herrera (Eds.), Working Notes Pape...
2024
-
[11]
L. Guo, W. Yang, L. Ma, J. Ruan, BLGAV: Generative AI Author Verification Model Based on BERT and BiLSTM, in: G. Faggioli, N. Ferro, P. Galuščáková, A. G. S. Herrera (Eds.), Working Notes Papers of the CLEF 2024 Evaluation Labs, CEUR-WS.org, 2024, pp. 2585–2592. URL: http: //c...
2024
-
[12]
Lorenz, F
L. Lorenz, F. Z. Aygüler, F. Schlatt, N. Mirzakhmedova, BaselineAvengers at PAN 2024: Often- Forgotten Baselines for LLM-Generated Text Detection, in: G. Faggioli, N. Ferro, P. Galuscáková, A. G. S. d. Herrera (Eds.), Working Notes of the Conference and Labs of the Evaluation ...
2024
-
[13]
Opara, StyloAI: Distinguishing AI-Generated Content with Stylometric Analysis, in: A
C. Opara, StyloAI: Distinguishing AI-Generated Content with Stylometric Analysis, in: A. M. Olney, I.-A. Chounta, Z. Liu, O. C. Santos, I. I. Bittencourt (Eds.), Artificial Intelligence in Education. Posters and Late Breaking Results, Workshops and Tutorials, Industry and Inno...
2024
-
[14]
doi:10.1007/978-3-031-64312-5_13
-
[15]
V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, S. Feizi, Can AI-Generated Text be Reliably Detected? Stress Testing AI Text Detectors Under Various Attacks, Transactions on Machine Learning Research (2025). URL: https://openreview.net/forum?id=OOgsAZdFOt
2025
-
[16]
Stiff, F
H. Stiff, F. Johansson, Detecting computer-generated disinformation, International Journal of Data Science and Analytics 13 (2022) 363–383. URL: https://doi.org/10.1007/s41060-021-00299-5. doi:10.1007/s41060-021-00299-5
2022 doi
-
[17]
M. M. Bhat, S. Parthasarathy, How Effectively Can Machines Defend Against Machine-Generated Fake News? An Empirical Study, in: A. Rogers, J. Sedoc, A. Rumshisky (Eds.), Proceedings of the First Workshop on Insights from Negative Results in NLP, Association for Computational Li...
2020
-
[18]
Crothers, N
E. Crothers, N. Japkowicz, H. Viktor, P. Branco, Adversarial Robustness of Neural-Statistical Features in Detection of Generative Transformers, in: 2022 International Joint Conference on Neural Networks (IJCNN), 2022, pp. 1–8. URL: https://ieeexplore.ieee.org/document/9892269....
2022
- [19]
- [20]
-
[21]
Przystalski, J
K. Przystalski, J. K. Argasiński, N. Lipp, D. Pacholczyk, Building Personality-Driven Lan- guage Models: How Neurotic is ChatGPT, Synthesis Lectures on Engineering, Science, and Technology, Springer Nature Switzerland, Cham, 2025. URL: https://link.springer.com/10.1007/ 978-3-...
2025 doi
-
[22]
A. M. Sarvazyan, J. À. González, P. Rosso, M. Franco-Salvador, Supervised Machine-Generated Text Detectors: Family and Scale Matters, in: A. Arampatzis, E. Kanoulas, T. Tsikrika, S. Vrochidis, A. Giachanou, D. Li, M. Aliannejadi, M. Vlachos, G. Faggioli, N. Ferro (Eds.), Exper...
2023 doi
-
[23]
Przystalski, J
K. Przystalski, J. K. Argasiński, I. Grabska-Gradzińska, J. Ochab, Stylometry recognizes human and llm-generated texts in short samples, 2025. Manuscript submitted for publication to *Expert Systems with Applications*
2025
-
[24]
A. M. Sarvazyan, J. À. González, M. Franco-Salvador, F. Rangel, B. Chulvi, P. Rosso, Overview of AuTexTification at IberLEF 2023: Detection and Attribution of Machine-Generated Text in Multiple Domains, in: Procesamiento del Lenguaje Natural, Jaén, Spain, 2023
2023
-
[25]
J. K. Argasiński, I. Grabska-Gradzińska, K. Przystalski, J. K. Ochab, T. Walkowiak, Stylomet- ric analysis of large language model-generated commentaries in the context of medical neuro- science, International Conference . . . (2024) 281–295. URL: https://link.springer.com/cha...
2024 doi
-
[26]
J. K. Ochab, T. Walkowiak, Implementing interpretable models in stylometric analysis, in: Digital Humanities 2024: Conference Abstracts, George Mason University (GMU), Washington, D.C., 2024
2024
-
[27]
Mikros, A
G. Mikros, A. Koursaris, D. Bilianos, ..., Ai-writing detection using an ensemble of transformers and stylometric features., IberLEF . . . (2023)
2023
-
[28]
P. Yu, J. Chen, X. Feng, Z. Xia, CHEAT: A Large-scale Dataset for Detecting CHatGPT-writtEn AbsTracts, IEEE Transactions on Big Data (2025) 1–9. URL: https://ieeexplore.ieee.org/abstract/ document/10858415. doi:10.1109/TBDATA.2025.3536929
2025
- [29]
-
[30]
Y. Li, Q. Li, L. Cui, W. Bi, Z. Wang, L. Wang, L. Yang, S. Shi, Y. Zhang, MAGE: Machine-generated Text Detection in the Wild, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...
2024 doi
-
[31]
Macko, R
D. Macko, R. Moro, A. Uchendu, J. Lucas, M. Yamashita, M. Pikuliak, I. Srba, T. Le, D. Lee, J. Simko, M. Bielikova, MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings of the 2023 Conference on Em...
2023 doi
-
[32]
Y. Wang, J. Mansurov, P. Ivanov, J. Su, A. Shelmanov, A. Tsvigun, C. Whitehouse, O. Mo- hammed Afzal, T. Mahmoud, T. Sasaki, T. Arnold, A. F. Aji, N. Habash, I. Gurevych, P. Nakov, M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text De- tectio...
2024
-
[33]
Okulska, D
I. Okulska, D. Stetsenko, A. Kołos, A. Karlińska, K. Głąbińska, A. Nowakowski, Stylometrix: An open-source multilingual tool for representing stylometric vectors, arXiv preprint arXiv:2309.12810 (2023)
2023 arXiv
-
[34]
M. Eder, M. Kestemont, J. Rybicki, Stylometry with R: A Package for Computational Text Analysis, The R Journal 8 (2016) 1–15. doi:10.32614/RJ-2016-007
2016 doi
-
[35]
Montani, M
I. Montani, M. Honnibal, M. Honnibal, A. Boyd, S. V. Landeghem, H. Peters, explosion/spaCy: v3.7.2: Fixes for APIs and requirements, 2023. URL: https://doi.org/10.5281/zenodo.10009823. doi:10.5281/zenodo.10009823
2023 doi
-
[36]
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, T.-Y. Liu, Lightgbm: A highly efficient gradient boosting decision tree, Advances in neural information processing systems 30 (2017) 3146–3154
2017
-
[37]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, E. Duchesnay, Scikit-learn: Machine learning in Python, Journal of Machine Learning Res...
2011
-
[38]
Fröbe, M
M. Fröbe, M. Wiegmann, N. Kolyada, B. Grahm, T. Elstner, F. Loebe, M. Hagen, B. Stein, M. Potthast, Continuous Integration for Reproducible Shared Tasks with TIRA.io, in: J. Kamps, L. Goeuriot, F. Crestani, M. Maistro, H. Joho, B. Davis, C. Gurrin, U. Kruschwitz, A. Caputo (Ed...
2023 doi
-
[39]
Macko, mdok of kinit: Robustly fine-tuned llm for binary and multiclass ai-generated text detection, 2025
D. Macko, mdok of kinit: Robustly fine-tuned llm for binary and multiclass ai-generated text detection, 2025. URL: https://arxiv.org/abs/2506.01702. arXiv:2506.01702
2025
-
[40]
Burrows, ‘Delta’: A Measure of Stylistic Difference and a Guide to Likely Authorship, Literary and Linguistic Computing 17 (2002) 267–287
J. Burrows, ‘Delta’: A Measure of Stylistic Difference and a Guide to Likely Authorship, Literary and Linguistic Computing 17 (2002) 267–287. doi:10.1093/llc/17.3.267
2002 doi
-
[41]
S. M. Lundberg, G. Erion, H. Chen, A. DeGrave, J. M. Prutkin, B. Nair, R. Katz, J. Himmelfarb, N. Bansal, S.-I. Lee, From local explanations to global understanding with explainable ai for trees, Nature Machine Intelligence 2 (2020) 2522–5839
2020
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.