REVIEW 2 major objections 4 minor 58 references
A Corpus of Persuasion Techniques in Slavic Languages
T0 review · 2 major / 4 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read A new Slavic corpus of 7,500 annotated spans lets machines spot 25 persuasion techniques in debates and social media.
desk verdict Solid new Slavic persuasion corpus with clear stats and baselines; missing IAA numbers are the main soft spot but do not kill usability. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The two-tier taxonomy of six rhetorical strategies (Attack on Reputation, Justification, Simplification, Distraction, Call, Manipulative Wording) that expand into 25 fine-grained techniques; every span and sentence is labelled according to this inventory, which then supports both multi-label classification and topic-technique correlation analysis.
What would settle it
Compute inter-annotator agreement (Cohen’s κ or Krippendorff’s α) on a held-out subset of the double-annotated documents; if agreement on the rarer techniques falls below conventional thresholds, the corpus cannot be treated as reliable gold for those classes.
Extended reading notes
Core claim
A publicly released multilingual corpus of approximately 7,500 text-span and sentence-level annotations of 25 persuasion techniques, drawn from 222 Bulgarian, Polish and Russian documents, supplies the first substantial training and evaluation resource for fine-grained persuasion detection in parliamentary debate and social-media genres for these languages.
Load-bearing premise
That dual annotation followed by curator correction yields labels reliable enough to serve as ground truth for a 25-class taxonomy, even though no inter-annotator agreement numbers are reported.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper releases a new multi-lingual corpus of persuasion techniques for three Slavic languages (Bulgarian, Polish, Russian), covering parliamentary debates and Telegram-style social media. Approximately 7 500 text-span annotations of 25 fine-grained techniques (grouped under six rhetorical strategies) are provided for 222 documents on contested national and international topics; sentence-level labels are derived by mapping. The authors describe acquisition, a multi-annotator + curator pipeline, topic distributions, co-occurrence heatmaps, and topic–technique correlations, then supply classical SVM, zero-shot LLM, and fine-tuned RoBERTa baselines for span- and sentence-level detection/classification. The central claim is that the resource is usable public ground truth that enables further research on these languages and genres.
Significance. If the labels are reliable, the corpus fills a clear gap: fine-grained span- and sentence-level persuasion annotations for Bulgarian and Polish parliamentary speech and for Russian social media have been scarce. The 25-technique taxonomy re-uses and modestly extends prior SemEval/CLEF work, the dual-level annotation format is practical, and the baselines (SVM micro-F1 0.11–0.35, LLM strategy F1 ~0.25–0.40, RoBERTa micro-F1 0.28–0.46) correctly illustrate task difficulty and give the community immediate reference numbers. Public release of data, guidelines, and code would therefore be a genuine service to Slavic NLP and disinformation research.
major comments (2)
- §3.3 (Annotation Process) and §3.5 (Statistics) never report any quantitative inter-annotator agreement (Cohen’s κ, Krippendorff’s α, pairwise F1, etc.) at either the 25-technique or the 6-strategy level, nor any curator-override rate. The authors themselves flag Whataboutism, Red Herring, Strawman, Causal/Consequential Oversimplification and Obfuscation as especially hard; without measured reliability it is impossible to know whether those classes are usable gold or largely noise. Process description alone does not substitute for IAA; this is load-bearing for the “usable ground-truth resource” claim.
- §5.1–5.2 baselines are trained and evaluated only on the newly released data with no external validation set, no comparison to prior English or multi-lingual persuasion corpora that share the same taxonomy, and no ablation of context length or multi-label handling. The reported numbers therefore demonstrate difficulty on this corpus but do not yet establish how transferable or competitive the resource is; a modest external or cross-corpus experiment would strengthen the benchmark claim.
minor comments (4)
- Figure 2 and Figure 3 captions and axis labels are dense; a short legend explaining the colour-coding of the six strategies would improve readability.
- Table 1 reports “AVG words/document” but does not state whether the average is tokenised or whitespace-split; a one-sentence clarification would help reproducibility.
- §3.4 annotation-format description is clear, yet an explicit statement of how multi-label sentences are encoded (order of labels, separator) would remove any remaining ambiguity for downstream users.
- A few typographical slips remain (e.g., “eWe believe” in §6, inconsistent capitalisation of technique names). A final proof-reading pass is warranted.
Circularity Check
No circularity: resource paper whose annotations, statistics and baselines are independent of any self-referential derivation.
full rationale
The paper’s central claims are the release of a new multi-lingual corpus (~7 500 text-span / sentence annotations of a 25-technique taxonomy across 222 documents) together with descriptive statistics, topic–technique co-occurrence heat-maps, and ordinary supervised / zero-shot baselines. There are no equations, no fitted constants re-used as “predictions,” no uniqueness theorems, and no ansatzes. The taxonomy itself is taken from earlier shared-task work by overlapping authors (Piskorski et al., 2023c; SlavicNLP 2025), but that is ordinary reuse of a publicly defined label set; the new contribution is the fresh dual-annotated data, not a re-derivation of the taxonomy. Baselines (character-n-gram SVM, zero-shot LLMs, language-specific RoBERTa fine-tunes) are trained and evaluated on the newly released splits; their performance numbers are therefore ordinary empirical results, not tautologies. Consequently the derivation chain is empty and the paper is free of the six circularity patterns.
Assumptions & free parameters
assumptions (2)
- domain assumption The 25-technique / 6-strategy taxonomy (extended from Piskorski et al. 2023c) is an adequate and transferable description of persuasion in parliamentary debates and Telegram posts.
- domain assumption Dual annotation followed by curator correction and cross-language meetings yields labels of sufficient quality to serve as gold standard.
Cite this review
Pith. "Pith review of A Corpus of Persuasion Techniques in Slavic Languages." pith.science (2026). https://pith.science/paper/V6NLPHQM
@misc{pith2026260710715,
author = {Pith},
title = {Pith review of: A Corpus of Persuasion Techniques in Slavic Languages},
year = {2026},
howpublished = {\url{https://pith.science/paper/V6NLPHQM}},
note = {Machine review of arXiv:2607.10715}
}
read the original abstract
Persuasion techniques are powerful rhetorical devices used to sway public opinion in a wide range of media. We present a new corpus of persuasion techniques, focusing on Slavic languages. The corpus contains documents in Bulgarian, Polish, and Russian, annotated with persuasion techniques at the coarse-grained text-span level and fine-grained sentence level. The techniques are drawn from a taxonomy of 25 fine-grained persuasion techniques, grouped under six broad categories of rhetorical persuasion strategies. The corpus contains approximately 7500 text spans from 222 documents that cover topics hotly debated at the national and international levels. We describe the corpus creation process, provide detailed statistics, and examine correlations between topics and persuasion techniques. We use classic ML-based and generative AI-based models to provide baselines and benchmark results for the detection and classification of persuasion techniques at the text-span level and sentence level.
Figures
Reference graph
Works this paper leans on
-
[1]
A rgotario: Computational Argumentation Meets Serious Games
Habernal, Ivan and Hannemann, Raffael and Pollak, Christian and Klamm, Christopher and Pauli, Patrick and Gurevych, Iryna. A rgotario: Computational Argumentation Meets Serious Games. EMNLP. 2017
2017
-
[2]
Modzelewski, Arkadiusz and Sosnowski, Witold and Labruna, Tiziano and Wierzbicki, Adam and Da San Martino, Giovanni. PC o T : Persuasion-Augmented Chain of Thought for Detecting Fake News and Social Media Disinformation. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. doi:10.18653/v1/2...
-
[3]
Fine-Grained Analysis of Propaganda in News Articles , booktitle =
Da San Martino, Giovanni and Yu, Seunghak and Barr\'. Fine-Grained Analysis of Propaganda in News Articles , booktitle =
-
[4]
MIPD : Exploring Manipulation and Intention In a Novel Corpus of P olish Disinformation
Modzelewski, Arkadiusz and Da San Martino, Giovanni and Savov, Pavel and Wilczy \'n ska, Magdalena Anna and Wierzbicki, Adam. MIPD : Exploring Manipulation and Intention In a Novel Corpus of P olish Disinformation. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.1103
-
[5]
Adapting Serious Game for Fallacious Argumentation to G erman: Pitfalls, Insights, and Best Practices
Habernal, Ivan and Pauli, Patrick and Gurevych, Iryna. Adapting Serious Game for Fallacious Argumentation to G erman: Pitfalls, Insights, and Best Practices. LREC: Conference on Language Resources and Evaluation. 2018
2018
-
[6]
The INCE p TION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation
Klie, Jan-Christoph and Bugert, Michael and Boullosa, Beto and Eckart de Castilho, Richard and Gurevych, Iryna. The INCE p TION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation. Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations. 2018
2018
-
[7]
2023 , author =
News Categorization, Framing and Persuasion Techniques: Annotation Guidelines , number =. 2023 , author =
2023
-
[8]
Cross-lingual Named Entity Corpus for S lavic Languages
Piskorski, Jakub and Marci \'n czuk, Micha and Yangarber, Roman. Cross-lingual Named Entity Corpus for S lavic Languages. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 2024
2024
Show all 58 references
-
[9]
Findings of the NLP 4 IF -2019 Shared Task on Fine-Grained Propaganda Detection
Da San Martino, Giovanni and Barr \'o n-Cede \ n o, Alberto and Nakov, Preslav. Findings of the NLP 4 IF -2019 Shared Task on Fine-Grained Propaganda Detection. Proceedings of the Second Workshop on Natural Language Processing for Internet Freedom: Censorship, Disinformation, ...
2019 doi
-
[10]
S em E val-2021 Task 6: Detection of Persuasion Techniques in Texts and Images
Dimitrov, Dimitar and Bin Ali, Bishr and Shaar, Shaden and Alam, Firoj and Silvestri, Fabrizio and Firooz, Hamed and Nakov, Preslav and Da San Martino, Giovanni. S em E val-2021 Task 6: Detection of Persuasion Techniques in Texts and Images. Proceedings of the 15th Internation...
2021 doi
-
[11]
S em E val-2020 Task 11: Detection of Propaganda Techniques in News Articles
Da San Martino, Giovanni and Barr \'o n-Cede \ n o, Alberto and Wachsmuth, Henning and Petrov, Rostislav and Nakov, Preslav. S em E val-2020 Task 11: Detection of Propaganda Techniques in News Articles. Proceedings of the Fourteenth Workshop on Semantic Evaluation. 2020. doi:1...
2020 doi
-
[12]
S em E val-2023 Task 3: Detecting the Category, the Framing, and the Persuasion Techniques in Online News in a Multi-lingual Setup
Piskorski, Jakub and Stefanovitch, Nicolas and Da San Martino, Giovanni and Nakov, Preslav. S em E val-2023 Task 3: Detecting the Category, the Framing, and the Persuasion Techniques in Online News in a Multi-lingual Setup. Proceedings of the 17th International Workshop on Sem...
2023 doi
-
[13]
The UNLP 2025 Shared Task on Detecting Social Media Manipulation
Kyslyi, Roman and Romanyshyn, Nataliia and Sydorskyi, Volodymyr. The UNLP 2025 Shared Task on Detecting Social Media Manipulation. Proceedings of the Fourth Ukrainian Natural Language Processing Workshop (UNLP 2025). 2025. doi:10.18653/v1/2025.unlp-1.12
2025 doi
-
[14]
Investigating Persuasion Techniques in
Abdurahmman Alzahrani and Eyad Babkier and Faisal Yanbaawi and Firas Yanbaawi and Hassan Alhuzali , year=. Investigating Persuasion Techniques in. 2405.12884 , archivePrefix=
-
[15]
S lavic NLP 2025 Shared Task: Detection and Classification of Persuasion Techniques in Parliamentary Debates and Social Media
Piskorski, Jakub and Dimitrov, Dimitar Iliyanov and Dobrani \'c , Filip and Ernst, Marina and Haneczok, Jacek and Koychev, Ivan and Ljube s i \'c , Nikola and Marcinczuk, Michal and Modzelewski, Arkadiusz and Moravski, Ivo and Yangarber, Roman. S lavic NLP 2025 Shared Task: De...
2025 doi
-
[16]
Multilingual Multifaceted Understanding of Online News in Terms of Genre, Framing, and Persuasion Techniques
Piskorski, Jakub and Stefanovitch, Nicolas and Nikolaidis, Nikolaos and Da San Martino, Giovanni and Nakov, Preslav. Multilingual Multifaceted Understanding of Online News in Terms of Genre, Framing, and Persuasion Techniques. Proceedings of the 61st Annual Meeting of the Asso...
2023 doi
-
[17]
Aho and Jeffrey D
Alfred V. Aho and Jeffrey D. Ullman , title =. 1972
1972
-
[18]
Publications Manual , year = "1983", publisher =
1983
-
[19]
Chandra and Dexter C
Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243
1981 doi
-
[20]
Scalable training of
Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of
-
[21]
Dan Gusfield , title =. 1997
1997
-
[22]
Tetreault , title =
Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =
2015
-
[23]
Overview of the
Jakub Piskorski and Nicolas Stefanovitch and Firoj Alam and Ricardo Campos and Dimitar Dimitrov and Al. Overview of the. Working Notes of the Conference and Labs of the Evaluation Forum. 2024 , url =
2024
-
[24]
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =
Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =
-
[25]
Proceedings of the 10th Workshop on Slavic Natural Language Processing 2025 (SlavicNLP 2025) , month =
Piskorski, Jakub and Dimitrov, Dimitar and Dobrani. Proceedings of the 10th Workshop on Slavic Natural Language Processing 2025 (SlavicNLP 2025) , month =. 2025 , address =
2025
-
[26]
S em E val-2024 Task 4: Multilingual Detection of Persuasion Techniques in Memes
Dimitrov, Dimitar and Alam, Firoj and Hasanain, Maram and Hasnat, Abul and Silvestri, Fabrizio and Nakov, Preslav and Da San Martino, Giovanni. S em E val-2024 Task 4: Multilingual Detection of Persuasion Techniques in Memes. Proceedings of the 18th International Workshop on S...
2024 doi
-
[27]
Overview of DIPROMATS 2023: automatic detection and characterization of propaganda techniques in messages from diplomats and authorities of world powers , volume=
Moral, Pablo and Marco, Guillermo and Gonzalo, Julio and Carrillo-de-Albornoz, Jorge and Gonzalo-Verdugo, Iv. Overview of DIPROMATS 2023: automatic detection and characterization of propaganda techniques in messages from diplomats and authorities of world powers , volume=. Pro...
2023
-
[28]
Overview of DIPROMATS 2024: Detection, Characterization and Tracking of Propaganda in Messages from Diplomats and Authorities of World Powers , volume=
Moral, Pablo and Fraile, Jes. Overview of DIPROMATS 2024: Detection, Characterization and Tracking of Propaganda in Messages from Diplomats and Authorities of World Powers , volume=. Procesamiento del lenguaje natural , publisher=. 2024 , pages=
2024
-
[29]
A r AIE val Shared Task: Persuasion Techniques and Disinformation Detection in A rabic Text
Hasanain, Maram and Alam, Firoj and Mubarak, Hamdy and Abdaljalil, Samir and Zaghouani, Wajdi and Nakov, Preslav and Da San Martino, Giovanni and Freihat, Abed. A r AIE val Shared Task: Persuasion Techniques and Disinformation Detection in A rabic Text. Proceedings of ArabicNL...
2023 doi
-
[30]
Arid and Ahmad, Fatema and Suwaileh, Reem and Biswas, Md
Hasanain, Maram and Hasan, Md. Arid and Ahmad, Fatema and Suwaileh, Reem and Biswas, Md. Rafiul and Zaghouani, Wajdi and Alam, Firoj. A r AIE val Shared Task: Propagandistic Techniques Detection in Unimodal and Multimodal A rabic Content. Proceedings of the Second Arabic Natur...
2024 doi
-
[31]
Overview of the WANLP 2022 Shared Task on Propaganda Detection in A rabic
Alam, Firoj and Mubarak, Hamdy and Zaghouani, Wajdi and Da San Martino, Giovanni and Nakov, Preslav. Overview of the WANLP 2022 Shared Task on Propaganda Detection in A rabic. Proceedings of the Seventh Arabic Natural Language Processing Workshop (WANLP). 2022. doi:10.18653/v1...
2022 doi
-
[32]
2025 , address =
Wang, Yutong and Nurbakova, Diana and Calabretto, Sylvie , booktitle =. 2025 , address =
2025
-
[33]
Multilabel Classification of Persuasion Techniques with self-improving
Sawi. Multilabel Classification of Persuasion Techniques with self-improving. Proceedings of the 10th Workshop on Slavic Natural Language Processing 2025 (SlavicNLP 2025) , editor =. 2025 , address =
2025
-
[34]
Robust Detection of Persuasion Techniques in
Ksi. Robust Detection of Persuasion Techniques in. Proceedings of the 10th Workshop on Slavic Natural Language Processing 2025 (SlavicNLP 2025) , editor =. 2025 , address =
2025
-
[35]
2025 , address =
Senichev, Sergey and Boriskin, Aleksandr and Krayko, Nikita and Galimzianova, Daria , booktitle =. 2025 , address =
2025
-
[36]
2025 , address =
Jose, Julia and Greenstadt, Rachel , booktitle =. 2025 , address =
2025
-
[37]
Empowering Persuasion Detection in
Xin, Zou and Chuhan, Wang and Dailin, Li and Yanan, Wang and Jian, Wang and Hongfei, Lin , booktitle=. Empowering Persuasion Detection in. 2025 , address=
2025
-
[38]
Hierarchical Classification of Propaganda Techniques in
Br. Hierarchical Classification of Propaganda Techniques in. Proceedings of the 10th Workshop on Slavic Natural Language Processing 2025 (SlavicNLP 2025) , editor =. 2025 , address =
2025
-
[39]
Fine‑Tuned Transformers for Detection and Classification of Persuasion Techniques in
Ekaterina Loginova , booktitle =. Fine‑Tuned Transformers for Detection and Classification of Persuasion Techniques in. 2025 , address =
2025
-
[40]
Fine-Tuned Transformer-Based Weighted Ensemble for Binary Classification in
Yahan, Mahshar and Sarker, Sakib and Amanul Islam, Mohammad , booktitle =. Fine-Tuned Transformer-Based Weighted Ensemble for Binary Classification in. 2025 , address =
2025
-
[41]
Language Resources and Evaluation , pages=
Erjavec, Toma. Language Resources and Evaluation , pages=. 2024 , publisher=
2024
-
[42]
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
Mochtak, Michal and Rupnik, Peter and Ljube. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
2024
-
[43]
Proceedings of machine translation summit x: papers , pages=
Europarl: A parallel corpus for statistical machine translation , author=. Proceedings of machine translation summit x: papers , pages=
-
[44]
International Conference on Speech and Computer , pages=
The parlaspeech collection of automatically generated speech and text datasets from parliamentary proceedings , author=. International Conference on Speech and Computer , pages=. 2024 , organization=
2024
-
[45]
Conference on Language Technologies & Digital Humanities 2022
The ParlaSpeech-HR benchmark for speaker profiling in Croatian , author=. Conference on Language Technologies & Digital Humanities 2022. 2002
2022
-
[46]
Leveraging Open Large Language Models for Multilingual Policy Topic Classification: The
Seb. Leveraging Open Large Language Models for Multilingual Policy Topic Classification: The. Social Science Computer Review , XXXpages=. 2024 , publisher=
2024
-
[47]
Jan-Christoph Klie and Michael Bugert and Beto Boullosa and Richard Eckart de Castilho and Iryna Gurevych , month =. The. Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations , url =. 2018 , location =
2018
-
[48]
Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation Campaign
Stefanovitch, Nicolas and Piskorski, Jakub. Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation Campaign. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. doi:10.18653/v1/2023....
2023 doi
-
[49]
Multi-source, multilingual information extraction and summarization , pages =
Information extraction: past, present and future , author =. Multi-source, multilingual information extraction and summarization , pages =
-
[50]
arXiv preprint arXiv:2004.14224 , year=
Exploiting structured knowledge in text via graph-guided representation learning , author=. arXiv preprint arXiv:2004.14224 , year=
2004 arXiv
-
[51]
arXiv preprint arXiv:2010.12688 , year=
Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training , author=. arXiv preprint arXiv:2010.12688 , year=
2010 arXiv
-
[52]
Multilingual Real-Time Event Extraction for Border Security Intelligence Gathering , booktitle =
Martin Atkinson and Jakub Piskorski and Erik van der Goot and Roman Yangarber , SSSauthor =. Multilingual Real-Time Event Extraction for Border Security Intelligence Gathering , booktitle =
-
[53]
Verification of Facts across Document Boundaries
Roman Yangarber , SSSauthor =. Verification of Facts across Document Boundaries. Proceedings of the International Workshop on Intelligent Information Access (
-
[54]
Diversity of Scenarios in Information Extraction
Silja Huttunen and Roman Yangarber and Ralph Grishman. Diversity of Scenarios in Information Extraction. Proceedings of the Third International Conference on Language Resources and Evaluation (LREC 2002)
2002
-
[55]
Evkoski, Bojan and Pollak, Senja , journal=
-
[56]
Journal of Computational Social Science , volume=
Sentiment and position-taking analysis of parliamentary debates: a systematic literature review , author=. Journal of Computational Social Science , volume=. 2020 , publisher=
2020
-
[57]
and Mikhailov, Vladislav and Fenogenova, Alena
Zmitrovich, Dmitry and Abramov, Aleksandr and Kalmykov, Andrey and Kadulin, Vitaly and Tikhonova, Maria and Taktasheva, Ekaterina and Astafurov, Danil and Baushenko, Mark and Snegirev, Artem and Shavrina, Tatiana and Markov, Sergei S. and Mikhailov, Vladislav and Fenogenova, A...
2024
-
[58]
Unsupervised Cross-lingual Representation Learning at Scale , journal =
Alexis Conneau and Kartikay Khandelwal and Naman Goyal and Vishrav Chaudhary and Guillaume Wenzek and Francisco Guzm. Unsupervised Cross-lingual Representation Learning at Scale , journal =. 2019 , url =. 1911.02116 , timestamp =
2019 arXiv
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.