REVIEW 1 major objections 1 minor 5 references
The Russian Legislative Corpus
T0 review · 1 major / 1 minor · reviewed 2026-05-23 · grok-4.3
Pith's one-line read A corpus compiles 304,382 Russian legislative texts from 1991 to 2025, available in plain and linguistically annotated forms.
desk verdict This is a data release of a Russian legislative corpus with UD annotations, but the collection method is undocumented enough that the size claims can't be verified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Russian Legislative Corpus itself, which supplies both raw texts and their conversion to Universal Dependencies CoNLL-U format for annotation.
What would settle it
Locating an official Russian law or regulation adopted between 1991 and 2025 that is absent from the corpus, or finding systematic duplicates or transcription mismatches when compared against government publication records.
Extended reading notes
Core claim
The authors have assembled and released a corpus of 304,382 Russian legislative documents (194,425,905 tokens) spanning 1991 to 2025, offered in a basic form with metadata and a detailed form that includes the original texts plus their CoNLL-U equivalents annotated for parts of speech, morphology, and syntactic dependencies.
Load-bearing premise
The collected documents accurately and exhaustively cover all primary and secondary Russian legislation in the period without major omissions, duplicates, or transcription errors.
Editorial extensions
If this is right
- Researchers can now run large-scale statistical or machine-learning studies on the language and structure of Russian statutes over three decades.
- The annotated version supports training or evaluation of dependency parsers and morphological analyzers on legal-domain Russian text.
- Temporal subsets of the corpus enable tracking of changes in legislative phrasing or complexity across years.
- The resource can serve as a test bed for information-extraction systems aimed at legal documents.
Reading between the lines
- Cross-referencing the corpus against official Russian government portals could quantify coverage gaps that the paper does not report.
- The same texts could be aligned with parallel corpora from other languages to study translation patterns in international law.
- Periodic updates to the corpus would allow longitudinal modeling of how new legislation interacts with prior texts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents the Russian Legislative Corpus, a collection of 304,382 primary and secondary legislative texts adopted between 1991 and 2025, totaling 194,425,905 tokens. Two versions are released: a basic version with simple metadata and a detailed version containing the original texts plus Universal Dependencies CoNLL-U annotations for POS tags, morphological features, and syntactic dependencies.
Significance. A verified exhaustive corpus of this scale would constitute a substantial resource for Russian NLP, enabling large-scale studies of legislative language, diachronic analysis, and training of domain-specific models. The inclusion of UD annotations further increases its utility for syntactic and morphological research. The significance, however, depends entirely on the undocumented collection and validation steps.
major comments (1)
- [Abstract / Data Collection (absent)] Abstract and entire manuscript: the headline claim that the 304,382 texts constitute a comprehensive, non-duplicative record of all Russian primary and secondary legislation 1991–2025 is unsupported by any description of data sources, scraping or API protocols, year-by-year benchmarking against official registries, deduplication methods, or sampling-based error audits. Without these, the numerical claims cannot be independently verified and the central contribution remains untestable.
minor comments (1)
- [Abstract] The abstract states token counts but provides no breakdown by year, document type, or source; adding such tables would improve transparency even if full methodology is added elsewhere.
Simulated Author's Rebuttal
Thank you for the thorough review. We agree that the data collection methodology requires more detailed documentation to support the claims of comprehensiveness. We will update the manuscript with a new section on corpus construction.
read point-by-point responses
-
Referee: Abstract and entire manuscript: the headline claim that the 304,382 texts constitute a comprehensive, non-duplicative record of all Russian primary and secondary legislation 1991–2025 is unsupported by any description of data sources, scraping or API protocols, year-by-year benchmarking against official registries, deduplication methods, or sampling-based error audits. Without these, the numerical claims cannot be independently verified and the central contribution remains untestable.
Authors: We will revise the manuscript to include a detailed 'Data Collection' section that describes the primary sources (official Russian Federation legal information portals), the automated collection protocols, deduplication procedures based on unique identifiers and content hashing, and cross-validation with annual official reports on legislative activity. This will substantiate the corpus's scope and allow for independent verification. revision: yes
Circularity Check
No circularity: data release paper with no derivations or fitted results
full rationale
The paper is a corpus presentation whose central claim is the existence, size (304,382 texts), and availability of the Russian legislative corpus. No equations, predictions, first-principles derivations, or fitted parameters are present. The collection is described as sourced from official portals but involves no mathematical reduction or self-citation chain that could be circular. The exhaustiveness claim is a factual assertion about data ingestion, not a derived result equivalent to its inputs by construction. This is the normal non-finding for a data-release manuscript.
Assumptions & free parameters
Cite this review
Pith. "Pith review of The Russian Legislative Corpus." pith.science (2026). https://pith.science/paper/2406.04855
@misc{pith2026240604855,
author = {Pith},
title = {Pith review of: The Russian Legislative Corpus},
year = {2026},
howpublished = {\url{https://pith.science/paper/2406.04855}},
note = {Machine review of arXiv:2406.04855}
}
read the original abstract
We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905 tokens). The corpus is available in two versions: the basic version contains texts with simple metadata, while the detailed version includes both the original texts and their equivalents converted to the Universal Dependencies CoNLL-U format, annotated with parts of speech, morphological features, and syntactic dependencies.
Figures
Reference graph
Works this paper leans on
-
[1]
Blinova, O. V. and N. A. Tarasov (2021). Complexity of Russian legal texts: assessment methods and language data. In Proceedings of the International Conference "Corpus Linguistics–2021" . https://events.spbu.ru/eventsContent/events/2021/corpora/Корпусная
work page 2021
-
[2]
Blinova, O. V. and N. A. Tarasov (2023). Language complexity across sub-styles and genres in legal Russian . Research Result. Theoretical and Applied Linguistics\/ 9\/ (2), 73--96. 10.18413/2313-8912-2023-9-2-0-5
-
[3]
Kuchakov, R. and D. Saveliev (2018). The complexity of the Russian legislation from the lexical and syntactic perspective ( Slozhnost' pravovyh aktov v Rossii : leksicheskoe i sintaksicheskoe kachestvo tekstov). Analytic report, IRL European University\/ . https://enforce.spb.ru/images/analit_zapiski/memo_readability_2018_web.pdf (in Russian)
work page 2018
-
[4]
Saveliev, D. (2018). On creating and using text of the Russian Federation corpus of legal acts acts as open dataset. Pravo. Zhurnal Vysshey shkoly ekonomiki\/ (1), 26--44. 10.17323/2072-8166.2018.1.26.44 (in Russian)
-
[5]
Saveliev, D. (2020). A study in complexity of sentences constituting Russian Federation legal acts. Pravo. Zhurnal Vysshey shkoly ekonomiki\/ (1), 50--74. 10.17323/2072-8166.2020.1.50.74 (in Russian)
Reviewed May 23, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.