REVIEW 1 cited by
Czech Text Document Corpus v 2.0
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and is freely available for research purposes at http://ctdc.kiv.zcu.cz/. This corpus was created in order to facilitate a straightforward comparison of the document classification approaches on Czech data. It is particularly dedicated to evaluation of multi-label document classification approaches, because one document is usually labelled with more than one label. Besides the information about the document classes, the corpus is also annotated at the morphological layer. This paper further shows the results of selected state-of-the-art methods on this corpus to offer the possibility of an easy comparison with these approaches.
Forward citations
Cited by 1 Pith paper
-
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
Circuits extracted from inflectionally varied Polish sentences are more robust to adversarial word attacks than circuits from syncretic or English variants, identifying layer-0 attention heads as inflection-specific.
Discussion (0). Continue with ORCID to comment.