Fine-tuned transformer models classify Swiss National Science Foundation grant peer review sentences into twelve content categories with an average macro F1 of 0.85, enabling large-scale analysis of review reports.
Automatic Analysis of Substantiation in Scientific Peer Reviews
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
With the increasing amount of problematic peer reviews in top AI conferences, the community is urgently in need of automatic quality control measures. In this paper, we restrict our attention to substantiation -- one popular quality aspect indicating whether the claims in a review are sufficiently supported by evidence -- and provide a solution automatizing this evaluation process. To achieve this goal, we first formulate the problem as claim-evidence pair extraction in scientific peer reviews, and collect SubstanReview, the first annotated dataset for this task. SubstanReview consists of 550 reviews from NLP conferences annotated by domain experts. On the basis of this dataset, we train an argument mining system to automatically analyze the level of substantiation in peer reviews. We also perform data analysis on the SubstanReview dataset to obtain meaningful insights on peer reviewing quality in NLP conferences over recent years.
fields
econ.EM 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Supervised Machine Learning Approach for Assessing Grant Peer Review Reports
Fine-tuned transformer models classify Swiss National Science Foundation grant peer review sentences into twelve content categories with an average macro F1 of 0.85, enabling large-scale analysis of review reports.