Pith. sign in

REVIEW 1 cited by

Automation of Citation Screening for Systematic Literature Reviews using Neural Networks: A Replicability Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.07534 v1 pith:QXD4U3N7 submitted 2022-01-19 cs.IR

classification cs.IR
keywords citationdatasetsscreeningapproachesdeepfirstlearningliterature
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the process of Systematic Literature Review, citation screening is estimated to be one of the most time-consuming steps. Multiple approaches to automate it using various machine learning techniques have been proposed. The first research papers that apply deep neural networks to this problem were published in the last two years. In this work, we conduct a replicability study of the first two deep learning papers for citation screening and evaluate their performance on 23 publicly available datasets. While we succeeded in replicating the results of one of the papers, we were unable to replicate the results of the other. We summarise the challenges involved in the replication, including difficulties in obtaining the datasets to match the experimental setup of the original papers and problems with executing the original source code. Motivated by this experience, we subsequently present a simpler model based on averaging word embeddings that outperforms one of the models on 18 out of 23 datasets and is, on average, 72 times faster than the second replicated approach. Finally, we measure the training time and the invariance of the models when exposed to a variety of input features and random initialisations, demonstrating differences in the robustness of these approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Reproducibility and Generalizability Study of Large Language Models for Query Generation

    cs.IR 2024-11 conditional novelty 5.0 of 10

    LLM-generated Boolean queries for systematic reviews are unstable across seeds, and the original ChatGPT results could not be reproduced with the documented setup.

Pith tools