Pith. sign in

REVIEW 2 cited by

Assessing the Ability of ChatGPT to Screen Articles for Systematic Reviews

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.06464 v1 pith:DFON5OT7 submitted 2023-07-12 cs.SE cs.CLcs.IR

classification cs.SEcs.CLcs.IR
keywords chatgptscreeningarticlesautomationfieldresearchreviewssystematic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

By organizing knowledge within a research field, Systematic Reviews (SR) provide valuable leads to steer research. Evidence suggests that SRs have become first-class artifacts in software engineering. However, the tedious manual effort associated with the screening phase of SRs renders these studies a costly and error-prone endeavor. While screening has traditionally been considered not amenable to automation, the advent of generative AI-driven chatbots, backed with large language models is set to disrupt the field. In this report, we propose an approach to leverage these novel technological developments for automating the screening of SRs. We assess the consistency, classification performance, and generalizability of ChatGPT in screening articles for SRs and compare these figures with those of traditional classifiers used in SR automation. Our results indicate that ChatGPT is a viable option to automate the SR processes, but requires careful considerations from developers when integrating ChatGPT into their SR tools.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI Simulation by Digital Twins: Systematic Survey, Reference Framework, and Mapping to a Standardized Architecture

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A systematic survey of digital twin enabled AI simulation produces the DT4AI reference framework and maps it to ISO 23247.

  2. LGAR: Zero-Shot LLM-Guided Neural Ranking for Abstract Screening in Systematic Literature Reviews

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LGAR combines zero-shot LLM graded relevance scoring with monoT5 re-ranking to rank abstracts for systematic reviews, outperforming QA-based baselines by 5-10 pp MAP on two benchmarks.

Pith tools