Pith. sign in

Zero-shot Generative Large Language Models for Systematic Review Screening Automation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Systematic reviews are crucial for evidence-based medicine as they comprehensively analyse published research findings on specific questions. Conducting such reviews is often resource- and time-intensive, especially in the screening phase, where abstracts of publications are assessed for inclusion in a review. This study investigates the effectiveness of using zero-shot large language models~(LLMs) for automatic screening. We evaluate the effectiveness of eight different LLMs and investigate a calibration technique that uses a predefined recall threshold to determine whether a publication should be included in a systematic review. Our comprehensive evaluation using five standard test collections shows that instruction fine-tuning plays an important role in screening, that calibration renders LLMs practical for achieving a targeted recall, and that combining both with an ensemble of zero-shot models saves significant screening time compared to state-of-the-art approaches.

fields

cs.AI 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

How Far Are AI Scientists from Changing the World?

cs.AI · 2025-07-31 · conditional · novelty 4.0

This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

citing papers explorer

Showing 1 of 1 citing paper.

  • How Far Are AI Scientists from Changing the World? cs.AI · 2025-07-31 · conditional · none · ref 173 · internal anchor

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.