PreScience: A Dataset and Benchmark for Scientific Forecasting

Amanpreet Singh; Anirudh Ajith; Austin C. Kozlowski; Daniel S. Weld; Doug Downey; James Evans; Jay DeYoung; Nadav Kunievsky; Oyvind Tafjord; Tom Hope

arxiv: 2602.20459 · v2 · pith:U2HK5TZ3new · submitted 2026-02-24 · 💻 cs.AI · cs.CL

PreScience: A Dataset and Benchmark for Scientific Forecasting

Anirudh Ajith , Amanpreet Singh , Jay DeYoung , Nadav Kunievsky , Austin C. Kozlowski , Oyvind Tafjord , James Evans , Daniel S. Weld

show 2 more authors

Tom Hope Doug Downey

This is my paper

classification 💻 cs.AI cs.CL

keywords presciencecitationdatasetforecastingpredictionscientificallenaiauthor

0 comments

read the original abstract

Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark for scientific forecasting built around 98K recent AI research papers, together with companion papers covering author publication histories and citation links, yielding 502K papers in total. The resulting paper records include titles, abstracts, disambiguated author identities, influential references, topic labels, citation trajectories, and metadata snapshotted to respect temporal cutoffs. We instantiate seven exemplar tasks: five paper-anchored tasks -- contribution generation, collaborator prediction, prior work selection, citation count prediction, and future combination prediction -- and two aggregate topic trend forecasting variants. We develop baselines ranging from simple heuristics and embedding methods to frontier language models and agentic systems, and introduce LACER, an LLM-based metric for evaluating similarity of generated contribution descriptions that agrees better with human judgments than existing metrics. Finally, we compose task models to generate a 12-month synthetic corpus and find that the resulting papers are systematically less diverse and less novel than human-authored research from the same period. We release the PreScience dataset (https://huggingface.co/datasets/allenai/prescience) and code (https://github.com/allenai/prescience).

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

GIANTS: Generative Insight Anticipation from Scientific Literature
cs.CL 2026-04 unverdicted novelty 8.0

GIANTS-4B, trained with RL on a new 17k-example benchmark of parent-to-child paper insights, achieves 34% relative improvement over gemini-3-pro in LM-judge similarity and is rated higher-impact by a citation predictor.
Forecasting Scientific Progress with Artificial Intelligence
cs.AI 2026-05 unverdicted novelty 7.0

Introduces the CUSP benchmark across 4760 events and finds frontier AI models can pick plausible directions but fail to predict whether or when scientific advances will occur, with performance varying by domain and in...
ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
cs.AI 2026-05 unverdicted novelty 5.0

ForeSci is a temporally controlled benchmark with 500 tasks for assessing LLM agents on forward-looking AI research judgments in four domains using cutoff-aligned knowledge bases.