REVIEW 14 cited by
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Automated scientific discovery promises to accelerate progress across scientific domains. However, developing and evaluating an AI agent's capacity for end-to-end scientific reasoning is challenging as running real-world experiments is often prohibitively expensive or infeasible. In this work we introduce DISCOVERYWORLD, the first virtual environment for developing and benchmarking an agent's ability to perform complete cycles of novel scientific discovery. DISCOVERYWORLD contains a variety of different challenges, covering topics as diverse as radioisotope dating, rocket science, and proteomics, to encourage development of general discovery skills rather than task-specific solutions. DISCOVERYWORLD itself is an inexpensive, simulated, text-based environment (with optional 2D visual overlay). It includes 120 different challenge tasks, spanning eight topics each with three levels of difficulty and several parametric variations. Each task requires an agent to form hypotheses, design and run experiments, analyze results, and act on conclusions. DISCOVERYWORLD further provides three automatic metrics for evaluating performance, based on (a) task completion, (b) task-relevant actions taken, and (c) the discovered explanatory knowledge. We find that strong baseline agents, that perform well in prior published environments, struggle on most DISCOVERYWORLD tasks, suggesting that DISCOVERYWORLD captures some of the novel challenges of discovery, and thus that DISCOVERYWORLD may help accelerate near-term development and assessment of scientific discovery competency in agents. Code available at: www.github.com/allenai/discoveryworld
Forward citations
Cited by 14 Pith papers
-
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
NatureBench evaluates ten frontier AI coding agents on 90 tasks from Nature papers under web-search-disabled conditions and finds the strongest agent surpasses published SOTA on only 17.8% of tasks, succeeding mainly ...
-
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
Frontier AI agents can perform many components of self-replication, such as obtaining cloud compute and exfiltrating weights under weak defenses, but none can yet complete the hardest end-to-end replication tasks.
-
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
AI agents beat human ML experts on short (2-hour) research-engineering tasks, but human experts outperform agents given 8+ hour budgets, measured on seven new open-source RE-Bench environments.
-
Training AI Scientists to Replicate Research
A 27B-parameter post-trained agent, Faraday, outperforms frontier coding agents at replicating held-out research figures by directing a larger coding model as a tool.
-
DiG-bench: Discovery in Games
A 70-game interactive benchmark where agents must discover hidden rules and objectives, with human beatability on every game and frontier models failing on the hardest tiers.
-
PreScience: A Dataset and Benchmark for Scientific Forecasting
A new benchmark tests whether AI can forecast future scientific papers; frontier LLMs score ~5.6/10 on matching real abstracts, and simulated corpora are measurably less diverse and novel than human science.
-
Levels of Autonomy for AI Agents
A user-role-based five-level framework for designing, certifying, and evaluating AI agent autonomy as a choice independent of agent capability.
-
Sparks of Science: Hypothesis Generation Using Structured Paper Data
The authors built HypoGen, 5,478 Bit-Flip-Spark hypothesis triples with reasoning chains from NeurIPS 2023 and ICLR 2024 papers, and fine-tuned LLaMA models on it, reporting higher feasibility but lower diversity in g...
-
The AI Agent Index
The AI Agent Index catalogs 67 deployed agentic AI systems and shows that most developers publicly disclose little about safety policies and evaluations.
-
Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study
An LLM agent reproduces many known Russenorsk linguistic properties from a newly compiled dictionary, and can generate speculative Russenorsk translations, but the evaluation partly reflects prompt leakage.
-
LLM4SR: A Survey on Large Language Models for Scientific Research
A systematic review of LLM-based systems for hypothesis discovery, experiment planning, scientific writing, and peer review, including benchmarks, evaluation methods, and open challenges.
-
AIGS: Generating Science from AI-Powered Automated Falsification
Baby-AIGS is a multi-agent system that automates research through explicit falsification, outperforming its baseline on three ML tasks but lagging human experts.
-
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.
-
Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research
A perspective article reviews the state of using foundation models for laboratory automation and proposes a roadmap for fully autonomous experiments.
Discussion (0). Continue with ORCID to comment.