Pith. sign in

REVIEW 14 cited by

DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06769 v2 pith:E5BYB6YG submitted 2024-06-10 cs.AI cs.CL

classification cs.AIcs.CL
keywords discoveryworlddiscoveryscientificagentagentsdevelopingenvironmentevaluating
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Automated scientific discovery promises to accelerate progress across scientific domains. However, developing and evaluating an AI agent's capacity for end-to-end scientific reasoning is challenging as running real-world experiments is often prohibitively expensive or infeasible. In this work we introduce DISCOVERYWORLD, the first virtual environment for developing and benchmarking an agent's ability to perform complete cycles of novel scientific discovery. DISCOVERYWORLD contains a variety of different challenges, covering topics as diverse as radioisotope dating, rocket science, and proteomics, to encourage development of general discovery skills rather than task-specific solutions. DISCOVERYWORLD itself is an inexpensive, simulated, text-based environment (with optional 2D visual overlay). It includes 120 different challenge tasks, spanning eight topics each with three levels of difficulty and several parametric variations. Each task requires an agent to form hypotheses, design and run experiments, analyze results, and act on conclusions. DISCOVERYWORLD further provides three automatic metrics for evaluating performance, based on (a) task completion, (b) task-relevant actions taken, and (c) the discovered explanatory knowledge. We find that strong baseline agents, that perform well in prior published environments, struggle on most DISCOVERYWORLD tasks, suggesting that DISCOVERYWORLD captures some of the novel challenges of discovery, and thus that DISCOVERYWORLD may help accelerate near-term development and assessment of scientific discovery competency in agents. Code available at: www.github.com/allenai/discoveryworld

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    NatureBench evaluates ten frontier AI coding agents on 90 tasks from Nature papers under web-search-disabled conditions and finds the strongest agent surpasses published SOTA on only 17.8% of tasks, succeeding mainly ...

  2. RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

    cs.CR 2025-04 conditional novelty 7.0 of 10

    Frontier AI agents can perform many components of self-replication, such as obtaining cloud compute and exfiltrating weights under weak defenses, but none can yet complete the hardest end-to-end replication tasks.

  3. RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

    cs.LG 2024-11 conditional novelty 7.0 of 10

    AI agents beat human ML experts on short (2-hour) research-engineering tasks, but human experts outperform agents given 8+ hour budgets, measured on seven new open-source RE-Bench environments.

  4. Training AI Scientists to Replicate Research

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A 27B-parameter post-trained agent, Faraday, outperforms frontier coding agents at replicating held-out research figures by directing a larger coding model as a tool.

  5. DiG-bench: Discovery in Games

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A 70-game interactive benchmark where agents must discover hidden rules and objectives, with human beatability on every game and frontier models failing on the hardest tiers.

  6. PreScience: A Dataset and Benchmark for Scientific Forecasting

    cs.AI 2026-02 conditional novelty 6.0 of 10

    A new benchmark tests whether AI can forecast future scientific papers; frontier LLMs score ~5.6/10 on matching real abstracts, and simulated corpora are measurably less diverse and novel than human science.

  7. Levels of Autonomy for AI Agents

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A user-role-based five-level framework for designing, certifying, and evaluating AI agent autonomy as a choice independent of agent capability.

  8. Sparks of Science: Hypothesis Generation Using Structured Paper Data

    cs.CL 2025-04 conditional novelty 6.0 of 10

    The authors built HypoGen, 5,478 Bit-Flip-Spark hypothesis triples with reasoning chains from NeurIPS 2023 and ICLR 2024 papers, and fine-tuned LLaMA models on it, reporting higher feasibility but lower diversity in g...

  9. The AI Agent Index

    cs.SE 2025-02 accept novelty 6.0 of 10

    The AI Agent Index catalogs 67 deployed agentic AI systems and shows that most developers publicly disclose little about safety policies and evaluations.

  10. Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study

    cs.CL 2025-05 conditional novelty 5.0 of 10

    An LLM agent reproduces many known Russenorsk linguistic properties from a newly compiled dictionary, and can generate speculative Russenorsk translations, but the evaluation partly reflects prompt leakage.

  11. LLM4SR: A Survey on Large Language Models for Scientific Research

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A systematic review of LLM-based systems for hypothesis discovery, experiment planning, scientific writing, and peer review, including benchmarks, evaluation methods, and open challenges.

  12. AIGS: Generating Science from AI-Powered Automated Falsification

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Baby-AIGS is a multi-agent system that automates research through explicit falsification, outperforming its baseline on three ML tasks but lagging human experts.

  13. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

  14. Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research

    cs.RO 2025-06 accept novelty 1.0 of 10

    A perspective article reviews the state of using foundation models for laboratory automation and proposes a roadmap for fully autonomous experiments.

Pith tools