Pith. sign in

REVIEW 3 major objections 8 minor 2 cited by

How Far Are AI Scientists from Changing the World?

T0 review · 3 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Current AI Scientist systems cannot yet produce papers that pass scientific muster; the strongest system averages 4.63/10 under an automated reviewer, and experimental weakness appears in every assessed paper.

desk verdict A useful, well-organized status survey of AI Scientist systems whose central 'not good enough' claim is right, but whose own new measurement leans on an uncalibrated AI reviewer and should be treated as illustrative, not definitive. read the letter →

arxiv 2507.23276 v2 pith:RY7EVEZT submitted 2025-07-31 cs.AI

classification cs.AI
keywords AIScientistlargelanguagemodelsautomatedscientificdiscoveryhypothesisgenerationverificationandfalsificationpeerreviewcapabilityframework
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that today's large-language-model-based AI Scientist systems are far from ready to reshape scientific research. The authors propose a four-level capability ladder—knowledge acquisition, idea generation, verification and falsification, and evolution—and use it to organize the field. Their new empirical evidence is a quality assessment of 28 publicly available papers produced by five leading systems, scored by an AI reviewer model. The best average rating is 4.63 out of 10, and every assessed paper shows experimental weakness. The paper concludes that current systems cannot independently produce artifacts meeting high-quality scientific communication standards, and that the bottleneck is verification and implementation, not idea generation alone.

What carries the argument

The organizing device is a four-level capability framework that defines what a mature AI Scientist must do: acquire knowledge from literature, generate feasible novel hypotheses, verify and falsify them through experiments, and evolve from feedback. The load-bearing empirical instrument is DeepReviewer-14B, an AI reviewer model that rates papers on soundness, presentation, and contribution and also lists defect categories; Table 4 and Table 5 use it to score 28 papers from five systems. The framework matters because it converts the vague question "how far are we?" into a checkable milestone list, and the reviewer scores supply the quantitative answer.

What would settle it

Have a panel of human experts blind-review the same 28 papers, along with human-written accepted papers, and compare their scores and defect categories with DeepReviewer-14B's; if humans rate the AI papers near parity with accepted human work, or if DeepReviewer-14B gives equally low scores to strong human papers, the central claim is unsupported.

Watch

Extended reading notes

Core claim

The central claim is that no current AI Scientist system can autonomously carry out the full research loop well enough to produce work that would pass genuine scientific scrutiny. The evidence is twofold: on implementation benchmarks, even the strongest language models score low at reproducing or executing research code, and on the paper's own evaluation, all five surveyed systems average below 4.63 out of 10, with "Experimental Weakness" flagged in 100% of the 28 papers. The authors attribute the gap to limits of the foundation models—hallucination, costly knowledge updating, and catastrophic forgetting—and to underdeveloped research abilities in feasibility assessment, rigorous experimentation, and long-term planning. Their proposed remedy is an explicit "evolution" capability, where systems improve through self-reflection, external feedback, and structured collaboration, before they can be expected to produce ground-breaking discoveries.

Load-bearing premise

The load-bearing premise is that DeepReviewer-14B's quality scores are valid and unbiased measures of scientific merit, since the paper does not calibrate the model against human expert judgments or test whether its harshness affects all systems equally.

Editorial extensions

If this is right

  • If the assessment is right, workshop acceptance of AI-generated papers is weak evidence of scientific maturity, since the same systems' full outputs rate far below typical accepted work.
  • Verification and implementation, not idea generation, are the binding constraints; improving code execution and experimental design should come before claims of autonomous discovery.
  • The four-level ladder implies that an AI system must demonstrate all levels—including evolution through feedback—before it is called a scientist, giving the field a concrete evaluation target.
  • Deploying current systems without safeguards would flood peer review with low-quality artifacts, so the paper's proposed detection, labeling, and human oversight mechanisms become prerequisites.
  • The review's category data, with 100% of papers showing experimental weakness and 96.4% showing methodological unclarity, provides a checklist that future systems can be measured against to track real progress.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's scores come from an AI reviewer, which is itself a language model; a natural next test is to check whether the same reviewer would give equally low percentiles to human-written papers, which would reveal whether the scale is harsh overall rather than specific to AI output.
  • One testable extension is to run the evaluation longitudinally: if iterative review-feedback cycles raise later-generation papers above the current 4.63/10 ceiling, that would support the authors' claim that evolution, not raw model scale, is the missing ingredient.
  • The framework suggests a practical benchmarking protocol—score any new AI Scientist on all four levels separately—so progress claims can be compared across systems instead of relying on anecdotal acceptance at workshops.
  • Because the 28 papers are publicly available and therefore likely curated toward higher quality, the true typical output of these systems may be even weaker than the already low average scores reported.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. This survey proposes a four-level capability framework for AI Scientist systems (knowledge acquisition, idea generation, verification and falsification, and evolution) and reviews representative methods and benchmarks for each level. It also presents a new empirical evaluation in Section 5.3, where 28 publicly available papers from five AI Scientist systems are scored by DeepReviewer-14B, yielding average ratings below 4.63/10 and a 100% incidence of 'Experimental Weakness' across the papers. On this basis, the paper concludes that current AI Scientist systems cannot independently produce scientific artifacts that meet established standards for high-quality scientific communication. The survey closes with limitations of foundation models, research-capability gaps, ethical considerations, and future directions.

Significance. The survey is comprehensive and timely, and the proposed capability framework provides a useful organizing structure for a rapidly evolving field. The paper compiles external benchmarks (MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, ML-Dev-Bench) that credibly demonstrate implementation and verification gaps in current systems. The new evaluation of full AI-generated manuscripts, if properly validated, would be an important contribution. However, the quantitative 'not good enough' claim currently rests on an uncalibrated AI reviewer with overlapping authorship, and the paper does not provide the statistical context needed to interpret the scores. With appropriate calibration and error analysis, the central conclusion could be made solid; as it stands, the evidence is not yet load-bearing in its reported form.

major comments (3)
  1. [§5.3, Tables 4–5] The central claim that current AI Scientist systems 'cannot independently produce scientific artifacts that meet established standards' rests entirely on ratings from DeepReviewer-14B (Zhu et al., 2025), yet the paper provides no calibration against human expert judgments, no inter-rater reliability statistics, no error bars, and no description of the reference distribution that defines the 'Percentile' column. Because DeepReviewer shares two authors with this manuscript, systematic bias cannot be ruled out. This is load-bearing, since Tables 4–5 are the only direct evaluation of complete AI-generated manuscripts in the paper.
  2. [§5.3, sample and statistics] The evaluation sample consists of only 28 papers, with 2–10 per system, and the paper itself notes in the Table 4 caption that publicly available papers 'may be curated and therefore may not fully represent the typical output of each system.' The paper does not report confidence intervals or significance tests, so the ranking across systems and the aggregate 'not good enough' conclusion are not robust to sampling variability. The authors should either enlarge the sample, report uncertainty, or soften the claims to match the evidentiary strength.
  3. [§4.2 vs. §5.3] The external benchmarks in Table 2 (MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, ML-Dev-Bench) directly support claims about weak implementation and verification capabilities, but they do not measure the holistic manuscript-quality dimensions (soundness, presentation, contribution) that Tables 4–5 assess. The paper should explicitly separate these two kinds of evidence and state that the manuscript-quality conclusion depends on the unvalidated DeepReviewer evaluation, not on the benchmarks.
minor comments (8)
  1. [§2.1] The model name 'CinicalBERT' should be 'ClinicalBERT' (Huang et al., 2019).
  2. [Figure 3] Figure 3 contains garbled text including long '/uni00000029/uni00000044/...' sequences and an unreadable table header ('Type Total citations Avg. citations'); this likely reflects a rendering or encoding error that must be fixed.
  3. [Throughout] The system name 'The AI Scientist' is typeset with non-standard spacing (e.g., 'A I Sc i e n t i s t'), making some sentences difficult to read; please use consistent formatting.
  4. [§5.3] In the sentence 'current AI Scientist systems not only struggle with scientific execution but also stuck with clearly articulating their research findings,' the phrase 'but also stuck' should be 'but also struggle.'
  5. [Table 4] The 'Percentile' column is undefined; the caption should state the reference distribution (e.g., percentile relative to what population of papers or scores).
  6. [§5.3 / Table 4] The paper does not describe how DeepReviewer-14B computes the overall 'Rating' from the three sub-scores (Soundness, Presentation, Contribution), nor whether the 0–10 scale permits non-integer intermediate values; please clarify the scoring protocol.
  7. [References] Several references have inconsistent formatting, including 'KABENAMUALU et al., 2023' in all caps and 'preprent' in Skarlinski et al. (2024); please proofread the bibliography.
  8. [§1] The phrase 'prospect-driven review' is used but never defined; consider adding one sentence explaining what makes the review 'prospect-driven.'

Circularity Check

1 steps flagged · score 4.0 of 10

Section 5.3's 'not good enough' result is carried by DeepReviewer-14B, a reviewer built by overlapping authors, used without independent calibration; the survey's broader gap claim retains independent support from external benchmarks.

  1. self citation load bearing [Section 5.3, Table 4 and Table 5]
    "To quantify these gaps, we employ DeepReviewer-14B (Zhu et al., 2025), an advanced AI reviewer model, to assess 28 publicly available research papers produced by 5 leading AI Scientist systems."

    The central claim of Section 5.3 — that current AI Scientist systems 'cannot independently produce scientific artifacts that meet established standards' — is operationalized entirely through DeepReviewer-14B's scores and defect categories. DeepReviewer-14B is cited only to Zhu et al. (2025), whose author list includes four present authors (Zhu, Weng, Yang, Zhang); no independent calibration against human reviewer ratings or inter-rater reliability is reported in this paper. The instrument also scores CycleResearcher-12B (Weng et al., 2025), another overlapping-author system in Table 4. The conclusion therefore reduces, for the full-manuscript-quality claim, to a self-cited evaluator rather than an externally validated standard.

full rationale

Most of the survey is a literature review and a capability taxonomy, which is not circular. Section 4.2 and Table 2 cite third-party benchmarks (MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, ML-Dev-Bench) that independently show implementation and verification gaps, so the broad 'considerable gap' conclusion does not reduce to the authors' own work. However, the specifically quantitative full-manuscript claim in Section 5.3 rests on DeepReviewer-14B, a model from Zhu et al. (2025) with four overlapping authors, and the paper provides no human calibration, no reference distribution for the Percentile column, and no external validation of the reviewer. Because that section is the only place where complete AI-generated manuscripts are directly evaluated, the strong 'cannot meet established standards' formulation is partially load-bearing on a self-citation. This is not a by-construction fit, so the score is moderate rather than high.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The survey itself introduces no fitted numerical parameters and no new physical or conceptual entities. Its load-bearing assumptions are the capability framework, the validity of DeepReviewer as a quality oracle, the representativeness of external benchmarks, and the representativeness of the 28-paper sample. None of these are independently established within the paper.

assumptions (4)
  • ad hoc to paper The four-level capability framework (knowledge acquisition, idea generation, verification and falsification, evolution) is a valid decomposition of the path to a mature AI Scientist.
    Introduced by the paper as its organizing structure (Section 1, Figure 1); no independent evidence is provided that these four levels are necessary or sufficient, making it a postulate for the survey's roadmap.
  • domain assumption DeepReviewer-14B scores are a valid measure of scientific paper quality.
    Section 5.3 uses DeepReviewer-14B as the sole evaluator of the 28 papers in Table 4 without calibrating it against human expert judgments or reporting agreement statistics.
  • domain assumption Benchmark results cited from MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, and ML-Dev-Bench are accurate and representative of current LLM verification capabilities.
    Table 2 reports second-hand performance numbers from multiple cited papers as evidence for the verification gap; the survey does not re-run any of these benchmarks.
  • domain assumption The publicly available papers from each AI Scientist system are not so unrepresentative as to invalidate cross-system comparisons.
    The note under Table 4 admits that public availability may bias toward higher-quality outputs; the analysis proceeds despite this acknowledged selection concern.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Far Are AI Scientists from Changing the World?." pith.science (2026). https://pith.science/paper/RY7EVEZT

@misc{pith2026250723276,
  author       = {Pith},
  title        = {Pith review of: How Far Are AI Scientists from Changing the World?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RY7EVEZT}},
  note         = {Machine review of arXiv:2507.23276}
}
read the original abstract

The emergence of large language models (LLMs) is propelling automated scientific discovery to the next level, with LLM-based Artificial Intelligence (AI) Scientist systems now taking the lead in scientific research. Several influential works have already appeared in the field of AI Scientist systems, with AI-generated research papers having been accepted at the ICLR 2025 workshop, suggesting that a human-level AI Scientist capable of uncovering phenomena previously unknown to humans, may soon become a reality. In this survey, we focus on the central question: How far are AI scientists from changing the world and reshaping the scientific research paradigm? To answer this question, we provide a prospect-driven review that comprehensively analyzes the current achievements of AI Scientist systems, identifying key bottlenecks and the critical components required for the emergence of a scientific agent capable of producing ground-breaking discoveries that solve grand challenges. We hope this survey will contribute to a clearer understanding of limitations of current AI Scientist systems, showing where we are, what is missing, and what the ultimate goals for scientific AI should be.

Figures

Figures reproduced from arXiv: 2507.23276 by the authors.

Figure 1
Figure 1. The capability level of an AI Scientist, illustrating the progression from foundational knowledge acquisition (Level 1), through idea generation (Level 2), rigorous hypothesis verification and falsification (Level 3), to continuous evolution (Level 4). We outline the core functions for each capability level. However, the research process is inherently constrained by human limitations such as limited time and cogniti… view at source ↗
Figure 2
Figure 2. The current capability landscape of AI Scientist systems across four progressive levels. We summarize the current achievements for each level and highlight critical gaps before AI Scientist systems can autonomously make ground-breaking scientific discoveries. 3 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. An analysis of the number of publications in the field of AI Scientist systems on arXiv. The upper panel displays the average number of citations up to now, cate￾gorized by containing implementation details. The lower panel shows the growth in the total number of these pa￾pers with the same categorization. Traditionally, scientific verification requires human scientists to design, implement, and analyze experi￾ments… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Three paradigms of AI reviewer systems with increasing complexity. The process begins with (1) classification & scoring systems that provide quantitative outputs (e.g., scores or accept/reject decisions). This paradigm gradually evolves into (2) generation Systems that…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction

    cs.AI 2026-07 accept novelty 7.0 of 10

    ZendoWorld shows that high labeling accuracy does not equal rule recovery, perception and induction are separate bottlenecks, and VLM agents propose near-uninformative experiments on active visual concept induction.

  2. Bridging the Gap on AI-Assisted Scientific Software Development Through Transparency and Traceability

    cs.SE 2026-05 conditional novelty 6.0 of 10

    Proposes guidance for responsible AI use in scientific software development under NQA-1 standards, illustrated with TMAP8 V&V cases to ensure accountability and auditability.

Reference graph

Works this paper leans on

214 extracted references · 14 canonical work pages · cited by 2 Pith papers

  1. [1]

    Litllm: A toolkit for scientific literature review

    Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H Laradji, Krishnamurthy DJ Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. Litllm: A toolkit for scientific literature review. arXiv preprint arXiv:2402.01788, 2024 a

  2. [2]

    Llms for literature review: Are we there yet? arXiv preprint arXiv:2412.15249, 2024 b

    Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H Laradji, Krishnamurthy DJ Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. Llms for literature review: Are we there yet? arXiv preprint arXiv:2412.15249, 2024 b

  3. [3]

    Litsearch: A retrieval benchmark for scientific literature search

    Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, and Tianyu Gao. Litsearch: A retrieval benchmark for scientific literature search. arXiv preprint arXiv:2407.18940, 2024

  4. [4]

    Automated literature review using nlp techniques and llm-based retrieval-augmented generation

    Nurshat Fateh Ali, Md Mahdi Mohtasim, Shakil Mosharrof, and T Gopi Krishna. Automated literature review using nlp techniques and llm-based retrieval-augmented generation. arXiv preprint arXiv:2411.18583, 2024

  5. [5]

    B io M -transformers: Building large biomedical language models with BERT , ALBERT and ELECTRA

    Sultan Alrowili and Vijay Shanker. B io M -transformers: Building large biomedical language models with BERT , ALBERT and ELECTRA . In Dina Demner-Fushman, Kevin Bretonnel Cohen, Sophia Ananiadou, and Junichi Tsujii, editors, Proceedings of the 20th Workshop on Biomedical Language Processing, pages 221--227, Online, June 2021. Association for Computationa...

  6. [6]

    Introducing claude 3.5 sonnet

    Anthropic. Introducing claude 3.5 sonnet. Anthropic Blog, 2024. URL https://www.anthropic.com/news/claude-3-5-sonnet

  7. [7]

    Constitutional ai: Harmlessness from ai feedback

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022

  8. [8]

    1,500 scientists lift the lid on reproducibility, 2016

    Monya Baker. 1,500 scientists lift the lid on reproducibility, 2016

Show all 214 references
  1. [9]

    Peerqa: A scientific question answering dataset from peer reviews

    Tim Baumg \"a rtner, Ted Briscoe, and Iryna Gurevych. Peerqa: A scientific question answering dataset from peer reviews. arXiv preprint arXiv:2502.13668, 2025

  2. [10]

    Joeran Beel, Min-Yen Kan, and Moritz Baumgart. Evaluating sakana's ai scientist for autonomous research: Wishful thinking or an emerging reality towards' artificial research intelligence'(ari)? arXiv preprint arXiv:2502.14297, 2025

  3. [11]

    Scibert: Pretrained language model for scientific text

    Iz Beltagy, Kyle Lo, and Arman Cohan. Scibert: Pretrained language model for scientific text. In EMNLP, 2019

  4. [12]

    Superintelligent agents pose catastrophic risks: Can scientist ai offer a safer path? arXiv preprint arXiv:2502.15657, 2025

    Yoshua Bengio, Michael Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, S \"o ren Mindermann, Adam Oberman, Jesse Richardson, Oliver Richardson, et al. Superintelligent agents pose catastrophic risks: Can scientist ai offer a safer path? arXiv preprin...

  5. [13]

    Has the machine learning review process become more arbitrary as the field has grown? the neurips 2021 consistency experiment

    Alina Beygelzimer, Yann N Dauphin, Percy Liang, and Jennifer Wortman Vaughan. Has the machine learning review process become more arbitrary as the field has grown? the neurips 2021 consistency experiment. arXiv preprint arXiv:2306.03262, 2023

  6. [14]

    Politepeer: does peer review hurt? a dataset to gauge politeness intensity in the peer reviews

    Prabhat Kumar Bharti, Meith Navlakha, Mayank Agarwal, and Asif Ekbal. Politepeer: does peer review hurt? a dataset to gauge politeness intensity in the peer reviews. Language Resources and Evaluation, 58 0 (4): 0 1291--1313, 2024

  7. [15]

    Iterative refinement of project-level code context for precise code generation with compiler feedback

    Zhangqian Bi, Yao Wan, Zheng Wang, Hongyu Zhang, Batu Guan, Fangxin Lu, Zili Zhang, Yulei Sui, Xuanhua Shi, and Hai Jin. Iterative refinement of project-level code context for precise code generation with compiler feedback. In Annual Meeting of the Association for Computationa...

  8. [16]

    Reflective multi-agent collaboration based on large language models

    Xiaohe Bo, Zeyu Zhang, Quanyu Dai, Xueyang Feng, Lei Wang, Rui Li, Xu Chen, and Ji-Rong Wen. Reflective multi-agent collaboration based on large language models. Advances in Neural Information Processing Systems, 37: 0 138595--138631, 2024

  9. [17]

    Chemcrow: Augmenting large-language models with chemistry tools

    Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Chemcrow: Augmenting large-language models with chemistry tools. arXiv preprint arXiv:2304.05376, 2023

  10. [18]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020

  11. [19]

    TLDR : Extreme summarization of scientific documents

    Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel Weld. TLDR : Extreme summarization of scientific documents. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4766--4777, Online, November 2020 a . Asso...

  12. [20]

    Tldr: Extreme summarization of scientific documents

    Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel S Weld. Tldr: Extreme summarization of scientific documents. arXiv preprint arXiv:2004.15011, 2020 b

  13. [21]

    Sang, Rahul K Arora, Robbie Kloosterman, Matthew Cecere, J

    Christian Cao, J. Sang, Rahul K Arora, Robbie Kloosterman, Matthew Cecere, J. Gorla, Richard Saleh, D. Chen, Ian Drennan, Bijan Teja, Michael Fehlings, P. Ronksley, Alexander A Leung, Dany E Weisz, Harriet Ware, Mairead Whelan, D. B. Emerson, Rahul K Arora, and Niklas Bobrovit...

  14. [22]

    Exploring scientific hypothesis generation with mamba

    Miaosen Chai, Emily Herron, Erick Cervantes, and Tirthankar Ghosal. Exploring scientific hypothesis generation with mamba. In Proceedings of the 1st Workshop on NLP for Science (NLP4Science), pages 197--207, 2024

  15. [23]

    Automated focused feedback generation for scientific writing assistance

    Eric Chamoun, Michael Schlichktrull, and Andreas Vlachos. Automated focused feedback generation for scientific writing assistance. arXiv preprint arXiv:2405.20477, 2024

  16. [24]

    Mle-bench: Evaluating machine learning agents on machine learning engineering

    Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al. Mle-bench: Evaluating machine learning agents on machine learning engineering. arXiv preprint arXiv:2410.07095, 2024

  17. [25]

    The toronto paper matching system: an automated paper-reviewer assignment system

    Laurent Charlin and Richard Zemel. The toronto paper matching system: an automated paper-reviewer assignment system. 2013

  18. [26]

    Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery

    Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, et al. Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery. arXiv preprint arXiv:2410.05080, 2024

  19. [27]

    React: A re view comment dataset for act ionability (and more)

    Gautam Choudhary, Natwar Modani, and Nitish Maurya. React: A re view comment dataset for act ionability (and more). In Web Information Systems Engineering--WISE 2021: 22nd International Conference on Web Information Systems Engineering, WISE 2021, Melbourne, VIC, Australia, Oc...

  20. [28]

    Relevai-reviewer: A benchmark on ai reviewers for survey paper relevance

    Paulo Henrique Couto, Quang Phuoc Ho, Nageeta Kumari, Benedictus Kent Rachmat, Thanh Gia Hieu Khuong, Ihsan Ullah, and Lisheng Sun-Hosoya. Relevai-reviewer: A benchmark on ai reviewers for survey paper relevance. arXiv preprint arXiv:2406.10294, 2024

  21. [29]

    Structured information extraction from scientific text with large language models

    John Dagdelen, Alexander Dunn, Sanghoon Lee, Nicholas Walker, Andrew S Rosen, Gerbrand Ceder, Kristin A Persson, and Anubhav Jain. Structured information extraction from scientific text with large language models. Nature Communications, 15 0 (1): 0 1418, 2024

  22. [30]

    Marg: Multi-agent review generation for scientific papers

    Mike D'Arcy, Tom Hope, Larry Birnbaum, and Doug Downey. Marg: Multi-agent review generation for scientific papers. arXiv preprint arXiv:2401.04259, 2024

  23. [31]

    ARIES : A corpus of scientific paper edits made in response to peer reviews

    Mike D ' Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, and Doug Downey. ARIES : A corpus of scientific paper edits made in response to peer reviews. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of ...

  24. [32]

    The general theory of relativity

    Albert Einstein. The general theory of relativity. In The meaning of relativity, pages 54--75. Springer, 1922

  25. [33]

    Introduction to artificial intelligence

    Wolfgang Ertel. Introduction to artificial intelligence. Springer Nature, 2024

  26. [34]

    Fox, Jennifer Meyer, and Emilie Aim \'e

    Charles W. Fox, Jennifer Meyer, and Emilie Aim \'e . Double-blind peer review affects reviewer ratings and editor decisions at an ecology journal. 37 0 (5): 0 1144--1157, May 2023. doi:10.1111/1365-2435.14259. Publisher Copyright: 2023 The Authors. Functional Ecology 2023 Brit...

  27. [35]

    Semantic scholar

    Suzanne Fricke. Semantic scholar. Journal of the Medical Library Association: JMLA, 106 0 (1): 0 145, 2018

  28. [36]

    Llm-ref: Enhancing reference handling in technical writing with large language models

    Kazi Ahmed Asif Fuad and Lizhong Chen. Llm-ref: Enhancing reference handling in technical writing with large language models. ArXiv, abs/2411.00294, 2024. URL https://api.semanticscholar.org/CorpusID:273798525

  29. [37]

    Reviewagents: Bridging the gap between human and ai-generated paper reviews

    Xian Gao, Jiacheng Ruan, Jingsheng Gao, Ting Liu, and Yuzhuo Fu. Reviewagents: Bridging the gap between human and ai-generated paper reviews. arXiv preprint arXiv:2503.08506, 2025

  30. [38]

    Retrieval-augmented generation for large language models: A survey

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997, 2 0 (1), 2023

  31. [39]

    Reviewer2: Optimizing review generation through prompt generation

    Zhaolin Gao, Kiant \'e Brantley, and Thorsten Joachims. Reviewer2: Optimizing review generation through prompt generation. arXiv preprint arXiv:2402.10886, 2024

  32. [41]

    Usefulness of llms as an author checklist assistant for scientific papers: Neurips'24 experiment

    Alexander Goldberg, Ihsan Ullah, Thanh Gia Hieu Khuong, Benedictus Kent Rachmat, Zhen Xu, Isabelle Guyon, and Nihar B Shah. Usefulness of llms as an author checklist assistant for scientific papers: Neurips'24 experiment. arXiv preprint arXiv:2411.03417, 2024 b

  33. [42]

    Towards an ai co-scientist

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2025

  34. [43]

    A survey on the rise of the ai scientists: Accelerating discovery and confronting ethical frontiers

    Muskaan Goyal. A survey on the rise of the ai scientists: Accelerating discovery and confronting ethical frontiers. World Journal of Advanced Engineering Technology and Sciences, 15: 0 564--569, 05 2025. doi:10.30574/wjaets.2025.15.2.0646

  35. [44]

    Agentic ai for scientific discovery: A survey of progress, challenges, and future directions

    Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack. Agentic ai for scientific discovery: A survey of progress, challenges, and future directions. arXiv preprint arXiv:2503.08979, 2025

  36. [45]

    Domain-specific language model pretraining for biomedical natural language processing, 2020

    Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. Domain-specific language model pretraining for biomedical natural language processing, 2020

  37. [46]

    Large language model based multi-agents: A survey of progress and challenges

    Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. Large language model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680, 2024

  38. [47]

    Automatic analysis of substantiation in scientific peer reviews

    Yanzhu Guo, Guokan Shang, Virgile Rennard, Michalis Vazirgiannis, and Chlo \'e Clavel. Automatic analysis of substantiation in scientific peer reviews. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023,...

  39. [48]

    Data extraction from polymer literature using large language models

    Sonakshi Gupta, Akhlak Mahmood, Pranav Shetty, Aishat Adeboye, and Rampi Ramprasad. Data extraction from polymer literature using large language models. Communications Materials, 5 0 (1): 0 269, 2024

  40. [49]

    Tanishq Gupta, Mohd Zaki, N. M. Anoop Krishnan, and Mausam. MatSciBERT : A materials domain language model for text mining and information extraction. npj Computational Materials, 8 0 (1): 0 102, May 2022. ISSN 2057-3960. doi:10.1038/s41524-022-00784-w. URL https://www.nature....

  41. [50]

    Pasa: An llm agent for comprehensive academic paper search, 2025

    Yichen He, Guanhua Huang, Peiyuan Feng, Yuan Lin, Yuchen Zhang, Hang Li, and Weinan E. Pasa: An llm agent for comprehensive academic paper search, 2025

  42. [51]

    The diminishing returns of masked language models to science, 2023

    Zhi Hong, Aswathy Ajith, Gregory Pauloski, Eamon Duede, Kyle Chard, and Ian Foster. The diminishing returns of masked language models to science, 2023

  43. [52]

    Automatic evaluation metrics for artificially generated scientific research

    Niklas H \"o pner, Leon Eshuijs, Dimitrios Alivanistos, Giacomo Zamprogno, and Ilaria Tiddi. Automatic evaluation metrics for artificially generated scientific research. arXiv preprint arXiv:2503.05712, 2025

  44. [53]

    Identification of tasks, datasets, evaluation metrics, and numeric scores for scientific leaderboards construction

    Yufang Hou, Charles Jochim, Martin Gleize, Francesca Bonin, and Debasis Ganguly. Identification of tasks, datasets, evaluation metrics, and numeric scores for scientific leaderboards construction. In Anna Korhonen, David Traum, and Llu \'i s M \`a rquez, editors, Proceedings o...

  45. [54]

    CHIME : LLM -assisted hierarchical organization of scientific studies for literature review support

    Chao-Chun Hsu, Erin Bransom, Jenna Sparks, Bailey Kuehl, Chenhao Tan, David Wadden, Lucy Wang, and Aakanksha Naik. CHIME : LLM -assisted hierarchical organization of scientific studies for literature review support. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Fi...

  46. [55]

    Nova: An iterative planning and search approach to enhance novelty and diversity of llm generated ideas

    Xiang Hu, Hongyu Fu, Jinge Wang, Yifeng Wang, Zhikun Li, Renjun Xu, Yu Lu, Yaochu Jin, Lili Pan, and Zhenzhong Lan. Nova: An iterative planning and search approach to enhance novelty and diversity of llm generated ideas. arXiv preprint arXiv:2410.14255, 2024

  47. [56]

    An overview of artificial intelligence ethics

    Changwu Huang, Zeqi Zhang, Bifei Mao, and Xin Yao. An overview of artificial intelligence ethics. IEEE Transactions on Artificial Intelligence, 4 0 (4): 0 799--819, 2022

  48. [57]

    Clinicalbert: Modeling clinical notes and predicting hospital readmission

    Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiv:1904.05342, 2019

  49. [58]

    Biomni: A general-purpose biomedical ai agent

    Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Yusuf Roohani, Ryan Li, Lin Qiu, Junze Zhang, Yin Di, et al. Biomni: A general-purpose biomedical ai agent. bioRxiv, pages 2025--05, 2025 a

  50. [59]

    A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Informatio...

  51. [60]

    Openreviewer: A specialized large language model for generating critical scientific paper reviews

    Maximilian Idahl and Zahra Ahmadi. Openreviewer: A specialized large language model for generating critical scientific paper reviews. arXiv preprint arXiv:2412.11948, 2024

  52. [61]

    Zochi technical report

    Intology. Zochi technical report. arXiv, 2025

  53. [62]

    Why most published research findings are false

    John PA Ioannidis. Why most published research findings are false. PLoS medicine, 2 0 (8): 0 e124, 2005

  54. [63]

    Discoveryworld: A virtual environment for developing and evaluating automated scientific discovery agents

    Peter Jansen, Marc-Alexandre C \^o t \'e , Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark. Discoveryworld: A virtual environment for developing and evaluating automated scientific discovery agents. Advances in Neur...

  55. [64]

    Codescientist: End-to-end semi-automated scientific discovery with code-based experimentation

    Peter Jansen, Oyvind Tafjord, Marissa Radensky, Pao Siangliulue, Tom Hope, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Daniel S Weld, and Peter Clark. Codescientist: End-to-end semi-automated scientific discovery with code-based experimentation. arXiv preprint arXiv:250...

  56. [65]

    Lgar: Zero-shot llm-guided neural ranking for abstract screening in systematic literature reviews

    Christian Jaumann, Andreas Wiedholz, and Annemarie Friedrich. Lgar: Zero-shot llm-guided neural ranking for abstract screening in systematic literature reviews. 2025. URL https://api.semanticscholar.org/CorpusID:279071051

  57. [66]

    Ai-researcher: Autonomous scientific innovation, 2025

    Tang Jiabin, Xia Lianghao, Li Zhonghang, and Huang Chao. Ai-researcher: Autonomous scientific innovation, 2025. URL https://arxiv.org/abs/2505.18705

  58. [67]

    Aide: Ai-driven exploration in the space of code

    Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, and Yuxiang Wu. Aide: Ai-driven exploration in the space of code. ArXiv, abs/2502.13138, 2025. URL https://api.semanticscholar.org/CorpusID:276421281

  59. [68]

    Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023. URL https://www.arxiv.org/abs/2310.06770

  60. [69]

    Probing biomedical embeddings from language models

    Qiao Jin, Bhuwan Dhingra, William Cohen, and Xinghua Lu. Probing biomedical embeddings from language models. In Proceedings of the 3rd Workshop on Evaluating Vector Space Representations for NLP, pages 82--89, 2019

  61. [70]

    A gent R eview: Exploring peer review dynamics with LLM agents

    Yiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen, Kaijie Zhu, Yijia Xiao, and Jindong Wang. A gent R eview: Exploring peer review dynamics with LLM agents. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in N...

  62. [71]

    Dsbench: How far are data science agents from becoming data science experts?, 2025

    Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu. Dsbench: How far are data science agents from becoming data science experts?, 2025. URL https://arxiv.org/abs/2409.07703

  63. [72]

    The global landscape of ai ethics guidelines

    Anna Jobin, Marcello Ienca, and Effy Vayena. The global landscape of ai ethics guidelines. Nature machine intelligence, 1 0 (9): 0 389--399, 2019

  64. [73]

    Cutting through the clutter: The potential of llms for efficient filtration in systematic literature reviews

    Lucas Joos, Daniel A Keim, and Maximilian T Fischer. Cutting through the clutter: The potential of llms for efficient filtration in systematic literature reviews. arXiv preprint arXiv:2407.10652, 2024

  65. [74]

    Highly accurate protein structure prediction with alphafold

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Z \' dek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. nature, 596 0 (7873): 0 583--589, 2021

  66. [75]

    Salomon Kabongo KABENAMUALU, Jennifer D’Souza, and S. Auer. Orkg-leaderboards: a systematic workflow for mining leaderboards as a knowledge graph. International Journal on Digital Libraries, pages 1--14, 2023. URL https://api.semanticscholar.org/CorpusID:258762176

  67. [76]

    Figureqa: An annotated figure dataset for visual reasoning

    Samira Ebrahimi Kahou, Adam Atkinson, Vincent Michalski, \'A kos K \'a d \'a r, Adam Trischler, and Yoshua Bengio. Figureqa: An annotated figure dataset for visual reasoning. ArXiv, abs/1710.07300, 2017. URL https://api.semanticscholar.org/CorpusID:3535069

  68. [77]

    A dataset of peer reviews (peerread): Collection, insights and nlp applications

    Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine Van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz. A dataset of peer reviews (peerread): Collection, insights and nlp applications. arXiv preprint arXiv:1804.09635, 2018

  69. [78]

    From who you know to what you read: Augmenting scientific recommendations with implicit social networks

    Hyeonsu B Kang, Rafal Kocielnik, Andrew Head, Jiangjiang Yang, Matt Latzke, Aniket Kittur, Daniel S Weld, Doug Downey, and Jonathan Bragg. From who you know to what you read: Augmenting scientific recommendations with implicit social networks. In Proceedings of the 2022 CHI Co...

  70. [79]

    Comlittee: Literature discovery with personal elected author committees

    Hyeonsu B Kang, Nouran Soliman, Matt Latzke, Joseph Chee Chang, and Jonathan Bragg. Comlittee: Literature discovery with personal elected author committees. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--20, 2023

  71. [80]

    AxCell : Automatic extraction of results from machine learning papers

    Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, and Robert Stojnic. AxCell : Automatic extraction of results from machine learning papers. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Con...

  72. [81]

    The perils of using mechanical turk to evaluate open-ended text generation

    Marzena Karpinska, Nader Akoury, and Mohit Iyyer. The perils of using mechanical turk to evaluate open-ended text generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1265--1285, 2021

  73. [82]

    Ethics of ai: A systematic literature review of principles and challenges

    Arif Ali Khan, Sher Badshah, Peng Liang, Muhammad Waseem, Bilal Khan, Aakash Ahmad, Mahdi Fahmideh, Mahmood Niazi, and Muhammad Azeem Akbar. Ethics of ai: A systematic literature review of principles and challenges. In Proceedings of the 26th international conference on evalua...

  74. [83]

    The automation of science

    Ross D King, Jem Rowland, Stephen G Oliver, Michael Young, Wayne Aubrey, Emma Byrne, Maria Liakata, Magdalena Markham, Pinar Pir, Larisa N Soldatova, et al. The automation of science. Science, 324 0 (5923): 0 85--89, 2009

  75. [84]

    Overcoming catastrophic forgetting in neural networks

    James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of scien...

  76. [85]

    Longeval: Guidelines for human evaluation of faithfulness in long-form summarization

    Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo. Longeval: Guidelines for human evaluation of faithfulness in long-form summarization. In Proceedings of the 17th Conference of the European Chapter of the Association for Comput...

  77. [86]

    The history of science

    Thomas Kuhn. The history of science. In Philosophy, Science, and History, pages 106--121. Routledge, 2014

  78. [87]

    Paperqa: Retrieval-augmented generative agent for scientific research

    Jakub L \'a la, Odhran O'Donoghue, Aleksandar Shtedritski, Sam Cox, Samuel G Rodriques, and Andrew D White. Paperqa: Retrieval-augmented generative agent for scientific research. arXiv preprint arXiv:2312.07559, 2023

  79. [88]

    Scientific literature: Information overload

    Esther Landhuis. Scientific literature: Information overload. Nature, 535 0 (7612): 0 457--458, 2016

  80. [89]

    Scientific discovery: Computational explorations of the creative processes

    P Langley. Scientific discovery: Computational explorations of the creative processes. MIT Press, 1987

  81. [90]

    Rizwan Parvez, Enamul Hoque, Shafiq R

    Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md. Rizwan Parvez, Enamul Hoque, Shafiq R. Joty, and Jimmy X. Huang. A systematic survey and critical review on evalua...

  82. [91]

    Deep learning

    Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521 0 (7553): 0 436--444, 2015

  83. [92]

    Biobert: a pre-trained biomedical language representation model for biomedical text mining

    Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36 0 (4): 0 1234--1240, 2020

  84. [93]

    P lag B ench: Exploring the duality of large language models in plagiarism generation and detection

    Jooyoung Lee, Toshini Agrawal, Adaku Uchendu, Thai Le, Jinghui Chen, and Dongwon Lee. P lag B ench: Exploring the duality of large language models in plagiarism generation and detection. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Proceedings of the 2025 Conference of...

  85. [94]

    Paperweaver: Enriching topical paper alerts by contextualizing recommended papers with user-collected papers

    Yoonjoo Lee, Hyeonsu B Kang, Matt Latzke, Juho Kim, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue. Paperweaver: Enriching topical paper alerts by contextualizing recommended papers with user-collected papers. In Proceedings of the 2024 CHI Conference on Human Factors ...

  86. [95]

    Matching papers and reviewers at large conferences

    Kevin Leyton-Brown, Yatin Nandwani, Hedayat Zarkoob, Chris Cameron, Neil Newman, Dinesh Raghu, et al. Matching papers and reviewers at large conferences. Artificial Intelligence, 331: 0 104119, 2024

  87. [96]

    Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models

    Lei Li, Yuqi Wang, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, and Qi Liu. Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models. In Annual Meeting of the Association for Computational Linguistics, 2024 a . URL https://api....

  88. [97]

    Chain of ideas: Revolutionizing research via novel idea development with llm agents

    Long Li, Weiwen Xu, Jiayan Guo, Ruochen Zhao, Xingxuan Li, Yuqian Yuan, Boqiang Zhang, Yuming Jiang, Yifei Xin, Ronghao Dang, et al. Chain of ideas: Revolutionizing research via novel idea development with llm agents. arXiv preprint arXiv:2410.13185, 2024 b

  89. [98]

    Peersum: a peer review dataset for abstractive multi-document summarization

    Miao Li, Jianzhong Qi, and Jey Han Lau. Peersum: a peer review dataset for abstractive multi-document summarization. arXiv preprint arXiv:2203.01769, 2022

  90. [99]

    Mlr-copilot: Autonomous machine learning research based on large language models agents, 2024 c

    Ruochen Li, Teerth Patel, Qingyun Wang, and Xinya Du. Mlr-copilot: Autonomous machine learning research based on large language models agents, 2024 c . URL https://arxiv.org/abs/2408.14033

  91. [100]

    Karlsson, Wei Shen, Manabu Okumura, and Chin-Yew Lin

    Yuhan Li, Jian Wu, Zhiwei Yu, B \"o rje F. Karlsson, Wei Shen, Manabu Okumura, and Chin-Yew Lin. All data on the table: Novel dataset and benchmark for cross-modality scientific information extraction. 2023. URL https://api.semanticscholar.org/CorpusID:265157731

  92. [101]

    Wilson, Woosang Lim, and William Yang Wang

    Zekun Li, Xianjun Yang, Kyuri Choi, Wanrong Zhu, Ryan Hsieh, HyeonJung Kim, Jin Hyuk Lim, Sungyoung Ji, Byungju Lee, Xifeng Yan, Linda Ruth Petzold, Stephen D. Wilson, Woosang Lim, and William Yang Wang. MMS ci: A dataset for graduate-level multi-discipline multimodal scientif...

  93. [102]

    Surveyx: Academic survey automation via large language models

    Xun Liang, Jiawei Yang, Yezhaohui Wang, Chen Tang, Zifan Zheng, Shichao Song, Zehao Lin, Yebin Yang, Simin Niu, Hanyu Wang, et al. Surveyx: Academic survey automation via large language models. arXiv preprint arXiv:2502.14776, 2025

  94. [103]

    Rouge: A package for automatic evaluation of summaries

    Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74--81, 2004

  95. [104]

    Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and X. Shi. Moprd: A multidisciplinary open peer review dataset. Neural Computing and Applications, 35: 0 24191--24206, 2022. URL https://api.semanticscholar.org/CorpusID:254535720

  96. [105]

    Moprd: A multidisciplinary open peer review dataset

    Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and Xiaodong Shi. Moprd: A multidisciplinary open peer review dataset. Neural Computing and Applications, 35 0 (34): 0 24191--24206, 2023

  97. [106]

    Autop2c: An llm-based agent framework for code repository generation from multimodal content in academic papers

    Zijie Lin, Yiqing Shen, Qilin Cai, He Sun, Jinrui Zhou, and Mingjun Xiao. Autop2c: An llm-based agent framework for code repository generation from multimodal content in academic papers. arXiv preprint arXiv:2504.20115, 2025

  98. [107]

    Ryan Liu and Nihar B. Shah. Reviewergpt? an exploratory study on using large language models for paper reviewing. ArXiv, abs/2306.00622, 2023. URL https://api.semanticscholar.org/CorpusID:258999338

  99. [108]

    Aigs: Generating science from ai-powered automated falsification

    Zijun Liu, Kaiming Liu, Yiqi Zhu, Xuanyu Lei, Zonghan Yang, Zhenhe Zhang, Peng Li, and Yang Liu. Aigs: Generating science from ai-powered automated falsification. arXiv preprint arXiv:2411.11910, 2024

  100. [109]

    Aaar-1.0: Assessing ai's potential to assist research, 2025

    Renze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Jian Xie, Yuxuan Sun, Yusen Zhang, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Wenchao Ma, Xi Li, Kai Zhang, Congying Xia, Lifu Huang, and Wenpeng Yin. Aaar-1.0: Assessing ai's potential to assist researc...

  101. [110]

    The ai scientist: Towards fully automated open-ended scientific discovery

    Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292v3, 2024. URL https://www.arxiv.org/abs/2408.06292v3

  102. [111]

    Llm4sr: A survey on large language models for scientific research

    Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du. Llm4sr: A survey on large language models for scientific research. arXiv preprint arXiv:2501.04306, 2025

  103. [112]

    Self-refine: Iterative refinement with self-feedback

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36: 0 46534--46594, 2023

  104. [113]

    Does my rebuttal matter? insights from a major nlp conference

    Yusuke Miyao. Does my rebuttal matter? insights from a major nlp conference. In Proceedings of NAACL-HLT, pages 1274--1290, 2019

  105. [114]

    Scigen: a dataset for reasoning-aware text generation from scientific tables

    Nafise Sadat Moosavi, Andreas R \"u ckl \'e , Dan Roth, and Iryna Gurevych. Scigen: a dataset for reasoning-aware text generation from scientific tables. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL http...

  106. [115]

    Literature-based discovery beyond the abc paradigm: a contrastive approach

    Erwan Moreau, Orla Hardiman, Mark Heverin, and Declan O’sullivan. Literature-based discovery beyond the abc paradigm: a contrastive approach. BioRxiv, pages 2021--09, 2021

  107. [116]

    Dora ai scientist: Multi-agent virtual research team for scientific exploration discovery and automated report generation

    Vladimir Naumov, Diana Zagirova, Sha Lin, Yupeng Xie, Wenhao Gou, Anatoly Urban, Nina Tikhonova, Khadija Alawi, Mike Durymanov, Fedor Galkin, et al. Dora ai scientist: Multi-agent virtual research team for scientific exploration discovery and automated report generation. bioRxiv, 2025

  108. [117]

    Philosophiae naturalis principia mathematica, volume 1

    Isaac Newton. Philosophiae naturalis principia mathematica, volume 1. G. Brookman, 1833

  109. [118]

    Newton's Principia: the mathematical principles of natural philosophy

    Isaac Newton and NW Chittenden. Newton's Principia: the mathematical principles of natural philosophy. Geo. P. Putnam, 1850

  110. [119]

    Alphaevolve: A coding agent for scientific and algorithmic discovery

    Alexander Novikov, Ng \^a n Vu, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al. Alphaevolve: A coding agent for scientific and algorithmic discovery. Technical report, Techn...

  111. [120]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 0 2...

  112. [121]

    Repograph: Enhancing ai software engineering with repository-level code graph

    Siru Ouyang, Wenhao Yu, Kaixin Ma, Zi-Qiang Xiao, Zhihan Zhang, Mengzhao Jia, Jiawei Han, Hongming Zhang, and Dong Yu. Repograph: Enhancing ai software engineering with repository-level code graph. ArXiv, abs/2410.14684, 2024. URL https://api.semanticscholar.org/CorpusID:273502041

  113. [122]

    Ml-dev-bench: Comparative analysis of ai agents on ml development workflows, 2025

    Harshith Padigela, Chintan Shah, and Dinkar Juyal. Ml-dev-bench: Comparative analysis of ai agents on ml development workflows, 2025. URL https://arxiv.org/abs/2502.00964

  114. [123]

    Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies

    Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies. Transactions of the Association for Computational Linguistics, 12: 0 4...

  115. [124]

    Enhancing repository-level code generation with integrated contextual information

    Zhiyuan Pan, Xing Hu, Xin Xia, and Xiaohu Yang. Enhancing repository-level code generation with integrated contextual information. ArXiv, abs/2406.03283, 2024 b . URL https://api.semanticscholar.org/CorpusID:270257850

  116. [125]

    Bleu: a method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, pages 311--318, 2002

  117. [126]

    Can llms help uncover insights about llms? a large-scale, evolving literature analysis of frontier llms

    Jungsoo Park, Junmo Kang, Gabriel Stanovsky, and Alan Ritter. Can llms help uncover insights about llms? a large-scale, evolving literature analysis of frontier llms. arXiv preprint arXiv:2502.18791, 2025

  118. [127]

    Automated research review support using machine learning, large language models, and natural language processing

    Vishnu S Pendyala, Karnavee Kamdar, and Kapil Mulchandani. Automated research review support using machine learning, large language models, and natural language processing. Electronics, 14 0 (2): 0 256, 2025

  119. [128]

    Review-llm: Harnessing large language models for personalized review generation

    Qiyao Peng, Hongtao Liu, Hongyan Xu, Qing Yang, Minglai Shao, and Wenjun Wang. Review-llm: Harnessing large language models for personalized review generation. ArXiv, abs/2407.07487, 2024. URL https://api.semanticscholar.org/CorpusID:271088888

  120. [129]

    The logic of scientific discovery

    Karl Popper. The logic of scientific discovery. Routledge, 2005

  121. [130]

    Citeme: Can language models accurately cite scientific claims? Advances in Neural Information Processing Systems, 37: 0 7847--7877, 2024

    Ori Press, Andreas Hochlehnert, Ameya Prabhu, Vishaal Udandarao, Ofir Press, and Matthias Bethge. Citeme: Can language models accurately cite scientific claims? Advances in Neural Information Processing Systems, 37: 0 7847--7877, 2024

  122. [131]

    Piflow: Principle-aware scientific discovery with multi-agent collaboration, 2025

    Yingming Pu, Tao Lin, and Hongyu Chen. Piflow: Principle-aware scientific discovery with multi-agent collaboration, 2025. URL https://arxiv.org/abs/2505.15047

  123. [132]

    Exploring jiu-jitsu argumentation for writing peer review rebuttals

    Sukannya Purkayastha, Anne Lauscher, and Iryna Gurevych. Exploring jiu-jitsu argumentation for writing peer review rebuttals. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14...

  124. [133]

    Large language models are zero shot hypothesis proposers

    Biqing Qi, Kaiyan Zhang, Haoxiang Li, Kai Tian, Sihang Zeng, Zhang-Ren Chen, and Bowen Zhou. Large language models are zero shot hypothesis proposers. arXiv preprint arXiv:2311.05965, 2023

  125. [134]

    Scaling large-language-model-based multi-agent collaboration

    Chen Qian, Zihao Xie, Yifei Wang, Wei Liu, Yufan Dang, Zhuoyun Du, Weize Chen, Cheng Yang, Zhiyuan Liu, and Maosong Sun. Scaling large-language-model-based multi-agent collaboration. arXiv preprint arXiv:2406.07155, 2024

  126. [135]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21 0 (140): 0 1--67, 2020. URL...

  127. [136]

    Towards scientific intelligence: A survey of llm-based scientific agents

    Shuo Ren, Pu Jian, Zhenjiang Ren, Chunlin Leng, Can Xie, and Jiajun Zhang. Towards scientific intelligence: A survey of llm-based scientific agents. arXiv preprint arXiv:2503.24047, 2025

  128. [137]

    Mathematical discoveries from program search with large language models

    Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical discoveries from program search with large language models. Nature, 625 0 (7...

  129. [138]

    Biodiscoveryagent: An ai agent for designing genetic perturbation experiments

    Yusuf Roohani, Andrew Lee, Qian Huang, Jian Vora, Zachary Steinhart, Kexin Huang, Alexander Marson, Percy Liang, and Jure Leskovec. Biodiscoveryagent: An ai agent for designing genetic perturbation experiments. arXiv preprint arXiv:2405.17631, 2024

  130. [139]

    Agentrxiv: Towards collaborative autonomous research

    Samuel Schmidgall and Michael Moor. Agentrxiv: Towards collaborative autonomous research. arXiv preprint arXiv:2503.18102, 2025

  131. [140]

    Agent laboratory: Using llm agents as research assistants

    Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using llm agents as research assistants. arXiv preprint arXiv:2501.04227, 2025

  132. [141]

    u rgen Schmidhuber. G \

    J \"u rgen Schmidhuber. G \"o del machines: Fully self-referential optimal universal self-improvers. In Artificial general intelligence, pages 199--226. Springer, 2007

  133. [142]

    Bigpatent: A large-scale dataset for abstractive and coherent summarization

    Eva Sharma, Chen Li, and Lu Wang. Bigpatent: A large-scale dataset for abstractive and coherent summarization. arXiv preprint arXiv:1906.03741, 2019

  134. [143]

    Shortcutsbench: A large-scale real-world benchmark for api-based agents, 2025

    Haiyang Shen, Yue Li, Desong Meng, Dongqi Cai, Sheng Qi, Li Zhang, Mengwei Xu, and Yun Ma. Shortcutsbench: A large-scale real-world benchmark for api-based agents, 2025. URL https://arxiv.org/abs/2407.00132

  135. [144]

    Detecting pretraining data from large language models

    Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. Detecting pretraining data from large language models. In The Twelfth International Conference on Learning Representations, 2024

  136. [145]

    B io M egatron: Larger biomedical domain language model

    Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani. B io M egatron: Larger biomedical domain language model. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empiric...

  137. [146]

    Reflexion: Language agents with verbal reinforcement learning

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36: 0 8634--8652, 2023

  138. [147]

    The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity

    Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar. The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity. 2025. URL https://api.semanticscholar.org/C...

  139. [148]

    Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers

    Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers. arXiv preprint arXiv:2409.04109, 2024

  140. [149]

    Core-bench: Fostering the credibility of published research through a computational reproducibility agent benchmark

    Zachary S Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, and Arvind Narayanan. Core-bench: Fostering the credibility of published research through a computational reproducibility agent benchmark. arXiv preprint arXiv:2409.11363, 2024

  141. [150]

    Legobench: Scientific leaderboard generation benchmark, 2024

    Shruti Singh, Shoaib Alam, Husain Malwat, and Mayank Singh. Legobench: Scientific leaderboard generation benchmark, 2024

  142. [151]

    Skarlinski, Sam Cox, Jon M

    Michael D. Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela Hinks, Michael J. Hammerling, Manvitha Ponnapati, Samuel G. Rodriques, and Andrew D. White. Language agents achieve superhuman synthesis of scientific knowledge. arXiv preprent arXiv:2409.13740, 2024. URL...

  143. [152]

    Paperbench: Evaluating ai's ability to replicate ai research

    Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, et al. Paperbench: Evaluating ai's ability to replicate ai research. arXiv preprint arXiv:2504.01848, 2025

  144. [153]

    Towards fair, equitable, and efficient peer review

    Ivan Stelmakh. Towards fair, equitable, and efficient peer review. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 15736--15737, 2021

  145. [154]

    Peerreview4all: Fair and accurate reviewer assignment in peer review

    Ivan Stelmakh, Nihar B Shah, and Aarti Singh. Peerreview4all: Fair and accurate reviewer assignment in peer review. In Algorithmic Learning Theory, pages 828--856. PMLR, 2019

  146. [155]

    Shah, Aarti Singh, and Hal Daum'e

    Ivan Stelmakh, Nihar B. Shah, Aarti Singh, and Hal Daum'e. Prior and prejudice. Proceedings of the ACM on Human-Computer Interaction, 5: 0 1 -- 17, 2020. URL https://api.semanticscholar.org/CorpusID:227227578

  147. [156]

    Reviewriter: Ai-generated instructions for peer review writing

    Xiaotian Su, Thiemo Wambsganss, Roman Rietsche, Seyed Parsa Neshaei, and Tanja K \"a ser. Reviewriter: Ai-generated instructions for peer review writing. In Workshop on Innovative Use of NLP for Building Educational Applications, 2025. URL https://api.semanticscholar.org/Corpu...

  148. [157]

    Towards table-to-text generation with numerical reasoning

    Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, and Hiroya Takamura. Towards table-to-text generation with numerical reasoning. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Asso...

  149. [158]

    Lu Sun, Stone Tao, Junjie Hu, and Steven P. Dow. Metawriter: Exploring the potential and perils of ai writing support in scientific peer review. Proceedings of the ACM on Human-Computer Interaction, 8: 0 1 -- 32, 2024. URL https://api.semanticscholar.org/CorpusID:269470548

  150. [159]

    Undiscovered public knowledge

    Don R Swanson. Undiscovered public knowledge. The Library Quarterly, 56 0 (2): 0 103--118, 1986

  151. [160]

    Agatha: automatic graph mining and transformer based hypothesis generation approach

    Justin Sybrandt, Ilya Tyagin, Michael Shtutman, and Ilya Safro. Agatha: automatic graph mining and transformer based hypothesis generation approach. In Proceedings of the 29th ACM international conference on information & knowledge management, pages 2757--2764, 2020

  152. [161]

    X-scitldr: cross-lingual extreme summarization of scholarly documents

    Sotaro Takeshita, Tommaso Green, Niklas Friedrich, Kai Eckert, and Simone Paolo Ponzetto. X-scitldr: cross-lingual extreme summarization of scholarly documents. In Proceedings of the 22nd ACM/IEEE Joint Conference on Digital Libraries, pages 1--12, 2022

  153. [162]

    Peer review as a multi-turn and long-context dialogue with role-based interactions

    Cheng Tan, Dongxin Lyu, Siyuan Li, Zhangyang Gao, Jingxuan Wei, Siqi Ma, Zicheng Liu, and Stan Z Li. Peer review as a multi-turn and long-context dialogue with role-based interactions. arXiv preprint arXiv:2406.05688, 2024

  154. [163]

    Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, Shengyan Liu, Di Luo, Yutao Ma, Hao Tong, Kha Trinh, Chenyu Tian, Zihan Wang, Bohao Wu, Yanyu Xiong, Shengzhu Yin, Min Zhu, Kilian Adriano Lieret,...

  155. [164]

    Unsupervised word embeddings capture latent knowledge from materials science literature

    Vahe Tshitoyan, John Dagdelen, Leigh Weston, Alexander Dunn, Ziqin Rong, Olga Kononova, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain. Unsupervised word embeddings capture latent knowledge from materials science literature. Nature, 571 0 (7763): 0 95--98, 2019

  156. [165]

    A new sociology of humans and machines

    Milena Tsvetkova, Taha Yasseri, Niccolo Pescetelli, and Tobias Werner. A new sociology of humans and machines. Nature Human Behaviour, 8 0 (10): 0 1864--1876, 2024

  157. [166]

    Openreviewer: Mitigating challenges in LLM reviewing, 2024

    Keith Tyser, Jason Lee, Avi Shporer, Madeleine Udell, Dov Te'eni, and Iddo Drori. Openreviewer: Mitigating challenges in LLM reviewing, 2024. URL https://openreview.net/forum?id=gqmELEp3lZ

  158. [167]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017

  159. [168]

    Teaching code llms to use autocompletion tools in repository-level code generation

    Chong Wang, Jian Zhang, Yebo Feng, Tianlin Li, Weisong Sun, Yang Liu, and Xin Peng. Teaching code llms to use autocompletion tools in repository-level code generation. ACM Transactions on Software Engineering and Methodology, 2024 a . URL https://api.semanticscholar.org/Corpus...

  160. [169]

    Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents

    Peisong Wang, Ruotian Ma, Bang Zhang, Xingyu Chen, Zhiwei He, Kang Luo, Qingsong Lv, Qingxuan Jiang, Zheng Xie, Shanyi Wang, et al. Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents. arXiv preprint arXiv:2507.03112, 2025

  161. [170]

    Paperrobot: Incremental draft generation of scientific ideas

    Qingyun Wang, Lifu Huang, Zhiying Jiang, Kevin Knight, Heng Ji, Mohit Bansal, and Yi Luan. Paperrobot: Incremental draft generation of scientific ideas. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1980--1991, 2019

  162. [171]

    Reviewrobot: Explainable paper review generation based on knowledge synthesis

    Qingyun Wang, Qi Zeng, Lifu Huang, Kevin Knight, Heng Ji, and Nazneen Fatema Rajani. Reviewrobot: Explainable paper review generation based on knowledge synthesis. arXiv preprint arXiv:2010.06119, 2020

  163. [172]

    S ci MON : Scientific inspiration machines optimized for novelty

    Qingyun Wang, Doug Downey, Heng Ji, and Tom Hope. S ci MON : Scientific inspiration machines optimized for novelty. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...

  164. [173]

    Shuai Wang, Harrisen Scells, Shengyao Zhuang, Martin Potthast, Bevan Koopman, and G. Zuccon. Zero-shot generative large language models for systematic review screening automation. ArXiv, abs/2401.06320, 2024 c . URL https://api.semanticscholar.org/CorpusID:266977226

  165. [174]

    Autosurvey: Large language models can automatically write surveys

    Yidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang, Xin Zhang, Zhen Wu, Meishan Zhang, Xinyu Dai, Qingsong Wen, Wei Ye, et al. Autosurvey: Large language models can automatically write surveys. Advances in Neural Information Processing Systems, 37: 0 115119--115145, 2024 d

  166. [175]

    Chain-of-thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35: 0 24824--24837, 2022

  167. [176]

    Browsecomp: A simple yet challenging benchmark for browsing agents

    Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. Browsecomp: A simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516, 2025

  168. [177]

    Large language models are better reasoners with self-verification

    Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao. Large language models are better reasoners with self-verification. arXiv preprint arXiv:2212.09561, 2022

  169. [178]

    Cycleresearcher: Improving automated research via automated review

    Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. Cycleresearcher: Improving automated research via automated review. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id...

  170. [179]

    WestlakeNLP. Airaxiv. https://airaxiv.com/, 2025

  171. [180]

    Pylabrobot: An open-source, hardware-agnostic interface for liquid-handling robots and accessories

    Rick P Wierenga, Stefan M Golas, Wilson Ho, Connor W Coley, and Kevin M Esvelt. Pylabrobot: An open-source, hardware-agnostic interface for liquid-handling robots and accessories. Device, 1 0 (4), 2023

  172. [181]

    Lag: Llm agents for leaderboard auto generation on demanding

    Jian Wu, Jiayu Zhang, Dongyuan Li, Linyi Yang, Aoxiao Zhong, Renhe Jiang, Qingsong Wen, and Yue Zhang. Lag: Llm agents for leaderboard auto generation on demanding. arXiv preprint arXiv:2502.18209, 2025

  173. [182]

    Automated review generation method based on large language models

    Shican Wu, Xiao Ma, Dehui Luo, Lulu Li, Xiangcheng Shi, Xin Chang, Xiaoyun Lin, Ran Luo, Chunlei Pei, Changying Du, et al. Automated review generation method based on large language models. arXiv preprint arXiv:2407.20906, 2024

  174. [183]

    Scireplicate-bench: Benchmarking llms in agent-driven algorithmic reproduction from research papers

    Yanzheng Xiang, Hanqi Yan, Shuyin Ouyang, Lin Gui, and Yulan He. Scireplicate-bench: Benchmarking llms in agent-driven algorithmic reproduction from research papers. arXiv preprint arXiv:2504.00255, 2025

  175. [184]

    An empirical analysis of uncertainty in large language model evaluations

    Qiujie Xie, Qingqiu Li, Zhuohao Yu, Yuejie Zhang, Yue Zhang, and Linyi Yang. An empirical analysis of uncertainty in large language model evaluations. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=J4xLuCt2kg

  176. [185]

    The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search

    Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search. arXiv preprint arXiv:2504.08066, 2025

  177. [186]

    TELIN : Table entity LIN ker for extracting leaderboards from machine learning publications

    Sean Yang, Chris Tensmeyer, and Curtis Wigington. TELIN : Table entity LIN ker for extracting leaderboards from machine learning publications. In Tirthankar Ghosal, Sergi Blanco-Cuaresma, Alberto Accomazzi, Robert M. Patton, Felix Grezes, and Thomas Allen, editors, Proceedings...

  178. [187]

    Ai becomes a masterbrain scientist

    Zijie Yang, Yukai Wang, and Lijing Zhang. Ai becomes a masterbrain scientist. bioRxiv, pages 2023--04, 2023

  179. [188]

    Airalogy: Ai-empowered universal data digitization for research automation, 2025 a

    Zijie Yang, Qiji Zhou, Fang Guo, Sijie Zhang, Yexun Xi, Jinglei Nie, Yudian Zhu, Liping Huang, Chou Wu, Yonghe Xia, Xiaoyu Ma, Yingming Pu, Panzhong Lu, Junshu Pan, Mingtao Chen, Tiannan Guo, Yanmei Dou, Hongyu Chen, Anping Zeng, Jiaxing Huang, Tian Xu, and Yue Zhang. Airalogy...

  180. [189]

    Large language models for automated open-domain scientific hypotheses discovery

    Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and Erik Cambria. Large language models for automated open-domain scientific hypotheses discovery. In Findings of the Association for Computational Linguistics ACL 2024, pages 13545--13565, 2024

  181. [190]

    MOOSE -chem: Large language models for rediscovering unseen chemistry scientific hypotheses

    Zonglin Yang, Wanhao Liu, Ben Gao, Tong Xie, Yuqiang Li, Wanli Ouyang, Soujanya Poria, Erik Cambria, and Dongzhan Zhou. MOOSE -chem: Large language models for rediscovering unseen chemistry scientific hypotheses. In The Thirteenth International Conference on Learning Represent...

  182. [191]

    Extracting knowledge from scientific texts on patient-derived cancer models using large language models: algorithm development and validation

    Jiarui Yao, Zinaida Perova, Tushar Mandloi, Elizabeth Lewis, Helen Parkinson, and Guergana Savova. Extracting knowledge from scientific texts on patient-derived cancer models using large language models: algorithm development and validation. bioRxiv, pages 2025--01, 2025

  183. [192]

    Are we there yet? revealing the risks of utilizing large language models in scholarly peer review

    Rui Ye, Xianghe Pang, Jingyi Chai, Jiaao Chen, Zhenfei Yin, Zhen Xiang, Xiaowen Dong, Jing Shao, and Siheng Chen. Are we there yet? revealing the risks of utilizing large language models in scholarly peer review. arXiv preprint arXiv:2412.01708, 2024

  184. [193]

    Mcx-llm: an experiment in bridging natural language problem descriptions with quantitative scientific simulations

    Fan-Yu Yen and Qianqian Fang. Mcx-llm: an experiment in bridging natural language problem descriptions with quantitative scientific simulations. Optica Biophotonics Congress: Biomedical Optics 2024 (Translational, Microscopy, OCT, OTS, BRAIN), 2024. URL https://api.semanticsch...

  185. [194]

    Knowledge integration and decision support for accelerated discovery of antibiotic resistance genes

    Jason Youn, Navneet Rai, and Ilias Tagkopoulos. Knowledge integration and decision support for accelerated discovery of antibiotic resistance genes. Nature Communications, 13 0 (1): 0 2360, 2022

  186. [195]

    Researchtown: Simulator of human research community

    Haofei Yu, Zhaochen Hong, Zirui Cheng, Kunlun Zhu, Keyang Xuan, Jinwei Yao, Tao Feng, and Jiaxuan You. Researchtown: Simulator of human research community. arXiv preprint arXiv:2412.17767, 2024

  187. [196]

    Is your paper being reviewed by an llm? a new benchmark dataset and approach for detecting ai text in peer review

    Sungduk Yu, Man Luo, Avinash Madusu, Vasudev Lal, and Phillip Howard. Is your paper being reviewed by an llm? a new benchmark dataset and approach for detecting ai text in peer review. arXiv preprint arXiv:2502.19614, 2025

  188. [197]

    Dolphin: Closed-loop open-ended auto-research through thinking, practice, and feedback

    Jiakang Yuan, Xiangchao Yan, Botian Shi, Tao Chen, Wanli Ouyang, Bo Zhang, Lei Bai, Yu Qiao, and Bowen Zhou. Dolphin: Closed-loop open-ended auto-research through thinking, practice, and feedback. arXiv preprint arXiv:2501.03916, 2025

  189. [198]

    Can we automate scientific reviewing? Journal of Artificial Intelligence Research, 75: 0 171--212, 2022

    Weizhe Yuan, Pengfei Liu, and Graham Neubig. Can we automate scientific reviewing? Journal of Artificial Intelligence Research, 75: 0 171--212, 2022

  190. [199]

    Appraising the potential uses and harms of llms for medical systematic reviews

    Hye Sun Yun, Iain J Marshall, Thomas A Trikalinos, and Byron C Wallace. Appraising the potential uses and harms of llms for medical systematic reviews. arXiv preprint arXiv:2305.11828, 2023

  191. [200]

    Star: Bootstrapping reasoning with reasoning

    Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35: 0 15476--15488, 2022

  192. [201]

    Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation

    Qi Zeng, Mankeerat Sidhu, Hou Pong Chan, Lu Wang, and Heng Ji. Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation. In AI4Research/DemocrAI@IJCAI, 2023. URL https://api.semanticscholar.org/CorpusID:258865977

  193. [202]

    Chen, Yue Zhang, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen

    Fengji Zhang, B. Chen, Yue Zhang, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. Repocoder: Repository-level code completion through iterative retrieval and generation. In Conference on Empirical Methods in Natural Language Processing, 2023. URL https://api.se...

  194. [203]

    Investigating fairness disparities in peer review: A language model enhanced approach

    Jiayao Zhang, Hongming Zhang, Zhun Deng, and Dan Roth. Investigating fairness disparities in peer review: A language model enhanced approach. ArXiv, abs/2211.06398, 2022. URL https://api.semanticscholar.org/CorpusID:253499393

  195. [204]

    Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges

    Kechi Zhang, Jia Li, Ge Li, Xianjie Shi, and Zhi Jin. Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges. In Annual Meeting of the Association for Computational Linguistics, 2024 a . URL https://api.semanticschol...

  196. [205]

    A comprehensive survey of scientific large language models and their applications in scientific discovery

    Yu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang, Shuiwang Ji, Wei Wang, and Jiawei Han. A comprehensive survey of scientific large language models and their applications in scientific discovery. arXiv preprint arXiv:2406.10833, 2024 b

  197. [206]

    Siren’s song in the ai ocean: A survey on hallucination in large language models

    Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. Siren’s song in the ai ocean: A survey on hallucination in large language models. Computational Linguistics, pages 1--45, 2025

  198. [207]

    Chain of agents: Large language models collaborating on long-context tasks

    Yusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister, Rui Zhang, and Sercan Arik. Chain of agents: Large language models collaborating on long-context tasks. Advances in Neural Information Processing Systems, 37: 0 132208--132237, 2024 c

  199. [208]

    Judging llm-as-a-judge with mt-bench and chatbot arena

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36: 0 46595--46623, 2023

  200. [209]

    From automation to autonomy: A survey on large language models in scientific discovery

    Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song. From automation to autonomy: A survey on large language models in scientific discovery. arXiv preprint arXiv:2505.13259, 2025

  201. [210]

    Large language models penetration in scholarly writing and peer review

    Li Zhou, Ruijie Zhang, Xunlian Dai, Daniel Hershcovich, and Haizhou Li. Large language models penetration in scholarly writing and peer review. ArXiv, abs/2502.11193, 2025 a . URL https://api.semanticscholar.org/CorpusID:276408109

  202. [211]

    Is llm a reliable reviewer? a comprehensive evaluation of llm on automatic paper reviewing tasks

    Ruiyang Zhou, Lu Chen, and Kai Yu. Is llm a reliable reviewer? a comprehensive evaluation of llm on automatic paper reviewing tasks. In International Conference on Language Resources and Evaluation, 2024 a . URL https://api.semanticscholar.org/CorpusID:269803977

  203. [212]

    Is LLM a reliable reviewer? a comprehensive evaluation of LLM on automatic paper reviewing tasks

    Ruiyang Zhou, Lu Chen, and Kai Yu. Is LLM a reliable reviewer? a comprehensive evaluation of LLM on automatic paper reviewing tasks. In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings of the 2024 Joint ...

  204. [213]

    From hypothesis to publication: A comprehensive survey of ai-driven research support systems

    Zekun Zhou, Xiaocheng Feng, Lei Huang, Xiachong Feng, Ziyun Song, Ruihan Chen, Liang Zhao, Weitao Ma, Yuxuan Gu, Baoxin Wang, et al. From hypothesis to publication: A comprehensive survey of ai-driven research support systems. arXiv preprint arXiv:2503.01424, 2025 b

  205. [214]

    Deepreview: Improving llm-based paper review with human-like deep thinking process

    Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang. Deepreview: Improving llm-based paper review with human-like deep thinking process. arXiv preprint arXiv:2503.08569, 2025

  206. [215]

    Large language models for automated scholarly paper review: A survey

    Zhenzhen Zhuang, Jiandong Chen, Hongfeng Xu, Yuwen Jiang, and Jialiang Lin. Large language models for automated scholarly paper review: A survey. arXiv preprint arXiv:2501.10326, 2025

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.