REVIEW 3 major objections 8 minor 2 cited by
How Far Are AI Scientists from Changing the World?
T0 review · 3 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Current AI Scientist systems cannot yet produce papers that pass scientific muster; the strongest system averages 4.63/10 under an automated reviewer, and experimental weakness appears in every assessed paper.
desk verdict A useful, well-organized status survey of AI Scientist systems whose central 'not good enough' claim is right, but whose own new measurement leans on an uncalibrated AI reviewer and should be treated as illustrative, not definitive. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a four-level capability framework that defines what a mature AI Scientist must do: acquire knowledge from literature, generate feasible novel hypotheses, verify and falsify them through experiments, and evolve from feedback. The load-bearing empirical instrument is DeepReviewer-14B, an AI reviewer model that rates papers on soundness, presentation, and contribution and also lists defect categories; Table 4 and Table 5 use it to score 28 papers from five systems. The framework matters because it converts the vague question "how far are we?" into a checkable milestone list, and the reviewer scores supply the quantitative answer.
What would settle it
Have a panel of human experts blind-review the same 28 papers, along with human-written accepted papers, and compare their scores and defect categories with DeepReviewer-14B's; if humans rate the AI papers near parity with accepted human work, or if DeepReviewer-14B gives equally low scores to strong human papers, the central claim is unsupported.
Extended reading notes
Core claim
The central claim is that no current AI Scientist system can autonomously carry out the full research loop well enough to produce work that would pass genuine scientific scrutiny. The evidence is twofold: on implementation benchmarks, even the strongest language models score low at reproducing or executing research code, and on the paper's own evaluation, all five surveyed systems average below 4.63 out of 10, with "Experimental Weakness" flagged in 100% of the 28 papers. The authors attribute the gap to limits of the foundation models—hallucination, costly knowledge updating, and catastrophic forgetting—and to underdeveloped research abilities in feasibility assessment, rigorous experimentation, and long-term planning. Their proposed remedy is an explicit "evolution" capability, where systems improve through self-reflection, external feedback, and structured collaboration, before they can be expected to produce ground-breaking discoveries.
Load-bearing premise
The load-bearing premise is that DeepReviewer-14B's quality scores are valid and unbiased measures of scientific merit, since the paper does not calibrate the model against human expert judgments or test whether its harshness affects all systems equally.
Editorial extensions
If this is right
- If the assessment is right, workshop acceptance of AI-generated papers is weak evidence of scientific maturity, since the same systems' full outputs rate far below typical accepted work.
- Verification and implementation, not idea generation, are the binding constraints; improving code execution and experimental design should come before claims of autonomous discovery.
- The four-level ladder implies that an AI system must demonstrate all levels—including evolution through feedback—before it is called a scientist, giving the field a concrete evaluation target.
- Deploying current systems without safeguards would flood peer review with low-quality artifacts, so the paper's proposed detection, labeling, and human oversight mechanisms become prerequisites.
- The review's category data, with 100% of papers showing experimental weakness and 96.4% showing methodological unclarity, provides a checklist that future systems can be measured against to track real progress.
Reading between the lines
- The paper's scores come from an AI reviewer, which is itself a language model; a natural next test is to check whether the same reviewer would give equally low percentiles to human-written papers, which would reveal whether the scale is harsh overall rather than specific to AI output.
- One testable extension is to run the evaluation longitudinally: if iterative review-feedback cycles raise later-generation papers above the current 4.63/10 ceiling, that would support the authors' claim that evolution, not raw model scale, is the missing ingredient.
- The framework suggests a practical benchmarking protocol—score any new AI Scientist on all four levels separately—so progress claims can be compared across systems instead of relying on anecdotal acceptance at workshops.
- Because the 28 papers are publicly available and therefore likely curated toward higher quality, the true typical output of these systems may be even weaker than the already low average scores reported.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey proposes a four-level capability framework for AI Scientist systems (knowledge acquisition, idea generation, verification and falsification, and evolution) and reviews representative methods and benchmarks for each level. It also presents a new empirical evaluation in Section 5.3, where 28 publicly available papers from five AI Scientist systems are scored by DeepReviewer-14B, yielding average ratings below 4.63/10 and a 100% incidence of 'Experimental Weakness' across the papers. On this basis, the paper concludes that current AI Scientist systems cannot independently produce scientific artifacts that meet established standards for high-quality scientific communication. The survey closes with limitations of foundation models, research-capability gaps, ethical considerations, and future directions.
Significance. The survey is comprehensive and timely, and the proposed capability framework provides a useful organizing structure for a rapidly evolving field. The paper compiles external benchmarks (MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, ML-Dev-Bench) that credibly demonstrate implementation and verification gaps in current systems. The new evaluation of full AI-generated manuscripts, if properly validated, would be an important contribution. However, the quantitative 'not good enough' claim currently rests on an uncalibrated AI reviewer with overlapping authorship, and the paper does not provide the statistical context needed to interpret the scores. With appropriate calibration and error analysis, the central conclusion could be made solid; as it stands, the evidence is not yet load-bearing in its reported form.
major comments (3)
- [§5.3, Tables 4–5] The central claim that current AI Scientist systems 'cannot independently produce scientific artifacts that meet established standards' rests entirely on ratings from DeepReviewer-14B (Zhu et al., 2025), yet the paper provides no calibration against human expert judgments, no inter-rater reliability statistics, no error bars, and no description of the reference distribution that defines the 'Percentile' column. Because DeepReviewer shares two authors with this manuscript, systematic bias cannot be ruled out. This is load-bearing, since Tables 4–5 are the only direct evaluation of complete AI-generated manuscripts in the paper.
- [§5.3, sample and statistics] The evaluation sample consists of only 28 papers, with 2–10 per system, and the paper itself notes in the Table 4 caption that publicly available papers 'may be curated and therefore may not fully represent the typical output of each system.' The paper does not report confidence intervals or significance tests, so the ranking across systems and the aggregate 'not good enough' conclusion are not robust to sampling variability. The authors should either enlarge the sample, report uncertainty, or soften the claims to match the evidentiary strength.
- [§4.2 vs. §5.3] The external benchmarks in Table 2 (MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, ML-Dev-Bench) directly support claims about weak implementation and verification capabilities, but they do not measure the holistic manuscript-quality dimensions (soundness, presentation, contribution) that Tables 4–5 assess. The paper should explicitly separate these two kinds of evidence and state that the manuscript-quality conclusion depends on the unvalidated DeepReviewer evaluation, not on the benchmarks.
minor comments (8)
- [§2.1] The model name 'CinicalBERT' should be 'ClinicalBERT' (Huang et al., 2019).
- [Figure 3] Figure 3 contains garbled text including long '/uni00000029/uni00000044/...' sequences and an unreadable table header ('Type Total citations Avg. citations'); this likely reflects a rendering or encoding error that must be fixed.
- [Throughout] The system name 'The AI Scientist' is typeset with non-standard spacing (e.g., 'A I Sc i e n t i s t'), making some sentences difficult to read; please use consistent formatting.
- [§5.3] In the sentence 'current AI Scientist systems not only struggle with scientific execution but also stuck with clearly articulating their research findings,' the phrase 'but also stuck' should be 'but also struggle.'
- [Table 4] The 'Percentile' column is undefined; the caption should state the reference distribution (e.g., percentile relative to what population of papers or scores).
- [§5.3 / Table 4] The paper does not describe how DeepReviewer-14B computes the overall 'Rating' from the three sub-scores (Soundness, Presentation, Contribution), nor whether the 0–10 scale permits non-integer intermediate values; please clarify the scoring protocol.
- [References] Several references have inconsistent formatting, including 'KABENAMUALU et al., 2023' in all caps and 'preprent' in Skarlinski et al. (2024); please proofread the bibliography.
- [§1] The phrase 'prospect-driven review' is used but never defined; consider adding one sentence explaining what makes the review 'prospect-driven.'
Circularity Check
Section 5.3's 'not good enough' result is carried by DeepReviewer-14B, a reviewer built by overlapping authors, used without independent calibration; the survey's broader gap claim retains independent support from external benchmarks.
-
self citation load bearing
[Section 5.3, Table 4 and Table 5]
"To quantify these gaps, we employ DeepReviewer-14B (Zhu et al., 2025), an advanced AI reviewer model, to assess 28 publicly available research papers produced by 5 leading AI Scientist systems."
The central claim of Section 5.3 — that current AI Scientist systems 'cannot independently produce scientific artifacts that meet established standards' — is operationalized entirely through DeepReviewer-14B's scores and defect categories. DeepReviewer-14B is cited only to Zhu et al. (2025), whose author list includes four present authors (Zhu, Weng, Yang, Zhang); no independent calibration against human reviewer ratings or inter-rater reliability is reported in this paper. The instrument also scores CycleResearcher-12B (Weng et al., 2025), another overlapping-author system in Table 4. The conclusion therefore reduces, for the full-manuscript-quality claim, to a self-cited evaluator rather than an externally validated standard.
full rationale
Most of the survey is a literature review and a capability taxonomy, which is not circular. Section 4.2 and Table 2 cite third-party benchmarks (MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, ML-Dev-Bench) that independently show implementation and verification gaps, so the broad 'considerable gap' conclusion does not reduce to the authors' own work. However, the specifically quantitative full-manuscript claim in Section 5.3 rests on DeepReviewer-14B, a model from Zhu et al. (2025) with four overlapping authors, and the paper provides no human calibration, no reference distribution for the Percentile column, and no external validation of the reviewer. Because that section is the only place where complete AI-generated manuscripts are directly evaluated, the strong 'cannot meet established standards' formulation is partially load-bearing on a self-citation. This is not a by-construction fit, so the score is moderate rather than high.
Assumptions & free parameters
assumptions (4)
- ad hoc to paper The four-level capability framework (knowledge acquisition, idea generation, verification and falsification, evolution) is a valid decomposition of the path to a mature AI Scientist.
- domain assumption DeepReviewer-14B scores are a valid measure of scientific paper quality.
- domain assumption Benchmark results cited from MLE-Bench, PaperBench, SciReplicate-Bench, CORE-Bench, and ML-Dev-Bench are accurate and representative of current LLM verification capabilities.
- domain assumption The publicly available papers from each AI Scientist system are not so unrepresentative as to invalidate cross-system comparisons.
Cite this review
Pith. "Pith review of How Far Are AI Scientists from Changing the World?." pith.science (2026). https://pith.science/paper/RY7EVEZT
@misc{pith2026250723276,
author = {Pith},
title = {Pith review of: How Far Are AI Scientists from Changing the World?},
year = {2026},
howpublished = {\url{https://pith.science/paper/RY7EVEZT}},
note = {Machine review of arXiv:2507.23276}
}
read the original abstract
The emergence of large language models (LLMs) is propelling automated scientific discovery to the next level, with LLM-based Artificial Intelligence (AI) Scientist systems now taking the lead in scientific research. Several influential works have already appeared in the field of AI Scientist systems, with AI-generated research papers having been accepted at the ICLR 2025 workshop, suggesting that a human-level AI Scientist capable of uncovering phenomena previously unknown to humans, may soon become a reality. In this survey, we focus on the central question: How far are AI scientists from changing the world and reshaping the scientific research paradigm? To answer this question, we provide a prospect-driven review that comprehensively analyzes the current achievements of AI Scientist systems, identifying key bottlenecks and the critical components required for the emergence of a scientific agent capable of producing ground-breaking discoveries that solve grand challenges. We hope this survey will contribute to a clearer understanding of limitations of current AI Scientist systems, showing where we are, what is missing, and what the ultimate goals for scientific AI should be.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 2 Pith papers
-
Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction
ZendoWorld shows that high labeling accuracy does not equal rule recovery, perception and induction are separate bottlenecks, and VLM agents propose near-uninformative experiments on active visual concept induction.
-
Bridging the Gap on AI-Assisted Scientific Software Development Through Transparency and Traceability
Proposes guidance for responsible AI use in scientific software development under NQA-1 standards, illustrated with TMAP8 V&V cases to ensure accountability and auditability.
Reference graph
Works this paper leans on
-
[1]
Litllm: A toolkit for scientific literature review
Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H Laradji, Krishnamurthy DJ Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. Litllm: A toolkit for scientific literature review. arXiv preprint arXiv:2402.01788, 2024 a
arXiv 2024
-
[2]
Llms for literature review: Are we there yet? arXiv preprint arXiv:2412.15249, 2024 b
Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H Laradji, Krishnamurthy DJ Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. Llms for literature review: Are we there yet? arXiv preprint arXiv:2412.15249, 2024 b
arXiv 2024
-
[3]
Litsearch: A retrieval benchmark for scientific literature search
Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, and Tianyu Gao. Litsearch: A retrieval benchmark for scientific literature search. arXiv preprint arXiv:2407.18940, 2024
arXiv 2024
-
[4]
Automated literature review using nlp techniques and llm-based retrieval-augmented generation
Nurshat Fateh Ali, Md Mahdi Mohtasim, Shakil Mosharrof, and T Gopi Krishna. Automated literature review using nlp techniques and llm-based retrieval-augmented generation. arXiv preprint arXiv:2411.18583, 2024
arXiv 2024
-
[5]
B io M -transformers: Building large biomedical language models with BERT , ALBERT and ELECTRA
Sultan Alrowili and Vijay Shanker. B io M -transformers: Building large biomedical language models with BERT , ALBERT and ELECTRA . In Dina Demner-Fushman, Kevin Bretonnel Cohen, Sophia Ananiadou, and Junichi Tsujii, editors, Proceedings of the 20th Workshop on Biomedical Language Processing, pages 221--227, Online, June 2021. Association for Computationa...
-
[6]
Introducing claude 3.5 sonnet
Anthropic. Introducing claude 3.5 sonnet. Anthropic Blog, 2024. URL https://www.anthropic.com/news/claude-3-5-sonnet
2024
-
[7]
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022
arXiv 2022
-
[8]
1,500 scientists lift the lid on reproducibility, 2016
Monya Baker. 1,500 scientists lift the lid on reproducibility, 2016
2016
Show all 214 references
-
[9]
Peerqa: A scientific question answering dataset from peer reviews
Tim Baumg \"a rtner, Ted Briscoe, and Iryna Gurevych. Peerqa: A scientific question answering dataset from peer reviews. arXiv preprint arXiv:2502.13668, 2025
2025 arXiv
-
[10]
Joeran Beel, Min-Yen Kan, and Moritz Baumgart. Evaluating sakana's ai scientist for autonomous research: Wishful thinking or an emerging reality towards' artificial research intelligence'(ari)? arXiv preprint arXiv:2502.14297, 2025
2025
-
[11]
Scibert: Pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. Scibert: Pretrained language model for scientific text. In EMNLP, 2019
2019
-
[12]
Superintelligent agents pose catastrophic risks: Can scientist ai offer a safer path? arXiv preprint arXiv:2502.15657, 2025
Yoshua Bengio, Michael Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, S \"o ren Mindermann, Adam Oberman, Jesse Richardson, Oliver Richardson, et al. Superintelligent agents pose catastrophic risks: Can scientist ai offer a safer path? arXiv preprin...
2025 arXiv
-
[13]
Has the machine learning review process become more arbitrary as the field has grown? the neurips 2021 consistency experiment
Alina Beygelzimer, Yann N Dauphin, Percy Liang, and Jennifer Wortman Vaughan. Has the machine learning review process become more arbitrary as the field has grown? the neurips 2021 consistency experiment. arXiv preprint arXiv:2306.03262, 2023
2021 arXiv
-
[14]
Politepeer: does peer review hurt? a dataset to gauge politeness intensity in the peer reviews
Prabhat Kumar Bharti, Meith Navlakha, Mayank Agarwal, and Asif Ekbal. Politepeer: does peer review hurt? a dataset to gauge politeness intensity in the peer reviews. Language Resources and Evaluation, 58 0 (4): 0 1291--1313, 2024
2024
-
[15]
Iterative refinement of project-level code context for precise code generation with compiler feedback
Zhangqian Bi, Yao Wan, Zheng Wang, Hongyu Zhang, Batu Guan, Fangxin Lu, Zili Zhang, Yulei Sui, Xuanhua Shi, and Hai Jin. Iterative refinement of project-level code context for precise code generation with compiler feedback. In Annual Meeting of the Association for Computationa...
2024
-
[16]
Reflective multi-agent collaboration based on large language models
Xiaohe Bo, Zeyu Zhang, Quanyu Dai, Xueyang Feng, Lei Wang, Rui Li, Xu Chen, and Ji-Rong Wen. Reflective multi-agent collaboration based on large language models. Advances in Neural Information Processing Systems, 37: 0 138595--138631, 2024
2024
-
[17]
Chemcrow: Augmenting large-language models with chemistry tools
Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Chemcrow: Augmenting large-language models with chemistry tools. arXiv preprint arXiv:2304.05376, 2023
2023 arXiv
-
[18]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020
1901
-
[19]
TLDR : Extreme summarization of scientific documents
Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel Weld. TLDR : Extreme summarization of scientific documents. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4766--4777, Online, November 2020 a . Asso...
2020 doi
-
[20]
Tldr: Extreme summarization of scientific documents
Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel S Weld. Tldr: Extreme summarization of scientific documents. arXiv preprint arXiv:2004.15011, 2020 b
2004 arXiv
-
[21]
Sang, Rahul K Arora, Robbie Kloosterman, Matthew Cecere, J
Christian Cao, J. Sang, Rahul K Arora, Robbie Kloosterman, Matthew Cecere, J. Gorla, Richard Saleh, D. Chen, Ian Drennan, Bijan Teja, Michael Fehlings, P. Ronksley, Alexander A Leung, Dany E Weisz, Harriet Ware, Mairead Whelan, D. B. Emerson, Rahul K Arora, and Niklas Bobrovit...
2024
-
[22]
Exploring scientific hypothesis generation with mamba
Miaosen Chai, Emily Herron, Erick Cervantes, and Tirthankar Ghosal. Exploring scientific hypothesis generation with mamba. In Proceedings of the 1st Workshop on NLP for Science (NLP4Science), pages 197--207, 2024
2024
-
[23]
Automated focused feedback generation for scientific writing assistance
Eric Chamoun, Michael Schlichktrull, and Andreas Vlachos. Automated focused feedback generation for scientific writing assistance. arXiv preprint arXiv:2405.20477, 2024
2024 arXiv
-
[24]
Mle-bench: Evaluating machine learning agents on machine learning engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al. Mle-bench: Evaluating machine learning agents on machine learning engineering. arXiv preprint arXiv:2410.07095, 2024
-
[25]
The toronto paper matching system: an automated paper-reviewer assignment system
Laurent Charlin and Richard Zemel. The toronto paper matching system: an automated paper-reviewer assignment system. 2013
2013
-
[26]
Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery
Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, et al. Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery. arXiv preprint arXiv:2410.05080, 2024
-
[27]
React: A re view comment dataset for act ionability (and more)
Gautam Choudhary, Natwar Modani, and Nitish Maurya. React: A re view comment dataset for act ionability (and more). In Web Information Systems Engineering--WISE 2021: 22nd International Conference on Web Information Systems Engineering, WISE 2021, Melbourne, VIC, Australia, Oc...
2021
-
[28]
Relevai-reviewer: A benchmark on ai reviewers for survey paper relevance
Paulo Henrique Couto, Quang Phuoc Ho, Nageeta Kumari, Benedictus Kent Rachmat, Thanh Gia Hieu Khuong, Ihsan Ullah, and Lisheng Sun-Hosoya. Relevai-reviewer: A benchmark on ai reviewers for survey paper relevance. arXiv preprint arXiv:2406.10294, 2024
2024
-
[29]
Structured information extraction from scientific text with large language models
John Dagdelen, Alexander Dunn, Sanghoon Lee, Nicholas Walker, Andrew S Rosen, Gerbrand Ceder, Kristin A Persson, and Anubhav Jain. Structured information extraction from scientific text with large language models. Nature Communications, 15 0 (1): 0 1418, 2024
2024
-
[30]
Marg: Multi-agent review generation for scientific papers
Mike D'Arcy, Tom Hope, Larry Birnbaum, and Doug Downey. Marg: Multi-agent review generation for scientific papers. arXiv preprint arXiv:2401.04259, 2024
2024 arXiv
-
[31]
ARIES : A corpus of scientific paper edits made in response to peer reviews
Mike D ' Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, and Doug Downey. ARIES : A corpus of scientific paper edits made in response to peer reviews. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of ...
2024 doi
-
[32]
The general theory of relativity
Albert Einstein. The general theory of relativity. In The meaning of relativity, pages 54--75. Springer, 1922
1922
-
[33]
Introduction to artificial intelligence
Wolfgang Ertel. Introduction to artificial intelligence. Springer Nature, 2024
2024
-
[34]
Fox, Jennifer Meyer, and Emilie Aim \'e
Charles W. Fox, Jennifer Meyer, and Emilie Aim \'e . Double-blind peer review affects reviewer ratings and editor decisions at an ecology journal. 37 0 (5): 0 1144--1157, May 2023. doi:10.1111/1365-2435.14259. Publisher Copyright: 2023 The Authors. Functional Ecology 2023 Brit...
2023
-
[35]
Semantic scholar
Suzanne Fricke. Semantic scholar. Journal of the Medical Library Association: JMLA, 106 0 (1): 0 145, 2018
2018
-
[36]
Llm-ref: Enhancing reference handling in technical writing with large language models
Kazi Ahmed Asif Fuad and Lizhong Chen. Llm-ref: Enhancing reference handling in technical writing with large language models. ArXiv, abs/2411.00294, 2024. URL https://api.semanticscholar.org/CorpusID:273798525
2024 arXiv
-
[37]
Reviewagents: Bridging the gap between human and ai-generated paper reviews
Xian Gao, Jiacheng Ruan, Jingsheng Gao, Ting Liu, and Yuzhuo Fu. Reviewagents: Bridging the gap between human and ai-generated paper reviews. arXiv preprint arXiv:2503.08506, 2025
2025 arXiv
-
[38]
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997, 2 0 (1), 2023
2023 arXiv
-
[39]
Reviewer2: Optimizing review generation through prompt generation
Zhaolin Gao, Kiant \'e Brantley, and Thorsten Joachims. Reviewer2: Optimizing review generation through prompt generation. arXiv preprint arXiv:2402.10886, 2024
2024 arXiv
-
[41]
Usefulness of llms as an author checklist assistant for scientific papers: Neurips'24 experiment
Alexander Goldberg, Ihsan Ullah, Thanh Gia Hieu Khuong, Benedictus Kent Rachmat, Zhen Xu, Isabelle Guyon, and Nihar B Shah. Usefulness of llms as an author checklist assistant for scientific papers: Neurips'24 experiment. arXiv preprint arXiv:2411.03417, 2024 b
2024 arXiv
-
[42]
Towards an ai co-scientist
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2025
2025 arXiv
-
[43]
A survey on the rise of the ai scientists: Accelerating discovery and confronting ethical frontiers
Muskaan Goyal. A survey on the rise of the ai scientists: Accelerating discovery and confronting ethical frontiers. World Journal of Advanced Engineering Technology and Sciences, 15: 0 564--569, 05 2025. doi:10.30574/wjaets.2025.15.2.0646
2025 doi
-
[44]
Agentic ai for scientific discovery: A survey of progress, challenges, and future directions
Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack. Agentic ai for scientific discovery: A survey of progress, challenges, and future directions. arXiv preprint arXiv:2503.08979, 2025
2025 arXiv
-
[45]
Domain-specific language model pretraining for biomedical natural language processing, 2020
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. Domain-specific language model pretraining for biomedical natural language processing, 2020
2020
-
[46]
Large language model based multi-agents: A survey of progress and challenges
Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. Large language model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680, 2024
2024 arXiv
-
[47]
Automatic analysis of substantiation in scientific peer reviews
Yanzhu Guo, Guokan Shang, Virgile Rennard, Michalis Vazirgiannis, and Chlo \'e Clavel. Automatic analysis of substantiation in scientific peer reviews. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023,...
2023 doi
-
[48]
Data extraction from polymer literature using large language models
Sonakshi Gupta, Akhlak Mahmood, Pranav Shetty, Aishat Adeboye, and Rampi Ramprasad. Data extraction from polymer literature using large language models. Communications Materials, 5 0 (1): 0 269, 2024
2024
-
[49]
Tanishq Gupta, Mohd Zaki, N. M. Anoop Krishnan, and Mausam. MatSciBERT : A materials domain language model for text mining and information extraction. npj Computational Materials, 8 0 (1): 0 102, May 2022. ISSN 2057-3960. doi:10.1038/s41524-022-00784-w. URL https://www.nature....
2022 doi
-
[50]
Pasa: An llm agent for comprehensive academic paper search, 2025
Yichen He, Guanhua Huang, Peiyuan Feng, Yuan Lin, Yuchen Zhang, Hang Li, and Weinan E. Pasa: An llm agent for comprehensive academic paper search, 2025
2025
-
[51]
The diminishing returns of masked language models to science, 2023
Zhi Hong, Aswathy Ajith, Gregory Pauloski, Eamon Duede, Kyle Chard, and Ian Foster. The diminishing returns of masked language models to science, 2023
2023
-
[52]
Automatic evaluation metrics for artificially generated scientific research
Niklas H \"o pner, Leon Eshuijs, Dimitrios Alivanistos, Giacomo Zamprogno, and Ilaria Tiddi. Automatic evaluation metrics for artificially generated scientific research. arXiv preprint arXiv:2503.05712, 2025
2025 arXiv
-
[53]
Identification of tasks, datasets, evaluation metrics, and numeric scores for scientific leaderboards construction
Yufang Hou, Charles Jochim, Martin Gleize, Francesca Bonin, and Debasis Ganguly. Identification of tasks, datasets, evaluation metrics, and numeric scores for scientific leaderboards construction. In Anna Korhonen, David Traum, and Llu \'i s M \`a rquez, editors, Proceedings o...
2019 doi
-
[54]
CHIME : LLM -assisted hierarchical organization of scientific studies for literature review support
Chao-Chun Hsu, Erin Bransom, Jenna Sparks, Bailey Kuehl, Chenhao Tan, David Wadden, Lucy Wang, and Aakanksha Naik. CHIME : LLM -assisted hierarchical organization of scientific studies for literature review support. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Fi...
2024 doi
-
[55]
Nova: An iterative planning and search approach to enhance novelty and diversity of llm generated ideas
Xiang Hu, Hongyu Fu, Jinge Wang, Yifeng Wang, Zhikun Li, Renjun Xu, Yu Lu, Yaochu Jin, Lili Pan, and Zhenzhong Lan. Nova: An iterative planning and search approach to enhance novelty and diversity of llm generated ideas. arXiv preprint arXiv:2410.14255, 2024
-
[56]
An overview of artificial intelligence ethics
Changwu Huang, Zeqi Zhang, Bifei Mao, and Xin Yao. An overview of artificial intelligence ethics. IEEE Transactions on Artificial Intelligence, 4 0 (4): 0 799--819, 2022
2022
-
[57]
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiv:1904.05342, 2019
1904 arXiv
-
[58]
Biomni: A general-purpose biomedical ai agent
Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Yusuf Roohani, Ryan Li, Lin Qiu, Junze Zhang, Yin Di, et al. Biomni: A general-purpose biomedical ai agent. bioRxiv, pages 2025--05, 2025 a
2025
-
[59]
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Informatio...
2025
-
[60]
Openreviewer: A specialized large language model for generating critical scientific paper reviews
Maximilian Idahl and Zahra Ahmadi. Openreviewer: A specialized large language model for generating critical scientific paper reviews. arXiv preprint arXiv:2412.11948, 2024
2024 arXiv
-
[61]
Zochi technical report
Intology. Zochi technical report. arXiv, 2025
2025
-
[62]
Why most published research findings are false
John PA Ioannidis. Why most published research findings are false. PLoS medicine, 2 0 (8): 0 e124, 2005
2005
-
[63]
Discoveryworld: A virtual environment for developing and evaluating automated scientific discovery agents
Peter Jansen, Marc-Alexandre C \^o t \'e , Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark. Discoveryworld: A virtual environment for developing and evaluating automated scientific discovery agents. Advances in Neur...
2024
-
[64]
Codescientist: End-to-end semi-automated scientific discovery with code-based experimentation
Peter Jansen, Oyvind Tafjord, Marissa Radensky, Pao Siangliulue, Tom Hope, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Daniel S Weld, and Peter Clark. Codescientist: End-to-end semi-automated scientific discovery with code-based experimentation. arXiv preprint arXiv:250...
2025 arXiv
-
[65]
Lgar: Zero-shot llm-guided neural ranking for abstract screening in systematic literature reviews
Christian Jaumann, Andreas Wiedholz, and Annemarie Friedrich. Lgar: Zero-shot llm-guided neural ranking for abstract screening in systematic literature reviews. 2025. URL https://api.semanticscholar.org/CorpusID:279071051
2025
-
[66]
Ai-researcher: Autonomous scientific innovation, 2025
Tang Jiabin, Xia Lianghao, Li Zhonghang, and Huang Chao. Ai-researcher: Autonomous scientific innovation, 2025. URL https://arxiv.org/abs/2505.18705
2025 arXiv
-
[67]
Aide: Ai-driven exploration in the space of code
Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, and Yuxiang Wu. Aide: Ai-driven exploration in the space of code. ArXiv, abs/2502.13138, 2025. URL https://api.semanticscholar.org/CorpusID:276421281
2025 arXiv
-
[68]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023. URL https://www.arxiv.org/abs/2310.06770
2023 arXiv
-
[69]
Probing biomedical embeddings from language models
Qiao Jin, Bhuwan Dhingra, William Cohen, and Xinghua Lu. Probing biomedical embeddings from language models. In Proceedings of the 3rd Workshop on Evaluating Vector Space Representations for NLP, pages 82--89, 2019
2019
-
[70]
A gent R eview: Exploring peer review dynamics with LLM agents
Yiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen, Kaijie Zhu, Yijia Xiao, and Jindong Wang. A gent R eview: Exploring peer review dynamics with LLM agents. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in N...
2024 doi
-
[71]
Dsbench: How far are data science agents from becoming data science experts?, 2025
Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu. Dsbench: How far are data science agents from becoming data science experts?, 2025. URL https://arxiv.org/abs/2409.07703
2025 arXiv
-
[72]
The global landscape of ai ethics guidelines
Anna Jobin, Marcello Ienca, and Effy Vayena. The global landscape of ai ethics guidelines. Nature machine intelligence, 1 0 (9): 0 389--399, 2019
2019
-
[73]
Cutting through the clutter: The potential of llms for efficient filtration in systematic literature reviews
Lucas Joos, Daniel A Keim, and Maximilian T Fischer. Cutting through the clutter: The potential of llms for efficient filtration in systematic literature reviews. arXiv preprint arXiv:2407.10652, 2024
2024 arXiv
-
[74]
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Z \' dek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. nature, 596 0 (7873): 0 583--589, 2021
2021
-
[75]
Salomon Kabongo KABENAMUALU, Jennifer D’Souza, and S. Auer. Orkg-leaderboards: a systematic workflow for mining leaderboards as a knowledge graph. International Journal on Digital Libraries, pages 1--14, 2023. URL https://api.semanticscholar.org/CorpusID:258762176
2023
-
[76]
Figureqa: An annotated figure dataset for visual reasoning
Samira Ebrahimi Kahou, Adam Atkinson, Vincent Michalski, \'A kos K \'a d \'a r, Adam Trischler, and Yoshua Bengio. Figureqa: An annotated figure dataset for visual reasoning. ArXiv, abs/1710.07300, 2017. URL https://api.semanticscholar.org/CorpusID:3535069
2017 arXiv
-
[77]
A dataset of peer reviews (peerread): Collection, insights and nlp applications
Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine Van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz. A dataset of peer reviews (peerread): Collection, insights and nlp applications. arXiv preprint arXiv:1804.09635, 2018
2018 arXiv
-
[78]
From who you know to what you read: Augmenting scientific recommendations with implicit social networks
Hyeonsu B Kang, Rafal Kocielnik, Andrew Head, Jiangjiang Yang, Matt Latzke, Aniket Kittur, Daniel S Weld, Doug Downey, and Jonathan Bragg. From who you know to what you read: Augmenting scientific recommendations with implicit social networks. In Proceedings of the 2022 CHI Co...
2022
-
[79]
Comlittee: Literature discovery with personal elected author committees
Hyeonsu B Kang, Nouran Soliman, Matt Latzke, Joseph Chee Chang, and Jonathan Bragg. Comlittee: Literature discovery with personal elected author committees. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--20, 2023
2023
-
[80]
AxCell : Automatic extraction of results from machine learning papers
Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, and Robert Stojnic. AxCell : Automatic extraction of results from machine learning papers. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Con...
2020 doi
-
[81]
The perils of using mechanical turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. The perils of using mechanical turk to evaluate open-ended text generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1265--1285, 2021
2021
-
[82]
Ethics of ai: A systematic literature review of principles and challenges
Arif Ali Khan, Sher Badshah, Peng Liang, Muhammad Waseem, Bilal Khan, Aakash Ahmad, Mahdi Fahmideh, Mahmood Niazi, and Muhammad Azeem Akbar. Ethics of ai: A systematic literature review of principles and challenges. In Proceedings of the 26th international conference on evalua...
2022
-
[83]
The automation of science
Ross D King, Jem Rowland, Stephen G Oliver, Michael Young, Wayne Aubrey, Emma Byrne, Maria Liakata, Magdalena Markham, Pinar Pir, Larisa N Soldatova, et al. The automation of science. Science, 324 0 (5923): 0 85--89, 2009
2009
-
[84]
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of scien...
2017
-
[85]
Longeval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo. Longeval: Guidelines for human evaluation of faithfulness in long-form summarization. In Proceedings of the 17th Conference of the European Chapter of the Association for Comput...
2023
-
[86]
The history of science
Thomas Kuhn. The history of science. In Philosophy, Science, and History, pages 106--121. Routledge, 2014
2014
-
[87]
Paperqa: Retrieval-augmented generative agent for scientific research
Jakub L \'a la, Odhran O'Donoghue, Aleksandar Shtedritski, Sam Cox, Samuel G Rodriques, and Andrew D White. Paperqa: Retrieval-augmented generative agent for scientific research. arXiv preprint arXiv:2312.07559, 2023
2023 arXiv
-
[88]
Scientific literature: Information overload
Esther Landhuis. Scientific literature: Information overload. Nature, 535 0 (7612): 0 457--458, 2016
2016
-
[89]
Scientific discovery: Computational explorations of the creative processes
P Langley. Scientific discovery: Computational explorations of the creative processes. MIT Press, 1987
1987
-
[90]
Rizwan Parvez, Enamul Hoque, Shafiq R
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md. Rizwan Parvez, Enamul Hoque, Shafiq R. Joty, and Jimmy X. Huang. A systematic survey and critical review on evalua...
2024
-
[91]
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521 0 (7553): 0 436--444, 2015
2015
-
[92]
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36 0 (4): 0 1234--1240, 2020
2020
-
[93]
P lag B ench: Exploring the duality of large language models in plagiarism generation and detection
Jooyoung Lee, Toshini Agrawal, Adaku Uchendu, Thai Le, Jinghui Chen, and Dongwon Lee. P lag B ench: Exploring the duality of large language models in plagiarism generation and detection. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Proceedings of the 2025 Conference of...
2025
-
[94]
Paperweaver: Enriching topical paper alerts by contextualizing recommended papers with user-collected papers
Yoonjoo Lee, Hyeonsu B Kang, Matt Latzke, Juho Kim, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue. Paperweaver: Enriching topical paper alerts by contextualizing recommended papers with user-collected papers. In Proceedings of the 2024 CHI Conference on Human Factors ...
2024
-
[95]
Matching papers and reviewers at large conferences
Kevin Leyton-Brown, Yatin Nandwani, Hedayat Zarkoob, Chris Cameron, Neil Newman, Dinesh Raghu, et al. Matching papers and reviewers at large conferences. Artificial Intelligence, 331: 0 104119, 2024
2024
-
[96]
Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models
Lei Li, Yuqi Wang, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, and Qi Liu. Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models. In Annual Meeting of the Association for Computational Linguistics, 2024 a . URL https://api....
2024
-
[97]
Chain of ideas: Revolutionizing research via novel idea development with llm agents
Long Li, Weiwen Xu, Jiayan Guo, Ruochen Zhao, Xingxuan Li, Yuqian Yuan, Boqiang Zhang, Yuming Jiang, Yifei Xin, Ronghao Dang, et al. Chain of ideas: Revolutionizing research via novel idea development with llm agents. arXiv preprint arXiv:2410.13185, 2024 b
-
[98]
Peersum: a peer review dataset for abstractive multi-document summarization
Miao Li, Jianzhong Qi, and Jey Han Lau. Peersum: a peer review dataset for abstractive multi-document summarization. arXiv preprint arXiv:2203.01769, 2022
2022 arXiv
-
[99]
Mlr-copilot: Autonomous machine learning research based on large language models agents, 2024 c
Ruochen Li, Teerth Patel, Qingyun Wang, and Xinya Du. Mlr-copilot: Autonomous machine learning research based on large language models agents, 2024 c . URL https://arxiv.org/abs/2408.14033
2024
-
[100]
Karlsson, Wei Shen, Manabu Okumura, and Chin-Yew Lin
Yuhan Li, Jian Wu, Zhiwei Yu, B \"o rje F. Karlsson, Wei Shen, Manabu Okumura, and Chin-Yew Lin. All data on the table: Novel dataset and benchmark for cross-modality scientific information extraction. 2023. URL https://api.semanticscholar.org/CorpusID:265157731
2023
-
[101]
Wilson, Woosang Lim, and William Yang Wang
Zekun Li, Xianjun Yang, Kyuri Choi, Wanrong Zhu, Ryan Hsieh, HyeonJung Kim, Jin Hyuk Lim, Sungyoung Ji, Byungju Lee, Xifeng Yan, Linda Ruth Petzold, Stephen D. Wilson, Woosang Lim, and William Yang Wang. MMS ci: A dataset for graduate-level multi-discipline multimodal scientif...
2025
-
[102]
Surveyx: Academic survey automation via large language models
Xun Liang, Jiawei Yang, Yezhaohui Wang, Chen Tang, Zifan Zheng, Shichao Song, Zehao Lin, Yebin Yang, Simin Niu, Hanyu Wang, et al. Surveyx: Academic survey automation via large language models. arXiv preprint arXiv:2502.14776, 2025
2025 arXiv
-
[103]
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74--81, 2004
2004
-
[104]
Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and X. Shi. Moprd: A multidisciplinary open peer review dataset. Neural Computing and Applications, 35: 0 24191--24206, 2022. URL https://api.semanticscholar.org/CorpusID:254535720
2022
-
[105]
Moprd: A multidisciplinary open peer review dataset
Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and Xiaodong Shi. Moprd: A multidisciplinary open peer review dataset. Neural Computing and Applications, 35 0 (34): 0 24191--24206, 2023
2023
-
[106]
Autop2c: An llm-based agent framework for code repository generation from multimodal content in academic papers
Zijie Lin, Yiqing Shen, Qilin Cai, He Sun, Jinrui Zhou, and Mingjun Xiao. Autop2c: An llm-based agent framework for code repository generation from multimodal content in academic papers. arXiv preprint arXiv:2504.20115, 2025
2025 arXiv
-
[107]
Ryan Liu and Nihar B. Shah. Reviewergpt? an exploratory study on using large language models for paper reviewing. ArXiv, abs/2306.00622, 2023. URL https://api.semanticscholar.org/CorpusID:258999338
2023 arXiv
-
[108]
Aigs: Generating science from ai-powered automated falsification
Zijun Liu, Kaiming Liu, Yiqi Zhu, Xuanyu Lei, Zonghan Yang, Zhenhe Zhang, Peng Li, and Yang Liu. Aigs: Generating science from ai-powered automated falsification. arXiv preprint arXiv:2411.11910, 2024
2024 arXiv
-
[109]
Aaar-1.0: Assessing ai's potential to assist research, 2025
Renze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Jian Xie, Yuxuan Sun, Yusen Zhang, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Wenchao Ma, Xi Li, Kai Zhang, Congying Xia, Lifu Huang, and Wenpeng Yin. Aaar-1.0: Assessing ai's potential to assist researc...
2025 arXiv
-
[110]
The ai scientist: Towards fully automated open-ended scientific discovery
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292v3, 2024. URL https://www.arxiv.org/abs/2408.06292v3
2024 arXiv
-
[111]
Llm4sr: A survey on large language models for scientific research
Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du. Llm4sr: A survey on large language models for scientific research. arXiv preprint arXiv:2501.04306, 2025
2025 arXiv
-
[112]
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36: 0 46534--46594, 2023
2023
-
[113]
Does my rebuttal matter? insights from a major nlp conference
Yusuke Miyao. Does my rebuttal matter? insights from a major nlp conference. In Proceedings of NAACL-HLT, pages 1274--1290, 2019
2019
-
[114]
Scigen: a dataset for reasoning-aware text generation from scientific tables
Nafise Sadat Moosavi, Andreas R \"u ckl \'e , Dan Roth, and Iryna Gurevych. Scigen: a dataset for reasoning-aware text generation from scientific tables. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL http...
2021
-
[115]
Literature-based discovery beyond the abc paradigm: a contrastive approach
Erwan Moreau, Orla Hardiman, Mark Heverin, and Declan O’sullivan. Literature-based discovery beyond the abc paradigm: a contrastive approach. BioRxiv, pages 2021--09, 2021
2021
-
[116]
Dora ai scientist: Multi-agent virtual research team for scientific exploration discovery and automated report generation
Vladimir Naumov, Diana Zagirova, Sha Lin, Yupeng Xie, Wenhao Gou, Anatoly Urban, Nina Tikhonova, Khadija Alawi, Mike Durymanov, Fedor Galkin, et al. Dora ai scientist: Multi-agent virtual research team for scientific exploration discovery and automated report generation. bioRxiv, 2025
2025
-
[117]
Philosophiae naturalis principia mathematica, volume 1
Isaac Newton. Philosophiae naturalis principia mathematica, volume 1. G. Brookman, 1833
-
[118]
Newton's Principia: the mathematical principles of natural philosophy
Isaac Newton and NW Chittenden. Newton's Principia: the mathematical principles of natural philosophy. Geo. P. Putnam, 1850
-
[119]
Alphaevolve: A coding agent for scientific and algorithmic discovery
Alexander Novikov, Ng \^a n Vu, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al. Alphaevolve: A coding agent for scientific and algorithmic discovery. Technical report, Techn...
2025
-
[120]
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 0 2...
2022
-
[121]
Repograph: Enhancing ai software engineering with repository-level code graph
Siru Ouyang, Wenhao Yu, Kaixin Ma, Zi-Qiang Xiao, Zhihan Zhang, Mengzhao Jia, Jiawei Han, Hongming Zhang, and Dong Yu. Repograph: Enhancing ai software engineering with repository-level code graph. ArXiv, abs/2410.14684, 2024. URL https://api.semanticscholar.org/CorpusID:273502041
-
[122]
Ml-dev-bench: Comparative analysis of ai agents on ml development workflows, 2025
Harshith Padigela, Chintan Shah, and Dinkar Juyal. Ml-dev-bench: Comparative analysis of ai agents on ml development workflows, 2025. URL https://arxiv.org/abs/2502.00964
2025 arXiv
-
[123]
Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies. Transactions of the Association for Computational Linguistics, 12: 0 4...
2024
-
[124]
Enhancing repository-level code generation with integrated contextual information
Zhiyuan Pan, Xing Hu, Xin Xia, and Xiaohu Yang. Enhancing repository-level code generation with integrated contextual information. ArXiv, abs/2406.03283, 2024 b . URL https://api.semanticscholar.org/CorpusID:270257850
2024
-
[125]
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, pages 311--318, 2002
2002
-
[126]
Can llms help uncover insights about llms? a large-scale, evolving literature analysis of frontier llms
Jungsoo Park, Junmo Kang, Gabriel Stanovsky, and Alan Ritter. Can llms help uncover insights about llms? a large-scale, evolving literature analysis of frontier llms. arXiv preprint arXiv:2502.18791, 2025
2025 arXiv
-
[127]
Automated research review support using machine learning, large language models, and natural language processing
Vishnu S Pendyala, Karnavee Kamdar, and Kapil Mulchandani. Automated research review support using machine learning, large language models, and natural language processing. Electronics, 14 0 (2): 0 256, 2025
2025
-
[128]
Review-llm: Harnessing large language models for personalized review generation
Qiyao Peng, Hongtao Liu, Hongyan Xu, Qing Yang, Minglai Shao, and Wenjun Wang. Review-llm: Harnessing large language models for personalized review generation. ArXiv, abs/2407.07487, 2024. URL https://api.semanticscholar.org/CorpusID:271088888
2024 arXiv
-
[129]
The logic of scientific discovery
Karl Popper. The logic of scientific discovery. Routledge, 2005
2005
-
[130]
Citeme: Can language models accurately cite scientific claims? Advances in Neural Information Processing Systems, 37: 0 7847--7877, 2024
Ori Press, Andreas Hochlehnert, Ameya Prabhu, Vishaal Udandarao, Ofir Press, and Matthias Bethge. Citeme: Can language models accurately cite scientific claims? Advances in Neural Information Processing Systems, 37: 0 7847--7877, 2024
2024
-
[131]
Piflow: Principle-aware scientific discovery with multi-agent collaboration, 2025
Yingming Pu, Tao Lin, and Hongyu Chen. Piflow: Principle-aware scientific discovery with multi-agent collaboration, 2025. URL https://arxiv.org/abs/2505.15047
2025
-
[132]
Exploring jiu-jitsu argumentation for writing peer review rebuttals
Sukannya Purkayastha, Anne Lauscher, and Iryna Gurevych. Exploring jiu-jitsu argumentation for writing peer review rebuttals. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14...
2023 doi
-
[133]
Large language models are zero shot hypothesis proposers
Biqing Qi, Kaiyan Zhang, Haoxiang Li, Kai Tian, Sihang Zeng, Zhang-Ren Chen, and Bowen Zhou. Large language models are zero shot hypothesis proposers. arXiv preprint arXiv:2311.05965, 2023
2023 arXiv
-
[134]
Scaling large-language-model-based multi-agent collaboration
Chen Qian, Zihao Xie, Yifei Wang, Wei Liu, Yufan Dang, Zhuoyun Du, Weize Chen, Cheng Yang, Zhiyuan Liu, and Maosong Sun. Scaling large-language-model-based multi-agent collaboration. arXiv preprint arXiv:2406.07155, 2024
2024 arXiv
-
[135]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21 0 (140): 0 1--67, 2020. URL...
2020
-
[136]
Towards scientific intelligence: A survey of llm-based scientific agents
Shuo Ren, Pu Jian, Zhenjiang Ren, Chunlin Leng, Can Xie, and Jiajun Zhang. Towards scientific intelligence: A survey of llm-based scientific agents. arXiv preprint arXiv:2503.24047, 2025
2025
-
[137]
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical discoveries from program search with large language models. Nature, 625 0 (7...
2024
-
[138]
Biodiscoveryagent: An ai agent for designing genetic perturbation experiments
Yusuf Roohani, Andrew Lee, Qian Huang, Jian Vora, Zachary Steinhart, Kexin Huang, Alexander Marson, Percy Liang, and Jure Leskovec. Biodiscoveryagent: An ai agent for designing genetic perturbation experiments. arXiv preprint arXiv:2405.17631, 2024
2024 arXiv
-
[139]
Agentrxiv: Towards collaborative autonomous research
Samuel Schmidgall and Michael Moor. Agentrxiv: Towards collaborative autonomous research. arXiv preprint arXiv:2503.18102, 2025
2025 arXiv
-
[140]
Agent laboratory: Using llm agents as research assistants
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using llm agents as research assistants. arXiv preprint arXiv:2501.04227, 2025
2025 arXiv
-
[141]
u rgen Schmidhuber. G \
J \"u rgen Schmidhuber. G \"o del machines: Fully self-referential optimal universal self-improvers. In Artificial general intelligence, pages 199--226. Springer, 2007
2007
-
[142]
Bigpatent: A large-scale dataset for abstractive and coherent summarization
Eva Sharma, Chen Li, and Lu Wang. Bigpatent: A large-scale dataset for abstractive and coherent summarization. arXiv preprint arXiv:1906.03741, 2019
1906 arXiv
-
[143]
Shortcutsbench: A large-scale real-world benchmark for api-based agents, 2025
Haiyang Shen, Yue Li, Desong Meng, Dongqi Cai, Sheng Qi, Li Zhang, Mengwei Xu, and Yun Ma. Shortcutsbench: A large-scale real-world benchmark for api-based agents, 2025. URL https://arxiv.org/abs/2407.00132
2025 arXiv
-
[144]
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. Detecting pretraining data from large language models. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[145]
B io M egatron: Larger biomedical domain language model
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani. B io M egatron: Larger biomedical domain language model. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empiric...
2020 doi
-
[146]
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36: 0 8634--8652, 2023
2023
-
[147]
The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar. The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity. 2025. URL https://api.semanticscholar.org/C...
2025
-
[148]
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers. arXiv preprint arXiv:2409.04109, 2024
2024 arXiv
-
[149]
Core-bench: Fostering the credibility of published research through a computational reproducibility agent benchmark
Zachary S Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, and Arvind Narayanan. Core-bench: Fostering the credibility of published research through a computational reproducibility agent benchmark. arXiv preprint arXiv:2409.11363, 2024
2024 arXiv
-
[150]
Legobench: Scientific leaderboard generation benchmark, 2024
Shruti Singh, Shoaib Alam, Husain Malwat, and Mayank Singh. Legobench: Scientific leaderboard generation benchmark, 2024
2024
-
[151]
Skarlinski, Sam Cox, Jon M
Michael D. Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela Hinks, Michael J. Hammerling, Manvitha Ponnapati, Samuel G. Rodriques, and Andrew D. White. Language agents achieve superhuman synthesis of scientific knowledge. arXiv preprent arXiv:2409.13740, 2024. URL...
-
[152]
Paperbench: Evaluating ai's ability to replicate ai research
Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, et al. Paperbench: Evaluating ai's ability to replicate ai research. arXiv preprint arXiv:2504.01848, 2025
2025 arXiv
-
[153]
Towards fair, equitable, and efficient peer review
Ivan Stelmakh. Towards fair, equitable, and efficient peer review. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 15736--15737, 2021
2021
-
[154]
Peerreview4all: Fair and accurate reviewer assignment in peer review
Ivan Stelmakh, Nihar B Shah, and Aarti Singh. Peerreview4all: Fair and accurate reviewer assignment in peer review. In Algorithmic Learning Theory, pages 828--856. PMLR, 2019
2019
-
[155]
Shah, Aarti Singh, and Hal Daum'e
Ivan Stelmakh, Nihar B. Shah, Aarti Singh, and Hal Daum'e. Prior and prejudice. Proceedings of the ACM on Human-Computer Interaction, 5: 0 1 -- 17, 2020. URL https://api.semanticscholar.org/CorpusID:227227578
2020
-
[156]
Reviewriter: Ai-generated instructions for peer review writing
Xiaotian Su, Thiemo Wambsganss, Roman Rietsche, Seyed Parsa Neshaei, and Tanja K \"a ser. Reviewriter: Ai-generated instructions for peer review writing. In Workshop on Innovative Use of NLP for Building Educational Applications, 2025. URL https://api.semanticscholar.org/Corpu...
2025
-
[157]
Towards table-to-text generation with numerical reasoning
Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, and Hiroya Takamura. Towards table-to-text generation with numerical reasoning. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Meeting of the Asso...
2021
-
[158]
Lu Sun, Stone Tao, Junjie Hu, and Steven P. Dow. Metawriter: Exploring the potential and perils of ai writing support in scientific peer review. Proceedings of the ACM on Human-Computer Interaction, 8: 0 1 -- 32, 2024. URL https://api.semanticscholar.org/CorpusID:269470548
2024
-
[159]
Undiscovered public knowledge
Don R Swanson. Undiscovered public knowledge. The Library Quarterly, 56 0 (2): 0 103--118, 1986
1986
-
[160]
Agatha: automatic graph mining and transformer based hypothesis generation approach
Justin Sybrandt, Ilya Tyagin, Michael Shtutman, and Ilya Safro. Agatha: automatic graph mining and transformer based hypothesis generation approach. In Proceedings of the 29th ACM international conference on information & knowledge management, pages 2757--2764, 2020
2020
-
[161]
X-scitldr: cross-lingual extreme summarization of scholarly documents
Sotaro Takeshita, Tommaso Green, Niklas Friedrich, Kai Eckert, and Simone Paolo Ponzetto. X-scitldr: cross-lingual extreme summarization of scholarly documents. In Proceedings of the 22nd ACM/IEEE Joint Conference on Digital Libraries, pages 1--12, 2022
2022
-
[162]
Peer review as a multi-turn and long-context dialogue with role-based interactions
Cheng Tan, Dongxin Lyu, Siyuan Li, Zhangyang Gao, Jingxuan Wei, Siqi Ma, Zicheng Liu, and Stan Z Li. Peer review as a multi-turn and long-context dialogue with role-based interactions. arXiv preprint arXiv:2406.05688, 2024
2024 arXiv
-
[163]
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, Shengyan Liu, Di Luo, Yutao Ma, Hao Tong, Kha Trinh, Chenyu Tian, Zihan Wang, Bohao Wu, Yanyu Xiong, Shengzhu Yin, Min Zhu, Kilian Adriano Lieret,...
2024 arXiv
-
[164]
Unsupervised word embeddings capture latent knowledge from materials science literature
Vahe Tshitoyan, John Dagdelen, Leigh Weston, Alexander Dunn, Ziqin Rong, Olga Kononova, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain. Unsupervised word embeddings capture latent knowledge from materials science literature. Nature, 571 0 (7763): 0 95--98, 2019
2019
-
[165]
A new sociology of humans and machines
Milena Tsvetkova, Taha Yasseri, Niccolo Pescetelli, and Tobias Werner. A new sociology of humans and machines. Nature Human Behaviour, 8 0 (10): 0 1864--1876, 2024
2024
-
[166]
Openreviewer: Mitigating challenges in LLM reviewing, 2024
Keith Tyser, Jason Lee, Avi Shporer, Madeleine Udell, Dov Te'eni, and Iddo Drori. Openreviewer: Mitigating challenges in LLM reviewing, 2024. URL https://openreview.net/forum?id=gqmELEp3lZ
2024
-
[167]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017
2017
-
[168]
Teaching code llms to use autocompletion tools in repository-level code generation
Chong Wang, Jian Zhang, Yebo Feng, Tianlin Li, Weisong Sun, Yang Liu, and Xin Peng. Teaching code llms to use autocompletion tools in repository-level code generation. ACM Transactions on Software Engineering and Methodology, 2024 a . URL https://api.semanticscholar.org/Corpus...
2024
-
[169]
Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents
Peisong Wang, Ruotian Ma, Bang Zhang, Xingyu Chen, Zhiwei He, Kang Luo, Qingsong Lv, Qingxuan Jiang, Zheng Xie, Shanyi Wang, et al. Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents. arXiv preprint arXiv:2507.03112, 2025
2025
-
[170]
Paperrobot: Incremental draft generation of scientific ideas
Qingyun Wang, Lifu Huang, Zhiying Jiang, Kevin Knight, Heng Ji, Mohit Bansal, and Yi Luan. Paperrobot: Incremental draft generation of scientific ideas. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1980--1991, 2019
1980
-
[171]
Reviewrobot: Explainable paper review generation based on knowledge synthesis
Qingyun Wang, Qi Zeng, Lifu Huang, Kevin Knight, Heng Ji, and Nazneen Fatema Rajani. Reviewrobot: Explainable paper review generation based on knowledge synthesis. arXiv preprint arXiv:2010.06119, 2020
2010 arXiv
-
[172]
S ci MON : Scientific inspiration machines optimized for novelty
Qingyun Wang, Doug Downey, Heng Ji, and Tom Hope. S ci MON : Scientific inspiration machines optimized for novelty. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...
2024 doi
-
[173]
Shuai Wang, Harrisen Scells, Shengyao Zhuang, Martin Potthast, Bevan Koopman, and G. Zuccon. Zero-shot generative large language models for systematic review screening automation. ArXiv, abs/2401.06320, 2024 c . URL https://api.semanticscholar.org/CorpusID:266977226
2024 arXiv
-
[174]
Autosurvey: Large language models can automatically write surveys
Yidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang, Xin Zhang, Zhen Wu, Meishan Zhang, Xinyu Dai, Qingsong Wen, Wei Ye, et al. Autosurvey: Large language models can automatically write surveys. Advances in Neural Information Processing Systems, 37: 0 115119--115145, 2024 d
2024
-
[175]
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35: 0 24824--24837, 2022
2022
-
[176]
Browsecomp: A simple yet challenging benchmark for browsing agents
Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. Browsecomp: A simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516, 2025
2025 arXiv
-
[177]
Large language models are better reasoners with self-verification
Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao. Large language models are better reasoners with self-verification. arXiv preprint arXiv:2212.09561, 2022
2022 arXiv
-
[178]
Cycleresearcher: Improving automated research via automated review
Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. Cycleresearcher: Improving automated research via automated review. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id...
2025
-
[179]
WestlakeNLP. Airaxiv. https://airaxiv.com/, 2025
2025
-
[180]
Pylabrobot: An open-source, hardware-agnostic interface for liquid-handling robots and accessories
Rick P Wierenga, Stefan M Golas, Wilson Ho, Connor W Coley, and Kevin M Esvelt. Pylabrobot: An open-source, hardware-agnostic interface for liquid-handling robots and accessories. Device, 1 0 (4), 2023
2023
-
[181]
Lag: Llm agents for leaderboard auto generation on demanding
Jian Wu, Jiayu Zhang, Dongyuan Li, Linyi Yang, Aoxiao Zhong, Renhe Jiang, Qingsong Wen, and Yue Zhang. Lag: Llm agents for leaderboard auto generation on demanding. arXiv preprint arXiv:2502.18209, 2025
2025
-
[182]
Automated review generation method based on large language models
Shican Wu, Xiao Ma, Dehui Luo, Lulu Li, Xiangcheng Shi, Xin Chang, Xiaoyun Lin, Ran Luo, Chunlei Pei, Changying Du, et al. Automated review generation method based on large language models. arXiv preprint arXiv:2407.20906, 2024
2024 arXiv
-
[183]
Scireplicate-bench: Benchmarking llms in agent-driven algorithmic reproduction from research papers
Yanzheng Xiang, Hanqi Yan, Shuyin Ouyang, Lin Gui, and Yulan He. Scireplicate-bench: Benchmarking llms in agent-driven algorithmic reproduction from research papers. arXiv preprint arXiv:2504.00255, 2025
2025 arXiv
-
[184]
An empirical analysis of uncertainty in large language model evaluations
Qiujie Xie, Qingqiu Li, Zhuohao Yu, Yuejie Zhang, Yue Zhang, and Linyi Yang. An empirical analysis of uncertainty in large language model evaluations. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=J4xLuCt2kg
2025
-
[185]
The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search. arXiv preprint arXiv:2504.08066, 2025
2025 arXiv
-
[186]
TELIN : Table entity LIN ker for extracting leaderboards from machine learning publications
Sean Yang, Chris Tensmeyer, and Curtis Wigington. TELIN : Table entity LIN ker for extracting leaderboards from machine learning publications. In Tirthankar Ghosal, Sergi Blanco-Cuaresma, Alberto Accomazzi, Robert M. Patton, Felix Grezes, and Thomas Allen, editors, Proceedings...
2022 doi
-
[187]
Ai becomes a masterbrain scientist
Zijie Yang, Yukai Wang, and Lijing Zhang. Ai becomes a masterbrain scientist. bioRxiv, pages 2023--04, 2023
2023
-
[188]
Airalogy: Ai-empowered universal data digitization for research automation, 2025 a
Zijie Yang, Qiji Zhou, Fang Guo, Sijie Zhang, Yexun Xi, Jinglei Nie, Yudian Zhu, Liping Huang, Chou Wu, Yonghe Xia, Xiaoyu Ma, Yingming Pu, Panzhong Lu, Junshu Pan, Mingtao Chen, Tiannan Guo, Yanmei Dou, Hongyu Chen, Anping Zeng, Jiaxing Huang, Tian Xu, and Yue Zhang. Airalogy...
2025 arXiv
-
[189]
Large language models for automated open-domain scientific hypotheses discovery
Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and Erik Cambria. Large language models for automated open-domain scientific hypotheses discovery. In Findings of the Association for Computational Linguistics ACL 2024, pages 13545--13565, 2024
2024
-
[190]
MOOSE -chem: Large language models for rediscovering unseen chemistry scientific hypotheses
Zonglin Yang, Wanhao Liu, Ben Gao, Tong Xie, Yuqiang Li, Wanli Ouyang, Soujanya Poria, Erik Cambria, and Dongzhan Zhou. MOOSE -chem: Large language models for rediscovering unseen chemistry scientific hypotheses. In The Thirteenth International Conference on Learning Represent...
2025
-
[191]
Extracting knowledge from scientific texts on patient-derived cancer models using large language models: algorithm development and validation
Jiarui Yao, Zinaida Perova, Tushar Mandloi, Elizabeth Lewis, Helen Parkinson, and Guergana Savova. Extracting knowledge from scientific texts on patient-derived cancer models using large language models: algorithm development and validation. bioRxiv, pages 2025--01, 2025
2025
-
[192]
Are we there yet? revealing the risks of utilizing large language models in scholarly peer review
Rui Ye, Xianghe Pang, Jingyi Chai, Jiaao Chen, Zhenfei Yin, Zhen Xiang, Xiaowen Dong, Jing Shao, and Siheng Chen. Are we there yet? revealing the risks of utilizing large language models in scholarly peer review. arXiv preprint arXiv:2412.01708, 2024
2024 arXiv
-
[193]
Mcx-llm: an experiment in bridging natural language problem descriptions with quantitative scientific simulations
Fan-Yu Yen and Qianqian Fang. Mcx-llm: an experiment in bridging natural language problem descriptions with quantitative scientific simulations. Optica Biophotonics Congress: Biomedical Optics 2024 (Translational, Microscopy, OCT, OTS, BRAIN), 2024. URL https://api.semanticsch...
2024
-
[194]
Knowledge integration and decision support for accelerated discovery of antibiotic resistance genes
Jason Youn, Navneet Rai, and Ilias Tagkopoulos. Knowledge integration and decision support for accelerated discovery of antibiotic resistance genes. Nature Communications, 13 0 (1): 0 2360, 2022
2022
-
[195]
Researchtown: Simulator of human research community
Haofei Yu, Zhaochen Hong, Zirui Cheng, Kunlun Zhu, Keyang Xuan, Jinwei Yao, Tao Feng, and Jiaxuan You. Researchtown: Simulator of human research community. arXiv preprint arXiv:2412.17767, 2024
2024 arXiv
-
[196]
Is your paper being reviewed by an llm? a new benchmark dataset and approach for detecting ai text in peer review
Sungduk Yu, Man Luo, Avinash Madusu, Vasudev Lal, and Phillip Howard. Is your paper being reviewed by an llm? a new benchmark dataset and approach for detecting ai text in peer review. arXiv preprint arXiv:2502.19614, 2025
2025
-
[197]
Dolphin: Closed-loop open-ended auto-research through thinking, practice, and feedback
Jiakang Yuan, Xiangchao Yan, Botian Shi, Tao Chen, Wanli Ouyang, Bo Zhang, Lei Bai, Yu Qiao, and Bowen Zhou. Dolphin: Closed-loop open-ended auto-research through thinking, practice, and feedback. arXiv preprint arXiv:2501.03916, 2025
2025 arXiv
-
[198]
Can we automate scientific reviewing? Journal of Artificial Intelligence Research, 75: 0 171--212, 2022
Weizhe Yuan, Pengfei Liu, and Graham Neubig. Can we automate scientific reviewing? Journal of Artificial Intelligence Research, 75: 0 171--212, 2022
2022
-
[199]
Appraising the potential uses and harms of llms for medical systematic reviews
Hye Sun Yun, Iain J Marshall, Thomas A Trikalinos, and Byron C Wallace. Appraising the potential uses and harms of llms for medical systematic reviews. arXiv preprint arXiv:2305.11828, 2023
2023 arXiv
-
[200]
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35: 0 15476--15488, 2022
2022
-
[201]
Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation
Qi Zeng, Mankeerat Sidhu, Hou Pong Chan, Lu Wang, and Heng Ji. Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation. In AI4Research/DemocrAI@IJCAI, 2023. URL https://api.semanticscholar.org/CorpusID:258865977
2023
-
[202]
Chen, Yue Zhang, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen
Fengji Zhang, B. Chen, Yue Zhang, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. Repocoder: Repository-level code completion through iterative retrieval and generation. In Conference on Empirical Methods in Natural Language Processing, 2023. URL https://api.se...
2023
-
[203]
Investigating fairness disparities in peer review: A language model enhanced approach
Jiayao Zhang, Hongming Zhang, Zhun Deng, and Dan Roth. Investigating fairness disparities in peer review: A language model enhanced approach. ArXiv, abs/2211.06398, 2022. URL https://api.semanticscholar.org/CorpusID:253499393
2022 arXiv
-
[204]
Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges
Kechi Zhang, Jia Li, Ge Li, Xianjie Shi, and Zhi Jin. Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges. In Annual Meeting of the Association for Computational Linguistics, 2024 a . URL https://api.semanticschol...
2024
-
[205]
A comprehensive survey of scientific large language models and their applications in scientific discovery
Yu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang, Shuiwang Ji, Wei Wang, and Jiawei Han. A comprehensive survey of scientific large language models and their applications in scientific discovery. arXiv preprint arXiv:2406.10833, 2024 b
2024 arXiv
-
[206]
Siren’s song in the ai ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. Siren’s song in the ai ocean: A survey on hallucination in large language models. Computational Linguistics, pages 1--45, 2025
2025
-
[207]
Chain of agents: Large language models collaborating on long-context tasks
Yusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister, Rui Zhang, and Sercan Arik. Chain of agents: Large language models collaborating on long-context tasks. Advances in Neural Information Processing Systems, 37: 0 132208--132237, 2024 c
2024
-
[208]
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36: 0 46595--46623, 2023
2023
-
[209]
From automation to autonomy: A survey on large language models in scientific discovery
Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song. From automation to autonomy: A survey on large language models in scientific discovery. arXiv preprint arXiv:2505.13259, 2025
2025
-
[210]
Large language models penetration in scholarly writing and peer review
Li Zhou, Ruijie Zhang, Xunlian Dai, Daniel Hershcovich, and Haizhou Li. Large language models penetration in scholarly writing and peer review. ArXiv, abs/2502.11193, 2025 a . URL https://api.semanticscholar.org/CorpusID:276408109
2025 arXiv
-
[211]
Is llm a reliable reviewer? a comprehensive evaluation of llm on automatic paper reviewing tasks
Ruiyang Zhou, Lu Chen, and Kai Yu. Is llm a reliable reviewer? a comprehensive evaluation of llm on automatic paper reviewing tasks. In International Conference on Language Resources and Evaluation, 2024 a . URL https://api.semanticscholar.org/CorpusID:269803977
2024
-
[212]
Is LLM a reliable reviewer? a comprehensive evaluation of LLM on automatic paper reviewing tasks
Ruiyang Zhou, Lu Chen, and Kai Yu. Is LLM a reliable reviewer? a comprehensive evaluation of LLM on automatic paper reviewing tasks. In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings of the 2024 Joint ...
2024
-
[213]
From hypothesis to publication: A comprehensive survey of ai-driven research support systems
Zekun Zhou, Xiaocheng Feng, Lei Huang, Xiachong Feng, Ziyun Song, Ruihan Chen, Liang Zhao, Weitao Ma, Yuxuan Gu, Baoxin Wang, et al. From hypothesis to publication: A comprehensive survey of ai-driven research support systems. arXiv preprint arXiv:2503.01424, 2025 b
2025
-
[214]
Deepreview: Improving llm-based paper review with human-like deep thinking process
Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang. Deepreview: Improving llm-based paper review with human-like deep thinking process. arXiv preprint arXiv:2503.08569, 2025
2025 arXiv
-
[215]
Large language models for automated scholarly paper review: A survey
Zhenzhen Zhuang, Jiandong Chen, Hongfeng Xu, Yuwen Jiang, and Jialiang Lin. Large language models for automated scholarly paper review: A survey. arXiv preprint arXiv:2501.10326, 2025
2025 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.