REVIEW 67 cited by
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advancements in large language models (LLMs) have sparked optimism about their potential to accelerate scientific discovery, with a growing number of works proposing research agents that autonomously generate and validate new ideas. Despite this, no evaluations have shown that LLM systems can take the very first step of producing novel, expert-level ideas, let alone perform the entire research process. We address this by establishing an experimental design that evaluates research idea generation while controlling for confounders and performs the first head-to-head comparison between expert NLP researchers and an LLM ideation agent. By recruiting over 100 NLP researchers to write novel ideas and blind reviews of both LLM and human ideas, we obtain the first statistically significant conclusion on current LLM capabilities for research ideation: we find LLM-generated ideas are judged as more novel (p < 0.05) than human expert ideas while being judged slightly weaker on feasibility. Studying our agent baselines closely, we identify open problems in building and evaluating research agents, including failures of LLM self-evaluation and their lack of diversity in generation. Finally, we acknowledge that human judgements of novelty can be difficult, even by experts, and propose an end-to-end study design which recruits researchers to execute these ideas into full projects, enabling us to study whether these novelty and feasibility judgements result in meaningful differences in research outcome.
Forward citations
Showing 60 of 67 Pith papers that cite this
-
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
A randomized execution study with 43 experts shows that LLM-generated research ideas lose more of their appeal than human ideas when actually implemented, reversing part of their ideation-stage advantage.
-
Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution
A citation-graph system that traces how research gaps evolve along branching literature trajectories can generate AI research ideas rated near human-paper quality on novelty and groundedness.
-
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Agent-safety benchmark scores are not interchangeable: R-Judge's F1 is matched by an always-unsafe baseline, rankings differ across benchmarks, and which held-out outcome you choose flips the capability–safety correlation.
-
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
Conference accept/reject outcomes yield 15 operational ideation patterns that, as an LLM skill suite, improve automated-judged research-proposal quality over no-skill and generic-skill baselines.
-
FARS: A Fully Automated Research System Deployed at Scale
FARS deployed at scale produced 166 AI/ML papers across 67 topics that received 282 structured human reviews indicating some review-worthy outputs alongside recurring failure modes.
-
Formalizing Learning from Language Feedback with Provable Guarantees
Introduces a formal framework for learning from language feedback, a transfer eluder dimension complexity measure, and HELiX, a no-regret algorithm whose regret scales with this dimension.
-
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
Model benchmark performance only weakly predicts how well people learn from AI explanations, with notable outliers across code and math.
-
Predicting Empirical AI Research Outcomes with Language Models
A fine-tuned, retrieval-augmented language model predicts which of two AI research ideas will perform better empirically, beating human experts and frontier models on a new contamination-controlled benchmark.
-
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
SPOT shows that state-of-the-art AI models detect fewer than one in five known errors in full scientific papers, with precision below 7%.
-
Echoes in AI: Quantifying lack of plot diversity in LLM outputs
The Sui Generis score quantifies plot-level uniqueness in LLM story generation and shows that GPT-4 and LLaMA-3 stories contain more repeated plot elements than human-written stories.
-
Training AI Scientists to Replicate Research
A 27B-parameter post-trained agent, Faraday, outperforms frontier coding agents at replicating held-out research figures by directing a larger coding model as a tool.
-
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
LigBench evaluates research ideas through formalization, pairwise comparison against 11k+ papers with debiased OpenReview scores, and Elo-based score aggregation, validated against human experts and paper acceptance outcomes.
-
Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods
Idea Search, a tree search variant that samples and updates a dynamically scored bank of atomic design ideas, modestly improves LLM-generated single-cell batch integration code over a pure tree search baseline.
-
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
In two adversarial real-world domains with time-delayed expert ground truth, frontier LLMs over-generate plausible ideas but under-select the specific solutions experts adopted, indicating a filtering and prioritization gap.
-
Can AI Follow In Einstein's Footsteps?
AI for physics has moved from explicit equation discovery to black-box prediction, a trajectory the authors argue reverses the historical progression of human physics, leaving the invention of new mathematical framewo...
-
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review
In a controlled test, three mid-2025 LLMs shared under 6% of literature references with physics experts, and 64% of their real references had at least one metadata error.
-
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
MetaEvolve trains LLMs with reinforcement learning on synthesized code-refinement trajectories, reporting large gains on coding and numerical-optimization benchmarks.
-
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
HypoArena is a 988-case benchmark that asks LLMs to generate hypothesis sets from conclusion-free reconstructed contexts and ranks 15 models via pairwise arena judgments.
-
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
A new benchmark (IG-Bench) reveals that LLM-based scientists fail at compositional lineage reasoning, with the best system reaching only 27.3% exact accuracy.
-
Beyond the Golden Record: Toward a Design Theory for Trustworthy Master Data Management with Self-Sovereign Identity
A design theory is derived for trustworthy master data management based on self-sovereign identity to support reliable, sovereign, and accountable data sharing in data ecosystems.
-
AI Can Learn Scientific Taste
Reinforcement learning on citation-preference pairs teaches a model to predict which papers will be cited more and to propose ideas that LLM judges rate as likely to be cited more—but "taste" here means citation impact.
-
HypoChainer: A Collaborative System Combining LLMs and Knowledge Graphs for Hypothesis-Driven Scientific Discovery
In a small user study and two case studies, a hypothesis-chain workflow grounded in knowledge graphs helped biomedical researchers construct and validate hypotheses from machine-learning predictions more effectively t...
-
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
LIMITGEN evaluates LLMs on identifying paper limitations and shows strong models still miss roughly half of obvious flaws, with retrieval augmentation giving modest but consistent gains.
-
THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning?
THE-Tree constructs causally-linked semantic evolution trees from surveys and literature, and the authors report improved graph completion, future prediction, and LLM-based paper evaluation.
-
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
LLM-based long-horizon event simulation, used as a reward signal, is claimed to improve safety alignment and indirect-harm detection, but evaluation confounds simulation with the capability of the external projector model.
-
Effective Red-Teaming of Policy-Adherent Agents
A policy-aware red-teaming system (CRAFT) induces policy violations in LLM customer service agents at much higher rates than generic jailbreak prompts, using a new security-focused benchmark (tau-break) built from tau-bench.
-
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
ScienceMeter evaluates language model knowledge updates across three axes, preservation of old scientific claims, acquisition of new claims, and projection to future findings, and finds all current methods fall short.
-
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
LLMs struggle to use passive observations for reverse engineering, but active intervention improves performance, largely through the process of generating queries rather than the data obtained.
-
Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis
A meta-analysis of 28 studies finds no average creativity gap between GenAI and humans, a small boost when humans collaborate with GenAI, and a large drop in idea diversity in those collaborations.
-
Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
A new benchmark (TruthHypo) and a knowledge-grounded hallucination detector (KnowHD) show that grounding scores can partially select truthful LLM-generated biomedical hypotheses, but the result is at risk from knowled...
-
Towards Automated Scoping of AI for Social Good Projects
A retrieval-augmented LLM pipeline, the Problem Scoping Agent, generates AI4SG project proposals that blind human reviewers score as comparable to expert-written proposals.
-
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
IRIS combines Monte Carlo Tree Search, fine-grained LLM review, and targeted literature retrieval in a human-in-the-loop system for generating research briefs.
-
Aspirational Affordances of AI
The paper defines aspirational affordances and aspirational harms, arguing that AI can reduce the interpretive resources people use to imagine alternative futures.
-
AI Idea Bench 2025: AI Research Idea Generation Benchmark
The paper builds a post-cutoff benchmark dataset and a multi-metric evaluation framework for scoring LLM-generated research ideas against real paper motivations and experiment plans.
-
Sparks of Science: Hypothesis Generation Using Structured Paper Data
The authors built HypoGen, 5,478 Bit-Flip-Spark hypothesis triples with reasoning chains from NeurIPS 2023 and ICLR 2024 papers, and fine-tuned LLaMA models on it, reporting higher feasibility but lower diversity in g...
-
Co-Writing with AI, on Human Terms: Aligning Research with User Demands Across the Writing Process
Writers' preferred level of AI control in co-writing splits by whether they contribute content or form, and the current literature often over-assists the very stages writers want to own.
-
Automatic Evaluation Metrics for Artificially Generated Scientific Research
A simple title-and-abstract model predicts citation counts better than review scores and outperforms LLM reviewers in matching human review scores, but remains below human consistency.
-
How do Humans and Language Models Reason About Creativity? A Comparative Analysis
When LLMs rate STEM solution originality, showing example solutions improves accuracy but sharply increases correlations among creativity facets to near 1, unlike human raters.
-
We're Different, We're the Same: Creative Homogeneity Across LLMs
Across three divergent-thinking tests, responses from seven LLM families were substantially more similar to one another than responses from 102 humans were to one another.
-
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
Dolphin closes the loop between idea generation, code implementation, and experimental feedback, improving on baselines and producing one 3D classification model comparable to a human-designed state-of-the-art.
-
MapExplorer: New Content Generation from Low-Dimensional Visualizations
A new task and benchmark that generate contextually aligned text for arbitrary coordinates in 2D projection maps, evaluated with an LLM-based metric.
-
Large Language Models show both individual and collective creativity comparable to humans
The best LLMs match the average human on overall creativity, excel in divergent thinking and problem solving, lag in creative writing, and when sampled repeatedly match a small human group.
-
Bridging AI and Science: Implications from a Large-Scale Literature Analysis of AI4Science
Using LLM-extracted labels from 162,656 papers, this analysis shows that AI4Science connections are sparse and uneven, with a few hub areas dominating while many scientific problems and AI methods remain weakly linked.
-
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
AI-generated one-page research proposals are scored about the same as human-written ones by human reviewers, but AI reviewers favor AI-written proposals by roughly one point and detect authorship perfectly.
-
Spacer: Towards Engineered Scientific Inspiration
Spacer proposes a graph-based keyword recombinator plus LLM pipeline that generates plausible scientific hypotheses, validated by reconstructing recent paper theses and embedding similarity to published work.
-
Deep Researcher with Test-Time Diffusion
A test-time 'denoising' loop that repeatedly revises a draft report using fresh web retrieval, plus a component-wise self-evolution step, beats existing deep research agents on several benchmarks.
-
AI-Researcher: Autonomous Scientific Innovation
AI-Researcher runs an end-to-end ML research pipeline with LLM agents, and Scientist-Bench measures how close the resulting papers come to human-authored publications.
-
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
A closed-loop LLM-agent framework that auto-generates research ideas and code, reported to improve baseline performance on all 12 tasks it was tested on.
-
Sparks: Multi-Agent Artificial Intelligence Model Discovers Protein Design Principles
An autonomous multi-agent AI claims it discovered that beta-sheet peptides become mechanically stronger than alpha-helices above roughly 80 residues, yet the tested range stops at 80 and the difference is not statisti...
-
AI Safety Should Prioritize the Future of Work
AI safety should treat the future of work as a core concern, with worker support, transparent training data, and collective licensing.
-
LLM4SR: A Survey on Large Language Models for Scientific Research
A systematic review of LLM-based systems for hypothesis discovery, experiment planning, scientific writing, and peer review, including benchmarks, evaluation methods, and open challenges.
-
LangYa: Revolutionizing Cross-Spatiotemporal Ocean Forecasting
LangYa is a single AI model that forecasts global ocean temperature, salinity, and currents for 1 to 7 days at 1/12° resolution, with reported RMSE improvements over numerical systems and the XiHe AI model.
-
On the Role of Model Prior in Real-World Inductive Reasoning
LLMs' hypotheses for real-world classification tasks are driven mostly by task priors, and in-context demonstrations, even with flipped labels, do little to change them.
-
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
A human-scored benchmark shows that even GPT-4o averages below 4/5 correctness and all tested models struggle with scientific diagram prompts that combine spatial, numeric, and attribute requirements.
-
AIGS: Generating Science from AI-Powered Automated Falsification
Baby-AIGS is a multi-agent system that automates research through explicit falsification, outperforming its baseline on three ML tasks but lagging human experts.
-
AI for Auto-Research: Roadmap & User Guide
The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.
-
How Far Are AI Scientists from Changing the World?
This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.
-
Agent Ideate: A Framework for Product Idea Generation from Patents Using Agentic AI
A multi-agent LLM framework with optional web search generates product ideas from patents and outperforms a single-prompt LLM on 150 patents, though results vary by domain.
-
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.
-
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
The paper advocates protecting and leveraging OpenReview's peer review corpus as a community asset for LLM-based review assistance, benchmarks, and alignment.
Discussion (0). Continue with ORCID to comment.