Pith. sign in

REVIEW 24 cited by

LLM4SR: A Survey on Large Language Models for Scientific Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.04306 v1 pith:UHUYKZA5 submitted 2025-01-08 cs.CL cs.DL

LLM4SR: A Survey on Large Language Models for Scientific Research

classification cs.CL cs.DL
keywords researchllmsscientificsurveyacrosslanguagelargellm4sr
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In recent years, the rapid advancement of Large Language Models (LLMs) has transformed the landscape of scientific research, offering unprecedented support across various stages of the research cycle. This paper presents the first systematic survey dedicated to exploring how LLMs are revolutionizing the scientific research process. We analyze the unique roles LLMs play across four critical stages of research: hypothesis discovery, experiment planning and implementation, scientific writing, and peer reviewing. Our review comprehensively showcases the task-specific methodologies and evaluation benchmarks. By identifying current challenges and proposing future research directions, this survey not only highlights the transformative potential of LLMs, but also aims to inspire and guide researchers and practitioners in leveraging LLMs to advance scientific inquiry. Resources are available at the following repository: https://github.com/du-nlp-lab/LLM4SR

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences

    cs.AI 2026-02 accept novelty 8.0

    ReplicatorBench evaluates LLM agents on replicating social and behavioral science claims across retrieval, computation, and interpretation stages, finding strength in experiment execution but weakness in resource retrieval.

  2. Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape

    cs.SE 2026-04 accept novelty 7.0

    A survey of 457 SE researchers finds widespread GenAI use concentrated in writing and ideation, with productivity gains but persistent concerns over accuracy, bias, and the need for clearer governance rules.

  3. AlphaEvolve: A coding agent for scientific and algorithmic discovery

    cs.AI 2025-06 unverdicted novelty 7.0

    AlphaEvolve is an LLM-orchestrated evolutionary coding agent that discovered a 4x4 complex matrix multiplication algorithm using 48 scalar multiplications, the first improvement over Strassen's algorithm in 56 years, ...

  4. Graph2Idea:Retrieval-Augmented Scientific Idea Generation with Graph-Structured Contexts

    cs.AI 2026-06 unverdicted novelty 6.0

    Graph2Idea builds dynamic knowledge graphs from retrieved literature to supply compact, relational contexts that guide LLMs in generating novel, feasible, and high-quality scientific ideas, outperforming flat-text bas...

  5. GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

    cs.IR 2026-06 unverdicted novelty 6.0

    GIScholarBench shows LLMs exhibit consistent overconfidence across three scholarly tasks in GIS, with different manifestations in factual retrieval, citation expansion, and idea generation.

  6. SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

    cs.AI 2026-05 unverdicted novelty 6.0

    SciResearcher automates creation of diverse scientific reasoning tasks from academic evidence to train an 8B model that sets new SOTA at 19.46% on HLE-Bio/Chem-Gold and gains 13-15% on SuperGPQA-Hard-Biology and TRQA-...

  7. When AI reviews science: Can we trust the referee?

    cs.AI 2026-04 unverdicted novelty 6.0

    AI peer review systems are vulnerable to prompt injections, prestige biases, assertion strength effects, and contextual poisoning, as demonstrated by a new attack taxonomy and causal experiments on real conference sub...

  8. DeepInflation: an AI agent for research and model discovery of inflation

    astro-ph.CO 2026-01 conditional novelty 6.0

    An LLM agent with symbolic regression finds simple inflation potentials that match target CMB observables, but the outputs are fitted to the targets rather than independently predicted.

  9. CacheClip: Accelerating RAG with Effective KV Cache Reuse

    cs.LG 2025-10 unverdicted novelty 6.0

    CacheClip accelerates RAG prefill by up to 3.33x via auxiliary-model-guided selective KV recomputation while retaining 85-91% of full-attention quality on NIAH and LongBench.

  10. RExBench: Can coding agents autonomously implement AI research extensions?

    cs.CL 2025-06 unverdicted novelty 6.0

    RExBench is a new benchmark showing that LLM coding agents fail to autonomously implement most realistic research extensions to prior AI papers.

  11. Scientific discovery as meta-optimization: a combinatorial optimization case study

    cs.AI 2026-06 unverdicted novelty 5.0

    Introduces consensus objective aggregation for meta-optimization of scientific discovery and reports improved scaling and speedup for 3-SAT algorithm discovery using digital MemComputing machines.

  12. PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality

    cs.CL 2026-06 unverdicted novelty 5.0

    PeerCheck finds that chain-of-thought prompting improves LLM academic reviews while retrieval-augmented generation sometimes lowers quality, and that LLMs and humans emphasize different aspects of papers.

  13. SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding

    cs.CL 2026-06 unverdicted novelty 5.0

    SciLens introduces an evidence-conditioned atomic entailment framework that grounds claims to modality-specific witnesses in tables and figures, achieving 79.2% macro-F1 on SciClaimEval.

  14. EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation

    cs.CL 2026-05 unverdicted novelty 5.0

    EvoGens uses rank-based mutation, semantic-aware crossover, and lightweight evaluation to evolve populations of LLM-generated scientific ideas, boosting novelty and diversity metrics.

  15. SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

    cs.AI 2026-05 unverdicted novelty 5.0

    SciResearcher is a new agentic data-construction framework that trains an 8B model via supervised fine-tuning and reinforcement learning to reach 19.46% on HLE-Bio/Chem-Gold and 13-15% gains on related biology and lit...

  16. MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts

    cs.CL 2026-04 unverdicted novelty 5.0

    MedConclusion is a 5.7M-instance benchmark dataset for generating biomedical conclusions from structured PubMed abstracts, with LLM evaluations showing conclusion writing differs from summarization and that judge choi...

  17. Towards Fully Automated Molecular Simulations: Multi-Agent Framework for Simulation Setup and Force Field Extraction

    cs.AI 2025-09 conditional novelty 5.0

    A multi-agent LLM system generates RASPA simulation inputs and extracts literature force field parameters with moderate to high accuracy on a small set of zeolite tasks.

  18. Hephaestus: Toward a Cybersecurity AI Scientist

    cs.CR 2026-06 unverdicted novelty 4.0

    The paper proposes the Cybersecurity AI Scientist as a modular multi-agent architecture for automating cybersecurity research, distinguished by its focus on non-stationary threats and anchored in a four-zeros risk-tru...

  19. LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

    cs.CL 2026-06 unverdicted novelty 4.0

    A survey synthesizing LLM methods for peer review critique generation and score prediction, including taxonomies, benchmark limitations, domain biases, and robustness risks such as prompt injection.

  20. AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

    cs.AI 2026-05 unverdicted novelty 4.0

    A survey organizing AI-powered research automation into five workflow stages, defining AutoResearch and Vibe Research, and proposing five evaluation dimensions while noting domain-conditioned limits on autonomy.

  21. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  22. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 conditional novelty 4.0

    AI can generate research artifacts faster than it can verify them, so across all eight lifecycle stages the credible deployment mode is human-governed collaboration rather than full autonomy.

  23. Conversational AI for Rapid Scientific Prototyping: A Case Study on ESA's ELOPE Competition

    cs.AI 2026-01 conditional novelty 4.0

    One engineer paired with ChatGPT and reached second place in ESA's ELOPE competition in about one week of work; the paper draws best-practice lessons from that experience.

  24. Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator

    cs.DL 2025-07 unverdicted novelty 4.0

    The paper proposes a four-role framework for LLMs in scientific innovation and reviews methods, benchmarks, and limitations across Assistant, Collaborator, Scientist, and Evaluator roles.