REVIEW 20 cited by
OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Scientific progress depends on researchers' ability to synthesize the growing body of literature. Can large language models (LMs) assist scientists in this task? We introduce OpenScholar, a specialized retrieval-augmented LM that answers scientific queries by identifying relevant passages from 45 million open-access papers and synthesizing citation-backed responses. To evaluate OpenScholar, we develop ScholarQABench, the first large-scale multi-domain benchmark for literature search, comprising 2,967 expert-written queries and 208 long-form answers across computer science, physics, neuroscience, and biomedicine. On ScholarQABench, OpenScholar-8B outperforms GPT-4o by 5% and PaperQA2 by 7% in correctness, despite being a smaller, open model. While GPT4o hallucinates citations 78 to 90% of the time, OpenScholar achieves citation accuracy on par with human experts. OpenScholar's datastore, retriever, and self-feedback inference loop also improves off-the-shelf LMs: for instance, OpenScholar-GPT4o improves GPT-4o's correctness by 12%. In human evaluations, experts preferred OpenScholar-8B and OpenScholar-GPT4o responses over expert-written ones 51% and 70% of the time, respectively, compared to GPT4o's 32%. We open-source all of our code, models, datastore, data and a public demo.
Forward citations
Cited by 20 Pith papers
-
FARS: A Fully Automated Research System Deployed at Scale
FARS deployed at scale produced 166 AI/ML papers across 67 topics that received 282 structured human reviews indicating some review-worthy outputs alongside recurring failure modes.
-
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
NatureBench evaluates ten frontier AI coding agents on 90 tasks from Nature papers under web-search-disabled conditions and finds the strongest agent surpasses published SOTA on only 17.8% of tasks, succeeding mainly ...
-
Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
Typed quantity verification exposes a canonical-equivalence blind spot in neural fact-checkers; Symbolic Augmentation fixes it (36.5%→98.2%) and transfers to SciFact-Open (+0.037 binary macro-F1).
-
From Queries to Criteria: Understanding How Astronomers Evaluate LLMs
A user study of an astronomy RAG bot identifies the question types and evaluation criteria astronomers actually use, and turns them into a 40-item benchmark.
-
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
SciVer is the first benchmark for multimodal scientific claim verification over full paper context, and current foundation models score about 16 points below expert humans.
-
Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
Scientific-memory leaderboards are misleading unless retrieval budget and modality are reported; under matched budgets, simple RAG baselines tie with structured memory systems.
-
Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis
Citation-faithfulness metrics for AI science agents are verifier-dependent (3–18% on identical outputs), and a split-conformal guard provides a finite-sample catch-rate guarantee anchored on human gold.
-
Multi-Turn Agentic Scientific Literature Search via Workflow Induction
PaperPilot induces executable DAG workflows for multi-turn literature search and trains via imitation plus preference optimization, raising Hit@5 from 58.0 to 77.0 over a baseline agent.
-
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
AISE-Bench is a real-user-query benchmark with annotated API trajectories and grounded answers that exposes LLM agents' weak performance in multi-step academic information seeking.
-
Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving
Pattern-aware speculative tool execution cuts agent end-to-end latency by roughly half and observed tool latency by about 1.8× by overlapping predicted tools with LLM generation.
-
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
Idea evaluation can be automated with LLM agents that retrieve heterogeneous online evidence, simulate diverse reviewers, and score ideas on multiple dimensions, outperforming existing judges on acceptance-label predi...
-
A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
A scalable literature-synthesis pipeline that retrieves, filters, extracts, summarizes, and converts AR-model papers into runnable training scripts, with F1 > 0.85 extraction and three reproduction case studies.
-
Characterizing Deep Research: A Benchmark and Formal Definition
Deep research is characterized by high search and reasoning intensity; the new LiveDRBench measures claim-level precision and recall, where the best current model scores 0.55 F1.
-
HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
HedraRAG uses a graph abstraction and dynamic transformations to pipeline generation and retrieval stages, achieving 1.5x to 5x speedups in heterogeneous RAG serving.
-
Scaffolding Recursive Divergence and Convergence in Story Ideation
Reverger scaffolds recursive divergence and multi-direction convergence in story ideation, with a user study whose headline significance statistics are internally impossible.
-
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
LIMITGEN evaluates LLMs on identifying paper limitations and shows strong models still miss roughly half of obvious flaws, with retrieval augmentation giving modest but consistent gains.
-
Compare: A Framework for Scientific Comparisons
Compare is a RAG-based system that generates qualitative, citation-supported comparisons of scientific contributions at institution and publication granularity.
-
Interaction as Intelligence: Deep Research With Human-AI Partnership
A human-in-the-loop deep research system with transparent, interruptible interaction is claimed to outperform commercial baselines, but the evidence is weakened by small samples and biased instructions.
-
SciSage: A Multi-Agent Framework for High-Quality Scientific Survey Generation
SciSage, a multi-agent reflection-based framework, is reported to outperform previous LLM survey generators on coherence and citation F1, while a new benchmark, SurveyScope, enables standardized evaluation.
-
AI4Research: A Survey of Artificial Intelligence for Scientific Research
A survey that organizes AI-for-research work into five tasks, comprehension, survey, discovery, writing, and peer review, and compiles associated tools and benchmarks.
Discussion (0). Sign in to comment.