Pith. sign in

REVIEW 4 cited by

LitSearch: A Retrieval Benchmark for Scientific Literature Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.18940 v2 pith:FTGKTRDT submitted 2024-07-10 cs.IR cs.AIcs.CLcs.DLcs.LG

classification cs.IRcs.AIcs.CLcs.DLcs.LG
keywords litsearchsearchquestionsretrievalresearchbenchmarkdenseliterature
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Literature search questions, such as "Where can I find research on the evaluation of consistency in generated summaries?" pose significant challenges for modern search engines and retrieval systems. These questions often require a deep understanding of research concepts and the ability to reason across entire articles. In this work, we introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers. LitSearch is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and (2) questions manually written by authors about their recently published papers. All LitSearch questions were manually examined or edited by experts to ensure high quality. We extensively benchmark state-of-the-art retrieval models and also evaluate two LLM-based reranking pipelines. We find a significant performance gap between BM25 and state-of-the-art dense retrievers, with a 24.8% absolute difference in recall@5. The LLM-based reranking strategies further improve the best-performing dense retriever by 4.4%. Additionally, commercial search engines and research tools like Google Search perform poorly on LitSearch, lagging behind the best dense retriever by up to 32 recall points. Taken together, these results show that LitSearch is an informative new testbed for retrieval systems while catering to a real-world use case.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

    cs.CL 2025-10 reject novelty 6.0 of 10

    A tool-augmented reward model trained with GRPO on 27K synthetic pairs beats existing reward models on long-form QA judgment and improves downstream alignment.

  2. NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question Answering

    cs.CL 2025-05 conditional novelty 6.0 of 10

    NeuSym-RAG combines SQL-based symbolic retrieval with neural vector search in an iterative LLM agent, using multi-view PDF parsing, and reports large gains over simple RAG baselines on full-paper QA.

  3. How Far Are AI Scientists from Changing the World?

    cs.AI 2025-07 conditional novelty 4.0 of 10

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

  4. No Stupid Questions: An Analysis of Question Query Generation for Citation Recommendation

    cs.IR 2025-06 conditional novelty 4.0 of 10

    LLM-generated questions are sometimes better retrieval queries than extractive keywords for citation recommendation, but selecting the best question at inference time remains largely unsolved.

Pith tools