A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.
LitSearch: A retrieval benchmark for scientific literature search
5 Pith papers cite this work, alongside 9 external citations. Polarity classification is still indexing.
representative citing papers
A new benchmark (IG-Bench) reveals that LLM-based scientists fail at compositional lineage reasoning, with the best system reaching only 27.3% exact accuracy.
HyBIRD adds point, cone, and factorized hyperbolic bridges over a frozen dense retriever for methodology inspiration retrieval, reaching 59.034 mAP on the MIR benchmark while producing query need profiles and evidence bundles.
Search-R3 trains LLMs to output search embeddings as a direct product of step-by-step reasoning via supervised pre-training and a specialized RL environment that avoids full corpus re-encoding.
PaSaMaster, a recursive self-evolving agentic literature retrieval system, reports 16.5× higher F1 than Google Scholar and 37.8% higher F1 than GPT-5.2 on PaSaMaster-Bench across 38 disciplines with zero source hallucination at ~1% cost.
citing papers explorer
-
Can AI Agents Synthesize Scientific Conclusions?
A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.
-
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
A new benchmark (IG-Bench) reveals that LLM-based scientists fail at compositional lineage reasoning, with the best system reaching only 27.3% exact accuracy.
-
HyBIRD: Hyperbolic Bridge Retrieval and Diagnosis for Methodology Inspiration Retrieval
HyBIRD adds point, cone, and factorized hyperbolic bridges over a frozen dense retriever for methodology inspiration retrieval, reaching 59.034 mAP on the MIR benchmark while producing query need profiles and evidence bundles.
-
Search-R3: Unifying Reasoning and Embedding in Large Language Models
Search-R3 trains LLMs to output search embeddings as a direct product of step-by-step reasoning via supervised pre-training and a specialized RL environment that avoids full corpus re-encoding.
-
Towards Recursive Self-Evolving Agentic Literature Retrieval
PaSaMaster, a recursive self-evolving agentic literature retrieval system, reports 16.5× higher F1 than Google Scholar and 37.8% higher F1 than GPT-5.2 on PaSaMaster-Bench across 38 disciplines with zero source hallucination at ~1% cost.