LoHoSearch is a new benchmark of 544 KG-constructed questions across 11 domains where the strongest search agent scores 34.74% and context strategies add at most 6.8%.
The Web as a Knowledge-Base for Answering Complex Questions
8 Pith papers cite this work. Polarity classification is still indexing.
years
2026 8representative citing papers
KG-Guard augments knowledge graphs with a virtual question node and uses a graph encoder plus MLP to classify LLM-proposed answers as hallucinations or not, reporting superior F1 scores and downstream improvements on three benchmarks.
DualGraph combines semantic textual KGs with symbolic KGs for semi-structured QA and introduces the SpecsQA benchmark, outperforming baselines on both open and specification questions.
CacheRAG is a cache-augmented architecture for LLM KGQA using ISR parsing, hierarchical MMR-based retrieval, and bounded subgraph expansion, claiming +13.2% accuracy gains on CRAG.
GraphWalker achieves state-of-the-art results on CWQ and WebQSP by training KGQA agents via synthetic random-walk trajectories in stage-wise SFT plus RL, with improved out-of-distribution generalization.
OCC-RAG develops task-specialized SLMs (0.6B and 1.7B) via a new synthetic data pipeline for multi-hop reasoning and context faithfulness, claiming to match or exceed 2-6x larger general models on HotpotQA, MuSiQue, TAT-QA, ConFiQA, and MuSiQue-Un.
KnowledgeBerg is a 4,800-question, 17-language benchmark showing LLMs fail at systematically enumerating bounded knowledge universes and performing compositional set-based reasoning over them.
STAR is a semantic-tuned and tail-adaptive retriever for GraphRAG that uses cross-attention interaction learning and path-weighted contrastive learning to mitigate Semantic Shortcut Bias and Long-Tail Path Bias, reporting 1.8% retrieval and 2.2% QA gains.
citing papers explorer
-
LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling
LoHoSearch is a new benchmark of 544 KG-constructed questions across 11 domains where the strongest search agent scores 34.74% and context strategies add at most 6.8%.
-
KG-Guard: Graph-Based Hallucination Detection for Knowledge Base Question Answering
KG-Guard augments knowledge graphs with a virtual question node and uses a graph encoder plus MLP to classify LLM-proposed answers as hallucinations or not, reporting superior F1 scores and downstream improvements on three benchmarks.
-
Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering
DualGraph combines semantic textual KGs with symbolic KGs for semi-structured QA and introduces the SpecsQA benchmark, outperforming baselines on both open and specification questions.
-
CacheRAG: A Semantic Caching System for Retrieval-Augmented Generation in Knowledge Graph Question Answering
CacheRAG is a cache-augmented architecture for LLM KGQA using ISR parsing, hierarchical MMR-based retrieval, and bounded subgraph expansion, claiming +13.2% accuracy gains on CRAG.
-
GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum
GraphWalker achieves state-of-the-art results on CWQ and WebQSP by training KGQA agents via synthetic random-walk trajectories in stage-wise SFT plus RL, with improved out-of-distribution generalization.
-
OCC-RAG: Optimal Cognitive Core for Faithful Question Answering
OCC-RAG develops task-specialized SLMs (0.6B and 1.7B) via a new synthetic data pipeline for multi-hop reasoning and context faithfulness, claiming to match or exceed 2-6x larger general models on HotpotQA, MuSiQue, TAT-QA, ConFiQA, and MuSiQue-Un.
-
KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models
KnowledgeBerg is a 4,800-question, 17-language benchmark showing LLMs fail at systematically enumerating bounded knowledge universes and performing compositional set-based reasoning over them.
-
STAR: Semantic-Tuned and Tail-Adaptive Retriever for Graph-Augmented Generation
STAR is a semantic-tuned and tail-adaptive retriever for GraphRAG that uses cross-attention interaction learning and path-weighted contrastive learning to mitigate Semantic Shortcut Bias and Long-Tail Path Bias, reporting 1.8% retrieval and 2.2% QA gains.