The paper delivers the first survey of abductive reasoning in LLMs, a unified two-stage taxonomy, a compact benchmark, and an analysis of gaps relative to deductive and inductive reasoning.
Detectiveqa: Evaluating long-context reasoning on detective novels.arXiv preprint arXiv:2409.02465
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Many-shot CoT-ICL improves when demonstrations are ordered for smooth conceptual progression, with CDS delivering up to 5.42 percentage-point gains on math tasks using 64 examples.
MemoryAgentBench is a multi-turn benchmark covering four memory competencies, and current memory agents fail at selective forgetting and long-range understanding.
citing papers explorer
-
Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs
The paper delivers the first survey of abductive reasoning in LLMs, a unified two-stage taxonomy, a compact benchmark, and an analysis of gaps relative to deductive and inductive reasoning.
-
Many-Shot CoT-ICL: Making In-Context Learning Truly Learn
Many-shot CoT-ICL improves when demonstrations are ordered for smooth conceptual progression, with CDS delivering up to 5.42 percentage-point gains on math tasks using 64 examples.
-
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
MemoryAgentBench is a multi-turn benchmark covering four memory competencies, and current memory agents fail at selective forgetting and long-range understanding.
- Trace Only What You Need: Structure-Aware On-Demand Hypergraph Memory for Long-Document Question Answering