REVIEW 7 cited by
A Survey of Large Language Model Agents for Question Answering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper surveys the development of large language model (LLM)-based agents for question answering (QA). Traditional agents face significant limitations, including substantial data requirements and difficulty in generalizing to new environments. LLM-based agents address these challenges by leveraging LLMs as their core reasoning engine. These agents achieve superior QA results compared to traditional QA pipelines and naive LLM QA systems by enabling interaction with external environments. We systematically review the design of LLM agents in the context of QA tasks, organizing our discussion across key stages: planning, question understanding, information retrieval, and answer generation. Additionally, this paper identifies ongoing challenges and explores future research directions to enhance the performance of LLM agent QA systems.
Forward citations
Cited by 7 Pith papers
-
Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution
An eight-agent question-asking system that front-loads intent clarification produced more complete prompts, higher-rated outputs, and single-turn task completion in a four-person pilot, with unstable effect sizes.
-
TopoTuner: Topological Finetuning of Large Language Models
TopoTuner uses Wasserstein distances between persistence diagrams of attention projection weights to build reusable freezing profiles that match or beat LoRA while updating ~1-3% of model parameters.
-
dgMARK: Decoding-Guided Watermarking for Diffusion Language Models
Steering the unmasking order of diffusion language models so that tokens at parity-matching positions get revealed first creates a detectable watermark with modest quality loss.
-
LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization
LeTS hybridizes process-level and outcome-level rewards for GRPO-based RAG training, improving accuracy and reducing redundant searches on multi-hop QA benchmarks.
-
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
PhantomCircuit traces knowledge overshadowing to attention circuits during training and prunes circuit edges to recover the overshadowed answer.
-
SHERPA: A Model-Driven Framework for Large Language Model Execution
A framework that executes LLM tasks through hierarchical state machines improves output quality in 12 of 15 comparisons, but the evaluation lacks error bars and includes test-set-informed design choices.
-
PAGE-RAG: Evidence-Grounded Adaptive Graph Retrieval for Long-Document Question Answering
PAGE-RAG combines an always-on textual retrieval floor with a graph 'skeleton', routes queries adaptively, and explicitly abstains when evidence is insufficient, yielding competitive accuracy with reliable refusal.
Discussion (0). Sign in to comment.