Pith. sign in

REVIEW 7 cited by

A Survey of Large Language Model Agents for Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.19213 v1 pith:FCOWCQG2 submitted 2025-03-24 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords agentsquestionansweringchallengesenvironmentslanguagelargemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper surveys the development of large language model (LLM)-based agents for question answering (QA). Traditional agents face significant limitations, including substantial data requirements and difficulty in generalizing to new environments. LLM-based agents address these challenges by leveraging LLMs as their core reasoning engine. These agents achieve superior QA results compared to traditional QA pipelines and naive LLM QA systems by enabling interaction with external environments. We systematically review the design of LLM agents in the context of QA tasks, organizing our discussion across key stages: planning, question understanding, information retrieval, and answer generation. Additionally, this paper identifies ongoing challenges and explores future research directions to enhance the performance of LLM agent QA systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution

    cs.MA 2026-08 conditional novelty 6.0 of 10

    An eight-agent question-asking system that front-loads intent clarification produced more complete prompts, higher-rated outputs, and single-turn task completion in a four-person pilot, with unstable effect sizes.

  2. TopoTuner: Topological Finetuning of Large Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    TopoTuner uses Wasserstein distances between persistence diagrams of attention projection weights to build reusable freezing profiles that match or beat LoRA while updating ~1-3% of model parameters.

  3. dgMARK: Decoding-Guided Watermarking for Diffusion Language Models

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Steering the unmasking order of diffusion language models so that tokens at parity-matching positions get revealed first creates a detectable watermark with modest quality loss.

  4. LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LeTS hybridizes process-level and outcome-level rewards for GRPO-based RAG training, improving accuracy and reducing redundant searches on multi-hop QA benchmarks.

  5. Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PhantomCircuit traces knowledge overshadowing to attention circuits during training and prunes circuit edges to recover the overshadowed answer.

  6. SHERPA: A Model-Driven Framework for Large Language Model Execution

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A framework that executes LLM tasks through hierarchical state machines improves output quality in 12 of 15 comparisons, but the evaluation lacks error bars and includes test-set-informed design choices.

  7. PAGE-RAG: Evidence-Grounded Adaptive Graph Retrieval for Long-Document Question Answering

    cs.IR 2026-07 conditional novelty 4.0 of 10

    PAGE-RAG combines an always-on textual retrieval floor with a graph 'skeleton', routes queries adaptively, and explicitly abstains when evidence is insufficient, yielding competitive accuracy with reliable refusal.

Pith tools