Pith. sign in

AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs

1 Pith paper cite this work, alongside 1 external citations. Polarity classification is still indexing.

1 Pith paper citing it
1 external citations · Pith
abstract

Generating diverse and effective clarifying questions is crucial for improving query understanding and retrieval performance in open-domain conversational search (CS) systems. We propose AGENT-CQ (Automatic GENeration, and evaluaTion of Clarifying Questions), an end-to-end LLM-based framework addressing the challenges of scalability and adaptability faced by existing methods that rely on manual curation or template-based approaches. AGENT-CQ consists of two stages: a generation stage employing LLM prompting strategies to generate clarifying questions, and an evaluation stage (CrowdLLM) that simulates human crowdsourcing judgments using multiple LLM instances to assess generated questions and answers based on comprehensive quality metrics. Extensive experiments on the ClariQ dataset demonstrate CrowdLLM's effectiveness in evaluating question and answer quality. Human evaluation and CrowdLLM show that the AGENT-CQ - generation stage, consistently outperforms baselines in various aspects of question and answer quality. In retrieval-based evaluation, LLM-generated questions significantly enhance retrieval effectiveness for both BM25 and cross-encoder models compared to human-generated questions.

fields

cs.IR 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Characterizing Web Search in The Age of Generative AI

cs.IR · 2025-10-13 · conditional · novelty 6.0

AI search engines vary greatly in how much they rely on web pages versus internal model knowledge, and these differences shift which sources and concepts users see.

citing papers explorer

Showing 1 of 1 citing paper.

  • Characterizing Web Search in The Age of Generative AI cs.IR · 2025-10-13 · conditional · none · ref 36 · internal anchor

    AI search engines vary greatly in how much they rely on web pages versus internal model knowledge, and these differences shift which sources and concepts users see.