Pith. sign in

REVIEW 8 cited by

CLAM: Selective Clarification for Ambiguous Questions with Generative Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.07769 v2 pith:K7GUDGLN submitted 2022-12-15 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords ambiguouslanguagemodelsclarificationquestionsclamusersquestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Users often ask dialogue systems ambiguous questions that require clarification. We show that current language models rarely ask users to clarify ambiguous questions and instead provide incorrect answers. To address this, we introduce CLAM: a framework for getting language models to selectively ask for clarification about ambiguous user questions. In particular, we show that we can prompt language models to detect whether a given question is ambiguous, generate an appropriate clarifying question to ask the user, and give a final answer after receiving clarification. We also show that we can simulate users by providing language models with privileged information. This lets us automatically evaluate multi-turn clarification dialogues. Finally, CLAM significantly improves language models' accuracy on mixed ambiguous and unambiguous questions relative to SotA.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Twelve coding LLMs resolve injected user-specific ambiguity more often on the first turn when given same-user session history (average FT-ES +15.6 pp), though shuffled history explains part of the benefit.

  2. The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Adding a structured list of six unknowable aspects of a user's life to an LLM prompt reduces sycophantic, harmful, and hallucinated advice in synthetic tests across five model families.

  3. Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?

    cs.CL 2025-08 conditional novelty 6.0 of 10

    An auxiliary LLM can diagnose whether an LLM's uncertainty comes from ambiguous questions or missing knowledge by analyzing patterns of disagreement among multiple sampled answers.

  4. Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A training method using reinforcement learning and answerability heuristics lets small language models actively ask for missing math details and then solve problems, raising accuracy on the new GSM-MC benchmark from 0...

  5. SG-CoT: An Ambiguity-Aware Robotic Planning Framework using Scene Graph Representations

    cs.RO 2026-03 reject novelty 5.0 of 10

    SG-CoT grounds an LLM planner's chain-of-thought in a scene graph via iterative API queries, improving ambiguity detection and clarification in simulated manipulation, though its success metric credits any clarifying ...

  6. Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Per the abstract, large reasoning models systematically fail to ask for missing information on under-specified math problems, a skill standard benchmarks never test.

  7. Demystifying Feature Requests: Leveraging LLMs to Refine Feature Requests in Open-Source Software

    cs.SE 2025-07 conditional novelty 5.0 of 10

    GPT-4o with in-context learning can flag ambiguity and incompleteness in GitHub feature requests and draft clarification questions, though moderate annotator agreement and a small sample limit the strength of the evidence.

  8. Referential ambiguity and clarification requests: comparing human and LLM behaviour

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Humans seldom ask clarification questions for referential ambiguity, while LLMs ask them more often, and reasoning prompts increase LLM question frequency and relevance.

Pith tools