Pith. sign in

REVIEW 13 cited by

AmbigQA: Answering Ambiguous Open-domain Questions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.10645 v2 pith:KT2QZG7C submitted 2020-04-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords ambigqaopen-domainquestionsambiguityansweringnq-openquestiontask
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Ambiguity is inherent to open-domain question answering; especially when exploring new topics, it can be difficult to ask questions that have a single, unambiguous answer. In this paper, we introduce AmbigQA, a new open-domain question answering task which involves finding every plausible answer, and then rewriting the question for each one to resolve the ambiguity. To study this task, we construct AmbigNQ, a dataset covering 14,042 questions from NQ-open, an existing open-domain QA benchmark. We find that over half of the questions in NQ-open are ambiguous, with diverse sources of ambiguity such as event and entity references. We also present strong baseline models for AmbigQA which we show benefit from weakly supervised learning that incorporates NQ-open, strongly suggesting our new task and data will support significant future research effort. Our data and baselines are available at https://nlp.cs.washington.edu/ambigqa.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning

    cs.CL 2026-04 conditional novelty 6.0 of 10

    RRPO formulates document reranking as a sequential MDP and optimizes a pointwise reranker with PPO using LLM generation rewards and a reference-anchored deterministic baseline.

  2. Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

    cs.CL 2025-07 conditional novelty 6.0 of 10

    HumorBench scores LLM explanations of cartoon jokes against expert-written objective elements and finds reasoning skills transfer from STEM benchmarks, while extra thinking tokens help only some models.

  3. PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    PRGB introduces a placeholder-based, fine-grained RAG benchmark that evaluates LLMs on filtering, combination, and multi-hop reasoning, with English and Chinese datasets.

  4. Direct Retrieval-augmented Optimization: Synergizing Knowledge Selection and Language Models

    cs.IR 2025-05 conditional novelty 6.0 of 10

    DRO jointly trains a generative document selector and an LLM generator by treating document order as a latent variable and using importance-sampled expectation-maximization, beating prior RAG systems on five benchmarks.

  5. MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents

    cs.CL 2025-02 conditional novelty 6.0 of 10

    MTPChat adds explicit date stamps and synthetic earlier responses to multimodal persona dialogues, defines two temporal retrieval tasks, and reports modest gains from a gated fusion module.

  6. Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation

    cs.IR 2025-01 conditional novelty 6.0 of 10

    K-RagRec improves LLM-based recommendation by retrieving and encoding knowledge graph subgraphs as soft prompts, outperforming existing retrieval-augmented LLM recommenders on three datasets.

  7. Acknowledging Focus Ambiguity in Visual Questions

    cs.CV 2025-01 conditional novelty 6.0 of 10

    VQ-FocusAmbiguity is a 5,500-example dataset annotating all plausible focus regions for visual questions, and state-of-the-art models perform poorly at recognizing and locating focus ambiguity.

  8. Unanswerability Evaluation for Retrieval Augmented Generation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    UAEval4RAG synthesizes six categories of unanswerable queries from any knowledge base and evaluates whether RAG systems reject them acceptably.

  9. Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Per the abstract, large reasoning models systematically fail to ask for missing information on under-specified math problems, a skill standard benchmarks never test.

  10. LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries

    cs.IR 2025-05 conditional novelty 5.0 of 10

    LLM-based query expansion can hurt retrieval when the LLM lacks knowledge of the query or the query is highly ambiguous.

  11. ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation

    cs.AI 2025-06 reject novelty 4.0 of 10

    ChemAU adds a position penalty to token-level uncertainty estimates so that flagged reasoning steps are corrected by a fine-tuned chemistry model, reporting improved accuracy on GPQA, MMLU-Pro, and SuperGPQA chemistry...

  12. Multiple Abstraction Level Retrieve Augment Generation

    cs.CL 2025-01 conditional novelty 4.0 of 10

    MAL-RAG retrieves document, section, paragraph, and multi-sentence chunks together and claims a 25.7% improvement in AI-judged answer correctness on glycoscience questions over single-level RAG.

  13. A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.

Pith tools