REVIEW 3 cited by
(QA)$^2$: Question Answering with Questionable Assumptions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Naturally occurring information-seeking questions often contain questionable assumptions -- assumptions that are false or unverifiable. Questions containing questionable assumptions are challenging because they require a distinct answer strategy that deviates from typical answers for information-seeking questions. For instance, the question "When did Marie Curie discover Uranium?" cannot be answered as a typical "when" question without addressing the false assumption "Marie Curie discovered Uranium". In this work, we propose (QA)$^2$ (Question Answering with Questionable Assumptions), an open-domain evaluation dataset consisting of naturally occurring search engine queries that may or may not contain questionable assumptions. To be successful on (QA)$^2$, systems must be able to detect questionable assumptions and also be able to produce adequate responses for both typical information-seeking questions and ones with questionable assumptions. Through human rater acceptability on end-to-end QA with (QA)$^2$, we find that current models do struggle with handling questionable assumptions, leaving substantial headroom for progress.
Forward citations
Cited by 3 Pith papers
-
HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning
A new benchmark shows search-augmented LLMs correctly refuse only up to 42.9% of unanswerable multi-hop questions, with failures split between hallucinated answers and search exhaustion.
-
Towards LLM Agents for Earth Observation
On a new 140-question Earth observation benchmark, the best LLM agent scores 33% accuracy with Google Earth Engine access because generated code fails to run over 58% of the time.
-
Unanswerability Evaluation for Retrieval Augmented Generation
UAEval4RAG synthesizes six categories of unanswerable queries from any knowledge base and evaluates whether RAG systems reject them acceptably.
Discussion (0). Continue with ORCID to comment.