Pith. sign in

REVIEW 3 cited by

Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12145 v1 pith:T3YRJ3P2 submitted 2024-03-18 cs.CL

classification cs.CL
keywords questionsfalseassumptionschallengingdatasetsdetectionnaturallyoccurring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Sensitivity to false assumptions (or false premises) in information-seeking questions is critical for robust question-answering (QA) systems. Recent work has shown that false assumptions in naturally occurring questions pose challenges to current models, with low performance on both generative QA and simple detection tasks (Kim et al. 2023). However, the focus of existing work on naturally occurring questions leads to a gap in the analysis of model behavior on the long tail of the distribution of possible questions. To this end, we introduce Syn-(QA)$^2$, a set of two synthetically generated QA datasets: one generated using perturbed relations from Wikidata, and the other by perturbing HotpotQA (Yang et al. 2018). Our findings from evaluating a range of large language models are threefold: (1) false assumptions in QA are challenging, echoing the findings of prior work, (2) the binary detection task is challenging even compared to the difficulty of generative QA itself, possibly due to the linguistic structure of the problem, and (3) the detection task is more challenging with long-tail questions compared to naturally occurring questions, highlighting the utility of our synthetic datasets and generation method.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication

    cs.CL 2026-01 unverdicted novelty 7.0 of 10

    LLMs often fail to redirect health questions containing misconceptions, unlike clinicians, exposing safety gaps in patient-facing medical AI.

  2. Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions

    cs.CL 2025-06 conditional novelty 6.0 of 10

    In German political questions, three LLMs frequently accommodated false presuppositions, and correct answers to direct questions did not guarantee rejection of the false assumption.

  3. MultiHoax: A Dataset of Multi-hop False-Premise Questions

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new multi-hop false-premise benchmark shows that leading large language models detect embedded falsehoods in only a minority of cases, with the best model reaching about 23% on the full two-stage protocol.

Pith tools