REVIEW 4 major objections 5 minor 59 references
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read When a loaded question presupposes a false political claim, GPT-4o, Mistral-7B-v03, and Llama-3-8B often confirm the falsehood instead of correcting it — even when their answers to direct factual questions show they know the truth.
desk verdict Loaded questions expose a genuine failure to reject false political presuppositions, but the knowledge measure needs validation and the analysis lacks significance tests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a paired-question design that separates knowledge from grounding. Each verified party stance yields a confirmatory direct question ('Is it true that F?'), a disconfirmatory direct question ('Is it true that not F?'), and a loaded question built with a factive verb such as 'find out' or 'resent' that presupposes the false claim. Six direct-question answers (three samples of each type) place each loaded question into a belief group — false belief, weak, moderate, or strong correct belief — while three annotated responses to the loaded question are scored 0 (accommodation), 1 (imprecise), or 2 (rejection) and summed into a grounding score from 0 to 6. The framing concepts are common ground and presupposition failure from linguistics: a false presupposition should be repaired by a grounding act, and silent accommodation is appropriate only when the speaker genuinely lacks the knowledge.
What would settle it
Run the same loaded questions with a single added instruction — 'you may state that the question rests on a false assumption' — and check whether rejection rates jump above 90 percent; if they do, the default accommodation is a policy or conversational norm rather than missing knowledge. Collecting human responses to the identical items would additionally show what rejection rate a cooperative speaker actually achieves.
Extended reading notes
Core claim
The paper's central claim is that instruction-tuned LLMs do not systematically reject misinformation, even when knowledge is present. Given a loaded question such as 'Did voters resent that the AfD is not in favour of permanent border controls?' — which presupposes a false party position — the models frequently accommodate the false presupposition by answering as if the question were well-founded, or issue vague, irrelevant responses, instead of performing the grounding act of pointing out that the presupposition fails. Direct-question accuracy is not a reliable predictor of rejection: only GPT-4o shifts markedly toward rejection when its belief is strong and correct (grounding score 6 in 52.69 percent of such cases), while Mistral and Llama show weak or inconsistent knowledge-dependence. The paper also finds a systematic agreement bias — confirmatory questions are answered correctly far more often than disconfirmatory ones — and party-dependent behavior, with GPT rejecting misinformation about the far-right AfD at high rates even when underlying knowledge is weak.
Load-bearing premise
The paper assumes that answering simple yes/no knowledge questions correctly means the model genuinely knows the fact, so that not rejecting a false presupposition in a loaded question counts as a grounding failure rather than some other conversational strategy.
Editorial extensions
If this is right
- A user who asks an LLM a question embedding a false political claim will often receive an answer that confirms the claim rather than corrects it, so the model can reinforce misinformation in ordinary conversation.
- Benchmarks that measure only direct factual question accuracy will miss this failure mode, since a model can score high on direct knowledge questions while systematically accommodating false presuppositions in loaded ones.
- Instruction-tuned dialogue systems intended for political education or election advice need grounding as an explicitly trained or instructed capability, not just factual knowledge, to avoid amplifying false beliefs.
- Evaluation suites should include negated and loaded question types alongside confirmatory ones, because the agreement bias means confirmatory accuracy overstates a model's true reliability.
Reading between the lines
- A natural extension the paper leaves untested is whether explicit grounding instructions ('if the question presupposes a false claim, say so') would raise rejection rates near ceiling; the authors deliberately prompted without such instructions, so capability versus default behavior remains open.
- Comparing the same false presupposition under different trigger types (factive verbs versus definite descriptions) would show whether the failure is tied to how the presupposition is packaged linguistically, an axis the companion study begins to explore.
- The absence of a human conversational baseline means the 'face-saving' interpretation is one reading among several; eliciting human responses to the same loaded items would show what rejection rate a cooperative human actually produces.
- A user-study follow-up measuring whether people's beliefs shift toward a false presupposition after the model accommodates it would quantify the misinformation-amplification risk the paper flags as its motivating concern.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether instruction-tuned LLMs engage in conversational grounding when asked loaded political questions that presuppose false party positions. Using verified stances from the 2024 German Wahl-O-Mat, the authors constructed confirmatory and disconfirmatory direct questions and loaded questions embedding false presuppositions, and evaluated GPT-4-o, Mistral-7B-v03, and Llama-3-8B on these items. Responses to loaded questions were manually annotated by seven raters, yielding a Fleiss kappa of 0.82, and the authors grouped items into belief groups based on the number of correct direct-question answers. The results show that all models frequently accommodate false presuppositions rather than reject them, with GPT rejecting more often when its direct-question accuracy is 6/6, while Mistral and Llama mostly accommodate or respond imprecisely. The paper further reports a systematic agreement bias in direct questions and a political bias in GPT's rejection behavior, particularly for the far-right AfD. The central conclusion is that the models do not systematically reject misinformation even when knowledge is present.
Significance. If the central claim holds, the paper makes a valuable contribution to understanding LLMs' pragmatic limitations in politically sensitive misinformation contexts. The public release of the FLEX Benchmark dataset, the careful construction of items from verified party positions, and the substantial annotation effort (Fleiss kappa 0.82) are clear strengths. The paper also usefully separates the notions of factual knowledge and conversational grounding, and its observation of a confirmatory bias in direct question answering is an important methodological caution for QA benchmarks. However, the conclusion rests critically on the operationalization of 'knowledge' as correct answers to direct polar questions, which appears to be confounded with a systematic agreement bias. The lack of inferential statistics and the absence of a true-presupposition control condition further limit the strength of the claims. The work is potentially important, but the central knowledge-grounding link is not yet cleanly established.
major comments (4)
- [§4.3] The operationalization of 'knowledge' via correct answers to direct polar questions is confounded by a systematic agreement bias. Table 3 shows a large confirmatory-disconfirmatory accuracy gap (GPT 89.5% vs. 63.6%; Mistral 80.1% vs. 43.1%), and Figure 3 reveals a strong preference for agreement. Under this measurement scheme, a model with a pure yes-bias and no factual knowledge would receive 3/6 correct answers and be assigned to the No/Weak Belief group. Consequently, the intermediate belief groups (WB and MB) do not cleanly represent increasing knowledge; they may instead reflect varying degrees of response bias. This is load-bearing because the central conclusion 'even when knowledge is present' depends on these groups. The authors should validate the knowledge measure independently, e.g., by testing with a separate set of forced-choice or fill-in-the-blank questions, or by restricting the 'knowledge present' conclusion to the Strong Correct Belief group (6/6 correct, which requires 'no' on disconfirmatory items).
- [§5] The evaluation of direct questions was performed manually by the two first authors, yet no inter-annotator reliability is reported for this judgment, and no confidence intervals or significance tests accompany the key quantitative comparisons. The entire results section (Tables 2 and 3, Figures 1 and 2) reports only raw percentages, making it impossible to assess whether the observed differences between belief groups or between models are statistically reliable. For instance, the claim that grounding scores increase with knowledge for Llama in Figure 2 is based on visual inspection of distributions. The authors should report measures of uncertainty (e.g., bootstrapped 95% confidence intervals) or perform appropriate significance tests (e.g., mixed-effects models with model and party as factors) to support the interpretive claims.
- [§7] The conclusion that 'the models do not systematically reject misinformation, even when knowledge is present' is not accurately supported by the data for GPT in the Strong Correct Belief group, where the modal grounding score is 6 (52.69%) and the paper itself states that 'Only GPT successfully rejected misinformation when equipped with strong and accurate beliefs.' The abstract and conclusion should be qualified to reflect this specificity: rejection rates remain substantially below the ideal even with high direct-question accuracy, and they are much lower in intermediate knowledge groups. As stated, the conclusion overgeneralizes beyond the actual results.
- [§5] Loaded questions are constructed only for false presuppositions, using factive verbs, and the paper acknowledges that responses to true presuppositions were collected but not analyzed. Without a true-presupposition control condition, the low rejection rates cannot be unambiguously attributed to the falsity of the presupposition or to an overall tendency to accommodate presuppositions in loaded questions. The face-saving interpretation in §5 ('Do LLMs save face?') requires such a contrast. Since the authors state that true-presupposition responses were already collected, analyzing them would strengthen the central claim and is feasible within the current scope.
minor comments (5)
- [§4.4] The sentence 'Consequently, all loaded questions embedding the claim ... weak belief.' appears twice in the text, and the placeholder 'Figure X' should be replaced with the actual figure number.
- [Appendix A.1] The example model responses reference 'the Greens' and 'Green Party', but the study only includes DIE LINKE, AfD, SPD, and CDU/CSU. This inconsistency suggests leftover content from an earlier version and should be corrected.
- [Throughout] The spelling of the Llama model is inconsistent ('LLaMA' vs. 'LLama' and 'Llama'); please unify to the official spelling.
- [Appendix] The annotation guidelines (Figures 4-6) are entirely in German, while the paper is in English. Providing an English translation or summary of the guidelines would make the annotation protocol accessible to international readers.
- [Figures 1 and 2] The figure captions do not fully describe the axes; please include clear labels such as 'Number of correct direct answers' and 'Grounding score' with units, and explicitly state that the percentages are relative frequencies within each belief group.
Circularity Check
No significant circularity: the paper's belief groups and grounding scores are independently measured, and the central claim is a transparent empirical operationalization rather than a derivation that reduces to its inputs.
full rationale
The paper's derivation chain is empirical rather than formal. Direct questions are used to measure factual knowledge, and loaded questions are used to measure grounding behavior; these are two separately constructed prompt sets with separately annotated responses. The belief groups in Section 4.4 are defined by the number of correct answers to six direct questions, while the grounding score is defined by annotations of three loaded-question responses. No equation or fitting procedure makes the loaded-question outcome a function of the direct-question score; the observed association between belief groups and grounding scores is an empirical finding, not a mathematical identity. The thresholds for belief groups (0-1, 2-3, 4-5, 6 correct) and grounding weights (0, 1, 2) are author-designed scoring conventions, but they are transparent and do not by construction force the conclusion that models fail to reject misinformation even when knowledge is present. The self-citations, including the concurrent FLEX Benchmark study by the same authors, are contextual references and are not load-bearing for the central result; no uniqueness theorem or unverified ansatz is imported from prior work to constrain the analysis. The paper's own Limitations section acknowledges validity caveats, such as the exclusion of true-presupposition loaded questions from analysis and the possibility that model knowledge cut-offs explain low direct-question accuracy; these are measurement-validity concerns, not evidence of circularity. The confirmatory versus disconfirmatory accuracy gap reported in Table 3 is an empirical result that supports the face-saving interpretation, but it is not a fitted input renamed as a prediction. In sum, the paper is self-contained against the Wahl-O-Mat ground truth and does not reduce any claimed result to its own inputs, so the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Belief group boundaries =
FB 0-1, WB 2-3, MB 4-5, SB 6
- Grounding score weights =
Accommodation 0, Imprecise 1, Rejection 2
- Sample count per prompt =
3
assumptions (5)
- domain assumption Wahl-O-Mat party positions from the 2024 European election constitute accurate ground truth for party stances.
- domain assumption Factive verbs such as 'find out', 'resent', and 'discover' reliably trigger presuppositions, and the negated claims embedded in loaded questions are false when the party holds the opposite stance.
- domain assumption The annotation categories Accommodated, Rejected, and Imprecise validly operationalize grounding behavior, with Rejection as the ideal outcome.
- domain assumption Accuracy on direct confirmatory and disconfirmatory questions approximates a model's knowledge of the fact.
- domain assumption Human annotation agreement (Fleiss kappa 0.82) is sufficient to treat the labels as reliable.
Cite this review
Pith. "Pith review of Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions." pith.science (2026). https://pith.science/paper/7OUWXME6
@misc{pith2026250608952,
author = {Pith},
title = {Pith review of: Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions},
year = {2026},
howpublished = {\url{https://pith.science/paper/7OUWXME6}},
note = {Machine review of arXiv:2506.08952}
}
read the original abstract
Communication among humans relies on conversational grounding, allowing interlocutors to reach mutual understanding even when they do not have perfect knowledge and must resolve discrepancies in each other's beliefs. This paper investigates how large language models (LLMs) manage common ground in cases where they (don't) possess knowledge, focusing on facts in the political domain where the risk of misinformation and grounding failure is high. We examine the ability of LLMs to answer direct knowledge questions and loaded questions that presuppose misinformation. We evaluate whether loaded questions lead LLMs to engage in active grounding and correct false user beliefs, in connection to their level of knowledge and their political bias. Our findings highlight significant challenges in LLMs' ability to engage in grounding and reject false user beliefs, raising concerns about their role in mitigating misinformation in political discourse.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
IU - S tudie zu D emokratie und B ildung
2024. IU - S tudie zu D emokratie und B ildung. https://www.iu.de/news/iu-studie-demokratie-und-bildung/. Accessed: 2024-9-17
work page 2024
-
[4]
Vevake Balaraman, Arash Eshghi, Ioannis Konstas, and Ioannis Papaioannou. 2023. https://doi.org/10.18653/v1/2023.sigdial-1.52 No that ' s not what I meant: Handling third position repair in conversational question answering . In Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 562--571, Prague, Czechia....
-
[5]
Yejin Bang, Delong Chen, Nayeon Lee, and Pascale Fung. 2024. https://aclanthology.org/2024.acl-long.600 Measuring P olitical B ias in L arge L anguage M odels: What is said and how it is said . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11142--11159, Bangkok, Thailand. Associat...
work page 2024
-
[6]
Beaver, Bart Geurts, and Kristie Denlinger
David I. Beaver, Bart Geurts, and Kristie Denlinger. 2024. Presupposition . In Edward N. Zalta and Uri Nodelman, editors, The Stanford Encyclopedia of Philosophy , F all 2024 edition. Metaphysics Research Lab, Stanford University
work page 2024
-
[7]
Emily M Bender and Alex Lascarides. 2019. Linguistic fundamentals for natural language processing II : 100 essentials from semantics and pragmatics. Synth. Lect. Hum. Lang. Technol., 12(3):1--268
work page 2019
-
[8]
Luciana Benotti and Patrick Blackburn. 2021. https://doi.org/10.18653/v1/2021.eacl-main.41 Grounding as a collaborative process . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 515--531, Online. Association for Computational Linguistics
Show all 59 references
-
[9]
Penelope Brown and Stephen C Levinson. 1987. Politeness: Some universals in language usage. 4. Cambridge university press
1987
-
[10]
Khyathi Raghavi Chandu, Yonatan Bisk, and Alan W Black. 2021. Grounding ‘grounding’ in NLP . In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4283--4305, Stroudsburg, PA, USA. Association for Computational Linguistics
2021
-
[11]
Herbert H. Clark. 1996. Using Language. “Using” Linguistic Books. Cambridge University Press
1996
-
[12]
Luigi Curini and Eugenio Pizzimenti. 2020. Searching for a unicorn: Fake news and electoral behaviour. Democracy and Fake News, pages 77--91
2020
-
[13]
Ashwin Daswani, Rohan Sawant, and Najoung Kim. 2024. https://arxiv.org/abs/2403.12145 Syn-qa2: Evaluating false assumptions in long-tail questions with synthetic qa datasets . Preprint, arXiv:2403.12145
2024 arXiv
-
[14]
Judith Degen and Judith Tonhauser. 2021. https://doi.org/10.1162/opmi_a_00042 Prior Beliefs Modulate Projection . Open Mind, 5:59--70
2021 doi
-
[15]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[16]
we’re running out of fuel!
Chi-Hé Elder and David Beaver. 2022. “we’re running out of fuel!”: When does miscommunication go unrepaired? Intercult. Pragmat., 19(5):541--570
2022
-
[17]
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. https://doi.org/10.18653/v1/2023.acl-long.656 From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models . In Proceedings of the 61st An...
2023 doi
-
[18]
Constanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau, and Anders S gaard. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.900 Defining knowledge: Bridging epistemology and large language models . In Proceedings of the 2024 Conference on Empirical Methods in Na...
2024 doi
-
[19]
Daniel Fried, Nicholas Tomlin, Jennifer Hu, Roma Patel, and Aida Nematzadeh. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.840 Pragmatics in language grounding: Phenomena, tasks, and modeling approaches . In Findings of the Association for Computational Linguistics: EM...
2023 doi
-
[20]
Suyash Fulay, William Brannon, Shrestha Mohanty, Cassandra Overney, Elinor Poole-Dayan, Deb Roy, and Jad Kabbara. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.508 On the relationship between truth and political bias in language models . In Proceedings of the 2024 Conferen...
2024 doi
-
[21]
Bart Geurts. 2024. Common Ground in Pragmatics . In Edward N. Zalta and Uri Nodelman, editors, The Stanford Encyclopedia of Philosophy , W inter 2024 edition. Metaphysics Research Lab, Stanford University
2024
-
[22]
Erving Goffman. 1955. On face-work: An analysis of ritual elements in social interaction. Psychiatry, 18(3):213--231
1955
-
[23]
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. The political ideology of conversational AI : Converging evidence on ChatGPT’s pro-environmental, left-libertarian orientation. SSRN Electron. J
2023
-
[24]
U ber nein. Zeitschrift f \
Wolfgang Imo. 2017. \"U ber nein. Zeitschrift f \"u r germanistische Linguistik , 45(1):40--72
2017
-
[25]
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams. 2020. https://doi.org/10.18653/v1/2020.acl-main.768 Are natural language inference models IMPPRESsive ? L earning IMPlicature and PRESupposition . In Proceedings of the 58th Annual Meeting of the Association f...
2020 doi
-
[26]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825
2023 arXiv
-
[27]
Nanjiang Jiang and Marie-Catherine de Marneffe. 2019. https://doi.org/10.18653/v1/P19-1412 Do you know that florence is packed with visitors? evaluating state-of-the-art models of speaker commitment . In Proceedings of the 57th Annual Meeting of the Association for Computation...
2019 doi
-
[28]
Lalitha Kameswari, Dama Sravani, and Radhika Mamidi. 2020. https://doi.org/10.18653/v1/2020.socialnlp-1.1 Enhancing bias detection in political news using pragmatic presupposition . In Proceedings of the Eighth International Workshop on Natural Language Processing for Social M...
2020 doi
-
[29]
Bowman, and Jackson Petty
Najoung Kim, Phu Mon Htut, Samuel R. Bowman, and Jackson Petty. 2023. https://doi.org/10.18653/v1/2023.acl-long.472 ( QA ) ^2 : Question answering with questionable assumptions . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume...
2023 doi
-
[30]
Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, and Deepak Ramachandran. 2021. https://doi.org/10.18653/v1/2021.acl-long.304 Which linguist invented the lightbulb? presupposition verification for question-answering . In Proceedings of the 59th Annual Meeting of the Association...
2021 doi
-
[31]
Clara Lachenmaier, Eleonore Lumer, Hendrik Buschmeier, and Sina Zarrie . 2024. Towards understanding the entanglement of human stereotypes and system biases in human-robot interaction. In Companion of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, pages...
2024
-
[32]
Staffan Larsson. 2018. Grounding as a side-effect of grounding. Top. Cogn. Sci., 10(2):389--408
2018
-
[33]
Seung-Hee Lee. 2016. Information and affiliation: Disconfirming responses to polar questions and what follows in third position. Journal of Pragmatics, 100:59--72
2016
-
[34]
Levinson
Stephen C. Levinson. 1983. Pragmatics. Cambridge Textbooks in Linguistics. Cambridge University Press
1983
-
[35]
Eleonore Lumer, Clara Lachenmaier, Sina Zarrie , and Hendrik Buschmeier. 2023. Indirect politeness of disconfirming answers to humans and robots. In 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), pages 1808--1815. IEEE
2023
-
[36]
Meta AI . 2024. Introducing Meta Llama 3: The most capable openly available LLM to date . https://ai.meta.com/blog/meta-llama-3/. Accessed: 2025-05-30
2024
-
[37]
Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau
Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022. https://doi.org/10.1162/tacl_a_00494 Reducing conversational agents ' overconfidence through linguistic calibration . Transactions of the Association for Computational Linguistics, 10:857--872
2022 doi
-
[38]
Mistral AI . 2024. Mistral-7B-Instruct-v0.3 . https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3. Accessed: 2025-05-30
2024
-
[39]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786
2022 arXiv
-
[40]
Jan Nehring, Aleksandra Gabryszak, Pascal J \"u rgens, Aljoscha Burchardt, Stefan Schaffer, Matthias Spielkamp, and Birgit Stark. 2024. https://aclanthology.org/2024.lrec-main.884/ Large language models are echo chambers . In Proceedings of the 2024 Joint International Confere...
2024
-
[41]
OpenAI . 2024. Hello GPT-4o . https://openai.com/index/hello-gpt-4o/. Accessed: 2025-05-30
2024
-
[42]
Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, D...
2023
-
[43]
Ildiko Pilan, Laurent Pr \'e vot, Hendrik Buschmeier, and Pierre Lison. 2024. https://doi.org/10.18653/v1/2024.sigdial-1.38 Conversational feedback in scripted versus spontaneous dialogues: A comparative analysis . In Proceedings of the 25th Annual Meeting of the Special Inter...
2024 doi
-
[44]
Laura Ruis, Akbir Khan, Stella Biderman, Sara Hooker, Tim Rockt\" a schel, and Edward Grefenstette. 2023. The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by LLMs . In Proceedings of the 37th International Conference on Neural ...
2023
-
[45]
Michel Schimpf, Anton Wyrowski, Roman Mayr, Sebastian Maier, and Robin Frasch. 2024. https://wahl.chat/ Wahl.chat – dein K I ‚- A ssistent für eine informierte wahl . Accessed: February 7, 2025
2024
-
[46]
Omar Shaikh, Kristina Gligoric, Ashna Khetan, Matthias Gerstgrasser, Diyi Yang, and Dan Jurafsky. 2024. https://aclanthology.org/2024.naacl-long.348 Grounding gaps in language model generations . In Proceedings of the 2024 Conference of the North American Chapter of the Associ...
2024
-
[47]
Judith Sieker, Oliver Bott, Torgrim Solstad, and Sina Zarrie . 2023. https://doi.org/10.18653/v1/2023.inlg-main.15 Beyond the bias: Unveiling the quality of implicit causality prompt continuations in language models . In Proceedings of the 16th International Natural Language G...
2023 doi
-
[48]
Judith Sieker, Simeon Junker, Ronja Utescher, Nazia Attari, Heiko Wersing, Hendrik Buschmeier, and Sina Zarrie . 2024. https://doi.org/10.18653/v1/2024.emnlp-main.1084 The illusion of competence: Evaluating the effect of explanations on users' mental models of visual question ...
2024 doi
-
[49]
Judith Sieker, Clara Lachenmaier, and Sina Zarrieß. 2025. https://arxiv.org/abs/2505.22354 LLMs struggle to reject false presuppositions when misinformation stakes are high . Preprint, arXiv:2505.22354. To appear in: Proceedings of CogSci 2025
2025
-
[50]
Judith Sieker and Sina Zarrie . 2023. https://doi.org/10.18653/v1/2023.blackboxnlp-1.14 When your language model cannot E ven do determiners right: Probing for anti-presuppositions and the maximize presupposition! principle . In Proceedings of the 6th BlackboxNLP Workshop: Ana...
2023 doi
-
[51]
Neha Srikanth, Rupak Sarkar, Heran Mane, Elizabeth Aparicio, Quynh Nguyen, Rachel Rudinger, and Jordan Boyd-Graber. 2024. https://aclanthology.org/2024.naacl-long.403 Pregnant questions: The importance of pragmatic awareness in maternal health question answering . In Proceedin...
2024
-
[52]
Robert Stalnaker. 1973. https://doi.org/10.1007/bf00262951 Presuppositions . Journal of Philosophical Logic, 2(4):447--457
1973 doi
-
[53]
Polina Tsvilodub, Michael Franke, Robert Hawkins, and Noah Goodman. 2023. https://escholarship.org/uc/item/1mb6p7gn Overinformative question answering by humans and machines . In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 45
2023
-
[54]
Kai von Fintel. 2008. https://doi.org/10.1111/j.1520-8583.2008.00144.x What is presupposition accommodation, again? Philosophical Perspectives, 22(1):137--170
2008
-
[55]
Richard von Maydell. 2024. Parteienpositionen: Ann \"a herung in der mitte und zunehmende distanz zur afd. Wirtschaftsdienst, 104(12):861--866
2024
-
[56]
Alice Xia, Roxana M Barbu, Kathleen Van Benthem, Daniel A Di Giovanni, Ida Toivonen, and Raj Singh. 2019. Detecting presupposition failure and accommodation with EEG . Proceedings of the Annual Meeting of the Cognitive Science Society, 41(0)
2019
-
[57]
Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.486 Knowledge conflicts for LLM s: A survey . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, page...
2024 doi
-
[58]
S Yablo. 2006. Non-catastrophic presupposition failure
2006
-
[59]
Xinyan Yu, Sewon Min, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.583 CREPE : Open-domain question answering with false presuppositions . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (...
2023 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.