REVIEW 2 major objections 6 minor 2 cited by
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
T0 review · 2 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Explanations increase reliance on LLM answers—correct or not—while sources and contradictions curb overreliance on wrong ones.
desk verdict Strong pre-registered study with a credible sources effect; the inconsistency sub-claim is real but currently confounded by question identity. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing apparatus is a 2 × 2 × 2 within-subjects experiment using 12 difficult binary factual questions, where each of 308 participants saw eight response types from a hypothetical LLM named Theta: {correct, incorrect} × {no explanation, explanation} × {no sources, clickable sources}. Explanations—supporting details that justify the answer—and answers were generated in advance with ChatGPT and Perplexity AI so that content could be controlled, and reliance was measured behaviorally as whether the participant's final answer agreed with Theta's answer, complemented by self-reported confidence, justification-quality and actionability ratings, source-clicking, and follow-up questions. In addition, the authors coded naturally occurring inconsistencies in the explanations (sets of statements that cannot both be true) and ran a pre-registered ANOVA comparing incorrect answers with no explanation, consistent explanation, and inconsistent explanation; a 16-person think-aloud study supplied the qualitative account of how users notice these cues.
What would settle it
Run the same 12 questions with matched explanation pairs that are identical except for one internal contradiction and randomly assign participants to versions; if agreement with incorrect answers does not drop when the contradiction is present, the paper's inconsistency claim is falsified.
Extended reading notes
Core claim
Users agree with an LLM's answer more often when the answer comes with an explanation, regardless of whether the answer is actually right; the paper shows this in a controlled setting where the same difficult questions are paired with correct or incorrect answers, with or without explanations and sources. Sources change the pattern: they raise agreement when the answer is correct and lower it when the answer is wrong, and they increase time on task and the odds of overriding an incorrect answer. Explanations that contain a logical inconsistency—for instance, an answer that contradicts the numbers cited to support it—produce significantly less agreement and higher accuracy than consistent explanations when the answer is wrong. The paper therefore concludes that explanation is not a single good thing: its effect depends on whether it invites verification (sources) and whether it contains visible cracks (inconsistencies).
Load-bearing premise
The inconsistency finding rests on comparing a small set of questions whose explanations happened to contain contradictions against the other questions' explanations, so the result holds only if the contradiction itself—not the content or difficulty of those three questions—is what changed participants' agreement and accuracy.
Editorial extensions
If this is right
- Explanations alone make users more likely to accept an answer, so adding them without other safeguards can actively increase overreliance on wrong LLM outputs.
- Providing clickable, accurate sources is a concrete corrective: in the no-explanation condition it raised agreement with correct answers from 67.2% to 73.4% and lowered agreement with incorrect answers from 78.2% to 68.2%.
- When the LLM answer is wrong, sources without explanation produce the highest user accuracy (31.8%), while explanation alone gives the lowest (17.2%); when the answer is right, explanation plus sources gives the highest accuracy (79.9%).
- Inconsistent explanations cut overreliance: agreement with wrong answers fell from 83.3% to 69.7% and accuracy rose from 16.7% to 30.3% compared with consistent explanations.
- Because the beneficial source effect was obtained with real, mostly accurate links, the authors expect that fake, broken, or irrelevant sources would not help and could even increase perceived credibility.
Reading between the lines
- Beyond the paper: if the causal story is that inconsistency triggers deeper scrutiny, then automatically detecting and highlighting contradictions (for example, by checking whether the answer matches the numbers in the explanation) should reproduce the effect without requiring users to spot the flaw themselves.
- Beyond the paper: the paper's observational comparison suggests a testable design rule—LLM answers whose supporting explanation is internally consistent should be treated as more reliable for answer selection, since consistency of the explanation is itself predictive of correctness in the paper's stimulus set.
- Beyond the paper: the source effect may depend on individual differences: users who click links (roughly 119 of 308 participants clicked in at least one task, while 189 never clicked) may be the main beneficiaries, so an interface that actively nudges source-checking could widen the benefit to the majority who do not click.
- Beyond the paper: the experiment used deliberately hard questions where lay users have little prior knowledge, so the findings may not transfer to domains where users can independently evaluate the answer; testing with familiar topics would show whether sources still carry the same corrective force.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates how features of LLM responses shape users' reliance when answering objective questions. Study 1 is a think-aloud study (N=16) that identifies explanations, inconsistencies, and sources as key features. Study 2 is a pre-registered, within-subjects experiment (N=308) with a 2x2x2 design (answer correctness, presence of explanation, presence of clickable sources) using realistic LLM-generated responses from ChatGPT and Perplexity AI. The main mixed-effects analyses show that explanations increase agreement with both correct and incorrect answers, while sources increase appropriate reliance on correct answers and reduce overreliance on incorrect answers. A secondary observational analysis suggests that inconsistent explanations are associated with reduced overreliance on incorrect answers. The authors discuss implications for designing LLM interfaces to foster appropriate reliance.
Significance. The randomized manipulation of explanations and sources, combined with mixed-effects models that include participant and question random effects, makes the main 2x2x2 findings credible and directly relevant to the HCI community. The pre-registration, power analysis, and use of realistic LLM-generated stimuli are notable strengths that increase confidence in the explanation and source effects. If the inconsistency finding were causally supported, the design recommendation to highlight inconsistencies would be of considerable practical value; however, as it stands, that sub-claim is not yet established because the analysis is observational and confounded.
major comments (2)
- [§4.3.1, Figure 5] The comparison of consistent vs. inconsistent explanations is confounded by question identity. Only 3 of the 12 task questions produced naturally occurring inconsistent explanations, all for incorrect answers, and these three questions may differ systematically in difficulty, answer plausibility, or specific content (e.g., the arithmetic error in the Brazil population item). The ANOVA does not include participant or question random effects, and it collapses across the sources manipulation, so the reported differences (agreement 69.7% vs. 83.3%; accuracy 30.3% vs. 16.7%) cannot be attributed to inconsistency per se. This confound is load-bearing because the abstract presents the inconsistency result as a finding and §5.1 uses it to recommend interventions that highlight inconsistencies.
- [Abstract and §5.1] The causal claim that inconsistencies reduce overreliance, and the design recommendation to highlight inconsistencies as an intervention, are stronger than the evidence supports. The independent variable was not manipulated; it was observed post hoc after the experiment. Even a mixed-effects reanalysis would address non-independence but would not address selection on question content or the small number of inconsistent items. The authors should either reframe the inconsistency result as an exploratory, hypothesis-generating finding or conduct a follow-up experiment that factorially manipulates inconsistency while holding content constant.
minor comments (6)
- [§4.2.2] In the sentence reporting the confidence effect, β = .96, SE = .10, p < .001, the formatting of the p-value is inconsistent with the rest of the paper; please standardize the formatting of regression results throughout.
- [§4.3.1] The text does not state the cell size for the inconsistent explanation condition (N=155); adding this number would help readers interpret the precision of the estimates.
- [Figure 5] The y-axis labels are not fully specified in the figure caption; please indicate the response scale or unit for each panel to improve readability.
- [§5.3] The limitations section does not mention the confound in the inconsistency analysis; given the prominence of this finding, it should be explicitly acknowledged as an observational analysis with limited internal validity.
- [§4.1.4] The coding of inconsistencies was performed by the authors without reporting inter-rater reliability; since this variable is central to the inconsistency sub-claim, a second coder or a reliability statistic would strengthen confidence in the coding.
- [Appendix A] There is a typo in the first sentence: the the relationship should be the relationship.
Circularity Check
No significant circularity: the paper's central claims are tested in a pre-registered randomized experiment, and the observational inconsistency analysis raises internal-validity concerns but does not reduce to its own inputs.
full rationale
The paper's derivation chain is empirical rather than definitional. Study 1 (think-aloud) is used only to generate hypotheses about explanations, inconsistencies, and sources; Study 2 then tests those hypotheses in a pre-registered 2x2x2 within-subjects experiment in which explanation presence and source presence are randomly manipulated. The main reliance and accuracy findings are estimated from mixed-effects models with participant and question random effects, so they are not forced by construction. The inconsistency analysis in Section 4.3.1 is observational: the authors explicitly state that 'the presence of inconsistencies is not something we control for or manipulate,' and they compare naturally occurring inconsistent explanations against consistent explanations. This creates a potential confound with question identity and content, which is a validity limitation that the paper partially acknowledges, but it is not circular reasoning because the inconsistency variable is defined independently of the outcome (as 'sets of statements that cannot be true at the same time') and the observed effect could plausibly have gone in either direction. The paper's self-citations (e.g., prior work on uncertainty expression and interpretability) are used for methodological precedent and literature positioning, not as load-bearing proof of the current findings. No fitted parameter is renamed as a prediction, and no result is equivalent to its inputs by definition.
Assumptions & free parameters
assumptions (4)
- domain assumption Participants' behavior in a single-response controlled experiment reflects how users rely on LLMs in multi-turn interactions.
- domain assumption The sources used in Study 2 are real, relevant, and accurate, so the effect of sources may depend on source quality.
- standard math Mixed-effects regression assumptions, such as linearity and normality of residuals, hold for the fitted models.
- domain assumption The coding of inconsistencies by the first author is accurate and reproducible.
Cite this review
Pith. "Pith review of Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies." pith.science (2026). https://pith.science/paper/YDA3ZISB
@misc{pith2026250208554,
author = {Pith},
title = {Pith review of: Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies},
year = {2026},
howpublished = {\url{https://pith.science/paper/YDA3ZISB}},
note = {Machine review of arXiv:2502.08554}
}
read the original abstract
Large language models (LLMs) can produce erroneous responses that sound fluent and convincing, raising the risk that users will rely on these responses as if they were correct. Mitigating such overreliance is a key challenge. Through a think-aloud study in which participants use an LLM-infused application to answer objective questions, we identify several features of LLM responses that shape users' reliance: explanations (supporting details for answers), inconsistencies in explanations, and sources. Through a large-scale, pre-registered, controlled experiment (N=308), we isolate and study the effects of these features on users' reliance, accuracy, and other measures. We find that the presence of explanations increases reliance on both correct and incorrect responses. However, we observe less reliance on incorrect responses when sources are provided or when explanations exhibit inconsistencies. We discuss the implications of these findings for fostering appropriate reliance on LLMs.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 2 Pith papers
-
Humans overrely on overconfident language models, across languages
LLMs produce overconfident-sounding answers in all five tested languages, and bilingual users show the highest overreliance risk in Japanese despite its frequent hedges.
-
ContextBuddy: AI-Enhanced Contextual Insights for Security Alert Investigation (Applied to Intrusion Detection)
ContextBuddy trains an imitation-learning assistant on RL-simulated analysts' context requests and shows its suggestions improve alert classification accuracy and speed in simulation and a small non-expert user study.
Reference graph
Works this paper leans on
-
[1]
Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju. 2024. Faithful- ness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models. arXiv:cs.CL/2402.04614 https://arxiv.org/abs/2402.04614
arXiv 2024
-
[2]
Hussam Alkaissi and Samy I. McFarlane. 2023. Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus 15, 2 (2023). https://doi. org/10.7759/cureus.35179
-
[3]
Ravinithesh Annapureddy, Alessandro Fornaroli, and Daniel Gatica-Perez. 2024. Generative AI Literacy: Twelve Defining Competencies. Digit. Gov.: Res. Pract. (aug 2024). https://doi.org/10.1145/3685680 Just Accepted
doi:10.1145/3685680 2024
-
[4]
Sara Aronowitz and Tania Lombrozo. 2020. Experiential Explanation. Topics in Cognitive Science 12, 4 (2020), 1321–1336. https://doi.org/10.1111/tops.12445 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/tops.12445
-
[5]
Pepa Atanasova, Oana-Maria Camburu, Christina Lioma, Thomas Lukasiewicz, Jakob Grue Simonsen, and Isabelle Augenstein. 2023. Faithfulness Tests for Natural Language Explanations. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . Association for Computational Linguistics, Toronto, Canada, ...
2023
-
[6]
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the Whole Exceed Its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Computing Machinery, New Yor...
arXiv 2021
-
[7]
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Ben- netot, Siham Tabik, Alberto Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera. 2020. Ex- plainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Informatio...
-
[8]
Christos Bechlivanidis, David A Lagnado, Jeffrey C Zemla, and Steven Sloman
Show all 147 references
-
[9]
Boren and J
T. Boren and J. Ramey. 2000. Thinking Aloud: Reconciling Theory and Practice. IEEE Transactions on Professional Communication 43, 3 (2000), 261–278. https: //doi.org/10.1109/47.867942
2000 doi
-
[10]
Richard E Boyatzis. 1998. Transforming Qualitative Information: Thematic Anal- ysis and Code Development . sage
1998
-
[11]
Gelman Brandy N
Susan A. Gelman Brandy N. Frazier and Henry M. Wellman. 2016. Young Children Prefer and Remember Satisfying Explanations. Journal of Cognition and Development 17, 5 (2016), 718–736. https://doi.org/10.1080/15248372.2015. 1098649 PMID: 28713222
2016
-
[12]
Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psy- chology. Qualitative Research in Psychology 3, 2 (2006), 77–101. https: //doi.org/10.1191/1478088706qp063oa
2006 doi
-
[13]
Sylvain Bromberger. 1966. Why-Questions. In Readings in the Philosophy of Science, Baruch A. Brody (Ed.). Prentice Hall, Inc., Englewood Cliffs, 66–84
1966
-
[14]
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 188 (apr 2021), 21 pages. https://doi.org/10.1145/3449287
2021 doi
-
[15]
Ben Buchanan, Andrew Lohn, Micah Musser, and Katerina Sedova. 2021. Truth, Lies, and Automation: How Language Models Could Change Disinformation . Re- port. Center for Security and Emerging Technology. https://doi.org/10.51593/ 2021CA003
2021
-
[16]
Zana Buçinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman. 2020. Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems. In Proceedings of the 25th International Conference on Intelligent User Interfaces. 454–464
2020
-
[17]
Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan. 2015. The Role of Explanations on Trust and Reliance in Clinical Decision Support Systems. In Proceedings of the 2015 International Conference on Healthcare Informatics (ICHI ’15). IEEE Computer Society, USA, 160–169. https...
2015 doi
-
[18]
Carrie J Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry. 2021. Onboarding Materials as Cross-functional Boundary Objects for Developing AI Assistants. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems (CHI EA ’21) . A...
2021
-
[19]
Shiye Cao and Chien-Ming Huang. 2022. Understanding User Reliance on AI in Assisted Decision-Making. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 471 (nov 2022), 23 pages. https://doi.org/10.1145/3555572
2022 doi
-
[20]
Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal
Valerie Chen, Q. Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal. 2023. Understanding the Role of Human Intuition on Reliance in Human-AI Decision- Making with Explanations. Proc. ACM Hum.-Comput. Interact. 7, CSCW2, Article 370 (oct 2023). https://doi.org/10.1145/3610219
2023 doi
-
[21]
Cheng-Han Chiang and Hung-yi Lee. 2024. Over-Reasoning and Redundant Calculation of Large Language Models. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers) , Yvette Graham and Matthew Purver...
2024
-
[22]
Leah Chong, Guanglu Zhang, Kosa Goucher-Lambert, Kenneth Kotovsky, and Jonathan Cagan. 2022. Human Confidence in Artificial Intelligence and in Themselves: The Evolution and Impact of Confidence on Adoption of AI Advice. Computers in Human Behavior 127 (2022), 107018. https://...
2022
-
[23]
Christiano, Jan Leike, Tom B
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep Reinforcement Learning from Human Preferences. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17). Curran Associates Inc., R...
2017
-
[24]
Olanubi, Joseph M
Michelle Cohn, Mahima Pushkarna, Gbolahan O. Olanubi, Joseph M. Moran, Daniel Padgett, Zion Mengesha, and Courtney Heldreth. 2024. Believing Anthro- pomorphism: Examining the Role of Anthropomorphic Cues on Trust in Large Language Models. In Extended Abstracts of the CHI Confe...
2024
-
[25]
Collins, Albert Q
Katherine M. Collins, Albert Q. Jiang, Simon Frieder, Lionel Wong, Miri Zilka, Umang Bhatt, Thomas Lukasiewicz, Yuhuai Wu, Joshua B. Tenenbaum, William Hart, Timothy Gowers, Wenda Li, Adrian Weller, and Mateja Jamnik. 2024. Evaluating Language Models for Mathematics through In...
2024 doi
-
[26]
Francisco Cruz and Tania Lombrozo. 2024. The Effect of Jargon on Perceptions of Explanation Quality: Reconciling Contradictory Findings. In Proceedings of the Annual Meeting of the Cognitive Science Society , Vol. 46. https://escholarship. org/uc/item/4ds9s5tj
2024
-
[27]
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho. 2024. Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models.Journal of Legal Analysis 16, 1 (06 2024), 64–93. https://doi.org/10.1093/jla/laae003
2024 doi
-
[28]
Rafferty, and Christopher D
Marie-Catherine de Marneffe, Anna N. Rafferty, and Christopher D. Manning
-
[29]
Igor Douven and Patricia Mirabile. 2018. Best, Second-Best, and Good-Enough Explanations: How They Matter to Reasoning. Journal of Experimental Psy- chology: Learning, Memory, and Cognition 44, 11 (2018), 1792–1813. https: //doi.org/10.1037/xlm0000545
2018 doi
-
[31]
Malin Eiband, Daniel Buschek, Alexander Kremer, and Heinrich Hussmann
-
[32]
Hovy, Hinrich Schütze, and Yoav Goldberg
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, and Yoav Goldberg. 2021. Measuring and Improving Consistency in Pretrained Language Models. Transactions of the Association for Computational Linguistics 9 (2021), 1012–1031. h...
2021
-
[33]
Raymond Fok and Daniel S. Weld. [n.d.]. In Search of Verifiability: Explanations Rarely Enable Complementary Performance in AI-advised Decision Making. AI Magazine ([n. d.]). https://doi.org/10.1002/aaai.12182
-
[34]
Mark C Fox, K Anders Ericsson, and Ryan Best. 2011. Do Procedures for Verbal Reporting of Thinking Have to be Reactive? A Meta-Analysis and Recommenda- tions for Best Reporting Methods. Psychological Bulletin 137, 2 (2011), 316–344. https://doi.org/10.1037/a0021663
2011 doi
-
[35]
Bas. C. van Fraassen. 1980. The Scientific Image . Oxford University Press. https://doi.org/10.1093/0198244274.001.0001
1980 doi
-
[36]
Frazier, Susan A
Brandy N. Frazier, Susan A. Gelman, and Henry M. Wellman. 2009. Preschoolers’ Search for Explanatory Information Within Adult–Child Conversation. Child Development 80, 6 (2009), 1592–1611. https://doi.org/10.1111/j.1467-8624.2009. 01356.x
2009
-
[37]
Gajos and Lena Mamykina
Krzysztof Z. Gajos and Lena Mamykina. 2022. Do People Engage Cognitively with AI? Impact of AI Assistance on Incidental Learning. In Proceedings of the 27th International Conference on Intelligent User Interfaces (IUI ’22) . Association for Computing Machinery, New York, NY, U...
2022
-
[38]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:cs.CL/2312.10997 https://arxiv.org/abs/2312.10997
2024 arXiv
-
[39]
Carly Giffin, Daniel Wilkenfeld, and Tania Lombrozo. 2017. The Explanatory Effect of a Label: Explanations with Named Categories are More Satisfying. Cognition 168 (2017), 357–369. https://doi.org/10.1016/j.cognition.2017.07.011
2017 doi
-
[40]
Ana Valeria González, Gagan Bansal, Angela Fan, Yashar Mehdad, Robin Jia, and Srinivasan Iyer. 2021. Do Explanations Help Users Detect Errors in Open- Domain QA? An Evaluation of Spoken vs. Visual Explanations. InFindings of the Association for Computational Linguistics: ACL-I...
2021 doi
-
[41]
Ben Green and Yiling Chen. 2019. The Principles and Limits of Algorithm-in- the-Loop Decision Making. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 50 (nov 2019), 24 pages. https://doi.org/10.1145/3359152
2019 doi
-
[42]
Peter Green and Catriona J. MacLeod. 2016. SIMR: An R package for Power Analysis of Generalized Linear Mixed Models by Simulation. Methods in Ecology and Evolution 7, 4 (2016), 493–498. https://doi.org/10.1111/2041-210X.12504
2016 doi
-
[43]
Gaole He, Stefan Buijsman, and Ujwal Gadiraju. 2023. How Stated Accuracy of an AI System and Analogies to Explain Accuracy Affect Human Reliance on the System. Proc. ACM Hum.-Comput. Interact. 7, CSCW2, Article 276 (oct 2023), 29 pages. https://doi.org/10.1145/3610067
2023 doi
-
[44]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring Mathematical Problem Solving With the MATH Dataset. In Proceedings of the Neural Information Pro- cessing Systems Track on Datasets and Benchmar...
2021
-
[45]
Hopkins, Deena Skolnick Weisberg, and Jordan C.V
Emily J. Hopkins, Deena Skolnick Weisberg, and Jordan C.V. Taylor. 2019. Does Expertise Moderate the Seductive Allure of Reductive Explanations? Acta Psychologica 198 (2019), 102890. https://doi.org/10.1016/j.actpsy.2019.102890
2019
-
[46]
Hopkins, Deena S
Emily J. Hopkins, Deena S. Weisberg, and Jordan C. V. Taylor. 2016. The Se- ductive Allure is a Reductive Allure: People Prefer Scientific Explanations that Contain Logically Irrelevant Reductive Information. Cognition 155 (2016), 67–76. https://doi.org/10.1016/j.cognition.2016.06.011
2016 doi
-
[47]
Jie Huang and Kevin Chen-Chuan Chang. 2023. Towards Reasoning in Large Language Models: A Survey. In Findings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toro...
2023 doi
-
[48]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. arXiv:cs.CL/2311....
2023 arXiv
-
[49]
Patrick J. Hurley. 2000. A Concise Introduction to Logic . Wadsworth, Belmont, CA
2000
-
[51]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 55, 12, Article 248 (mar 2023), 38 pages. https://doi.org/10.1145/3571730
2023 doi
-
[52]
Daniel Kahneman. 2003. A Perspective on Judgment and Choice: Mapping Bounded Rationality. The American Psychologist 58, 9 (2003), 697–720. https: //doi.org/10.1037/0003-066X.58.9.697
2003 doi
-
[53]
Daniel Kahneman. 2011. Thinking, Fast and Slow . Farrar, Straus and Giroux
2011
-
[54]
Ho, Percy Liang, and Arvind Narayanan
Sayash Kapoor, Rishi Bommasani, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Peter Cihon, Aspen Hopkins, Kevin Bankston, Stella Biderman, Miranda Bogen, Rumman Chowdhury, Alex Engler, Peter Henderson, Yacine Jer- nite, Seth Lazar, Stefano Maffulli, Alondra Nelson, Joelle Pi...
2024
-
[55]
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. 2020. Interpreting Interpretability: Under- standing Data Scientists’ Use of Interpretability Tools for Machine Learning. In Proceedings of the 2020 CHI Conference on Huma...
2020
-
[56]
Frank C. Keil. 2006. Explanation and Understanding. Annual Review of Psychol- ogy 57 (2006), 227–254. https://doi.org/10.1146/annurev.psych.57.102904.190100
2006
-
[57]
Deborah Kelemen, Joshua Rottman, and Rebecca Seston. 2013. Professional Physical Scientists Display Tenacious Teleological Tendencies: Purpose-based Reasoning as a Cognitive Default. Journal of Experimental Psychology: General 142, 4 (2013), 1074–1083. https://doi.org/10.1037/a0030399
2013 doi
-
[58]
National Geographic Kids. 2017. Weird But True! Human Body: 300 Outrageous Facts about Your A wesome Anatomy. National Geographic. https://books.google. com/books?id=dJo8DgAAQBAJ
2017
-
[59]
2023.Weird But True World 2024
National Geographic Kids. 2023.Weird But True World 2024. National Geographic. https://books.google.com/books?id=JMiPzwEACAAJ
2023
-
[60]
I’m Not Sure, But
Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan. 2024. "I’m Not Sure, But... ": Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and Trust. In Proceedings of the 2024 ACM Conference on Fa...
2024
-
[61]
Sunnie S. Y. Kim, Nicole Meister, Vikram V. Ramaswamy, Ruth Fong, and Olga Russakovsky. 2022. HIVE: Evaluating the Human Interpretability of Visual Explanations. In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Proceedings, Part...
2022 doi
-
[62]
Help Me Help the AI
Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. "Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (...
2023
-
[63]
Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. Humans, AI, and Context: Understanding End-Users’ Trust in a Real-World Computer Vision Application. In Proceedings of the 2023 ACM Conference on Fairness, Accountability,...
2023
-
[64]
Yoonsu Kim, Jueon Lee, Seoyoung Kim, Jaehyuk Park, and Juho Kim. 2024. Un- derstanding Users’ Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level. InProceedings of the 29th International Conference on Intelligent User Interfaces ...
2024
-
[65]
Kurkul and Kathleen H
Katelyn E. Kurkul and Kathleen H. Corriveau. 2018. Question, Explanation, Follow-Up: A Mechanism for Learning From Others? Child Development 89, 1 (2018), 280–294. https://doi.org/10.1111/cdev.12726
2018 doi
-
[66]
Philippe Laban, Lidiya Murakhovs’ka, Caiming Xiong, and Chien-Sheng Wu
-
[67]
Bennett, and Marti A
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-Visiting NLI-based Models for Incon- sistency Detection in Summarization. Transactions of the Associ- ation for Computational Linguistics 10 (02 2022), 163–177. https: //doi.org/10.1162/tac...
2022 doi
-
[68]
Why is ‘Chicago’ deceptive?
Vivian Lai, Han Liu, and Chenhao Tan. 2020. "Why is ‘Chicago’ deceptive?" Towards Building Model-Driven Tutorials for Humans. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20) . Association for Computing Machinery, New York, NY, USA, 1–13...
2020
-
[69]
Vivian Lai and Chenhao Tan. 2019. On Human Predictions with Explanations and Predictions of Machine Learning Models: A Case Study on Deception Detection. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19). Association for Computing Machin...
2019
-
[70]
Placebic
Ellen Langer, Arthur Black, and Benzio Chanowitz. 1978. The Mindlessness of Ostensibly Thoughtful Action: The Role of "Placebic" Information in Inter- personal Interaction. Journal of Personality and Social Psychology 36, 6 (1978), 635–642. https://doi.org/10.1037/0022-3514.36.6.635
1978 doi
-
[71]
Yoonjoo Lee, Kihoon Son, Tae Soo Kim, Jisu Kim, John Joon Young Chung, Eytan Adar, and Juho Kim. 2024. One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations. In Proceedings of the 2024 ACM Conference on Fairness, Accountabilit...
2024
-
[72]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Gener- ation for Knowledge-Intensive NLP Tasks. In Advances...
2020
-
[73]
Q Vera Liao and S Shyam Sundar. 2022. Designing for Responsible Trust in AI Systems: A Communication Perspective. Proceedings of the 2022 Conference on Fairness, Accountability, and Transparency (2022)
2022
-
[74]
Q Vera Liao and Kush R Varshney. 2021. Human-Centered Explainable AI (XAI): From Algorithms to User Experiences. arXiv preprint arXiv:2110.10790 (2021)
2021 arXiv
-
[75]
Liquin and Tania Lombrozo
Emily G. Liquin and Tania Lombrozo. 2022. Motivated to Learn: An Account of Explanatory Satisfaction. Cognitive Psychology 132 (2022), 101453. https: //doi.org/10.1016/j.cogpsych.2021.101453
2022
-
[76]
Han Liu, Vivian Lai, and Chenhao Tan. 2021. Understanding the Effect of Out- of-distribution Examples and Interactive Explanations on Human-AI Decision Making. Proc. ACM Hum.-Comput. Interact. 5, CSCW2, Article 408 (oct 2021), 45 pages. https://doi.org/10.1145/3479552
2021 doi
-
[77]
Nelson Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating Verifiability in Generative Search Engines. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singa...
2023 doi
-
[78]
Tania Lombrozo. 2006. The Structure and Function of Explanations. Trends in Cognitive Sciences 10, 10 (2024/09/10 2006), 464–470
2006
-
[79]
Tania Lombrozo. 2007. Simplicity and Probability in Causal Explanation. Cog- nitive Psychology 55, 3 (2007), 232–257. https://doi.org/10.1016/j.cogpsych.2006. 09.006
2007 doi
-
[80]
Tanya Lombrozo. 2012. Explanation and Abductive Inference. In The Oxford Handbook of Thinking and Reasoning . Oxford University Press. https://doi.org/ 10.1093/oxfordhb/9780199734689.013.0014
2012
-
[81]
Tania Lombrozo. 2016. Explanatory Preferences Shape Learning and Inference. Trends in Cognitive Sciences 20, 10 (2024/09/10 2016), 748–759
2016
-
[82]
Tania Lombrozo and Emily G. Liquin. 2023. Explanation Is Effective Because It Is Selective. Current Directions in Psychological Science 32, 3 (2023), 212–219. https://doi.org/10.1177/09637214231156106
2023 doi
-
[83]
Duri Long and Brian Magerko. 2020. What is AI Literacy? Competencies and Design Considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20) . Association for Computing Machinery, New York, NY, USA, 1–16. https://doi.org/10.1145/331...
2020
-
[84]
Zhuoran Lu and Ming Yin. 2021. Human Reliance on Machine Learning Models When Performance Feedback is Limited: Heuristics and Risks. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21) . Association for Computing Machinery, New York, NY, U...
2021
-
[85]
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Mari- anna Apidianaki, and Chris Callison-Burch. 2023. Faithful Chain-of-Thought Reasoning. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of...
2023
-
[86]
B. F. Malle and Joshua Knobe. 1997. Which Behaviors Do People Explain? A Basic Actor–Observer Asymmetry. Journal of Personality and Social Psychology 72, 2 (1997), 288
1997
-
[87]
Ana Marasović, Iz Beltagy, Doug Downey, and Matthew E. Peters. 2021. Few- Shot Self-Rationalization with Natural Language Prompts. In NAACL-HLT. https://api.semanticscholar.org/CorpusID:244130199
2021
-
[88]
Mills, Judith H
Candice M. Mills, Judith H. Danovitch, Sydney P. Rowles, and Ian L. Campbell
-
[89]
Sina Mohseni, Fan Yang, Shiva Pentyala, Mengnan Du, Yi Liu, Nic Lupfer, Xia Hu, Shuiwang Ji, and Eric Ragan. 2021. Machine Learning Explanations to Prevent Overtrust in Fake News Detection. Proceedings of the International AAAI Conference on Web and Social Media 15, 1 (May 202...
2021 doi
-
[90]
Hansen Morten Hertzum and Hans H.K
Kristin D. Hansen Morten Hertzum and Hans H.K. Andersen. 2009. Scrutinis- ing Usability Evaluation: Does Thinking Aloud Affect Behaviour and Men- tal Workload? Behaviour & Information Technology 28, 2 (2009), 165–181. https://doi.org/10.1080/01449290701773842
2009 doi
-
[91]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2024
-
[92]
Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, and Jiawei Han. 2023. The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions. In Proceedings of the 2023 Conference on Empirical Method...
2023 doi
-
[93]
Psychonomic Bulletin & Review 24, 5 (2017), 1465–1477
Children’s Success at Detecting Circular Explanations and Their Interest in Future Learning. Psychonomic Bulletin & Review 24, 5 (2017), 1465–1477
2017
-
[94]
Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Yang Wang. 2023. On the Risk of Misinformation Pollution with Large Language Models. In The 2023 Conference on Empirical Methods in Natural Language Processing. https://openreview.net/forum?id=voBhcwDyPt
2023
-
[95]
Samir Passi, Shipi Dhanorkar, and Mihaela Vorvoreanu. 2024. Appropriate reliance on Generative AI: Research synthesis . Technical Report MSR-TR-2024-7. Microsoft. https://www.microsoft.com/en-us/research/publication/appropriate- reliance-on-generative-ai-research-synthesis/
2024
-
[96]
2022.Overreliance on AI: Literature Review
Samir Passi and Mihaela Vorvoreanu. 2022.Overreliance on AI: Literature Review. Technical Report MSR-TR-2022-12. Microsoft. https://www.microsoft.com/en- us/research/publication/overreliance-on-ai-literature-review/ CHI ’25, April 26-May 1, 2025, Yokohama, Japan Kim, Vaughan, ...
2022
-
[97]
Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wort- man Wortman Vaughan, and Hanna Wallach. 2021. Manipulating and Measuring Model Interpretability. In Proceedings of the 2021 CHI Conference on Human Fac- tors in Computing Systems (CHI ’21). Associatio...
2021
-
[98]
Marvin Pafla, Kate Larson, and Mark Hancock. 2024. Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24) . Association f...
2024
-
[99]
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. 2022. Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges. Statistics Surveys 16, none (2022), 1 – 85. https: //doi.org/10.1214/21-SS133
2022 doi
-
[100]
Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023. Ver- bosity Bias in Preference Labeling by Large Language Models. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following . https://openreview. net/forum?id=magEgFpK1y
2023
-
[101]
Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2023. A Missing Piece in the Puzzle: Considering the Role of Task Complexity in Human-AI Decision Making. In Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization (UMAP ’23). Association for Compu...
2023
-
[102]
Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2024. Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24)...
2024
-
[103]
Danielson
Soo Young Rieh and David R. Danielson. 2007. Credibility: A Multidisciplinary Framework. Annual Review of Information Science and Technology 41, 1 (2007), 307–364. https://doi.org/10.1002/aris.2007.1440410114
2007 arXiv
-
[104]
Max Schemmer, Patrick Hemmer, Maximilian Nitsche, Niklas Kühl, and Michael Vössing. 2022. A Meta-Analysis of the Utility of Explainable Artificial Intelli- gence in Human-AI Decision-Making. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’22) ....
2022
-
[105]
Murray Shanahan. 2024. Talking about Large Language Models. Commun. ACM 67, 2 (Jan 2024), 68–79. https://doi.org/10.1145/3624724
2024 doi
-
[106]
Vera Liao, and Ziang Xiao
Nikhil Sharma, Q. Vera Liao, and Ziang Xiao. 2024. Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information Seeking. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York,...
2024
-
[107]
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023. Large Language Models can be Easily Distracted by Irrelevant Context. In Proceedings of the 40th International Conference on Machine Learning (ICML’23) . JMLR.o...
2023
-
[108]
Sashank Santhanam, Behnam Hedayatnia, Spandana Gella, Aishwarya Pad- makumar, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tur. 2022. Rome was built in 1776: A Case Study on Factual Correctness in Knowledge-Grounded Response Generation. arXiv:cs.CL/2110.05456 https://arxiv.org/ab...
2022 arXiv
-
[109]
Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, and Jordan Boyd-Graber. 2024. Large Language Models Help Humans Ver- ify Truthfulness – Except When They Are Convincingly Wrong. In Proceed- ings of the 2024 Conference of the North American Chapter ...
2024
-
[110]
Rothschild, Daniel G
Sofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, and Jake M. Hofman. 2023. Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment. arXiv:cs.HC/2307.03744
2023 arXiv
-
[111]
J.D. Trout. 2008. Seduction Without Cause: Uncovering Explanatory Neurophilia. Trends in Cognitive Sciences 12, 8 (2008), 281–282. https://doi.org/10.1016/j.tics. 2008.05.004
2008 doi
-
[112]
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. Lan- guage Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. InProceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’...
2023
-
[113]
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott We...
2021 doi
-
[114]
Bernstein, and Ranjay Krishna
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023. Explanations Can Reduce Overreliance on AI Systems During Decision-Making. Proc. ACM Hum.-Comput. Interact. 7, CSCW1, Article 129 (apr 2023), 38 ...
2023 doi
-
[115]
Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, Yidong Wang, Linyi Yang, Jindong Wang, Xing Xie, Zheng Zhang, and Yue Zhang
-
[116]
Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In The Eleventh International Conference on Learning Representations . https://open...
2023
-
[117]
Xinru Wang and Ming Yin. 2021. Are Explanations Helpful? A Compara- tive Study of the Effects of Explanations in AI-Assisted Decision-Making. In Proceedings of the 26th International Conference on Intelligent User Interfaces (IUI ’21). Association for Computing Machinery, New ...
2021
-
[118]
Vera Liao, and Jen- nifer Wortman Vaughan
Helena Vasconcelos, Gagan Bansal, Adam Fourney, Q. Vera Liao, and Jen- nifer Wortman Vaughan. 2024. Generation Probabilities Are Not Enough: Exploring the Effectiveness of Uncertainty Highlighting in AI-Powered Code Completions. ACM Transactions on Computer-Human Interaction (2024)
2024
-
[119]
Yeo Wei Jie, Ranjan Satapathy, Rick Goh, and Erik Cambria. 2024. How Inter- pretable are Reasoning Explanations from Prompting Large Language Models?. In Findings of the Association for Computational Linguistics: NAACL 2024 , Kevin Duh, Helena Gomez, and Steven Bethard (Eds.)....
2024 doi
-
[120]
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Court- ney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, Willi...
2022
-
[121]
Deena Skolnick Weisberg, Jordan C.V Taylor, and Emily J Hopkins. 2015. De- constructing the Seductive Allure of Neuroscience Explanations. Judgment and Decision Making 10, 5 (2015), 429–441
2015
-
[122]
Benjamin Weiser and Nate Schweber. 2023. The ChatGPT Lawyer Explains Himself. New York Times (June 2023)
2023
-
[123]
Henry M. Wellman. 2011. Reinvigorating Explanations for the Study of Early Cognitive Development. Child Development Perspectives 5, 1 (2011), 33–38. https://doi.org/10.1111/j.1750-8606.2010.00154.x
2011
-
[124]
Nadine Wathen and Jacquelyn Burkell
C. Nadine Wathen and Jacquelyn Burkell. 2002. Believe It or Not: Factors Influ- encing Credibility on the Web. Journal of the American Society for Information Science and Technology 53, 2 (2002), 134–144. https://doi.org/10.1002/asi.10016
2002 doi
-
[125]
Jennifer Wortman Vaughan and Hanna Wallach. 2021. A Human-Centered Agenda for Intelligible Machine Learning. In Machines We Trust: Perspectives on Dependable AI, Marcello Pelillo and Teresa Scantamburlo (Eds.). MIT Press
2021
-
[126]
Roy Xie, Chengxuan Huang, Junlin Wang, and Bhuwan Dhingra. 2024. Adver- sarial Math Word Problem Generation. arXiv:cs.CL/2402.17916
2024 arXiv
-
[127]
Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi. 2023. A Critical Evaluation of Evaluations for Long-form Question Answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Gra...
2023 doi
-
[128]
Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019. Understanding the Effect of Accuracy on Trust in Machine Learning Models. In Proceedings of the 2019 ACM CHI Conference on Human Factors in Computing Systems
2019
-
[129]
Kun Yu, Shlomo Berkovsky, Ronnie Taib, Jianlong Zhou, and Fang Chen. 2019. Do I Trust My Machine Teammate? An Investigation from Perception to De- cision. In Proceedings of the 24th International Conference on Intelligent User Interfaces (IUI ’19) . Association for Computing M...
2019
-
[130]
Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi
-
[131]
Zemla, Steven Sloman, Christos Bechlivanidis, and David A
Jeffrey C. Zemla, Steven Sloman, Christos Bechlivanidis, and David A. Lagnado
-
[132]
Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. 2020. Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-assisted Decision Making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. 295–305
2020
-
[133]
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for Large Language Models: A Survey. ACM Trans. Intell. Syst. Technol. 15, 2, Article 20 (feb 2024), 38 pages. https://doi.org/10.1145/3639372
2024 doi
-
[134]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2024. Judging LLM-as-a-Judge with MT- bench and Chatbot Arena. In Proceedings of the 37th Interna...
2024
-
[135]
Hwang, Xiang Ren, and Maarten Sap
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, and Maarten Sap. 2024. Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty. arXiv:cs.CL/2401.06730
2024 arXiv
-
[136]
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji rong Wen. 2023. Large Language Models for Information Retrieval: A Survey. ArXiv abs/2308.07107 (2023). https://api. semanticscholar.org/CorpusID:260887838
2023
-
[137]
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending Against Neural Fake News. In Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R....
2019
-
[139]
Psychonomic Bulletin & Review 24, 5 (2017), 1488–1500
Evaluating Everyday Explanations. Psychonomic Bulletin & Review 24, 5 (2017), 1488–1500. Fostering Appropriate Reliance on Large Language Models CHI ’25, April 26-May 1, 2025, Yokohama, Japan
2017
-
[145]
Which animal was sent to space first, cockroach or moon jellyfish?
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-Tuning Language Models from Human Preferences. arXiv preprint arXiv:1909.08593 (2019). https: //arxiv.org/abs/1909.08593 CHI ’25, April 26-...
2019 arXiv
-
[146]
https://science.nasa.gov/moon/moon- walkers/ 3
https://simple.wikipedia.org/wiki/List_of_people_who_have_ walked_on_the_Moon 2. https://science.nasa.gov/moon/moon- walkers/ 3. https://www.discovermagazine.com/planet-earth/what- has-been-found-in-the-deep-waters-of-the-mariana-trench CHI ’25, April 26-May 1, 2025, Yokohama,...
2025
-
[147]
https://www.healthline.com/health/hair-density • Incorrect: Yes, gorillas have twice as many hairs per square inch as humans
https://www.nationalgeographic.com/science/article/the-semi- naked-ape-or-why-peach-fuzz-makes-it-harder-for-parasites 3. https://www.healthline.com/health/hair-density • Incorrect: Yes, gorillas have twice as many hairs per square inch as humans. Gorillas have a significantly...
2011
-
[148]
https://midwesteyecenter
https://2020visioncare.com/the-eye-a-marvel-of-complexity- with-over-2-million-working-parts/ 2. https://midwesteyecenter. com/what-are-the-makings-of-the-human-eye/ 3. https://www. optometrists.org/general-practice-optometry/guide-to-eye-health/ how-does-the-eye-work/ • Incor...
2025
-
[149]
As of recent estimates, Brazil’s popula- tion is over 213 million people, which constitutes a significant ma- jority of South America’s total population of around 430 million
https://www.worldometers.info/world-population/south-america- population/ • Incorrect: Yes, more than two-thirds of South America’s popula- tion live in Brazil because Brazil is the largest and most populous country on the continent. As of recent estimates, Brazil’s popula- ti...
2025
-
[2008]
In Proceedings of ACL-08: HLT , Johanna D
Finding Contradictions in Text. In Proceedings of ACL-08: HLT , Johanna D. Moore, Simone Teufel, James Allan, and Sadaoki Furui (Eds.). Association for Computational Linguistics, Columbus, Ohio, 1039–1047. https://aclanthology. org/P08-1118
-
[2017]
Psychonomic Bulletin & Review 24, 5 (2017), 1451–1464
Concreteness and Abstraction in Everyday Explanation. Psychonomic Bulletin & Review 24, 5 (2017), 1451–1464. https://doi.org/10.3758/s13423-017- 1299-3
2017 doi
-
[2019]
In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (CHI EA ’19)
The Impact of Placebic Explanations on Trust in Intelligent Systems. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (CHI EA ’19) . Association for Computing Machinery, New York, NY, USA, 1–6. https://doi.org/10.1145/3290607.3312787
2019
-
[2022]
Reframing Human-AI Collaboration for Generating Free-Text Explana- tions. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Marine Carpuat, Marie-Catherine de Marneffe, and Ivan V...
2022 doi
-
[2023]
arXiv:cs.CL/2310.07521
Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity. arXiv:cs.CL/2310.07521
-
[2024]
arXiv:cs.CL/2311.08596 https://arxiv.org/abs/2311.08596
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment. arXiv:cs.CL/2311.08596 https://arxiv.org/abs/2311.08596
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.