{"id":"107cbdad-c656-490e-83bb-9ca6243f90e5","arxiv_id":"2502.08554","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Explanations increase user reliance on both correct and incorrect LLM answers, while sources and inconsistent explanations reduce overreliance on incorrect answers in a controlled experiment.","lead":"This paper reports two studies on how people decide whether to trust answers from large language models (LLMs). It finds that adding explanations makes people more likely to accept both right and wrong answers, while adding clickable sources or pointing out contradictions makes them less likely to accept wrong ones. The findings give interface designers a concrete lever: verify your LLM's sources and highlight inconsistencies to reduce overreliance.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Inconsistency sub-claim rests on 3 confounded questions; §4.3.1 lacks question random effects, so the causal design recommendation to highlight inconsistencies is not yet supported.","rationale":"I agree with the reader's verdict. The explanation and source effects are well-designed: within-subjects 2x2x2, pre-registered, mixed-effects models with question random effects, and the source effect is further supported by source-clicking and behavioral data. The inconsistency sub-claim is a small part of the paper but appears in the abstract and motivates a concrete design intervention, so it is not merely exploratory spin. The core problem is that inconsistency is a property of the stimulus, not a randomly assigned condition, and it is nested within only three questions. The ANOVA in §4.3.1 ignores this nesting. A mixed-model re-analysis would be informative but cannot fully separate question identity from inconsistency with only three inconsistent items; a follow-up experiment that yokes consistent and inconsistent versions of the same explanations is the decisive check. If the effect survives such a test, the paper's recommendation is strengthened; if not, the inconsistency claim should be downgraded to a hypothesis for future work. Since the reader already assigned CONDITIONAL and my concern matches theirs, the verdict should remain unchanged.","tokens_in":43706,"tokens_out":6515,"duration_ms":78236,"concrete_test":"Run a pre-registered follow-up experiment using the same 12 task questions and incorrect-answer explanations. For each of the 3 naturally inconsistent items, create a yoked consistent version by editing only the contradictory sentence(s) so the cited facts match the stated answer; for matched consistent items, create yoked inconsistent versions by inserting a contradictory sentence. Randomly assign participants to version, keep all other response features (answer, sources, formatting) identical, and analyze Agreement/Accuracy with Inconsistent as a within-question factor plus (1|participant) and (1|question). If agreement with incorrect answers no longer differs between versions, the §4.3.1 effect is a question-content artifact rather than an inconsistency effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's main 2x2x2 findings are well supported: explanations and sources are randomly manipulated, and the §4.1.3 analyses include (1|question) and (1|participant) random effects, so the claims that explanations increase reliance and that sources reduce overreliance on incorrect answers are credible. The inconsistency claim is the exception. In §4.1.4, only 3 of 12 incorrect-answer explanations naturally contained inconsistencies, and none of the correct-answer explanations did. Section 4.3.1 then compares 'No explanation', 'Consistent explanation', and 'Inconsistent explanation' using ANOVA without participant or question as random effects. The inconsistent cell (N=155) is thus fully confounded with the three specific questions (Brazil population, lungs vs. skin, mammals excluding platypus). The observed reduction in agreement (69.7% vs. 83.3%) and increase in accuracy (30.3% vs. 16.7%) may reflect question difficulty, answer plausibility, or the particular arithmetic error in the Brazil item rather than the general psychological mechanism that inconsistency reduces overreliance. Because the abstract states this as a finding and Section 5.1 recommends interventions that highlight inconsistencies, a causal claim is load-bearing. The paper itself describes the analysis as using 'natural inconsistencies' but does not flag this confound. A randomized manipulation, or at minimum a mixed-effects reanalysis, is required before the practical recommendation can be accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how features of LLM responses shape users' reliance when answering objective questions. Study 1 is a think-aloud study (N=16) that identifies explanations, inconsistencies, and sources as key features. Study 2 is a pre-registered, within-subjects experiment (N=308) with a 2x2x2 design (answer correctness, presence of explanation, presence of clickable sources) using realistic LLM-generated responses from ChatGPT and Perplexity AI. The main mixed-effects analyses show that explanations increase agreement with both correct and incorrect answers, while sources increase appropriate reliance on correct answers and reduce overreliance on incorrect answers. A secondary observational analysis suggests that inconsistent explanations are associated with reduced overreliance on incorrect answers. The authors discuss implications for designing LLM interfaces to foster appropriate reliance.","tokens_in":1092,"tokens_out":2770,"duration_ms":81704,"significance":"The randomized manipulation of explanations and sources, combined with mixed-effects models that include participant and question random effects, makes the main 2x2x2 findings credible and directly relevant to the HCI community. The pre-registration, power analysis, and use of realistic LLM-generated stimuli are notable strengths that increase confidence in the explanation and source effects. If the inconsistency finding were causally supported, the design recommendation to highlight inconsistencies would be of considerable practical value; however, as it stands, that sub-claim is not yet established because the analysis is observational and confounded.","major_comments":[{"comment":"The comparison of consistent vs. inconsistent explanations is confounded by question identity. Only 3 of the 12 task questions produced naturally occurring inconsistent explanations, all for incorrect answers, and these three questions may differ systematically in difficulty, answer plausibility, or specific content (e.g., the arithmetic error in the Brazil population item). The ANOVA does not include participant or question random effects, and it collapses across the sources manipulation, so the reported differences (agreement 69.7% vs. 83.3%; accuracy 30.3% vs. 16.7%) cannot be attributed to inconsistency per se. This confound is load-bearing because the abstract presents the inconsistency result as a finding and §5.1 uses it to recommend interventions that highlight inconsistencies.","section":"§4.3.1, Figure 5"},{"comment":"The causal claim that inconsistencies reduce overreliance, and the design recommendation to highlight inconsistencies as an intervention, are stronger than the evidence supports. The independent variable was not manipulated; it was observed post hoc after the experiment. Even a mixed-effects reanalysis would address non-independence but would not address selection on question content or the small number of inconsistent items. The authors should either reframe the inconsistency result as an exploratory, hypothesis-generating finding or conduct a follow-up experiment that factorially manipulates inconsistency while holding content constant.","section":"Abstract and §5.1"}],"minor_comments":[{"comment":"In the sentence reporting the confidence effect, β = .96, SE = .10, p < .001, the formatting of the p-value is inconsistent with the rest of the paper; please standardize the formatting of regression results throughout.","section":"§4.2.2"},{"comment":"The text does not state the cell size for the inconsistent explanation condition (N=155); adding this number would help readers interpret the precision of the estimates.","section":"§4.3.1"},{"comment":"The y-axis labels are not fully specified in the figure caption; please indicate the response scale or unit for each panel to improve readability.","section":"Figure 5"},{"comment":"The limitations section does not mention the confound in the inconsistency analysis; given the prominence of this finding, it should be explicitly acknowledged as an observational analysis with limited internal validity.","section":"§5.3"},{"comment":"The coding of inconsistencies was performed by the authors without reporting inter-rater reliability; since this variable is central to the inconsistency sub-claim, a second coder or a reliability statistic would strengthen confidence in the coding.","section":"§4.1.4"},{"comment":"There is a typo in the first sentence: the the relationship should be the relationship.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for CHI and the main 2x2x2 results are solid and clearly presented. My main concern is that the inconsistency sub-claim is overgeneralized relative to the evidence; I would urge the authors to either add a randomized inconsistency manipulation or limit the claim to exploratory. The paper would also benefit from more complete reporting of the inconsistency ANOVA, though that is secondary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a genuinely useful, well-run study on how response features shape reliance on LLMs, and the sources result is the real contribution. The explanation finding—explanations increase reliance on both correct and incorrect answers—replicates prior work, including Si et al., but the paper isolates it in a pre-registered N=308 design with mixed models and realistic stimuli. The sources effect is the new thing: clickable sources increase appropriate reliance on correct answers and reduce overreliance on incorrect ones, with a significant accuracy interaction. That part is credible and practically actionable. The related work is thorough and gives prior work proper credit.\n\nThe soft spot is exactly where the stress-test note lands. The inconsistency claim in the abstract—less reliance on incorrect responses when explanations exhibit inconsistencies—rests on a comparison of only 3 of 12 questions that naturally produced inconsistent explanations, analyzed with ANOVA without question or participant random effects. The inconsistent cell is confounded with question identity and content. The paper acknowledges the analysis is observational but does not flag the question confound, and then Section 5.1 recommends highlighting inconsistencies as a design intervention. That recommendation is not yet supported by the data as analyzed. A mixed-effects reanalysis or a follow-up that actually manipulates inconsistency would fix it.\n\nA couple of smaller points: the paper reports no data or code release, which makes independent verification harder, and Theta's 50% accuracy plus hard factual questions limits generalization, though the authors acknowledge that. None of that undercuts the main 2x2x2 findings.\n\nWho is it for? Anyone designing LLM-infused search or QA interfaces, and researchers working on overreliance. The sources finding alone is worth a serious referee; the inconsistency issue is addressable in revision. I would send it out.","headline":"Strong pre-registered study with a credible sources effect; the inconsistency sub-claim is real but currently confounded by question identity.","tokens_in":44495,"tokens_out":2117,"would_cite":true,"duration_ms":23755,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Explanations increase reliance on LLM answers—correct or not—while sources and contradictions curb overreliance on wrong ones.","keywords":["large language models","overreliance","appropriate reliance","explanations","sources","inconsistencies","human-AI interaction","question answering"],"falsifier":"Run the same 12 questions with matched explanation pairs that are identical except for one internal contradiction and randomly assign participants to versions; if agreement with incorrect answers does not drop when the contradiction is present, the paper's inconsistency claim is falsified.","tokens_in":43486,"feed_emoji":"🔗","tokens_out":7453,"duration_ms":71620,"temperature":0.7,"pith_summary":"Everyday users of large language models face a hard question: when should a fluent, confident answer be trusted? The paper tackles this through two empirical studies—a think-aloud session with 16 users and a pre-registered experiment with 308 participants—and claims that three response features govern reliance: explanations (supporting details), inconsistencies inside explanations, and clickable sources. Its central finding is that explanations increase reliance on both correct and incorrect answers, while sources improve reliance selectively (more agreement with correct answers, less with incorrect ones) and inconsistent explanations reduce overreliance on incorrect answers. If these findings hold, interface designers have concrete levers—provide accurate sources, surface contradictions—to help users get the benefit of LLM assistance without blind deference.","feed_headline":"Sources and contradictions curb blind trust in LLM answers","feed_subtitle":"Explanations boost trust even when wrong; sources and contradictions help users resist wrong answers.","key_machinery":"The load-bearing apparatus is a 2 × 2 × 2 within-subjects experiment using 12 difficult binary factual questions, where each of 308 participants saw eight response types from a hypothetical LLM named Theta: {correct, incorrect} × {no explanation, explanation} × {no sources, clickable sources}. Explanations—supporting details that justify the answer—and answers were generated in advance with ChatGPT and Perplexity AI so that content could be controlled, and reliance was measured behaviorally as whether the participant's final answer agreed with Theta's answer, complemented by self-reported confidence, justification-quality and actionability ratings, source-clicking, and follow-up questions. In addition, the authors coded naturally occurring inconsistencies in the explanations (sets of statements that cannot both be true) and ran a pre-registered ANOVA comparing incorrect answers with no explanation, consistent explanation, and inconsistent explanation; a 16-person think-aloud study supplied the qualitative account of how users notice these cues.","core_discovery":"Users agree with an LLM's answer more often when the answer comes with an explanation, regardless of whether the answer is actually right; the paper shows this in a controlled setting where the same difficult questions are paired with correct or incorrect answers, with or without explanations and sources. Sources change the pattern: they raise agreement when the answer is correct and lower it when the answer is wrong, and they increase time on task and the odds of overriding an incorrect answer. Explanations that contain a logical inconsistency—for instance, an answer that contradicts the numbers cited to support it—produce significantly less agreement and higher accuracy than consistent explanations when the answer is wrong. The paper therefore concludes that explanation is not a single good thing: its effect depends on whether it invites verification (sources) and whether it contains visible cracks (inconsistencies).","pith_inferences":["Beyond the paper: if the causal story is that inconsistency triggers deeper scrutiny, then automatically detecting and highlighting contradictions (for example, by checking whether the answer matches the numbers in the explanation) should reproduce the effect without requiring users to spot the flaw themselves.","Beyond the paper: the paper's observational comparison suggests a testable design rule—LLM answers whose supporting explanation is internally consistent should be treated as more reliable for answer selection, since consistency of the explanation is itself predictive of correctness in the paper's stimulus set.","Beyond the paper: the source effect may depend on individual differences: users who click links (roughly 119 of 308 participants clicked in at least one task, while 189 never clicked) may be the main beneficiaries, so an interface that actively nudges source-checking could widen the benefit to the majority who do not click.","Beyond the paper: the experiment used deliberately hard questions where lay users have little prior knowledge, so the findings may not transfer to domains where users can independently evaluate the answer; testing with familiar topics would show whether sources still carry the same corrective force."],"forward_implications":["Explanations alone make users more likely to accept an answer, so adding them without other safeguards can actively increase overreliance on wrong LLM outputs.","Providing clickable, accurate sources is a concrete corrective: in the no-explanation condition it raised agreement with correct answers from 67.2% to 73.4% and lowered agreement with incorrect answers from 78.2% to 68.2%.","When the LLM answer is wrong, sources without explanation produce the highest user accuracy (31.8%), while explanation alone gives the lowest (17.2%); when the answer is right, explanation plus sources gives the highest accuracy (79.9%).","Inconsistent explanations cut overreliance: agreement with wrong answers fell from 83.3% to 69.7% and accuracy rose from 16.7% to 30.3% compared with consistent explanations.","Because the beneficial source effect was obtained with real, mostly accurate links, the authors expect that fake, broken, or irrelevant sources would not help and could even increase perceived credibility."],"supporting_citations":[{"why":"Supplies the closest prior finding that LLM explanations increase reliance and that inconsistencies reduce overreliance, which this paper replicates and extends at larger scale.","marker":"[109]"},{"why":"Provides the one-response-per-task experimental paradigm and documents how inconsistent LLM outputs shape comprehension, grounding the stimulus design.","marker":"[71]"},{"why":"Contributes the controlled single-response method and the observation that users rarely click provided source links, which the source-clicking analysis builds on.","marker":"[60]"},{"why":"Establishes that AI explanations can increase overreliance even when the model is wrong, the baseline pattern the explanation result confirms for LLMs.","marker":"[6]"},{"why":"Shows that confidence and explanation affect accuracy and trust calibration in AI-assisted decisions, supporting the choice of reliance measures.","marker":"[132]"},{"why":"Documents that LLM-infused search engines generate inaccurate sources and unsupported statements, motivating the paper's caveat that source quality is load-bearing.","marker":"[77]"}],"fun_headline_variants":["Explanations boost trust in LLMs even when wrong","Sources and contradictions curb blind trust in LLM answers","Explanations raise trust in LLMs, but sources and logic gaps reduce it","Explanations increase reliance on LLMs unless sources or inconsistencies appear","Explanations boost trust even when wrong; sources and contradictions help"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The inconsistency finding rests on comparing a small set of questions whose explanations happened to contain contradictions against the other questions' explanations, so the result holds only if the contradiction itself—not the content or difficulty of those three questions—is what changed participants' agreement and accuracy.","fun_headline_variants_meta":{"raw":{"variants":["Explanations boost trust in LLMs even when wrong","Sources and contradictions curb blind trust in LLM answers","Explanations raise trust in LLMs, but sources and logic gaps reduce it","Explanations increase reliance on LLMs unless sources or inconsistencies appear","Explanations boost trust even when wrong; sources and contradictions help"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000782,"raw_usage":{"total_tokens":3414,"prompt_tokens":865,"completion_tokens":2549,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":2460}},"tokens_in":481,"tokens_out":2549,"duration_ms":18510,"temperature":1.0,"reasoning_tokens":2460,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T04:38:36.271186+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 12 questions with matched explanation pairs that are identical except for one internal contradiction and randomly assign participants to versions; if agreement with incorrect answers does not drop when the contradiction is present, the paper's inconsistency claim is falsified.","supporting_citations":[],"review_version":1}