{"id":"a12e2faa-5260-4d28-a12b-b796a585e3ab","arxiv_id":"2505.01192","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"High AI confidence increases reliance and reduces cognitive load, counterfactual explanations show mixed accuracy effects, and Need for Cognition did not differentiate users in a loan approval task.","lead":"This paper reports a 288-person online experiment on how AI predictions, confidence scores, and four explanation styles affect loan approval decisions. It finds that high AI confidence boosts reliance and cuts cognitive load, and that counterfactual explanations have mixed effects on accuracy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Confidence levels are confounded with instance difficulty: high- and low-confidence conditions use different loan applications, so H1a/H1b may reflect easier cases rather than the displayed confidence value.","rationale":"The reader's weakest assumption pointed to the hand-picked instances and the mismatch between stated and observed AI accuracy, framing it as a representativeness problem. I agree that this is fragile, but the more load-bearing issue is internal validity: the independent variable 'AI confidence' is not manipulated independently of the loan applicant's content. Because high and low confidence come from different instances, any observed effect on reliance or cognitive load could be driven by instance difficulty, typicality, or other features. The paper's own No-AI condition provides a natural control, but that comparison is not reported. This does not require rejecting the paper outright; the existing data could settle it. Therefore the conditional verdict stands, now explicitly conditioned on the no-AI per-instance baseline analysis. The rest of the empirical contributions—especially the null NFC results and the exploratory counterfactual analyses—are less affected by this confound, since they are either between-subjects or do not rely on the same causal reading. The concern is concrete, testable, and central to the abstract's strongest claim, so it should be addressed before the confidence finding is taken as established.","tokens_in":31904,"tokens_out":6501,"duration_ms":73331,"concrete_test":"Use the existing No-AI condition data (n≈48) to compare per-instance human accuracy and SEQ cognitive load between the four model-high-confidence and four model-low-confidence main-session instances (rows 1, 3, 5, 7 vs rows 2, 4, 6, 8 in Table 1). If No-AI participants are already more accurate and report lower load on the high-confidence instances, H1a/H1b are confounded by instance difficulty. Conversely, if No-AI users show no accuracy or load difference across these instance subsets, the confidence-display interpretation is supported. Also report per-instance accuracy within AI-assisted conditions to check whether 'reliance' simply tracks objectively correct answers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that high AI confidence increases reliance and reduces cognitive load—rests on a within-subjects comparison of eight hand-picked instances (Section 3.2.3). High confidence was not manipulated within the same instance; it is a property of each selected applicant. The instances were balanced on AI correctness, predicted class, and true class, but not on human-perceived difficulty or feature prototypicality. Model confidence from a Random Forest likely correlates with how typical or separable an applicant is, and such applicants are probably easier for humans too. If high-confidence instances are objectively easier, users would naturally agree with the AI more often and report lower cognitive load, producing the observed H1a (Log-Odds=1.22) and H1b (Log-Odds=-0.41) effects without any causal role for the confidence display. The paper does not report any per-instance difficulty control or a no-AI baseline comparison by confidence subgroup, so the central claim is not yet internally validated. The reader's concern about representativeness is related, but the sharper issue is this instance-identity confound.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an online user study (N = 288) on AI-assisted loan approval decisions, comparing six AI assistance conditions (No AI, AI with prediction/confidence/accuracy, and the latter plus example-based, feature-based, rule-based, or counterfactual explanations). The authors test hypotheses about AI confidence effects on reliance and cognitive load (H1a, H1b), the accuracy benefit of feature-based explanations (H1c), NFC-based differences in interface-element prioritization (H2), and NFC effects on accuracy and cognitive load (H3). They report that high AI confidence increases reliance and reduces cognitive load, that feature-based explanations do not improve accuracy, that both low and high NFC users rank explanations above AI information, and that no NFC differences in accuracy or cognitive load were found. Post hoc analyses suggest counterfactual explanations increase reliance and reduce cognitive load when AI predictions are correct, and the abstract claims they also enhance overall accuracy.","tokens_in":32180,"tokens_out":4591,"duration_ms":46165,"significance":"If the central confidence effects were causal, the paper would make a useful contribution to human-AI decision-making and XAI personalization. The study has concrete strengths: it uses an a priori power analysis, attention checks, mixed-effects and GEE models, a publicly available dataset, and a stated open pipeline for data processing and statistical analysis. The null results for feature-based explanations and for NFC differences are also informative in a literature with mixed findings. However, the main causal claims about confidence are threatened by a stimulus-level confound, and the counterfactual accuracy claim in the abstract goes beyond what the pre-specified analyses support.","major_comments":[{"comment":"H1a and H1b are tested by comparing high- and low-confidence loan applications, but AI confidence is a property of the instance, not a factor manipulated within the same instance. The eight main-session instances were balanced on AI correctness, predicted class, and true class, but not on human-perceived difficulty or prototypicality. If high-confidence instances are also easier for humans, the observed increases in reliance (Log-Odds = 1.22) and decreases in cognitive load (Log-Odds = -0.41) could occur without any causal role for the displayed confidence. The manuscript already contains data for a concrete test: the No AI condition saw the same instances without confidence information, so the authors can compare human accuracy and cognitive load on high- versus low-confidence instances in that condition; if no-AI performance also differs by confidence subgroup, the H1 effects are confounded. This issue is load-bearing for the paper's central claim and should be addressed directly.","section":"3.2.3 and Table 1"},{"comment":"The statistical models include a random intercept for participant but no by-item random effects for the eight main-session loan applications, and the instances were not randomly sampled from the test set. With only one or two instances per confidence-by-correctness combination, the reported H1 effects may be driven by idiosyncratic features of the selected applicants rather than by the confidence construct. I ask the authors to report per-instance estimates (e.g., reliance and SEQ means for each of the eight instances) and, if feasible, an item-level analysis or a sensitivity analysis that excludes the most extreme instances.","section":"4.2 and 3.2.3"},{"comment":"The claim that counterfactual explanations 'enhanced overall accuracy' is not supported by the pre-specified analysis. H1c was evaluated at alpha = .01, and the feature-based comparison did not reach this threshold (Log-Odds = 0.34, p = .0349); the footnote reports a counterfactual effect at p = .0149, which also fails the stated threshold. The post hoc model in Section 5.3 reports a positive main effect for counterfactual explanations (Log-Odds = 0.87, p = .0015) but simultaneously reports a negative interaction with correct AI predictions (Log-Odds = -0.84, p < .0133), and the text acknowledges a 'trend' of decreased accuracy in that context. An exploratory main effect in a model with a significant interaction does not warrant the unqualified statement in the abstract and contribution list that counterfactuals 'enhanced overall accuracy'; the claim should be reworded as an exploratory finding with the negative interaction reported in the same sentence.","section":"Abstract and 5.3"}],"minor_comments":[{"comment":"The last sentence says 'we fail to reject the null hypothesis for H3c'; this should refer to H3b.","section":"5.2.3"},{"comment":"In the second paragraph, 'delining' should be 'delineating'.","section":"5.3"},{"comment":"Describing p = .0149 as showing an effect conflicts with the stated alpha = .01 threshold; the text should either call it a trend or explicitly note that it does not meet the threshold.","section":"5.2.1 and footnote 7"},{"comment":"The caption mentions 'lower and higher confidence intervals based on standard errors' where 'error bars' would be clearer, and the H1a panel does not appear to show the ticks referenced in the caption.","section":"Figure 3 caption"},{"comment":"The sentence reporting H2a says users prioritize the explanation second and AI information third; this is the pattern predicted by H2b, so the wording should clarify that H2a is rejected precisely because the observed order differs from H2a's prediction.","section":"5.2.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: worth a serious read, but don't take the abstract at face value. The study is competently run and the NFC null is a useful negative result, but the two headline findings rest on a design that can't separate confidence from instance identity.\n\nWhat's new: a six-condition between-subjects comparison of explanation styles in a loan task, with confidence from a real model (not simulated), N=288, and a median split on NFC. The NFC null in a high-stakes tabular task is a useful addition to a literature that mostly shows NFC effects in low-stakes settings. They also report the kind of practical detail (sample size computations, attention checks, practice session, incentives) that makes replication possible.\n\nWhere it wobbles. The central H1a/H1b claim is that high displayed confidence increases reliance and lowers cognitive load. But high- and low-confidence instances are different loan applications (Table 1). The eight main-session cases are balanced on AI correctness, predicted class, and true class, not on human-perceived difficulty or prototypicality. A Random Forest's confidence is likely correlated with how separable/typical an applicant is, and those applicants are probably easier for people too. So the higher agreement and lower SEQ scores on high-confidence cases may have nothing to do with the confidence display. The No AI condition saw the same eight instances, so the authors could check whether humans also find high-confidence instances easier; they didn't report that analysis. That is a load-bearing gap, not a nitpick.\n\nThe counterfactual accuracy claim is another soft spot. The pre-registered threshold was alpha=0.01; the feature-based H1c comparison for counterfactual was p=0.0149, and the accuracy boost appears in a post-hoc model (p=0.0015) with a negative interaction for correct AI predictions. Calling this 'enhanced overall accuracy' in the abstract is generous.\n\nCognitive load is measured with a single SEQ item, which is fine for a pilot but thin for a headline claim. The reproducibility link is mentioned but not visible in the text I read. There is also the usual post-hoc median split on NFC, which the authors themselves flag.\n\nNet: the paper is honest about several limitations and the discussion is measured, but the abstract and conclusion outrun the experimental design. The instance confound should be fixed either by a per-instance human difficulty control from the No AI condition or by repositioning the confidence results as descriptive. I would send it out for review, but with a letter asking for the reanalysis or an explicit downgrade of the claim.","headline":"Well-run empirical study whose headline confidence effects are confounded with instance difficulty; the counterfactual accuracy claim in the abstract is also a post-hoc stretch.","tokens_in":32652,"tokens_out":3329,"would_cite":true,"duration_ms":31460,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports an online experiment showing that high AI confidence significantly increases users' reliance on AI and reduces their cognitive load, while feature-based explanations do not improve accuracy.","keywords":["Loan approval prediction","AI-assisted decisions","Explainable AI","Reliance","Accuracy","Need for Cognition","Cognitive load","Counterfactual explanations"],"falsifier":"Run the same six-condition loan task with instances selected so the AI's observed accuracy matches its stated accuracy (about 83%), keeping confidence levels balanced; if high confidence still raises reliance and lowers cognitive load, the confidence effect survives the stated-observed accuracy mismatch. Alternatively, use a two-stage paradigm where users decide first and then see the AI's confidence; if the effect disappears, the result is anchoring rather than confidence-based delegation.","tokens_in":31743,"feed_emoji":"🤖","tokens_out":6453,"duration_ms":56143,"temperature":0.7,"pith_summary":"The paper reports an online experiment with 288 participants who decided whether to approve loan applications with varying levels of AI assistance: no AI, AI prediction/confidence/accuracy alone, or the same AI information plus one of four explanation styles (example-based, feature-based, rule-based, counterfactual). Its central claim is that the confidence score, not the explanation, is the main driver of user behavior: high AI confidence significantly increases reliance on the AI and reduces self-reported cognitive load, while feature-based explanations fail to improve accuracy over other conditions. The paper also claims that counterfactual explanations, although rated as less understandable, can increase overall accuracy and reduce cognitive load when the AI's prediction is correct. A third claim is that Need for Cognition, a personality trait widely used to split users, did not produce significant differences in accuracy, cognitive load, or how people ranked interface elements, suggesting the trait's effects may be task- and context-specific. A sympathetic reader would care because these results bear directly on how to design explainable AI interfaces that neither overburden users nor push them into blind agreement.","feed_headline":"High AI confidence makes users agree more and think less","feed_subtitle":"A 288-person loan study finds confidence scores, not explanations, drive reliance; personality traits mattered little.","key_machinery":"The experimental engine is a mixed-factorial online study built on a Random Forest classifier trained on a public loan-prediction dataset with about 83% test accuracy. AI confidence is defined from test-set quartiles of an entropy-based epistemic uncertainty estimate, with low below 44.3 and high above 61.6, and each participant sees eight main-session instances balanced across AI correctness, confidence, predicted class, and true class. Explanation styles are generated with nearest-neighbor examples, SHAP feature contributions, Anchors rules, and DiCE counterfactuals. The statistical machinery is a set of mixed-effects logistic regressions, generalized estimation equations, and Friedman/Nemenyi ranking tests, with a power analysis used to set the sample size and reduced alpha thresholds.","core_discovery":"On its own terms, the paper establishes that the information accompanying an AI suggestion changes behavior, and that the most consequential piece is the confidence value. Participants shown high-confidence AI predictions agreed with the AI significantly more often than those shown low-confidence predictions (log-odds 1.22), and they reported lower cognitive load (log-odds -0.41). The paper frames this as evidence that users treat confidence as a tiebreaker and that an uncalibrated or overconfident score can lead to over-reliance. It also finds that feature-based explanations, despite their dominance in XAI practice, did not outperform other assistance conditions on accuracy; counterfactual explanations, which users found harder to understand, improved accuracy overall and reduced load when paired with correct AI advice. Finally, the paper reports that low- and high-Need-for-Cognition participants both ranked loan attributes first, explanations second, and AI information last, and did not differ in accuracy or cognitive load, a null result the authors attribute to task complexity.","pith_inferences":["A testable extension: measure decision quality under high confidence when the AI is systematically wrong; if users are delegating rather than deliberating, accuracy should collapse exactly on those instances, which would make confidence displays a risk factor rather than a transparency feature.","The fixed 83% stated accuracy with 62.5% observed accuracy probably amplified both reliance and distrust; a replication that matches stated and observed accuracy would clarify whether the confidence effect is about the number itself or about resolving uncertainty.","The NFC null result suggests median-split personality measurement may be too coarse for complex tasks; traits like epistemic curiosity, or state measures like self-reported confidence, could capture variance this design missed.","The authors' one-stage paradigm leaves open that high confidence simply anchors the final answer; a two-stage design separating independent judgment from AI-informed judgment would test whether the cognitive-load reduction is genuine offloading."],"forward_implications":["Interfaces that display confidence scores will systematically nudge users toward the AI's recommendation, so deployments should calibrate confidence estimates to the true likelihood of correctness rather than showing raw model output.","Feature-based (Shapley-contribution) explanations, the current default in many AI systems, offer no measurable accuracy advantage over other explanation styles in a complex tabular task.","Counterfactual explanations, despite being rated less understandable, can raise overall decision accuracy and lower cognitive load when the AI advice is correct, making them a candidate for hybrid explanation designs.","Need for Cognition may not be a reliable personalization axis in complex, high-stakes tasks; adaptive interfaces may need other user characteristics or direct confidence calibration.","When the AI displays high confidence, users rank the AI information above the explanation, whereas low confidence blurs that ordering, so the confidence display changes which interface elements people actually use."],"supporting_citations":[{"why":"Established that high AI confidence increases trust and agreement; the paper's H1a directly extends this to a loan task.","marker":"(Zhang et al, 2020)"},{"why":"Showed that stated accuracy and confidence jointly shape trust; motivates presenting confidence alongside the fixed 83% accuracy.","marker":"(Rechkemmer and Yin, 2022)"},{"why":"Showed stated accuracy on held-out data increases trust; justifies the accuracy value shown in the interface.","marker":"(Yin et al, 2019)"},{"why":"Provided the nearest-neighbor example-based explanation design and evidence on overreliance that the example condition builds on.","marker":"(Chen et al, 2023)"},{"why":"Argued feature contribution explanations satisfy more desiderata; this is the premise behind the feature-based hypothesis H1c.","marker":"(Wang and Yin, 2022)"},{"why":"Found NFC differences under cognitive forcing functions; the comparison baseline for the paper's null NFC results.","marker":"(Bucinca et al, 2021)"},{"why":"Showed explanation-only designs benefit high-NFC users; one of the few NFC contrasts in the paper's discussion.","marker":"(Gajos and Mamykina, 2022)"},{"why":"Supplied the six-item Need for Cognition scale (NCS-6) used to split participants by the median.","marker":"(de Holanda Coelho et al, 2020)"},{"why":"Provides the task-complexity framing used to classify the loan scenario as high-complexity.","marker":"(Salimzadeh et al, 2023)"},{"why":"Frames cognitive load as an understudied outcome in AI-assisted decision-making, motivating H1b.","marker":"(Steyvers and Kumar, 2024)"}],"fun_headline_variants":["AI confidence sways users more than any explanation style","High confidence AI advice: users comply, think less","Counterfactual explanations improve AI decisions despite being less clear","Need for Cognition doesn't alter AI decision outcomes","Feature-based AI explanations fail to boost user accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the eight hand-picked instances and the quartile split capture how people respond to confidence itself, rather than how they respond to the unusual gap between the AI's stated 83% accuracy and its observed 62.5% accuracy on those instances.","fun_headline_variants_meta":{"raw":{"variants":["AI confidence sways users more than any explanation style","High confidence AI advice: users comply, think less","Counterfactual explanations improve AI decisions despite being less clear","Need for Cognition doesn't alter AI decision outcomes","Feature-based AI explanations fail to boost user accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1779,"prompt_tokens":1018,"completion_tokens":761,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":634,"completion_tokens_details":{"reasoning_tokens":686}},"tokens_in":634,"tokens_out":761,"duration_ms":7274,"temperature":1.0,"reasoning_tokens":686,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:23:30.959960+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same six-condition loan task with instances selected so the AI's observed accuracy matches its stated accuracy (about 83%), keeping confidence levels balanced; if high confidence still raises reliance and lowers cognitive load, the confidence effect survives the stated-observed accuracy mismatch. Alternatively, use a two-stage paradigm where users decide first and then see the AI's confidence; if the effect disappears, the result is anchoring rather than confidence-based delegation.","supporting_citations":[],"review_version":1}