{"id":"e13286e0-3a30-4f27-8732-d2f41a2d4e34","arxiv_id":"2505.10427","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Being emotionally aroused while an AI explains its reasoning is associated with lower understanding of the explanation, even though memory for the explained features is unaffected.","lead":"This study tested whether emotions change how people remember and understand explanations from an AI decision support system. It found that emotional arousal during an explanation was linked to poorer understanding, while induced moods before the task had little effect on recall of the explained features.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central RQ3b result rests on an unvalidated arousal detector; EmoNet is a speech model applied to faces, and the k=2.5/500 ms threshold has no sensitivity analysis.","rationale":"The reader's weakest assumption is exactly the right target. This is not a disagreement with consensus; it is a construct-validity threat to the central associative claim. The manuscript collects HRV and SAM and full EmoNet output vectors but never uses them to validate the binary reaction flags, and the detector threshold and window are unparameterized. The study has strengths: a concrete interaction design, clearly described dependent measures, and a mixed model with a random intercept for participants. The pattern of a null retention effect and a significant understanding effect is coherent and interesting. However, the central result is modest in strength (p = .015, N = 240 observations nested in 24 participants), so measurement validity is decisive. A threshold and window sensitivity grid plus convergent validation against HRV and SAM would settle whether the effect is real. Since the reader has already marked the paper CONDITIONAL with this exact caveat, my stress-test does not alter the verdict; I would make the sensitivity analysis a required condition of acceptance rather than an optional improvement.","tokens_in":8965,"tokens_out":4175,"duration_ms":45890,"concrete_test":"Recompute the Section 4.6 GLMM under a grid of detector settings, varying the z-threshold k over {2.0, 2.5, 3.0} and the rolling window over {250, 500, 1000} ms, and record the estimate, SE, and p-value for 'emotional reaction' in each cell. If the coefficient loses significance or changes sign in any plausible cell, the headline result is threshold-dependent; if it remains negative and significant across all cells, the objection is weakened. As a convergent validity check, compute the point-biserial correlation between the binary reaction flag and concurrently measured HRV-derived arousal and per-participant SAM arousal; absence of any positive correlation would indicate the detector is not measuring the intended construct.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central RQ3b result (Section 4.6: beta = -0.45, SE = 0.19, p = .015) depends entirely on the binary 'emotional reaction' detector defined in Section 3.1.1, Equations 1 and 2: a rolling z-score of EmoNet arousal with threshold k = 2.5 and a 500 ms window. That detector is unvalidated in two ways. First, EmoNet as cited is trained for speech emotion recognition, yet here it is applied to facial frames; the paper reports no check that its arousal output is meaningful for faces. Second, both k and the window are arbitrary, and no sensitivity analysis is reported. A 500 ms window can flag blinks, head movements, or speech-related facial motion as 'emotional reactions,' and the binary threshold turns any brief deviation into a reaction. The paper itself (Section 4.4) concedes that the causes of detected arousal remain unclear. Since p = .015 is not extremely robust, plausible changes in the detector definition or a failure to converge with the concurrently recorded HRV, SAM, or mDES measures would make the headline association a measurement artifact rather than evidence that arousal during explanation reduces understanding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an empirical HCI study (N=24) in which an embodied decision support system explains ten feature contributions to a risk estimate, and the authors test whether prior task-unrelated emotions (fear vs. happiness) and task-generated emotional reactions influence retention and understanding of those explanations. Emotional reactions are detected from EmoNet arousal using a rolling z-score threshold (k = 2.5, 500 ms window), while retention and understanding are measured via verbal recall and a direction-sorting task. The main reported findings are that prior emotion did not significantly affect retention (and only marginally affected understanding), and that detected emotional reactions were significantly negatively associated with understanding in a GLMM (Section 4.6, β = -0.45, SE = 0.19, p = .015), interpreted as arousal during explanation reducing comprehension. The paper also reports feature-level variation in arousal and acknowledges a systematic explanation error for one feature.","tokens_in":9173,"tokens_out":6115,"duration_ms":57526,"significance":"If the RQ3b effect is reliable, the paper makes a useful contribution to XAI evaluation by connecting affective state during explanation to comprehension outcomes, with a concrete implication that explanation systems may need to monitor and pace content in response to arousal. The study has genuine strengths: it separates prior task-unrelated emotions from task-related reactions, uses a mixed-effects model with a participant-level random intercept, and collects multimodal measures (HRV, facial data, self-report). However, the central result currently rests on an unvalidated arousal detector, so the contribution is conditional on resolving the measurement concerns below.","major_comments":[{"comment":"The task-related emotion detector is not validated. EmoNet, as cited in reference [6], is a speech emotion recognition model, but it is applied here to facial frames; no check is reported that its arousal output is meaningful for faces in this setup. Equations (1)-(2) introduce an arbitrary threshold k = 2.5 and a 500 ms rolling window, and no sensitivity analysis is provided. Since the headline RQ3b result (Section 4.6, β = -0.45, SE = 0.19, p = .015) is computed from this detector, plausible changes in k or the window, or contamination from blinks, head movements, or speech-related facial motion, could turn the association into a measurement artifact. The manuscript itself states in Section 4.4 that the causes of the detected arousal remain unclear. I request a sensitivity analysis over thresholds and window lengths, a validity check against concurrently recorded HRV, SAM, or human-coded video data, and a careful statement of what the binary reaction variable actually captures.","section":"Section 3.1.1 / 4.6"},{"comment":"The emotion manipulation check is descriptive only. Section 4.1 reports median SAM valence and arousal per condition and a 'trend' in the expected direction, but no inferential test is reported. The conclusions on RQ1a (null retention effect) and RQ1b (marginal understanding effect) depend on the induction having actually worked; without an inferential manipulation check (e.g., Mann-Whitney U or mixed ANOVA with effect sizes), a null or marginal result could simply reflect a failed or weak manipulation. Please add formal tests and effect sizes for the SAM data, and ideally for the mDES and HRV measures as well.","section":"Section 4.1"},{"comment":"The systematically erroneous explanation of the feature 'Einstellung bzgl. Zukunft' is acknowledged in Section 4.2 and shown to produce differential arousal between the Fear and Happy groups in Section 4.4, but this feature is retained in the GLMMs of Sections 4.5 and 4.6 without any control or exclusion. Because the error plausibly drives both emotional reactions and failures of understanding, the significant negative association between arousal and understanding may be confounded by this known stimulus artifact. I ask for a robustness analysis excluding this feature or including an error-flag covariate, and for reporting whether the arousal-understanding effect survives that analysis.","section":"Sections 4.2 / 4.4 / 4.6"},{"comment":"The definition of the emotional-reaction variable is inconsistent across sections. Section 4.4 states that a feature segment with one or more arousal bouts is counted as an emotional reaction, which is a binary variable; however, Figures 10 and 11 and the phrase 'number of emotional reactions' imply a count variable, and the model descriptions in Sections 4.5 and 4.6 do not state whether the fixed effect is binary or a count. The interpretation in Section 4.6 ('higher levels of positive emotional reactions') is also misleading because the detector captures arousal, not positive valence, and the variable as defined is binary. Please clarify the exact coding, provide the model formula, and align the wording with the actual variable definition.","section":"Sections 4.4 / 4.5 / 4.6"}],"minor_comments":[{"comment":"The abstract contains typos: 'ratantion' should be 'retention', and 'do not affected' should be 'did not affect'.","section":"Abstract"},{"comment":"'Hear Rate Variability' should be 'Heart Rate Variability'.","section":"Section 3.1.3"},{"comment":"The phrase 'cf. Fig. 3.4' should refer to the correct figure number (likely Figure 2).","section":"Section 3.4"},{"comment":"The caption placeholder 'Enter Caption' must be replaced with an actual descriptive caption.","section":"Figure 9"},{"comment":"The coefficient 'β = ˘2.31' appears to be a rendering error for 'β = -2.31'; please correct it.","section":"Section 4.6"},{"comment":"There are typographical issues such as 'ANOV A' and 'numberemotionalreactions' with missing spacing; a careful proofread is needed.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chris, quick read of Richter et al. What's genuinely new is the separation between task-unrelated prior emotions and explanation-generated arousal, measured via facial/HRV signals, with outcome variables being retention and understanding of XAI feature relevance. The central finding is that detected 'emotional reactions' during an explanation were significantly negatively associated with understanding of that feature (β=-0.45, SE=0.19, p=.015). That is a plausible and useful preliminary result for XAI design.\n\nThe paper does several things right. The study uses a concrete decision support system with a Holt-Laury risk task, applies GLMMs with random intercepts for the binary outcomes, and is admirably transparent about a programming error affecting one feature and about not knowing what causes the detected arousal. The writing is clear, the research questions are explicit, and the analysis is appropriate for the data size.\n\nThe soft spots are concentrated where the stress-test note points. The arousal detector uses EmoNet, which is a speech emotion recognition model, applied to facial frames. No evidence is given that its arousal output behaves sensibly on faces. The rolling z-score threshold (k=2.5, 500 ms window) is arbitrary and has no sensitivity analysis. At p=.015 the headline association is not robust enough to shrug off plausible detector variations. The emotion induction check is only descriptive (median comparisons, no inferential test), so the null effects for prior emotion could reflect a weak manipulation. The sample is small (24 participants, 15 fear / 9 happy) and imbalanced, and the confirmation-bias story for the wrong 'attitude towards future' feature is explicitly speculative. These are real limitations, not fatal ones.\n\nMy bottom line: this is a serious exploratory study worth sending to reviewers, but the central claim should be framed as preliminary. The authors need to validate the arousal detector or run sensitivity analyses, report an inferential manipulation check, and ideally increase the sample before anyone treats the negative association as established. I'd want a reviewer who knows affective computing and XAI, so the detector issue gets real scrutiny.","headline":"A plausible but measurement-sensitive result: explanation-triggered arousal (via an unvalidated EmoNet-on-faces detector) may reduce XAI understanding; worth reviewing as a preliminary study, not yet strong evidence.","tokens_in":9683,"tokens_out":3072,"would_cite":false,"duration_ms":29969,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Emotional arousal triggered during a feature's explanation is significantly associated with lower understanding of that feature's role in the AI's decision.","keywords":["explainable AI","emotional arousal","explanation retention","explanation understanding","decision support system","facial expression analysis","emotion induction","human-AI interaction"],"falsifier":"Re-run the analysis on the same 240 feature-explanation observations with arousal thresholds varied (for example, $k = 1.5$, $2.0$, $3.0$; window 250–1000 ms), or replace the facial arousal detector with HRV-derived or self-reported arousal for each feature. If the significant negative association between emotional reaction and understanding disappears, changes sign, or fails to replicate under any reasonable alternative detection setting, the claim that task-generated arousal hinders understanding would not survive.","tokens_in":8752,"feed_emoji":"📉","tokens_out":6340,"duration_ms":62138,"temperature":0.7,"pith_summary":"This paper asks whether emotions—both those a user brings into an interaction and those triggered by the explanation itself—change what people take away from an AI's feature-relevance explanations. In a study with 24 participants receiving risk-advice from an embodied agent, the authors find that task-unrelated induced emotions (fear vs happiness) did not affect how many explained features people could recall, and at most marginally affected their understanding of feature influence. The central result is that emotional arousal detected during a feature's explanation is significantly negatively associated with understanding of that feature's contribution: in a mixed model, $\\beta = -0.45$, $SE = 0.19$, $p = .015$. The authors interpret this as arousal during explanation hindering comprehension, and argue that XAI systems should seek the right level of arousal rather than assuming rational, emotion-free processing.","feed_headline":"Arousal during AI explanation predicts worse understanding","feed_subtitle":"In a risk-advice study, emotional spikes during feature explanations cut comprehension while recall held steady","key_machinery":"The load-bearing mechanism is a thresholded anomaly detector on arousal. For each facial frame, EmoNet supplies an arousal value; a rolling z-score with $k = 2.5$ over a 500 ms window (Equations 1 and 2) flags an emotional reaction whenever the current arousal exceeds the recent mean by 2.5 standard deviations. Each explained feature segment is then coded as having caused an emotional reaction or not, and that binary predictor enters a binomial GLMM (logit link) with emotion induction and explained feature as fixed effects and participant ID as random intercept. Retention is measured by verbal recall mapped to variable names; understanding is measured by whether the participant correctly sorts each feature as increasing or decreasing risk. This arrangement lets the authors separate prior emotion, content salience, and momentary arousal as predictors of the two outcomes.","core_discovery":"The paper's central claim is that task-generated emotional reactions, measured as brief arousal peaks during each feature's explanation, reduce the explainee's understanding of that feature's influence on the system's decision, while neither prior induced emotion nor arousal predicts whether the feature is verbally recalled. The evidence is a generalized linear mixed model on 240 feature-explanation observations: emotional reaction has a significant negative fixed effect on understanding ($\\beta = -0.45$, $SE = 0.19$, $p = .015$), with a positive intercept indicating generally good baseline understanding. The authors also report that some features (gender, current health status, political orientation) significantly predict retention, that feature categories differ in how often they trigger arousal, and that the wrongly explained 'attitude towards future' feature produced a suggestive emotion-congruent pattern in understanding: 80% of fearful participants accepted the system's wrong explanation versus 44% of happy participants.","pith_inferences":["Beyond the paper: the single wrongly explained feature functions as an accidental experiment—comparing understanding for correct versus erroneous explanations under fear versus happiness could isolate error detection from arousal, a contrast the paper notes but does not model.","Beyond the paper: because the arousal detector threshold ($k = 2.5$, 500 ms) was not varied, the reported coefficient could partly reflect measurement noise; testing the same GLMM with $k$ from 1.5 to 3.5 and with HRV-based or self-reported arousal would show whether the negative association is robust.","Beyond the paper: if the effect replicates, the practical target is not removing emotion from XAI but regulating arousal—explanations could be paced, chunked, or preceded by a calming step, with user state fed back into explanation generation."],"forward_implications":["If the central result is correct, XAI evaluation should treat recall and understanding as separate constructs, because arousal can lower comprehension without lowering memory.","Explanation interfaces that monitor arousal could adapt timing or presentation to keep the explainee in a comprehension-friendly arousal range.","Designers should expect some features (gender, health, political orientation) to be inherently more memorable and more arousal-inducing, so per-feature statistics are needed rather than aggregate explanation metrics.","Prior emotions may shape how users judge a feature's influence in an emotion-congruent way, so systems should verify understanding rather than relying on users' self-reports.","The absence of a retention effect suggests emotionally salient explanations do not improve memory, contrary to a common assumption that arousal aids encoding."],"supporting_citations":[{"why":"Supplies the arousal values from which emotional reactions are detected via rolling z-score.","marker":"[6]"},{"why":"Provides the same decision task and advice scheme (Flobi, Holt-and-Laury lottery) used in the study.","marker":"[10]"},{"why":"Supplies the biographic event-recall emotion induction and the earlier finding that prior emotions affect advice taking.","marker":"[11]"},{"why":"Supplies the Self-Assessment Manikin used for the emotion-induction manipulation check.","marker":"[3]"}],"fun_headline_variants":["Task-generated arousal during AI explanations hurts understanding, not recall","Emotional spikes while reading AI explanations predict worse comprehension","Arousal from AI feature explanations reduces understanding, leaves recall intact","Emotional arousal to AI explanations impairs understanding, spares recall"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The core result depends on treating a brief spike above a fixed threshold in a facial-expression arousal score as an emotional reaction; the paper does not validate that threshold or that the model, originally built for speech, works on faces, so the measured 'emotional reactions' could be an artifact of the detector.","fun_headline_variants_meta":{"raw":{"variants":["Task-generated arousal during AI explanations hurts understanding, not recall","Emotional spikes while reading AI explanations predict worse comprehension","Arousal from AI feature explanations reduces understanding, leaves recall intact","Emotional arousal to AI explanations impairs understanding, spares recall"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001228,"raw_usage":{"total_tokens":5058,"prompt_tokens":967,"completion_tokens":4091,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":4021}},"tokens_in":583,"tokens_out":4091,"duration_ms":30823,"temperature":1.0,"reasoning_tokens":4021,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:08:27.016400+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis on the same 240 feature-explanation observations with arousal thresholds varied (for example, $k = 1.5$, $2.0$, $3.0$; window 250–1000 ms), or replace the facial arousal detector with HRV-derived or self-reported arousal for each feature. If the significant negative association between emotional reaction and understanding disappears, changes sign, or fails to replicate under any reasonable alternative detection setting, the claim that task-generated arousal hinders understanding would not survive.","supporting_citations":[{"cited_title":"Gerczuk, S","cited_arxiv_id":null,"evidence_quote":"Supplies the arousal values from which emotional reactions are detected via rolling z-score."},{"cited_title":"Schütze, O","cited_arxiv_id":null,"evidence_quote":"Provides the same decision task and advice scheme (Flobi, Holt-and-Laury lottery) used in the study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Self-Assessment Manikin used for the emotion-induction manipulation check."}],"review_version":1}