{"id":"ccce904c-d5f3-4e3c-8aa1-10b296ba0432","arxiv_id":"2502.08920","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An emotion-sensitive LLM chatbot was rated as more trustworthy and competent than a neutral one, with no difference in problem-solving success.","lead":"Researchers tested a customer service chatbot that adjusts its tone to the user's emotions against a neutral one. Users rated the emotion-aware bot as more trustworthy and competent, even though both solved problems equally often.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The emotion-sensitive condition is likely confounded with response length and informativeness; the VADER-only manipulation check cannot rule this out.","rationale":"The reader's weakest assumption is that the VADER-based sentiment classification accurately captures emotional tone, so that the emotion-sensitive condition truly differs emotionally. My concern is distinct: even granting perfect VADER classification, the prompt manipulation is not isolated to emotional tone. The two conditions differ in the entire conversational style demanded by the system prompt, and the paper provides no check that the control responses are matched to the emotional responses on length, concreteness, or helpfulness. The central claim is about perceived competence and trust, and response substance is a direct alternative explanation for those ratings. The paper's own limitations section acknowledges VADER's coarseness and the small convenience sample, but it does not acknowledge this confound. Because the system prompts and interaction logs are not included, the confound cannot currently be ruled out. This does not overturn the paper's pilot-level finding, but it strengthens the need for replication with a cleaner manipulation. The reader's CONDITIONAL verdict already captures the need for more rigorous validation, so I recommend no change to the verdict. The concrete test of comparing response length and informativeness across conditions would settle whether the competence gap is really about emotion or about conversational substance.","tokens_in":5871,"tokens_out":4136,"duration_ms":46371,"concrete_test":"Ask the authors to release the exact system prompts and run both pipelines on a fixed set of, say, 40 scripted user messages covering the same emotion/issue combinations used in the study. Compare the two conditions on response length (tokens), number of sentences, number of task-relevant actions or steps offered, and reading ease. If the emotion-sensitive responses are systematically longer or more informative, then the perceived-competence difference in Section 4.3 cannot be attributed to emotional sensitivity alone, and the paper should be revised to report the confound or to match conditions on these dimensions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 describes two conditions that differ in system prompt: the emotion-sensitive bot selects one of three emotion-specified prompts based on VADER score, while the control is always instructed to be 'stoic and problem-focused.' The manipulation check in Section 4.1 confirms only that the sentiment polarity of the bot's output varies with user input (or stays near zero for the control). It does not establish that the two conditions are matched on response length, clarity, informativeness, or number of actionable troubleshooting steps. In customer-service text, a 'stoic and problem-focused' prompt can easily yield terse, unelaborated answers, whereas prompts that ask the model to acknowledge emotion tend to produce longer replies with empathetic framing. If the emotion-sensitive bot's answers were systematically longer or contained more troubleshooting detail, the higher ratings on 'capable of handling complex queries' and 'has the necessary knowledge' (Appendix A) would be explained by perceived effort or response substance, not by emotional sensitivity. This is a confound in the independent variable, not a statistics issue, and it survives even if VADER classifies sentiment perfectly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subjects experiment (N=30) comparing two LLM-based customer-service chatbots: an emotion-sensitive version whose system prompt is selected by VADER sentiment analysis of the user's input, and an emotion-insensitive control that is always instructed to be \"stoic and problem-focused.\" Participants interacted with one of the chatbots in an IT-support scenario and then rated their impressions. The authors report that the emotion-sensitive chatbot was rated higher on capability, knowledge, trustworthiness, willingness to use again, and supportiveness/understanding, while self-reported problem resolution rates and pre-post changes in PANAS negative affect did not differ between conditions. The paper concludes that emotional sensitivity enhances perceived trust and competence even though it does not affect objective issue resolution.","tokens_in":6061,"tokens_out":5605,"duration_ms":55611,"significance":"If the findings hold, they provide preliminary evidence that emotional sensitivity in LLM-based chatbots can improve users' trust and perceived competence without altering actual problem-solving performance, which is relevant both to affective computing and to customer-service applications. The study has several strengths: it uses a controlled manipulation of emotional sensitivity, a validated measure of affect (PANAS), an independent sentiment-analysis check on the bot's outputs, and a clear acknowledgment of its limitations and ethical considerations. The between-subjects design and the use of real LLM-based interactions are appropriate for the research question. However, the current evidence is weakened by a plausible confound between emotional sensitivity and response substance, by statistical reporting issues in the appendix, and by the very small convenience sample; these issues must be resolved before the central claim can be accepted as strong.","major_comments":[{"comment":"The two conditions differ not only in emotional sensitivity but also, very plausibly, in response length and informativeness. The emotion-sensitive condition selects among three emotion-specified system prompts, whereas the control is always instructed to be \"stoic and problem-focused\"; the manipulation check in Section 4.1 verifies only the sentiment polarity of the bot's responses via VADER. It does not establish that the two conditions produce responses matched on length, clarity, number of actionable troubleshooting steps, or other substance-related properties. Because Appendix A asks participants to rate the chatbot's capability and knowledge, higher ratings in the emotion-sensitive condition could reflect perceived effort or response detail rather than emotional sensitivity. The authors should report response-level metrics (e.g., token counts, number of steps, or a content analysis) or add a control condition that holds informativeness constant while varying only emotional tone.","section":"3.1, Figure 1, 4.1"},{"comment":"The statistical reporting is internally inconsistent. The item \"I feel that the chatbot was supportive and understanding\" has p = 0.053, which is not significant at the conventional 0.05 threshold, yet Section 4.3 states that \"significant differences\" were found on all listed items. Additionally, five separate ANOVAs are conducted without any correction for multiple comparisons; at a Bonferroni-corrected alpha of 0.01, only the trust item (p = 0.007) clearly survives, with the knowledge item (p = 0.010) borderline. Finally, the standard deviations reported for \"I would use this chatbot again\" (Std = 0.280 and 0.262) are implausibly low for 5-point Likert items with means around 4.3 and 3.3, and they more closely resemble standard errors; if so, the table's statistics for that item need to be corrected. These issues do not necessarily invalidate the overall competence/trust findings, but they must be corrected before the pattern of results can be taken at face value.","section":"Appendix A, 4.3"},{"comment":"The claim that emotional sensitivity does not affect issue resolution is stated as an equivalence (χ2(1, N=30)=0.268, p=0.605), but a null result from a sample of 30 has low statistical power and cannot support the conclusion that the two chatbots are equally effective. The authors should either report a power analysis or a confidence interval for the resolution-rate difference, and they should temper the wording in the abstract and conclusion from \"did not affect\" to something like \"no significant difference was detected.\"","section":"4.2, Abstract, 5"}],"minor_comments":[{"comment":"The word \"addiitonal\" should be \"additional.\"","section":"3.2"},{"comment":"The paper consistently writes \"V ADER\" with a space; this appears to be a rendering issue and should be corrected to \"VADER.\"","section":"Throughout"},{"comment":"The heatmaps lack axis labels and a legend for the color scale; in grayscale it is impossible to infer the distributions shown, so the figure should be made self-explanatory or be supplemented with descriptive statistics.","section":"Figure 2"},{"comment":"The phrase \"emotional and unemotional\" is inconsistent with the rest of the paper; use \"emotion-sensitive\" and \"emotion-insensitive\" for clarity.","section":"4.3"},{"comment":"The reference for Devlin (2018) is missing the co-authors and should be updated to the full citation.","section":"References"},{"comment":"The exact system prompts used for the three emotion conditions are not provided; including them in an appendix would substantially improve reproducibility.","section":"3.1"}],"recommendation":"major_revision","confidential_remarks":"This is a small pilot study from a class project, and the authors are transparent about that fact. The topic is of interest to the human-computer interaction and affective computing communities, and the core idea is compelling. My recommendation of major revision is driven by the confound in the experimental manipulation (emotional tone vs. response substance) and by the statistical reporting issues in Appendix A. With careful revisions—such as reporting response-level properties, correcting the statistical inconsistencies, and tempering the equivalence claim—the paper could become a useful empirical contribution. The editors may also wish to consider whether the venue's expectations for sample size and statistical rigor are compatible with a pilot study of this scale."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read for you on arXiv:2502.08920. It is a small pilot, clearly labeled as a class project, comparing an LLM customer-service chatbot that adapts its tone to VADER-detected user sentiment against a stoic control. The result: no difference in problem resolution, but the emotion-sensitive bot is rated higher on trust, knowledge, capability, and willingness to use again. That is a modest, plausible extension of emotional-labor work to LLM agents, and it is honestly reported as preliminary.\n\nWhat it does well: the writing is candid about the tiny convenience sample, the ethics statement is genuinely careful about manipulation risks, and the basic finding is consistent with prior work on emotional labor. The manipulation check at least confirms the bot outputs differed in sentiment polarity as intended. For a pilot, it is a clean enough setup.\n\nSoft spots, in order of severity. First, the confound the stress-test note flags is real and under-discussed: the two conditions differ in more than emotional tone. A \"stoic and problem-focused\" prompt invites terse, unelaborated answers; an emotion-aware prompt tends to produce longer, more informative responses. The perception items — \"capable of handling complex queries,\" \"has the necessary knowledge\" — could easily be driven by response length or perceived effort rather than emotional sensitivity. The VADER check confirms sentiment, not informational equivalence. That is not a fatal flaw for a pilot, but it should be named as the main threat to the causal claim.\n\nSecond, the statistics are slightly overclaimed. One headline item, \"supportive and understanding,\" is reported at p = 0.053 and still called significant in the appendix pattern. With five ANOVAs and no correction, that one should be reported as marginal. Third, no code, data, or prompts are provided, which limits what a replication-minded reader can do.\n\nThe reader's take is about right: conditional, plausible, not paradigm-shifting. I would add that the confound is more important than the multiple-comparisons issue, because a confound is a design problem, not a reporting fix. The paper itself acknowledges the sample and VADER limitations, so the authors aren't hiding the obvious.\n\nWho is this for? Practitioners building customer-service bots who want a sanity check that emotional tone matters for perceptions, and researchers working on affective computing who need a strawman or a pilot citation. It is not a paper that settles anything, but it is an honest, readable pilot that deserves a serious referee — an editor should send it out rather than desk-reject, with the expectation of heavy revision: larger sample, matched response length, preregistered analysis, and dropping the p = 0.053 claim.\n\nYes, bring it to reading group as a useful example of how to report a small pilot with clear limitations, and yes, I'd cite it as initial evidence that emotional-tone adaptations shift perceived competence without moving objective outcomes. Treat it as suggestive, not confirmatory.","headline":"A tiny, honest pilot study showing a plausible trust/competence boost from emotion-sensitive LLM chatbots, but the manipulation is confounded with response substance and the stats are overclaimed.","tokens_in":6565,"tokens_out":723,"would_cite":true,"duration_ms":9507,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An emotion-sensitive LLM chatbot is rated as more trustworthy, knowledgeable, and competent than an emotion-insensitive one, even though both versions resolve users' issues at the same rate.","keywords":["conversational AI","customer service chatbots","emotional sensitivity","large language models","sentiment analysis","VADER","user trust","perceived competence"],"falsifier":"A direct test would run the same two-condition design with a larger, more representative sample and a validated sentiment classifier; the central claim would fail if the emotion-sensitive bot no longer gets significantly higher trust and competence ratings, or if a manipulation check shows its responses are not reliably more emotionally variable than the stoic bot's. A second test would tell participants that the emotional responses are scripted: if the rating gap disappears, the effect is driven by perceived authenticity rather than emotional sensitivity itself.","tokens_in":5699,"feed_emoji":"🤖","tokens_out":7255,"duration_ms":67418,"temperature":0.7,"pith_summary":"This paper asks whether making a customer-service chatbot emotionally sensitive changes how users perceive it, and it answers yes for perceptions but not for outcomes. In a between-subjects experiment with 30 participants, an LLM-based chatbot that adjusted its tone to the user's sentiment was rated significantly higher on capability, knowledge, trustworthiness, supportiveness, and willingness to reuse, compared with an otherwise identical stoic chatbot. Problem-resolution rates were the same across conditions, and both chatbots reduced negative affect; the difference appeared only in user impressions. The authors take this as evidence that emotional sensitivity is a distinct pathway to customer satisfaction and raise the question of whether users are overtrusting emotionally expressive systems.","feed_headline":"Emotion-sensitive chatbots win trust without fixing more issues","feed_subtitle":"Users rated the empathetic bot more competent and trustworthy, even though both bots resolved problems equally often.","key_machinery":"The experimental mechanism is a prompt-selection loop built on VADER, a rule-based sentiment model. Each user message is scored; scores beyond ±0.1 from zero are labeled positive or negative, and this label selects a system prompt that tells ChatGPT-3.5 which emotional tone to use, while the control condition always receives the stoic, problem-focused prompt. The same VADER model then scores the bot's replies to verify the manipulation worked. This isolates emotional tone as the only difference between the two conditions, leaving the underlying LLM and problem-solving logic identical.","core_discovery":"The paper's central claim is that, in an IT customer-service context, users judge an emotionally sensitive LLM chatbot as more competent and trustworthy than an emotionally neutral one, despite no difference in how often problems are actually solved. The statistical results show significantly higher agreement for the emotional bot on 'capable of handling complex queries', 'has necessary knowledge', 'trust to assist', 'would use again', and 'supportive and understanding' (the last at p = 0.053). The paper also reports that negative emotion decreased after either interaction, so the benefit is not better mood repair; it is perception. The authors interpret this through emotional labor theory: meeting customers' emotional needs during service increases their trust and satisfaction, even when the functional outcome is unchanged.","pith_inferences":["A pre-registered replication with a representative sample and a conversation-level sentiment model would test whether the effect size survives; the current convenience sample of friends and family makes the magnitude uncertain.","The paper's design leaves open whether the trust gain would persist if users were told the emotions are generated by an algorithm; the authors' contrast with the backfire effect suggests this is the key boundary condition.","In real high-stakes support, inflated perceived competence could lead users to follow a bot's advice on serious issues without independent verification; measuring trust calibration directly would be a natural next step.","Using a stronger affect model than VADER might produce larger or cleaner differences, since rule-based sentiment classification can miss ironic or context-dependent emotion; that is an engineering extension the paper hints at but does not test."],"forward_implications":["Customer-service chatbots can raise perceived competence and trust without changing the underlying problem-solving system, so emotional tone is a distinct design lever.","Adding emotional sensitivity does not by itself improve or harm resolution rates, so it should complement, not replace, functional improvements.","Users may overtrust emotionally expressive bots: confidence in capability can rise while actual capability stays flat, so trust calibration becomes a safety issue.","Since both bots reduced negative affect equally, emotional sensitivity seems to pay off in evaluation rather than in mood repair, at least in short simulated interactions.","Qualitative responses suggest users experience the emotional bot as more personable ('friendly', 'great personality'), pointing to social presence as the mechanism behind the ratings."],"supporting_citations":[{"why":"Supplies the VADER sentiment model used to classify user input and to verify that bot responses match the intended emotional tone.","marker":"Hutto and Gilbert, 2014"},{"why":"Supplies the PANAS scale used to measure participants' positive and negative affect before and after the chatbot interaction.","marker":"Watson et al., 1988"},{"why":"Provides the emotional labor theory that motivates the prediction that meeting customers' emotional needs increases trust and satisfaction.","marker":"Hochschild, 1979"},{"why":"Documents the contrasting result that emotion expression by AI agents can backfire when users know it is generated by AI.","marker":"Han et al., 2023"},{"why":"Supplies the customer-emotion response tactics that the emotion-sensitive system prompt instructs the LLM to follow.","marker":"Magids et al., 2015"},{"why":"Supports the interpretation that human-like interaction qualities increase user acceptance of chatbots.","marker":"Rapp et al., 2021"},{"why":"Raises the overtrust concern the discussion applies to emotionally expressive chatbots.","marker":"Hancock et al., 2020"}],"fun_headline_variants":["Empathetic chatbots win trust, but not more fixes","Emotion-sensitive AI: trusted more, resolves same","Users see empathetic bots as more competent, not more effective","Emotion-aware chatbots boost perceived competence, not outcomes","Feeling bots: higher trust, unchanged issue resolution"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that VADER's ±0.1-neutrality threshold correctly identifies the emotional tone of user messages, so the emotion-sensitive chatbot actually receives and delivers distinct emotional prompts rather than responding to noise.","fun_headline_variants_meta":{"raw":{"variants":["Empathetic chatbots win trust, but not more fixes","Emotion-sensitive AI: trusted more, resolves same","Users see empathetic bots as more competent, not more effective","Emotion-aware chatbots boost perceived competence, not outcomes","Feeling bots: higher trust, unchanged issue resolution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1295,"prompt_tokens":797,"completion_tokens":498,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":413,"completion_tokens_details":{"reasoning_tokens":420}},"tokens_in":413,"tokens_out":498,"duration_ms":5416,"temperature":1.0,"reasoning_tokens":420,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T23:11:56.415627+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would run the same two-condition design with a larger, more representative sample and a validated sentiment classifier; the central claim would fail if the emotion-sensitive bot no longer gets significantly higher trust and competence ratings, or if a manipulation check shows its responses are not reliably more emotionally variable than the stoic bot's. A second test would tell participants that the emotional responses are scripted: if the rating gap disappears, the effect is driven by perceived authenticity rather than emotional sensitivity itself.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the PANAS scale used to measure participants' positive and negative affect before and after the chatbot interaction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the customer-emotion response tactics that the emotion-sensitive system prompt instructs the LLM to follow."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the interpretation that human-like interaction qualities increase user acceptance of chatbots."}],"review_version":1}