{"id":"aafa1d5c-45ab-464d-8b6e-8f9bf510d5cd","arxiv_id":"2411.12877","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In chat conversations, users rate AI chatbots as less empathetic than humans but still give the chatbot conversations higher quality ratings.","lead":"This study analyzed 155 chat conversations and found that people rated AI chatbots as less empathetic than human partners, even though they rated the chatbot conversations higher in overall quality. The result matters because it suggests that making chatbots sound empathetic is not enough, users' perceptions and expectations shape whether empathy actually lands.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Identity cue confound remains the decisive issue: since all chatbot turns were labeled 'Bot:', the empathy gap may be driven by labeling rather than language; the paper's own blind third-party annotations (p=0.07) fail to show an empathy gap.","rationale":"The reader's weakest assumption correctly identifies the identity cue as the load-bearing threat to the central claim. My independent reading of the paper and its supplement strengthens that concern: the paper's own third-party turn-level annotations, which were blind to identity, found no significant empathy gap between humans and chatbots (p=0.07, with bots slightly higher), whereas user self-reports collected with visible 'Bot:' labels showed a large, robust gap. This asymmetry is exactly what would be expected if the label itself, rather than the language, drives perceived empathy. The paper explicitly acknowledges the design limitation in its Limitations section, so the concern is not hidden, but acknowledging it does not remove the risk to the causal conclusion. The secondary issue about conversation quality being measured only in WASSA 2024 is also real but less central, because the main empathy claim does not depend on that measure. Given the direct self-report data and the plausible real-world relevance, the reader's CONDITIONAL verdict remains appropriate: the findings should be accepted only if framed as perception differences in a labeled interaction, not as evidence that empathetic language is insufficient. I therefore recommend no change to the reader's verdict.","tokens_in":14463,"tokens_out":4085,"duration_ms":44995,"concrete_test":"Run a follow-up controlled experiment using identical chatbot-generated transcripts (or the existing WASSA transcripts) with identity labels manipulated: present the same conversation to three groups labeled 'Bot:', 'Human:', or no label, holding language constant. If empathy ratings differ significantly across label conditions for identical language, the observed empathy gap is attributable to the identity cue rather than to chatbot language, and the central conclusion needs to be reframed. If no label effect appears, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that chatbots are perceived as less empathetic than humans despite using comparable empathetic language, so 'more than simply embedding empathetic language' is needed. The weakest link is the inseparable identity cue. In every human-chatbot conversation, participants saw bot utterances prefixed with 'Bot:' (Data section). Participants thus knew the partner was a machine, and prior work cited in the paper shows such identity cues alter judgments. The paper's own supplementary third-party annotations, where annotators were not told identity, found no significant empathy difference between humans and chatbots; turn-level empathy means were 2.36 vs 2.74, t=1.82, p=0.07, with bots slightly higher. If the identity label alone causes lower ratings, the data do not support the language-insufficiency conclusion. The Limitations section admits 'absence of a participant group that was unaware they were interacting with a chatbot' and that the impact of identity awareness cannot be directly assessed. This acknowledged limitation is load-bearing because the main user-rating result cannot be separated from labeling.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 155 crowd-worker conversations from the WASSA 2023, WASSA 2024, and Empathic Conversations datasets to compare how users perceive empathy and conversation quality in human-human versus human-chatbot interactions. The authors report that GPT-based chatbots were rated significantly lower than humans on general empathy and on affective, cognitive, and associative state empathy, while being rated higher on overall conversation quality. These user self-reports are supplemented by GPT-4o annotations, a model trained on the Empathic Conversations dataset, and four off-the-shelf empathy models. The paper concludes that achieving high-quality human-AI interaction requires more than embedding empathetic language, because users perceive chatbot empathy as lower despite comparable empathetic language.","tokens_in":14712,"tokens_out":5362,"duration_ms":55050,"significance":"The paper's main strength is its user-centered measurement: empathy ratings come directly from conversation participants using established psychological scales, and the lower-empathy finding is consistent across several dimensions and replicates in a within-subject subset. The authors also provide their code and prompts, which supports reproducibility. If the result holds, it is a useful corrective to studies that rely on third-party annotators and may show chatbots as more empathetic than humans. However, the significance of the central interpretive claim is currently limited by a visible identity cue that is inseparable from the chatbot's language, a problem the authors themselves acknowledge. The paper is valuable as a descriptive comparison of user perceptions when identity is known, but the stronger conclusion about 'more than simply embedding empathetic language' needs either new evidence or a substantial narrowing.","major_comments":[{"comment":"The identity-cue confound is load-bearing for the paper's main interpretive claim. The Data section states that bot utterances began with 'Bot:' and that participants were not explicitly told they were interacting with a chatbot, yet the visual cue indicated the presence of a bot. The Limitations section then concedes that the impact of chatbot identity awareness 'cannot be directly assessed.' Because the abstract concludes that high-quality human-AI interaction 'requires more than simply embedding empathetic language,' the user-rating results must be separable from the label. They are not. Supplementary S1 is directly relevant here: third-party annotators, who were blind to identity, found no significant empathy difference between humans and chatbots (t=1.82, p=0.07), with chatbot turns slightly higher on average. This is not a peripheral caveat; it removes key support for the 'language insufficiency' conclusion. I recommend either adding a condition without the 'Bot:' label, or substantially qualifying the central claim to 'when users are aware that they are interacting with a chatbot.'","section":"Data (Chatbot Implementation) and Limitations"},{"comment":"The claim that the lower-empathy finding is 'supported by ... a pre-trained empathy model' is inconsistent with Table 2. The perceived empathy model trained by the authors shows no human-chatbot difference (t=0.45, p=0.65), and the paper itself states that 'the empathetic language of humans and the empathetic language of bots are equivalent.' Among the off-the-shelf pre-trained models, the Interpret model favors humans while the Emo-React model favors chatbots. Only the GPT-4o annotations support the direction of the empathy gap. Please either remove 'pre-trained empathy model' from the supporting evidence in the Abstract and Conclusion, or identify which specific model is meant and explain why that model's direction is credited while the mixed and null results from the other models are not.","section":"Abstract and Discussion"},{"comment":"The GPT-4o annotation results are over-interpreted as an independent confirmation. Although GPT-4o was not given explicit identity labels, Supplementary S2 shows that GPT-4o correctly inferred speaker identity with 61% accuracy (F1=0.60), so the annotation contrast is not a clean blind comparison of language; it is partly a second opinion from a model that can detect LLM-like patterns. In addition, the per-conversation correlations with user ratings are weak (r=0.20 for humans, r=0.06 for chatbots, r=0.07 combined), so the statement that GPT-4o ratings 'aligned with user ratings' is true only at the level of aggregate means. The text should state both qualifications wherever the LLM-judge results are presented.","section":"Experiments and Results (LLM Judgement of Perceived Empathy) and Supplementary S2"},{"comment":"The conversation-quality conclusion is based on a single 5-point Likert item collected only in WASSA 2024, yet the paper states it as a general finding ('Chatbots receive higher ratings for conversation quality'). The Results and Conclusion should prominently state this dataset restriction and the single-item nature of the quality measure. The mixed-model analysis should also be presented with the number of observations per participant, since a random intercept for participant ID may be weakly identified if most participants contributed only one quality rating. Without this context, readers cannot assess the robustness of the claimed quality advantage.","section":"Experiments and Results (Psychological Ratings)"}],"minor_comments":[{"comment":"The abbreviation 'WASSA' appears with extra spaces in several places (e.g., 'W ASSA 2023'); please format it consistently.","section":"Data"},{"comment":"Please report effect sizes and confidence intervals for the t-tests in Table 2, since several p-values are close to conventional thresholds and the paper intentionally avoids effect-size thresholds.","section":"Psychological Ratings"},{"comment":"The text reports the interaction coefficient for cognitive state empathy as β=0.77 without specifying a sign; from the prose and Figure 2 it is unclear whether this is a positive or negative interaction. Please provide a full model table with standard errors.","section":"Experiments and Results (Psychological Ratings)"},{"comment":"The GPT-4o annotation prompt and rating scale are not described in the main text; please add them so readers know whether the 1-7 scale or another scale was used.","section":"Experiments and Results (LLM Judgement of Perceived Empathy)"},{"comment":"The sentence 'we intentionally avoided setting arbitrary thresholds for effect sizes' is confusing: reporting effect sizes does not require setting thresholds. Please clarify or remove this sentence.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid descriptive finding that users rate chatbots as less empathetic when they can see the 'Bot:' label, but the central interpretive claim about empathetic language being insufficient is not supported by the blind third-party annotations in S1. I believe the paper is salvageable by substantially qualifying the claims to the known-identity setting, or by adding an experiment that removes or manipulates the identity cue. As it stands, the gap between the abstract's causal-sounding conclusion and the acknowledged confound is too large for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline is that this is a solid, useful user-centered study with a real split result: users rate chatbot conversations higher in quality but lower in empathy. That pattern is new relative to the third-party annotation literature. The paper does several things right: four self-report empathy dimensions, a within-subject replication, blind GPT-4o annotations, and it ships code and data. The lower-empathy finding replicates internally.\n\nThe soft spot is the identity confound, and it is exactly where the stress-test note lands. Every chatbot turn was visibly labeled 'Bot:' so participants knew the partner was a machine. The paper's own supplementary third-party annotations, where annotators were blind to identity, found no significant empathy gap (t=1.82, p=0.07), with bots slightly higher. That matters because the abstract's stronger claim—that empathetic language is not enough—rests on the gap being about language, not just the label. The GPT-4o blind judge did find a gap, and the authors note it may have detected LLM patterns, so there is some blind evidence, but it is weaker and not from naive users. The limitation section acknowledges the missing unaware group honestly.\n\nOther smaller gaps: effect sizes are not reported; the conversation quality measure is a single 5-point item from WASSA 2024 only; and the pre-trained empathy model results are mixed, with some models showing humans lower and others higher. None of these kill the paper, but they temper the general claim.\n\nFor whom: this is useful for affective computing and human-AI interaction design, and for anyone working on evaluation methodology for chatbots. It deserves a serious referee: the data collection is real, the split result is interesting, and the limitations are acknowledged. I'd send it to review, with a request for effect sizes, a sensitivity analysis around the identity cue if possible, and a more careful wording of the conclusion.","headline":"A genuinely user-centered empathy comparison that splits quality from empathy, but the identity-label confound means the 'language insufficiency' conclusion is not yet proven.","tokens_in":15201,"tokens_out":1912,"would_cite":true,"duration_ms":18901,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Users rate chatbots as less empathetic than humans, even when they rate the conversation higher in quality.","keywords":["perceived empathy","chatbot","conversation quality","user-centered evaluation","state empathy","large language models","human-AI interaction","empathy perception"],"falsifier":"Present the exact same chatbot utterances to a new group of participants without any bot label (or with a false human label) and compare empathy ratings to the labeled condition. If the empathy gap disappears or reverses, the paper's conclusion that empathetic language is insufficient would be undercut; if the gap persists, the conclusion survives.","tokens_in":14312,"feed_emoji":"💬","tokens_out":4925,"duration_ms":43083,"temperature":0.7,"pith_summary":"The paper tries to establish that, from the user's point of view, empathy and conversation quality are not the same thing, and that current chatbots get them out of sync. Analyzing 155 crowd-worker conversations, the authors find that GPT-based chatbots are rated significantly higher in overall conversation quality than human partners, while being rated significantly lower on four self-report empathy measures. The same pattern appears when a large language model (GPT-4o) rates the conversations blind, and partially in one pre-trained empathy model, but not in models trained on human-human text. The authors argue this means high-quality human-AI interaction requires more than embedding empathetic language; it requires addressing how users interpret and experience empathy with a machine partner.","feed_headline":"Users find chatbots less empathetic yet rate chats higher","feed_subtitle":"Across 155 conversations, even GPT-based bots praised for smoothness were judged as worse at empathy than humans.","key_machinery":"The argument is carried by a multi-method comparison of perceived empathy. The primary instrument is state empathy, a transactional, three-component construct (affective, cognitive, associative) measured by six self-report questions, alongside a single-item general empathy rating and a closeness scale. These user ratings are cross-checked against three complementary automated views: GPT-4o annotations of whole conversations, a unigram-based regression model trained on human-human conversations, and four pre-trained empathy models from prior work. The divergence between what these measures see is itself the mechanism: user and LLM ratings detect the empathy gap, while models trained on human-human text do not, isolating the gap as a matter of perception rather than measurable language.","core_discovery":"The central discovery is a perception gap: users consistently rate chatbots as less empathetic than humans across general empathy, overall state empathy, and the affective, cognitive, and associative components of state empathy, while simultaneously rating chatbot conversations as higher in quality. Statistical models show that perceived empathy improves conversation quality for both partner types, but the association is stronger for humans; chatbots receive high quality ratings even when perceived empathy is low to moderate. The authors interpret this as evidence that users adjust their expectations for machines, and that a chatbot's empathetic language does not translate into perceived empathy.","pith_inferences":["A blind condition—hiding whether the partner is a bot—would test whether the visible 'Bot:' label alone drives the empathy gap; if it does, the claim that empathetic language is insufficient would need qualification.","The results imply that chatbot designers should target user expectations and identity signaling (e.g., framing, disclosure, personality cues) as much as response wording.","The finding that third-party annotators and text-trained models do not see the gap suggests future empathy benchmarks for AI should include the user's own perspective rather than only expert labels.","If the empathy gap comes from identity rather than language, then chatbots that convincingly mimic human identity—or are presented without disclosure—might close the gap, raising ethical questions about deception."],"forward_implications":["If users perceive chatbots as less empathetic despite empathetic wording, then adding more empathetic phrases to a chatbot will not by itself close the empathy gap.","Conversation quality ratings should not be read as proxies for empathy: a chatbot can be judged a better conversation partner while being judged a less empathetic one.","Third-party or offline empathy evaluations (human annotators or text-trained models) can miss the user's experience, so user-centered self-reports remain the essential measure for this question.","The weaker human-chatbot correlation between predicted and perceived empathy suggests that language-based empathy detectors need to be recalibrated for human-AI interactions."],"supporting_citations":[{"why":"Supplies the Empathic Conversations dataset of human-human text-based conversations, used to train the perceived-empathy language model.","marker":"Omitaomu et al. 2022"},{"why":"Supplies the 2023 shared-task dataset whose human-chatbot conversations are compared against human-human ones.","marker":"Barriere et al. 2023"},{"why":"Supplies the 2024 shared-task dataset, the only one with overall conversation-quality ratings, and the chatbot implementation for that year.","marker":"Giorgi et al. 2024"},{"why":"Provides the six-question state-empathy scale (affective, cognitive, associative) the paper revises for self-reports.","marker":"Shen 2010"},{"why":"Provides the classic empathy/distress distinction that underlies the general-empathy rating items.","marker":"Batson, Fultz, and Schoenrade 1987"},{"why":"Supplies the pre-trained emotional-reaction, interpretation, and exploration empathy models applied to conversation turns.","marker":"Sharma et al. 2020"},{"why":"Supplies the pre-trained Batson-empathy model used for turn-level empathy estimates.","marker":"Lahnala, Welch, and Flek 2022"},{"why":"Provides the contrasting third-party finding that chatbot responses were rated far more empathetic than physician responses, which this user-centered study says misses direct user perception.","marker":"Ayers et al. 2023"}],"fun_headline_variants":["Bots win quality but lose empathy in user eyes","Chatbots: High marks for talk, low marks for empathy","Empathy gap: Users prefer chatbots' conversation quality","Users rate AI chats higher, but find bots less empathetic","Empathetic chatbot words don't fool users, study finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The key assumption is that participants' lower empathy ratings for chatbots stem from the chatbot's actual conversational behavior, not just from the visible 'Bot:' label that told them they were talking to a machine.","fun_headline_variants_meta":{"raw":{"variants":["Bots win quality but lose empathy in user eyes","Chatbots: High marks for talk, low marks for empathy","Empathy gap: Users prefer chatbots' conversation quality","Users rate AI chats higher, but find bots less empathetic","Empathetic chatbot words don't fool users, study finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1256,"prompt_tokens":772,"completion_tokens":484,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":388,"completion_tokens_details":{"reasoning_tokens":403}},"tokens_in":388,"tokens_out":484,"duration_ms":5061,"temperature":1.0,"reasoning_tokens":403,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:05:12.534395+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Present the exact same chatbot utterances to a new group of participants without any bot label (or with a false human label) and compare empathy ratings to the labeled condition. If the empathy gap disappears or reverses, the paper's conclusion that empathetic language is insufficient would be undercut; if the gap persists, the conclusion survives.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 2023 shared-task dataset whose human-chatbot conversations are compared against human-human ones."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 2024 shared-task dataset, the only one with overall conversation-quality ratings, and the chatbot implementation for that year."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the six-question state-empathy scale (affective, cognitive, associative) the paper revises for self-reports."},{"cited_title":"D.; Fultz, J.; and Schoenrade, P","cited_arxiv_id":null,"evidence_quote":"Provides the classic empathy/distress distinction that underlies the general-empathy rating items."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the pre-trained Batson-empathy model used for turn-level empathy estimates."}],"review_version":1}