REVIEW 4 major objections 5 minor 56 references
The Illusion of Empathy: How AI Chatbots Shape Conversation Perception
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Users rate chatbots as less empathetic than humans, even when they rate the conversation higher in quality.
desk verdict A genuinely user-centered empathy comparison that splits quality from empathy, but the identity-label confound means the 'language insufficiency' conclusion is not yet proven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a multi-method comparison of perceived empathy. The primary instrument is state empathy, a transactional, three-component construct (affective, cognitive, associative) measured by six self-report questions, alongside a single-item general empathy rating and a closeness scale. These user ratings are cross-checked against three complementary automated views: GPT-4o annotations of whole conversations, a unigram-based regression model trained on human-human conversations, and four pre-trained empathy models from prior work. The divergence between what these measures see is itself the mechanism: user and LLM ratings detect the empathy gap, while models trained on human-human text do not, isolating the gap as a matter of perception rather than measurable language.
What would settle it
Present the exact same chatbot utterances to a new group of participants without any bot label (or with a false human label) and compare empathy ratings to the labeled condition. If the empathy gap disappears or reverses, the paper's conclusion that empathetic language is insufficient would be undercut; if the gap persists, the conclusion survives.
Extended reading notes
Core claim
The central discovery is a perception gap: users consistently rate chatbots as less empathetic than humans across general empathy, overall state empathy, and the affective, cognitive, and associative components of state empathy, while simultaneously rating chatbot conversations as higher in quality. Statistical models show that perceived empathy improves conversation quality for both partner types, but the association is stronger for humans; chatbots receive high quality ratings even when perceived empathy is low to moderate. The authors interpret this as evidence that users adjust their expectations for machines, and that a chatbot's empathetic language does not translate into perceived empathy.
Load-bearing premise
The key assumption is that participants' lower empathy ratings for chatbots stem from the chatbot's actual conversational behavior, not just from the visible 'Bot:' label that told them they were talking to a machine.
Editorial extensions
If this is right
- If users perceive chatbots as less empathetic despite empathetic wording, then adding more empathetic phrases to a chatbot will not by itself close the empathy gap.
- Conversation quality ratings should not be read as proxies for empathy: a chatbot can be judged a better conversation partner while being judged a less empathetic one.
- Third-party or offline empathy evaluations (human annotators or text-trained models) can miss the user's experience, so user-centered self-reports remain the essential measure for this question.
- The weaker human-chatbot correlation between predicted and perceived empathy suggests that language-based empathy detectors need to be recalibrated for human-AI interactions.
Reading between the lines
- A blind condition—hiding whether the partner is a bot—would test whether the visible 'Bot:' label alone drives the empathy gap; if it does, the claim that empathetic language is insufficient would need qualification.
- The results imply that chatbot designers should target user expectations and identity signaling (e.g., framing, disclosure, personality cues) as much as response wording.
- The finding that third-party annotators and text-trained models do not see the gap suggests future empathy benchmarks for AI should include the user's own perspective rather than only expert labels.
- If the empathy gap comes from identity rather than language, then chatbots that convincingly mimic human identity—or are presented without disclosure—might close the gap, raising ethical questions about deception.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes 155 crowd-worker conversations from the WASSA 2023, WASSA 2024, and Empathic Conversations datasets to compare how users perceive empathy and conversation quality in human-human versus human-chatbot interactions. The authors report that GPT-based chatbots were rated significantly lower than humans on general empathy and on affective, cognitive, and associative state empathy, while being rated higher on overall conversation quality. These user self-reports are supplemented by GPT-4o annotations, a model trained on the Empathic Conversations dataset, and four off-the-shelf empathy models. The paper concludes that achieving high-quality human-AI interaction requires more than embedding empathetic language, because users perceive chatbot empathy as lower despite comparable empathetic language.
Significance. The paper's main strength is its user-centered measurement: empathy ratings come directly from conversation participants using established psychological scales, and the lower-empathy finding is consistent across several dimensions and replicates in a within-subject subset. The authors also provide their code and prompts, which supports reproducibility. If the result holds, it is a useful corrective to studies that rely on third-party annotators and may show chatbots as more empathetic than humans. However, the significance of the central interpretive claim is currently limited by a visible identity cue that is inseparable from the chatbot's language, a problem the authors themselves acknowledge. The paper is valuable as a descriptive comparison of user perceptions when identity is known, but the stronger conclusion about 'more than simply embedding empathetic language' needs either new evidence or a substantial narrowing.
major comments (4)
- [Data (Chatbot Implementation) and Limitations] The identity-cue confound is load-bearing for the paper's main interpretive claim. The Data section states that bot utterances began with 'Bot:' and that participants were not explicitly told they were interacting with a chatbot, yet the visual cue indicated the presence of a bot. The Limitations section then concedes that the impact of chatbot identity awareness 'cannot be directly assessed.' Because the abstract concludes that high-quality human-AI interaction 'requires more than simply embedding empathetic language,' the user-rating results must be separable from the label. They are not. Supplementary S1 is directly relevant here: third-party annotators, who were blind to identity, found no significant empathy difference between humans and chatbots (t=1.82, p=0.07), with chatbot turns slightly higher on average. This is not a peripheral caveat; it removes key support for the 'language insufficiency' conclusion. I recommend either adding a condition without the 'Bot:' label, or substantially qualifying the central claim to 'when users are aware that they are interacting with a chatbot.'
- [Abstract and Discussion] The claim that the lower-empathy finding is 'supported by ... a pre-trained empathy model' is inconsistent with Table 2. The perceived empathy model trained by the authors shows no human-chatbot difference (t=0.45, p=0.65), and the paper itself states that 'the empathetic language of humans and the empathetic language of bots are equivalent.' Among the off-the-shelf pre-trained models, the Interpret model favors humans while the Emo-React model favors chatbots. Only the GPT-4o annotations support the direction of the empathy gap. Please either remove 'pre-trained empathy model' from the supporting evidence in the Abstract and Conclusion, or identify which specific model is meant and explain why that model's direction is credited while the mixed and null results from the other models are not.
- [Experiments and Results (LLM Judgement of Perceived Empathy) and Supplementary S2] The GPT-4o annotation results are over-interpreted as an independent confirmation. Although GPT-4o was not given explicit identity labels, Supplementary S2 shows that GPT-4o correctly inferred speaker identity with 61% accuracy (F1=0.60), so the annotation contrast is not a clean blind comparison of language; it is partly a second opinion from a model that can detect LLM-like patterns. In addition, the per-conversation correlations with user ratings are weak (r=0.20 for humans, r=0.06 for chatbots, r=0.07 combined), so the statement that GPT-4o ratings 'aligned with user ratings' is true only at the level of aggregate means. The text should state both qualifications wherever the LLM-judge results are presented.
- [Experiments and Results (Psychological Ratings)] The conversation-quality conclusion is based on a single 5-point Likert item collected only in WASSA 2024, yet the paper states it as a general finding ('Chatbots receive higher ratings for conversation quality'). The Results and Conclusion should prominently state this dataset restriction and the single-item nature of the quality measure. The mixed-model analysis should also be presented with the number of observations per participant, since a random intercept for participant ID may be weakly identified if most participants contributed only one quality rating. Without this context, readers cannot assess the robustness of the claimed quality advantage.
minor comments (5)
- [Data] The abbreviation 'WASSA' appears with extra spaces in several places (e.g., 'W ASSA 2023'); please format it consistently.
- [Psychological Ratings] Please report effect sizes and confidence intervals for the t-tests in Table 2, since several p-values are close to conventional thresholds and the paper intentionally avoids effect-size thresholds.
- [Experiments and Results (Psychological Ratings)] The text reports the interaction coefficient for cognitive state empathy as β=0.77 without specifying a sign; from the prose and Figure 2 it is unclear whether this is a positive or negative interaction. Please provide a full model table with standard errors.
- [Experiments and Results (LLM Judgement of Perceived Empathy)] The GPT-4o annotation prompt and rating scale are not described in the main text; please add them so readers know whether the 1-7 scale or another scale was used.
- [Limitations] The sentence 'we intentionally avoided setting arbitrary thresholds for effect sizes' is confusing: reporting effect sizes does not require setting thresholds. Please clarify or remove this sentence.
Circularity Check
No circularity: the central result is direct user self-report measurement, and the supporting models are external or transparently reported.
full rationale
The paper's central empirical finding—that chatbots are rated lower in perceived empathy but higher in conversational quality—comes directly from participant self-reports collected during the WASSA 2023 and 2024 extensions. This is measurement, not derivation, and no equation or construction reduces any prediction to its own input. The supporting analyses are not fitted to the target outcome. GPT-4o annotations were generated without explicit identity labels and correlate only weakly with the user labels (r = 0.07 combined), so they cannot be a renamed version of the user ratings. The perceived-empathy model was trained on the separate EC dataset using 10-fold cross-validation and then applied to the WASSA conversations; its null result (t = 0.45, p = 0.65) is an external check rather than circular confirmation. The pre-trained empathy models come from prior work (Lahnala et al. 2022; Sharma et al. 2020) and are used off-the-shelf; even where a prior author overlaps, their outputs are not recomputed from this paper's fitted values. The acknowledged identity-cue confound—all chatbot turns were visibly labeled 'Bot:'—is a genuine validity threat and an alternative explanation for the user-rating gap, but it does not make the empathy rating definitionally dependent on the label. The paper transparently reports its own third-party blind annotations (t = 1.82, p = 0.07), which weaken rather than circularly reinforce the central claim. No load-bearing self-citation chain is present, and the self-citations that do appear are to datasets and shared-task resources that are used as data sources, not as derivational premises. The paper is self-contained against external benchmarks, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- Ridge regularization lambda =
10,000
- Unigram frequency threshold =
5%
assumptions (4)
- domain assumption Self-reported general empathy (single 1-7 item) is a valid measure of perceived empathy.
- domain assumption The adapted six-item state empathy scale from Shen (2010) measures affective, cognitive, and associative empathy as experienced during the conversation.
- domain assumption The single 5-point conversation quality question captures overall conversation quality.
- ad hoc to paper The visible 'Bot:' label does not, by itself, account for the full empathy gap.
Cite this review
Pith. "Pith review of The Illusion of Empathy: How AI Chatbots Shape Conversation Perception." pith.science (2026). https://pith.science/paper/KK5AJI2U
@misc{pith2026241112877,
author = {Pith},
title = {Pith review of: The Illusion of Empathy: How AI Chatbots Shape Conversation Perception},
year = {2026},
howpublished = {\url{https://pith.science/paper/KK5AJI2U}},
note = {Machine review of arXiv:2411.12877}
}
read the original abstract
As AI chatbots increasingly incorporate empathy, understanding user-centered perceptions of chatbot empathy and its impact on conversation quality remains essential yet under-explored. This study examines how chatbot identity and perceived empathy influence users' overall conversation experience. Analyzing 155 conversations from two datasets, we found that while GPT-based chatbots were rated significantly higher in conversational quality, they were consistently perceived as less empathetic than human conversational partners. Empathy ratings from GPT-4o annotations aligned with user ratings, reinforcing the perception of lower empathy in chatbots compared to humans. Our findings underscore the critical role of perceived empathy in shaping conversation quality, revealing that achieving high-quality human-AI interactions requires more than simply embedding empathetic language; it necessitates addressing the nuanced ways users interpret and experience empathy in conversations with chatbots.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Alam, F.; Danieli, M.; and Riccardi, G. 2018. Annotating and modeling empathy in spoken conversations. Computer Speech & Language, 50: 40--61
work page 2018
-
[4]
Aron, A.; Aron, E. N.; and Smollan, D. 1992. Inclusion of other in the self scale and the structure of interpersonal closeness. Journal of personality and social psychology, 63(4): 596
work page 1992
-
[5]
Atkinson, J. M.; and Heritage, J. 1984. Structures of social action. Cambridge University Press
work page 1984
-
[6]
W.; Poliak, A.; Dredze, M.; Leas, E
Ayers, J. W.; Poliak, A.; Dredze, M.; Leas, E. C.; Zhu, Z.; Kelley, J. B.; Faix, D. J.; Goodman, A. M.; Longhurst, C. A.; Hogarth, M.; et al. 2023. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA internal medicine, 183(6): 589--596
2023
-
[7]
Barriere, V.; Sedoc, J.; Tafreshi, S.; and Giorgi, S. 2023. Findings of WASSA 2023 Shared Task on Empathy, Emotion and Personality Detection in Conversation and Reactions to News Articles. In Barnes, J.; De Clercq, O.; and Klinger, R., eds., Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis , ...
work page 2023
-
[8]
D.; Fultz, J.; and Schoenrade, P
Batson, C. D.; Fultz, J.; and Schoenrade, P. A. 1987. Distress and empathy: Two qualitatively distinct vicarious emotions with different motivational consequences. Journal of personality, 55(1): 19--39
work page 1987
Show all 56 references
-
[9]
Brave, S.; Nass, C.; and Hutchinson, K. 2005. Computers that care: investigating the effects of orientation of emotion exhibited by an embodied computer agent. International journal of human-computer studies, 62(2): 161--178
2005
-
[10]
A.; and Cudr \'e -Mauroux, P
Casas, J.; Spring, T.; Daher, K.; Mugellini, E.; Khaled, O. A.; and Cudr \'e -Mauroux, P. 2021. Enhancing conversational agents with empathic abilities. In Proceedings of the 21st ACM International Conference on Intelligent Virtual Agents, 41--47
2021
-
[11]
C.; and Curry, A
Curry, A. C.; and Curry, A. C. 2023. Computer says “no”: The case against empathetic conversational AI. In Findings of the Association for Computational Linguistics: ACL 2023, 8123--8130
2023
-
[12]
Decety, J.; and Jackson, P. L. 2004. The functional architecture of human empathy. Behavioral and cognitive neuroscience reviews, 3(2): 71--100
2004
-
[13]
Gao, J.; Liu, Y.; Deng, H.; Wang, W.; Cao, Y.; Du, J.; and Xu, R. 2021. Improving empathetic response generation by recognizing emotion cause in conversations. In Findings of the association for computational linguistics: EMNLP 2021, 807--819
2021
-
[14]
Giorgi, S.; Sedoc, J.; Barriere, V.; and Tafreshi, S. 2024. Findings of WASSA 2024 Shared Task on Empathy and Personality Detection in Interactions. In De Clercq, O.; Barriere, V.; Barnes, J.; Klinger, R.; Sedoc, J.; and Tafreshi, S., eds., Proceedings of the 14th Workshop on ...
2024
-
[15]
Go, E.; and Sundar, S. S. 2019. Humanizing chatbots: The effects of visual, identity and conversational cues on humanness perceptions. Computers in human behavior, 97: 304--316
2019
-
[16]
M.; Gray, K.; and Wegner, D
Gray, H. M.; Gray, K.; and Wegner, D. M. 2007. Dimensions of mind perception. Science, 315(5812): 619--619
2007
-
[17]
Herlin, I.; and Visap \"a \"a , L. 2016. Dimensions of empathy in relation to language. Nordic Journal of linguistics, 39(2): 135--157
2016
-
[18]
Hosseini, M.; and Caragea, C. 2021. Distilling knowledge for empathy detection. In Findings of the Association for Computational Linguistics: EMNLP 2021, 3713--3724
2021
-
[19]
Inan Nur, B.; Santoso, B.; and Putra, O. H. 2021. The method and metric of user experience evaluation: A systematic literature review. In Proceedings of ICSCA 2021
2021
-
[20]
Jain, G.; Pareek, S.; and Carlbring, P. 2024. Revealing the source: How awareness alters perceptions of AI and human-generated mental health responses. Internet Interventions, 36: 100745
2024
-
[21]
Jefferson, G. 1984. On stepwise transition from talk about a trouble to inappropriately next-positioned matters. Structures of social action: Studies in conversation analysis, 191: 222
1984
-
[22]
J.; and Sundar, S
Koh, Y. J.; and Sundar, S. S. 2010. Heuristic versus systematic processing of specialist versus generalist sources in online media. Human Communication Research, 36(2): 103--124
2010
-
[23]
Lahnala, A.; Welch, C.; and Flek, L. 2022. CAISA at WASSA 2022: Adapter-tuning for empathy prediction. In Proceedings of the 12th Workshop on Computational Approaches to Subjectivity, Sentiment & Social Media Analysis, 280--285
2022
-
[24]
Lee, A.; Kummerfeld, J.; Ann, L.; and Mihalcea, R. 2024 a . A Comparative Multidimensional Analysis of Empathetic Systems. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 179--189
2024
-
[25]
K.; Suh, J.; Zhan, H.; Li, J
Lee, Y. K.; Suh, J.; Zhan, H.; Li, J. J.; and Ong, D. C. 2024 b . Large language models produce responses perceived to be empathic. arXiv preprint arXiv:2403.18148
2024 arXiv
-
[26]
Lindstr \"o m, A.; and Sorjonen, M.-L. 2012. Affiliation in conversation. The handbook of conversation analysis, 250--369
2012
-
[27]
Majumder, N.; Hong, P.; Peng, S.; Lu, J.; Ghosal, D.; Gelbukh, A.; Mihalcea, R.; and Poria, S. 2020. MIME: MIMicking emotions for empathetic response generation. arXiv preprint arXiv:2010.01454
2020 arXiv
-
[28]
F.; and Kageki, N
Mori, M.; MacDorman, K. F.; and Kageki, N. 2012. The uncanny valley [from the field]. IEEE Robotics & automation magazine, 19(2): 98--100
2012
-
[29]
Nass, C.; Steuer, J.; and Tauber, E. R. 1994. Computers are social actors. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, 72--78. ACM
1994
-
[30]
Neumann, R.; and Chan, E. 2015. Measures of empathy: Self-report, behavioral, and neuroscientific approaches. In Measures of Personality and Social Psychological Constructs, 257--289. Academic Press
2015
-
[31]
B.; Schutz, A.; Lopes, P.; and Smith, C
Nezlek, J. B.; Schutz, A.; Lopes, P.; and Smith, C. V. 2007. Naturally occurring variability in state empathy. Empathy in mental illness, 187--200
2007
-
[32]
Omitaomu, D.; Tafreshi, S.; Liu, T.; Buechel, S.; Callison-Burch, C.; Eichstaedt, J.; Ungar, L.; and Sedoc, J. 2022. Empathic conversations: A multi-level dataset of contextualized conversations. arXiv preprint arXiv:2205.12698
2022 arXiv
-
[33]
R.; and Feng, S
Panickssery, A.; Bowman, S. R.; and Feng, S. 2024. Llm evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076
2024 arXiv
-
[34]
Per \"a kyl \"a , A. 2012. Conversation analysis in psychotherapy. The handbook of conversation analysis, 551--574
2012
-
[35]
Pfeiffer, J.; Vuli \'c , I.; Gurevych, I.; and Ruder, S. 2020. MAD-X : A n A dapter- B ased F ramework for M ulti- T ask C ross- L ingual T ransfer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 7654--7673. Online: Associati...
2020
-
[36]
D.; and De Waal, F
Preston, S. D.; and De Waal, F. B. M. 2002. Empathy: Its ultimate and proximate bases. Behavioral and Brain Sciences, 25(1): 1--20
2002
-
[37]
Qian, Y.; Zhang, W.-N.; and Liu, T. 2023. Harnessing the power of large language models for empathetic response generation: Empirical investigations and improvements. arXiv preprint arXiv:2310.05140
2023 arXiv
-
[38]
S.; and Yang, Y
Raamkumar, A. S.; and Yang, Y. 2022. Empathetic conversational systems: a review of current advances, gaps, and opportunities. IEEE Transactions on Affective Computing, 14(4): 2722--2739
2022
-
[39]
M.; Li, M.; and Boureau, Y.-L
Rashkin, H.; Smith, E. M.; Li, M.; and Boureau, Y.-L. 2018. Towards empathetic open-domain conversation models: A new benchmark and dataset. arXiv preprint arXiv:1811.00207
2018 arXiv
-
[40]
Schmidmaier, R.; Rupp, L.; et al. 2024. Perceived Empathy of Technology Scale (PETS): Measuring Empathy of Systems Toward the User. In Proceedings of the CHI Conference
2024
-
[41]
A.; Giorgi, S.; Sap, M.; Crutchley, P.; Ungar, L.; and Eichstaedt, J
Schwartz, H. A.; Giorgi, S.; Sap, M.; Crutchley, P.; Ungar, L.; and Eichstaedt, J. 2017. Dlatk: Differential language analysis toolkit. In Proceedings of the 2017 conference on empirical methods in natural language processing: System demonstrations, 55--60
2017
-
[42]
Shafaei, R.; Bahmani, Z.; Bahrami, B.; and Vaziri-Pashkam, M. 2020. Effect of perceived interpersonal closeness on the joint Simon effect in adolescents and adults. Scientific Reports, 10(1): 18107
2020
-
[43]
S.; Atkins, D
Sharma, A.; Miner, A. S.; Atkins, D. C.; and Althoff, T. 2020. A computational approach to understanding empathy expressed in text-based mental health support. arXiv preprint arXiv:2009.08441
2020 arXiv
-
[44]
Shen, L. 2010. On a scale of state empathy during message processing. Western Journal of Communication, 74(5): 504--524
2010
-
[45]
A.; Durbin, S.; Weyrich, M
Shetty, V. A.; Durbin, S.; Weyrich, M. S.; Mart \' nez, A. D.; Qian, J.; and Chin, D. L. 2024. A scoping review of empathy recognition in text using natural language processing. Journal of the American Medical Informatics Association, 31(3): 762--775
2024
-
[46]
J.; Zhang, J.; Sahay, S.; and Yu, Z
Shi, W.; Wang, X.; Oh, Y. J.; Zhang, J.; Sahay, S.; and Yu, Z. 2020. Effects of persuasive dialogues: testing bot identities and inquiry strategies. In Proceedings of the 2020 CHI conference on human factors in computing systems, 1--13
2020
-
[47]
Sorin, V.; Brin, D.; Barash, Y.; Konen, E.; Charney, A.; Nadkarni, G.; and Klang, E. 2023. Large language models (llms) and empathy-a systematic review. medRxiv, 2023--08
2023
-
[48]
S.; Bellur, S.; Oh, J.; Jia, H.; and Kim, H.-S
Sundar, S. S.; Bellur, S.; Oh, J.; Jia, H.; and Kim, H.-S. 2016. Theoretical importance of contingency in human-computer interaction: Effects of message interactivity on user engagement. Communication Research, 43(5): 595--625
2016
-
[49]
A.; Sutthithatip, S.; and Park, S
Urakami, J.; Moore, B. A.; Sutthithatip, S.; and Park, S. 2019. Users' perception of empathic expressions by an advanced intelligent system. In Proceedings of the 7th international conference on human-agent interaction, 11--18
2019
-
[50]
K.; Ferdiana, R.; and Hidayah, I
Wardhana, A. K.; Ferdiana, R.; and Hidayah, I. 2021. Empathetic chatbot enhancement and development: A literature review. In 2021 International Conference on Artificial Intelligence and Mechatronics Systems (AIMS), 1--6. IEEE
2021
-
[51]
Welivita, A.; and Pu, P. 2024 a . Are Large Language Models More Empathetic than Humans? arXiv preprint arXiv:2406.05063
2024 arXiv
-
[52]
Welivita, A.; and Pu, P. 2024 b . Is ChatGPT More Empathetic than Humans? arXiv preprint arXiv:2403.05572
2024 arXiv
-
[53]
Westman, M.; Shadach, E.; and Keinan, G. 2013. The crossover of positive and negative emotions: The role of state empathy. International Journal of Stress Management, 20(2): 116
2013
-
[54]
Xu, Z.; and Jiang, J. 2024. Multi-dimensional Evaluation of Empathetic Dialog Responses. arXiv preprint arXiv:2402.11409
2024 arXiv
-
[55]
Yin, Y.; Jia, N.; and Wakslak, C. J. 2024. AI can help people feel heard, but an AI label diminishes this impact. Proceedings of the National Academy of Sciences, 121(14): e2319112121
2024
-
[56]
M.; Scepanovic, S.; Quercia, D.; and Konrath, S
Zhou, K.; Aiello, L. M.; Scepanovic, S.; Quercia, D.; and Konrath, S. 2021. The language of situational empathy. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1): 1--19
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.