Pith. sign in

REVIEW 4 major objections 5 minor 76 references

E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read E-THER: A therapy-grounded multimodal dataset that trains AI to notice when a client's words and face disagree, and to respond with more genuine empathy.

desk verdict E-THER is a genuinely new dataset for verbal-visual incongruence in therapy conversations, but the claimed empathy gains rest on unvalidated lexical metrics and contradictory ablations. read the letter →

arxiv 2509.02100 v2 pith:AEI5AT4W submitted 2025-09-02 cs.HC cs.CL

classification cs.HCcs.CL
keywords multimodalempathydatasetverbal-visualincongruencePerson-CenteredTherapyvision-languagemodelsempathicresponsegenerationtherapeuticdialogueincongruence-awaretrainingannotation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces E-THER, a multimodal dataset built from about five hours of real therapeutic conversations, in which every client utterance is paired with a synchronized video frame and annotated for whether the person's words and facial expression match. The mismatches are labelled as 'minimizing' (the face shows more distress than the words admit), 'contradiction' (face and words oppose each other), or absent. The authors' central claim is that training vision-language models to detect this verbal-visual incongruence — instead of only recognizing surface emotions — produces replies that their Person-Centered Therapy-based metrics score as more authentic, more engaged, and more therapeutically appropriate, with fine-tuned IDEFICS2 and VideoLLaVA beating a GPT-4V baseline. If the claim holds, E-THER is a first resource for training empathic AI on a theoretically grounded signal that current empathy datasets ignore.

What carries the argument

The carrying object is the E-THER dataset plus its incongruence-focused training loss. Each training instance synchronizes a facial-expression frame with a transcribed client utterance and carries five annotation dimensions: incongruence type (minimizing, contradiction, or none), engagement level, and valence-arousal-dominance. During fine-tuning, a continuous incongruence score—combining VAD mismatch and cross-modal embedding distance—reweights each sample's loss between 1 and 2, so mismatched turns dominate learning; a 30% context dropout removes explicit emotional cues, and VAD labels are masked to force intrinsic affect inference. Evaluation uses four metrics built from hand-picked conve

What would settle it

Have licensed counsellors, blind to which model produced each response, rate the same generated replies on a validated clinical empathy scale. If the incongruence-trained models do not beat GPT-4V in those human judgments while still beating it on the paper's PCT metrics, the claimed empathy gains are an artifact of the keyword-based metric design.

Watch

Extended reading notes

Core claim

The central claim is that verbal-visual incongruence is a learnable signal for empathic AI, not a nuisance to be averaged away. A model trained to attend to turns where the client's words and facial expression disagree is claimed to respond in ways that score higher on Empathic Authenticity, Responsive Engagement, Therapeutic Concision, and PCT Adherence than a strong general-purpose baseline. The dataset grounds this signal in Person-Centered Therapy's core conditions, and the training pipeline amplifies the loss on incongruent turns, randomly drops explicit empathy context, and hides valence-arousal-dominance labels so the model must infer affect from vision and text together. In the paper

Load-bearing premise

The author-defined PCT metrics in Section V-A—built from hand-picked keyword markers and a word-embedding similarity step, without human ratings or clinical scales—must actually measure empathic communication quality for the central claim to hold.

Editorial extensions

If this is right

  • Incongruence-aware training is claimed to be architecture-agnostic: both IDEFICS2 and VideoLLaVA improve over GPT-4V on the authors' PCT metrics.
  • Removing visual input during training consistently lowers PCT Adherence, so the paper concludes that genuinely empathic response generation requires multimodal integration, not text alone.
  • The dense annotation scheme (789 annotation instances per hour) makes a modest 5-hour corpus sufficient to show training effects, offering a quality-over-quantity template.
  • The metric suite offers an alternative to BLEU/ROUGE-style evaluation for empathy, which the paper argues fails to capture therapeutic communication quality.
  • The authors explicitly scope the claim to supportive, non-clinical AI applications and do not claim clinical efficacy or AI-delivered therapy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The keyword-marker metrics could be gamed by models that learn to insert phrases like 'you feel' or 'I'm wondering'; a human-rater study with a validated empathy scale on the same outputs would show whether the gains are real.
  • The engagement annotations may serve better as analytical labels than as training weights, since the paper's own ablations show inconsistent benefits from engagement-based weighting.
  • The same annotation scheme could transfer to other modalities and settings—telehealth visits, spoken-language assistants, or non-English therapy contexts—where verbal-nonverbal mismatches also carry clinical meaning.
  • With only 20.4% of turns incongruent and 18 sessions, scaling to more sessions or augmenting incongruent turns could sharpen the training signal further; the authors leave this as future work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces E-THER, a multimodal dataset of 18 filmed therapy-style sessions (789 dialogue pairs) with annotations of verbal-visual incongruence, engagement level, and VAD dimensions, grounded in Person-Centered Therapy (PCT). It proposes LoRA-based fine-tuning of IDEFICS2, VideoLLaVA, and BLIP2 with an incongruence-weighted loss, context dropout, and self-supervised VAD prediction. Evaluation uses four author-defined PCT-based metrics (Empathic Authenticity, Responsive Engagement, Therapeutic Concision, PCT Adherence) plus BERTScore. The paper reports that E-THER-trained VideoLLaVA and IDEFICS2 outperform GPT-4V on several of these metrics, and that ablations show contributions from multimodal input and incongruence weighting.

Significance. The dataset is potentially valuable: to my knowledge, it is the first multimodal therapeutic dialogue corpus with expert-validated verbal-visual incongruence annotations, and the authors report inter-rater reliability and expert consensus (Table II). The training pipeline is clearly described and uses open models, which supports reproducibility. However, the paper's central empirical claim—that the trained models show 'notable gains in empathic and therapeutic conversational qualities'—is not supported by the evidence presented. The four evaluation metrics are unvalidated lexical counting rules; no human or clinical validation is reported; and no direct incongruence-detection accuracy is given. Because the metrics reward language patterns that the fine-tuning procedure directly encourages, the empirical results are uninterpretable as evidence of empathy. The dataset itself may merit a datasets-and-resources paper, but the current manuscript's claims exceed what the evaluation demonstrates.

major comments (4)
  1. [§V-A, Tables III–IV] The four PCT metrics are lexical heuristics: Empathic Authenticity counts words such as 'right', 'okay', and 'actually'; Responsive Engagement counts 'given', 'considering', and 'in your situation'; Therapeutic Concision counts 'specifically' and 'what I hear'; PCT Adherence counts 'you feel', 'I'm wondering', and 'that makes sense' and penalizes 'you should'/'you need to', with a vaguely specified GloVe-similarity step. No human rating study, no validated instrument (e.g., CARE, Jefferson Scale), and no convergent validity evidence is reported; Section VIII explicitly defers human expert evaluation to future work. Because the models are LoRA-fine-tuned on real therapist transcripts, they are directly pushed toward the very n-grams these metrics reward. The improvements over GPT-4V in Tables III–IV may therefore reflect distributional mimicry rather than empathy. I would need at least on
  2. [§III-C, §IV, §VI–VII] The paper's unique contribution is 'verbal-visual incongruence detection' (abstract and §IX), but no direct incongruence-detection accuracy is reported anywhere. There is no classification accuracy, precision/recall, F1, or ROC for detecting incongruence on the held-out test set. The experiments in Sections VI–VII report only response-generation metrics. The incongruence-weighted training loss in §IV-A is motivated by the value of detecting incongruence, yet the paper never measures whether the models actually detect it. Without such numbers, the central claim that E-THER 'enables artificial agents to develop empathic capabilities through detection of verbal-visual emotional incongruence' is untested.
  3. [§VII, Tables VI–VIII] The ablation results are internally inconsistent. Table VII reports IDEFICS2 'No Incongruence Weighting' Responsive Engagement change as −7.8%, while Table VIII reports the same condition as −17.8%. Additionally, the p-values for Empathic Authenticity in Tables VI and VII are identical across the two independently trained models (0.811, 0.419, 0.821), which is implausible and suggests copied or erroneous statistics. These inconsistencies undermine the ablation conclusions in §VII-B, including the claim of 'universal multimodal dependency' and 'architecture-specific vulnerabilities.'
  4. [§IV-A, Algorithm 1] The training objective is described with two different weighting formulas: Equation (4) uses w_i = 1 + s_i^γ, while Algorithm 1 uses w_i = λ_base + α·I_i. The paper does not state which formulation was used for the reported experiments, and the relationship between s_i and the binary I_i is not made precise. The hyperparameters γ, λ, τ_e, and p_dropout are fixed empirically with no sensitivity analysis. Given the small evaluation set (60 dialogue pairs from 2 conversations), the stability of the claimed improvements is unclear. The authors should report results across these choices and make the exact weighting scheme explicit.
minor comments (5)
  1. [§VI opening] Typo: 'semantic vlaidation' should be 'semantic validation'.
  2. [Table I and §III-D/E] The dataset is said to have '5 annotation dimensions' in Table I but is later described as a 'four-dimensional framework' and 'four-dimensional annotation' in §III-D/E. Please reconcile.
  3. [§IV-A vs Algorithm 1] The notation for the loss weight differs between the equation and the algorithm; unify the notation to avoid ambiguity about the actual training procedure.
  4. [§III-A/Data availability] No repository URL, license, or data-access statement is provided. For a dataset paper, this should be included.
  5. [Figure 3] The heatmaps are described qualitatively; no quantitative definition of 'information travel' is provided, so the figure is illustrative only and should be labeled as such.

Circularity Check

1 steps flagged · score 6.0 of 10

Empathy-gain claim is evaluated with lexical metrics built from the same PCT therapist language used as training targets; without external validation the improvement is partly by construction.

  1. fitted input called prediction [§I Contributions (Aligned evaluation); §IV Training Methodology; §V-A Core Evaluation Metrics; §VIII Limitations]
    "'Aligned evaluation: an automatic evaluation framework tailored to our annotations, i.e., conversational authenticity, responsive engagement, and alignment with Rogers’ core conditions' (Sec. I). 'PCT Adherence ... Empathic Understanding: Combines traditional empathy markers ("you feel", "you're experiencing") with genuine curiosity indicators ("I'm wondering", "what's that like") ... Unconditional Positive Regard: Measures non-judgmental acceptance language ("that makes sense", "that's understandable") while applying judgment penalties for directive or prescriptive responses ("you should", "y"

    The evaluation framework is explicitly 'tailored to our annotations,' and the annotations come from PCT therapist sessions that are the same data used for LoRA fine-tuning ('Training was performed on 16 conversations from our dataset'). The PCT Adherence metric counts the very phrases that saturate those therapist transcripts ('you feel', 'I'm wondering', 'that makes sense', etc.). A model that imitates the therapist transcript style will therefore mechanically receive higher Empathic Authenticity, Responsive Engagement, Therapeutic Concision, and PCT Adherence scores. The paper's headline 'notable gains' over GPT-4V are thus largely built into the choice of an evaluation ruler calibrated to the training distribution. Section VIII explicitly defers convergent validation against the Jeffers

full rationale

The central empirical claim—that E-THER-trained models show 'notable gains in empathic and therapeutic conversational qualities'—rests entirely on the four PCT metrics defined in §V-A. Each metric is a lexical scoring rule (e.g., PCT Adherence counts 'you feel', 'I'm wondering', 'that makes sense', penalizes 'you should'). The models are fine-tuned on 16 PCT therapist sessions, so the evaluation rewards reproduction of the training distribution's phrasing; the claimed empathy gain is at least in part a distributional-mimicry artifact rather than evidence of genuine empathic competence. The paper itself concedes in §VIII that no comparison against established clinical instruments (Jefferson Scale of Empathy, CARE) or human expert evaluation has been conducted, so the metric has no external anchor. Internal inconsistencies (IDEFICS2 Responsive Engagement reported as −7.8% in Table VII vs −17.8% in Table VIII; identical p-values across independent model/condition cells) further undermine the reliability of these scores. The dataset construction and incongruence annotations are independent contributions and not circular; the circularity is confined to the evaluation of the empathy-gain claim. No same-author self-citations are load-bearing.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The empirical claims depend on hand-set weighting coefficients, a small expert-validated annotation subsample, and author-defined evaluation metrics. The dataset and code are not released, and the relationship between the continuous incongruence score and the binary scheme actually used in training is unclear.

free parameters (6)
  • gamma (incongruence weighting exponent) = range [0.8, 1.2], exact value not stated
    Eq. (4) defines wi = 1 + s_i^gamma; the paper does not report which gamma was used, and Algorithm 1 instead uses a binary weighting.
  • lambda (cosine-distance weight) = 0.5
    Eq. (3) sets lambda=0.5 in the continuous incongruence score s_i.
  • tau_e (VAD mismatch temperature) = batch median VAD mismatch
    Eq. (3) uses tau_e as a scaling factor, set as the batch median rather than a fixed constant.
  • Context dropout probability p_dropout = 0.3
    Section IV-A2 states explicit empathy context is randomly removed during 30% of training iterations.
  • Binary incongruence weight (lambda_base, alpha) = lambda_base=1, alpha=1
    Algorithm 1 line 14 uses wi = lambda_base + alpha*Ii, giving a hand-set 2:1 emphasis at high incongruence, which conflicts with the continuous score equations.
  • Engagement weighting coefficient = wi = 1.0 + Ei
    Section VII-A1b introduces an inverse engagement weighting scheme wi = 1.0 + Ei for the ablation study.
assumptions (5)
  • domain assumption PCT constructs can be operationalized as keyword and GloVe-similarity scores
    Section V-A defines all four evaluation metrics from hand-picked lexical markers; no validation against human judgments or clinical instruments is provided.
  • domain assumption A single RGB frame at turn boundary captures the visual emotional state relevant to incongruence
    Section III-B extracts one frame at the moment of client verbal response; nonverbal affect is treated as a static facial expression.
  • domain assumption Expert review of 100 dialogue pairs validates all 789 annotations
    Section III-D reports 83% consensus on a 100-pair sample, yet 17 disagreements include presence/absence errors; full-dataset reliability is not established.
  • domain assumption Raters' distinction between minimizing and contradiction subtypes is reliable enough for training
    Section III-D notes subtype disagreements in 9 of 17 expert disagreements, but subtype annotations are still used as training labels.
  • domain assumption Fine-tuning on 16 sessions transfers to unseen clients and conversations
    Only two conversations are held out; no cross-validation or client-level split is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness." pith.science (2026). https://pith.science/paper/AEI5AT4W

@misc{pith2026250902100,
  author       = {Pith},
  title        = {Pith review of: E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AEI5AT4W}},
  note         = {Machine review of arXiv:2509.02100}
}
read the original abstract

A prevalent shortfall among current empathic AI systems is their inability to recognize when verbal expressions may not fully reflect underlying emotional states. This is because the existing datasets, used for the training of these systems, focus on surface-level emotion recognition without addressing the complex verbal-visual incongruence (mismatch) patterns useful for empathic understanding. In this paper, we present E-THER, the first Person-Centered Therapy-grounded multimodal dataset with multidimensional annotations for verbal-visual incongruence detection, enabling training of AI systems that develop genuine rather than performative empathic capabilities. The annotations included in the dataset are drawn from humanistic approach, i.e., identifying verbal-visual emotional misalignment in client-counsellor interactions - forming a framework for training and evaluating AI on empathy tasks. Additional engagement scores provide behavioral annotations for research applications. Notable gains in empathic and therapeutic conversational qualities are observed in state-of-the-art vision-language models (VLMs), such as IDEFICS and VideoLLAVA, using evaluation metrics grounded in empathic and therapeutic principles. Empirical findings indicate that our incongruence-trained models outperform general-purpose models in critical traits, such as sustaining therapeutic engagement, minimizing artificial or exaggerated linguistic patterns, and maintaining fidelity to PCT theoretical framework.

Figures

Figures reproduced from arXiv: 2509.02100 by the authors.

Figure 1
Figure 1. A snippet from one of the recorded conversations that depicts the verbal and visual content of the conversations in the [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Training pipeline showing dataset preparation (18 conversations - 16 Train + 2 Eval), VLM fine-tuning with LoRA [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Layer-wise information travel during empathy-enhanced training. Each heatmap shows relative cross-modal flow (rows: [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Effect sizes and significance levels across all ablation conditions for both VideoLLaVA and IDEFICS2. Points represent [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

76 extracted references · 70 canonical work pages

  1. [1]

    Towards em- pathetic open-domain conversation models: A new benchmark and dataset,

    H. Rashkin, E. M. Smith, M. Li, and Y .-L. Boureau, “Towards em- pathetic open-domain conversation models: A new benchmark and dataset,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2018, pp. 5370–5381

  2. [2]

    Are large language models more empathetic than humans?

    A. Welivita and P. Pu, “Are large language models more empathetic than humans?” arXiv preprint arXiv:2406.05063 , 2024

  3. [3]

    In principle obstacles for empathic ai: why we can’t replace human empathy in healthcare,

    C. Montemayor, J. Halpern, and A. Fairweather, “In principle obstacles for empathic ai: why we can’t replace human empathy in healthcare,” AI & society , vol. 37, no. 4, pp. 1353–1359, 2022

  4. [4]

    The necessary and sufficient conditions of therapeutic personality change,

    C. R. Rogers, “The necessary and sufficient conditions of therapeutic personality change,” Journal of consulting psychology , vol. 21, no. 2, pp. 95–103, 1957

  5. [5]

    Decoding of inconsistent communica- tions,

    A. Mehrabian and M. Wiener, “Decoding of inconsistent communica- tions,” Journal of personality and social psychology , vol. 6, no. 1, pp. 109–114, 1967

  6. [6]

    Modeling empathy and distress in reaction to news stories,

    S. Buechel, A. Buffone, B. Slaff, L. Ungar, and J. Sedoc, “Modeling empathy and distress in reaction to news stories,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018, pp. 4758–4765

  7. [7]

    Moel: Mixture of empathetic listeners,

    Z. Lin, A. Madotto, J. Shin, P. Xu, and P. Fung, “Moel: Mixture of empathetic listeners,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing , 2019, pp. 121– 132

  8. [8]

    Iemocap: Interactive emotional dyadic motion capture database,

    C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language Resources and Evaluation , vol. 42, no. 4, pp. 335–359, 2008

Show all 76 references
  1. [9]

    Meld: A multimodal multi-party dataset for emotion recognition in conversations,

    S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihal- cea, “Meld: A multimodal multi-party dataset for emotion recognition in conversations,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 527–536

  2. [10]

    Medic: A multimodal empathy dataset in counseling,

    Z. Zhu, X. Li, J. Pan, Y . Xiao, Y . Chang, A. Zhou, Y . Zheng, Y . Zhang, and S. Wang, “Medic: A multimodal empathy dataset in counseling,” in Proceedings of the 31st ACM International Conference on Multimedia . ACM, 2023, pp. 2340–2350

  3. [11]

    Towards multimodal emotional support conversation systems,

    Y . Chen, H. Liu, Y . Wang, J. Li, F. Zhang, H. Wu, J. Liu, and M. Zhang, “Towards multimodal emotional support conversation systems,” arXiv preprint arXiv:2408.03650, 2024

  4. [12]

    W. R. Miller and S. Rollnick, Motivational interviewing: Helping people change. Guilford press, 2012

  5. [13]

    The role of socio- emotional attributes in enhancing human-ai collaboration,

    M. Kolomaznik, V . Petrik, M. Slama, and V . Jurik, “The role of socio- emotional attributes in enhancing human-ai collaboration,” Frontiers in psychology, vol. 15, p. 1369957, 2024

  6. [14]

    The role of empathy in promoting change,

    J. C. Watson, P. L. Steckley, and E. J. McMullen, “The role of empathy in promoting change,” Psychotherapy Research, vol. 24, no. 3, pp. 286– 298, 2014

  7. [15]

    C. R. Rogers, On becoming a person: A therapist’s view of psychother- apy. Houghton Mifflin, 1961

  8. [16]

    A theory of therapy, personality, and interpersonal relationships: As developed in the client-centered framework,

    ——, “A theory of therapy, personality, and interpersonal relationships: As developed in the client-centered framework,” Psychology: A study of a science, vol. 3, pp. 184–256, 1959

  9. [17]

    The friendly relationship between thera- peutic empathy and person-centred care,

    D. Hardman and J. Howick, “The friendly relationship between thera- peutic empathy and person-centred care,” European Journal for Person Centered Healthcare, vol. 7, no. 2, pp. 351–357, 2019

  10. [18]

    J. D. Bozarth, Person-centered therapy: A revolutionary paradigm . PCCS Books, 1998

  11. [19]

    Does it matter if empathic ai has no empathy?

    E. Stevenson and Collaborators, “Does it matter if empathic ai has no empathy?” Nature Machine Intelligence , 2024, discusses risks and ethical concerns of empathic AI implementations. [Online]. Available: https://www.nature.com/articles/s42256-024-00841-7

  12. [20]

    Towards emotional support dialog systems,

    S. Liu, C. Zheng, O. Demasi, S. Sabour, Y . Li, Z. Yu, Y . Jiang, and M. Huang, “Towards emotional support dialog systems,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Languag...

  13. [21]

    Stickerconv: Generating multimodal empathetic responses from scratch,

    Y . Zhang, F. Kong, P. Wang, S. Sun, S. SWangLing, S. Feng, D. Wang, Y . Zhang, and K. Song, “Stickerconv: Generating multimodal empathetic responses from scratch,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Paper...

  14. [22]

    A large-scale dataset for empathetic response generation,

    A. Welivita, Y . Xie, and P. Pu, “A large-scale dataset for empathetic response generation,” in Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing . Association for Computational Linguistics, 2021, pp. 1251–1264

  15. [23]

    Emotionx-ar: Affective empathetic response generation with emotion regulation,

    Z. Wang, Y . Liu, and T. Zhao, “Emotionx-ar: Affective empathetic response generation with emotion regulation,” in Proceedings of EMNLP 2023, 2023, pp. 4532–4545

  16. [24]

    A computational framework for understanding empathy in conversational ai,

    S. Sabour, C. Zheng, and M. Huang, “A computational framework for understanding empathy in conversational ai,” in Proceedings of ACL 2022, 2022, pp. 2537–2549

  17. [25]

    Harnessing the power of large language models for empathetic response generation: Empirical investigations and improvements,

    Y . Qian, W. Zhang, and T. Liu, “Harnessing the power of large language models for empathetic response generation: Empirical investigations and improvements,” in Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, 2...

  18. [26]

    Multi-dimensional evaluation of empathetic di- alogue responses,

    Z. Xu and J. Jiang, “Multi-dimensional evaluation of empathetic di- alogue responses,” in Findings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics, 2024, pp. 2066–2087

  19. [27]

    Computational approaches to empathy: A survey,

    A. Sharma, A. S. Miner, D. C. Atkins, and T. Althoff, “Computational approaches to empathy: A survey,” ACM Computing Surveys , vol. 53, no. 5, pp. 1–36, 2020

  20. [28]

    A multi-modal open dataset for mental- disorder analysis,

    H. Cai, Z. Yuan, Y . Gao et al., “A multi-modal open dataset for mental- disorder analysis,” Scientific Data, vol. 9, no. 1, p. 178, 2022

  21. [29]

    An empirical study of clinical note generation from doctor-patient encounters,

    A. Ben Abacha, W.-w. Yim, Y . Fan, and T. Lin, “An empirical study of clinical note generation from doctor-patient encounters,” in Proceedings of EACL, 2023

  22. [30]

    Multimodal emotion recognition: A survey of methods, datasets, and challenges,

    T. Zhang, S. Sclaroff, and M. Betke, “Multimodal emotion recognition: A survey of methods, datasets, and challenges,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 3, pp. 3704– 3724, 2023

  23. [31]

    Instructblip: Towards general-purpose vision- language models with instruction tuning,

    W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision- language models with instruction tuning,” Advances in Neural Infor- mation Processing Systems , vol. 36, 2023

  24. [32]

    A multimodal investigation of emotional responding in alexithymia,

    R. M. Bagby and G. J. Taylor, “A multimodal investigation of emotional responding in alexithymia,” Cognition and Emotion , vol. 18, no. 6, pp. 781–799, 2004

  25. [33]

    Identifying complex emotions in alexithymia affected adolescents using machine learning techniques,

    S. Gannouni et al. , “Identifying complex emotions in alexithymia affected adolescents using machine learning techniques,” Diagnostics, vol. 12, no. 12, p. 3188, 2022

  26. [34]

    Contextual emotion recognition using large vision language models,

    Y . Etesam et al. , “Contextual emotion recognition using large vision language models,” arXiv preprint arXiv:2405.08992 , 2024

  27. [35]

    Leveraging vision transformers and entropy-based attention for accurate micro-expression recognition,

    Z. Wang et al. , “Leveraging vision transformers and entropy-based attention for accurate micro-expression recognition,” Scientific Reports, vol. 15, 2025

  28. [36]

    Multimodal emotion recognition: A comprehensive review, trends, and challenges,

    M. P. A. Ramaswamy and S. Palaniswamy, “Multimodal emotion recognition: A comprehensive review, trends, and challenges,” WIREs Data Mining and Knowledge Discovery , vol. 14, no. 6, p. e1563, 2024

  29. [37]

    Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, C. Xiong, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 730–19 742

  30. [38]

    Multimodal emotion recognition in social media: Integrating visual and textual cues with vision-language models,

    A. Gandhi, R. Sharma, N. Patel, and S. Kumar, “Multimodal emotion recognition in social media: Integrating visual and textual cues with vision-language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . Association for Computat...

  31. [39]

    Vision-language models for mental health assessment: Analyzing visual and textual indicators of depression and anxiety,

    S. Yoon, J. Kim, S. Lee, and C. Park, “Vision-language models for mental health assessment: Analyzing visual and textual indicators of depression and anxiety,” Journal of Medical Internet Research , vol. 25, p. e45678, 2023

  32. [40]

    Enhancing healthcare communication assessment with vision-language models: Applications in patient-provider interactions,

    L. Chen, J. Williams, C. Martinez, and D. Thompson, “Enhancing healthcare communication assessment with vision-language models: Applications in patient-provider interactions,” npj Digital Medicine , vol. 6, no. 1, p. 234, 2023

  33. [41]

    Towards understanding and mitigating social biases in language models for conversation ai,

    A. Sharma, I. W. Lin, A. S. Miner, D. C. Atkins, and T. Althoff, “Towards understanding and mitigating social biases in language models for conversation ai,” Nature Machine Intelligence, vol. 5, no. 7, pp. 656– 670, 2023

  34. [42]

    Multi-dimensional evaluation of empathetic di- alogue responses,

    Z. Xu and J. Jiang, “Multi-dimensional evaluation of empathetic di- alogue responses,” in Findings of the Association for Computational Linguistics: EMNLP 2024 . Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 2066–2087

  35. [43]

    How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation,

    C.-W. Liu, R. Lowe, I. Serban, M. Noseworthy, L. Charlin, and J. Pineau, “How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation,” Pro- ceedings of the 2016 Conference on Empirical Methods in Natural Lan...

  36. [44]

    Psychological metrics for dialog system evaluation,

    S. Giorgi, J. Sedoc, L. Ungar, S. Buechel, A. Buffone, H. A. Schwartz, S. Dill, A. Ramakrishna, and V . Ganesan, “Psychological metrics for dialog system evaluation,” arXiv preprint arXiv:2305.14757 , 2023

  37. [45]

    Perceived empathy of technology scale (pets): Measuring empathy of systems toward the user,

    “Perceived empathy of technology scale (pets): Measuring empathy of systems toward the user,” in Proceedings of the CHI Conference on Human Factors in Computing Systems . ACM, 2024

  38. [46]

    Third-party evaluators perceive ai as more compassionate than expert humans,

    D. Ovsyannikova, V . de Mello, and M. Inzlicht, “Third-party evaluators perceive ai as more compassionate than expert humans,” Communica- tions Psychology, vol. 3, no. 4, 2025

  39. [47]

    Algorithms for empathy: Using ma- chine learning to categorize common empathetic traits across profes- sional and peer-based conversations,

    S. Provence and A. Forcehimes, “Algorithms for empathy: Using ma- chine learning to categorize common empathetic traits across profes- sional and peer-based conversations,” PMC, 2024

  40. [48]

    K. K. Fitzpatrick, A. Darcy, and M. Vierhile, “Delivering cognitive be- havior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (woebot): a randomized controlled trial,” JMIR mHealth and uHealth , vol. 5, no. 6, p. e4...

  41. [49]

    Mental health surveillance over social media with digital cohorts,

    S. Chancellor, Z. Lin, E. Goodman, S. Zerwas, and M. De Choudhury, “Mental health surveillance over social media with digital cohorts,” in Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems. ACM, 2016, pp. 4306–4317

  42. [50]

    Empathy toward artificial intelligence versus human experiences and the role of transparency in mental health and social support chatbot design: Comparative study,

    J. Shen, D. DiPaola, S. Ali, M. Sap, H. W. Park, C. Breazeal et al. , “Empathy toward artificial intelligence versus human experiences and the role of transparency in mental health and social support chatbot design: Comparative study,” JMIR Mental Health , vol. 11, no. 1, p. e...

  43. [51]

    The role of empathy in psychoanalytic psychother- apy: A historical exploration,

    G. Kaluzeviciute, “The role of empathy in psychoanalytic psychother- apy: A historical exploration,” Cogent Psychology , vol. 7, no. 1, p. 1748792, 2020

  44. [52]

    Evidence based relationships 4: Empathy, congruence, unconditional positive regard, and real relationship,

    D. Mahon, “Evidence based relationships 4: Empathy, congruence, unconditional positive regard, and real relationship,” in Evidence based counselling & psychotherapy for the 21st century practitioner. Emerald Publishing Limited, 2023, pp. 71–83

  45. [53]

    Silent signals: New review highlights the impor- tance of nonverbal signals for perceived responsiveness,

    R. Ramadurai, “Silent signals: New review highlights the impor- tance of nonverbal signals for perceived responsiveness,” https://www. evidencebasedmentoring.org, 2024, accessed Sep 2025

  46. [54]

    Emotion in psychotherapy: Affect, cognition, and the process of change,

    L. S. Greenberg and J. D. Safran, “Emotion in psychotherapy: Affect, cognition, and the process of change,” Psychotherapy, vol. 24, no. 3, pp. 253–264, 1987

  47. [55]

    Empathic accuracy,

    W. Ickes, “Empathic accuracy,” Journal of personality , vol. 61, no. 4, pp. 587–610, 1993

  48. [56]

    The generalizability of the psychoanalytic concept of the working alliance,

    E. S. Bordin, “The generalizability of the psychoanalytic concept of the working alliance,” Psychotherapy: Theory, research & practice, vol. 16, no. 3, pp. 252–260, 1979

  49. [57]

    Alliance in individual psychotherapy,

    A. O. Horvath, A. Del Re, C. Fl ¨uckiger, and D. Symonds, “Alliance in individual psychotherapy,” Psychotherapy, vol. 48, no. 1, pp. 9–16, 2011

  50. [58]

    A circumplex model of affect,

    J. A. Russell, “A circumplex model of affect,” Journal of personality and social psychology , vol. 39, no. 6, pp. 1161–1178, 1980

  51. [59]

    Measuring emotion: the self-assessment manikin and the semantic differential,

    M. M. Bradley and P. J. Lang, “Measuring emotion: the self-assessment manikin and the semantic differential,” Journal of behavior therapy and experimental psychiatry, vol. 25, no. 1, pp. 49–59, 1994

  52. [60]

    Towards em- pathetic open-domain conversation models: A new benchmark and dataset,

    H. Rashkin, E. M. Smith, M. Li, and Y .-L. Boureau, “Towards em- pathetic open-domain conversation models: A new benchmark and dataset,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 5370–5381

  53. [61]

    What matters when building vision-language models?

    H. Laurenc ¸on, L. Saulnier, L. Tronchon, V . Saulnier, C. Akiki, A. Villegas, M. Haddad, L. Barrault, X. Bresson, A. F. Aji et al. , “What matters when building vision-language models?” arXiv preprint arXiv:2405.02246, 2024

  54. [62]

    Video-llava: Learning united visual representation by alignment before projection,

    B. Lin, B. Zhu, Y . Ye, M. Ning, P. Jin, and L. Yuan, “Video-llava: Learning united visual representation by alignment before projection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 2262–2272

  55. [63]

    Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, C. Xiong, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” in International conference on machine learning , 2023, pp. 19 730–19 742. IEEE TRANSACTIONS ON [JOURNAL NAME], VOL. X, NO. Y , [MONTH...

  56. [64]

    Mm-llms: Recent advances in multimodal large language models,

    D. Zhang, Y . Yu, J. Dong, C. Li, D. Su, C. Chu, and D. Yu, “Mm-llms: Recent advances in multimodal large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2401.13601

  57. [65]

    Lora: Low-rank adaptation of large language models,

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2021

  58. [66]

    Empathy,

    R. Elliott, A. C. Bohart, J. C. Watson, and L. S. Greenberg, “Empathy,” Psychotherapy, vol. 48, no. 1, pp. 43–49, 2011

  59. [67]

    Level of emotional awareness and mean length of utterance,

    J. A. Hall, S. Carter, M. C. Jimenez, and N. A. Frost, “Level of emotional awareness and mean length of utterance,” Emotion, vol. 1, no. 4, pp. 325–333, 2001

  60. [68]

    Cognitive load during problem solving: Effects on learning,

    J. Sweller, “Cognitive load during problem solving: Effects on learning,” Cognitive science, vol. 12, no. 2, pp. 257–285, 1988

  61. [69]

    Language features for automated evaluation of cognitive behavior psychotherapy sessions,

    N. Flemotomos, V . R. Martinez, J. Gibson, D. C. Atkins, T. A. Creed, and S. Narayanan, “Language features for automated evaluation of cognitive behavior psychotherapy sessions,” in Interspeech, 2018, pp. 1908–1912

  62. [70]

    Bertscore: Evaluating text generation with bert,

    T. Zhang, V . Kishore, F. Wu, K. Q. Weinberger, and Y . Artzi, “Bertscore: Evaluating text generation with bert,” in International Conference on Learning Representations, 2020

  63. [71]

    The empathy cycle: Refinement of a nuclear concept,

    G. T. Barrett-Lennard, “The empathy cycle: Refinement of a nuclear concept,” Journal of Counseling Psychology, vol. 28, no. 2, pp. 91–100, 1981

  64. [72]

    Artificial empathy in healthcare chatbots: Does it feel authentic?

    J. Luo, J. Huang, and H. Li, “Artificial empathy in healthcare chatbots: Does it feel authentic?” Computers in Human Behavior Reports , vol. 6, p. 100347, 2024

  65. [73]

    Egan, The skilled helper: A problem-management and opportunity- development approach to helping , 10th ed

    G. Egan, The skilled helper: A problem-management and opportunity- development approach to helping , 10th ed. Brooks/Cole, 2014

  66. [74]

    Toward effective counseling and psychotherapy: Training and practice,

    C. B. Truax and R. R. Carkhuff, “Toward effective counseling and psychotherapy: Training and practice,” 1967

  67. [75]

    The jefferson scale of empathy: development and preliminary psychometric data,

    M. Hojat, S. Mangione, T. J. Nasca, M. J. Cohen, J. S. Gonnella, J. B. Erdmann, J. J. Veloski, and M. Magee, “The jefferson scale of empathy: development and preliminary psychometric data,” Educational and Psychological Measurement , vol. 61, no. 2, pp. 349–365, 2001

  68. [76]

    The development and preliminary validation of the consultation and relational empathy (care) scale for use in primary care,

    S. W. Mercer and W. J. Reynolds, “The development and preliminary validation of the consultation and relational empathy (care) scale for use in primary care,” Family Practice, vol. 21, no. 6, pp. 699–705, 2004

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.