Pith. sign in

REVIEW 4 major objections 5 minor 88 references

Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read People disclose the same themes to a social robot as to a human therapist, and the robot's replies semantically match therapist responses.

desk verdict The cross-agent comparison is new and honestly discussed, but both headline numbers rest on thresholds whose null behavior is untested, so the alignment claim is not yet established. read the letter →

arxiv 2506.16473 v1 pith:FSHBBMIM submitted 2025-06-19 cs.HC cs.AIcs.CL

classification cs.HCcs.AIcs.CL
keywords human-robotinteractionemotionalsupportsemanticalignmentself-disclosurelargelanguagemodelstherapeuticdialoguesentenceembeddingsclusteranalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether supportive conversations with a large-language-model-powered social robot look like real therapy sessions, both in what people talk about and in how the helper responds. Comparing a public text-therapy corpus with logs from a five-session robot support intervention with university students, it reports that 90.88% of the disclosures made to the robot fit inside themes derived from the therapy data, and that for those matched themes, subject disclosures and agent responses score between about 0.74 and 0.91 in semantic similarity under Word2Vec and BERT embeddings, measured against a 0.5 chance baseline. The authors read this as evidence that people bring similar concerns to robots and therapists and that LLM-driven robots can produce semantically therapist-like replies, while also reporting that the reverse mapping is low (28.04%) and that most robot disclosures fall into one broad therapy cluster. If the finding holds, social robots could be evaluated against therapy-derived topic and language baselines and deployed as a complementary first layer of mental-health support.

What carries the argument

The load-bearing mechanism is a cluster-fit-then-compare pipeline. Each corpus of disclosures is embedded with a sentence-transformer model and clustered with K-means (seven clusters for therapy, six for the robot), and every cluster receives a capture radius equal to twice the average Euclidean distance of its own points to its centroid; a disclosure from the other corpus counts as aligned when it falls inside that radius. Matched clusters are then compared by mean pairwise semantic similarity across three embedding families—Word2Vec, BERT, and a sequence Transformer—and the LLM-generated cluster labels are validated by checking that each label description is closest to its own centroid. The cluster geometry supplies the units of 'topic,' and the embedding similarity supplies the numbers for 'responding alike.'

What would settle it

Run the identical pipeline on an age- and prompt-matched corpus—for example, the same participants speaking with a licensed therapist under the same structured prompts, or mid-life adults doing the robot intervention with open-ended prompts—and check whether the 90.88% mapping rate and the Word2Vec/BERT similarity scores above 0.74 survive. If they fall toward the 0.5 chance baseline or the 28.04% reverse-mapping level, the alignment is an artifact of demographics or interaction structure rather than of agent type.

Watch

Extended reading notes

Core claim

The central claim is that content-level alignment between robot-led and therapist-led support is real and measurable in both directions: people's disclosures to the robot and to the therapist organize into shared themes, and the robot's replies semantically echo those of human therapists. With the per-cluster acceptance threshold set to twice the average member-to-centroid distance, 90.88% of the 560 robot-corpus disclosures map onto clusters built from the therapy data, versus 28.04% in the other direction; for the matched clusters, mean pairwise semantic similarity under Word2Vec (roughly 0.76 to 0.91) and BERT (roughly 0.74 to 0.82) is significantly above the 0.5 chance baseline at p<0.001, while a sequence Transformer stays near chance. The authors attribute the asymmetry to the robot intervention's structured prompts, the students' developmental stage, and the breadth of the largest therapy cluster, and they explicitly frame the result as evidence about linguistic alignment rather than about therapeutic equivalence.

Load-bearing premise

The load-bearing premise is that the measured overlap between robot and therapy conversations comes from the shared substance of supportive talk, not from the accidental difference that the robot data were collected from university students in a structured wellbeing exercise while the therapy data came from mid-life adults in open-ended counselling.

Editorial extensions

If this is right

  • Robot support programs can be benchmarked against a therapy-derived taxonomy of disclosure themes rather than against user satisfaction alone.
  • The high semantic similarity of robot and therapist responses supports using an LLM-powered social robot as an accessible first tier of emotional support in settings where human therapists are scarce.
  • The low reverse mapping rate (28.04%) means current robot interventions elicit a narrower thematic range than open therapy, so future designs should use broader or adaptive prompts to cover the full space of concerns.
  • The model-dependence of the similarity scores (Word2Vec and BERT above chance, sequence Transformer near chance) locates the alignment at the level of word-distribution and contextual meaning rather than surface syntax.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural check the paper does not run is to shrink the acceptance radius from twice to 1.5 or 1.0 times the average member distance; if the 90.88% mapping collapses quickly, most of the claimed overlap is the breadth of the largest therapy cluster rather than genuine topical fit.
  • The paper's own diagnosis predicts an experiment: hold the participant population fixed and vary only the prompt structure; if the mapping rate tracks prompt breadth, the thematic overlap is partly an artifact of the intervention design rather than a property of talking to a robot.
  • The authors' framing implies a boundary test: give identical disclosures to a therapist and a robot and measure both semantic similarity and felt empathy; the prediction would be that language aligns while relational quality diverges.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper compares human-to-human (H2H) therapy conversations from the CounselChat corpus with human-to-robot (H2R) supportive conversations from a structured PERMA-based intervention with a QTrobot powered by GPT-3.5. The authors embed disclosures with MiniLM, cluster each dataset with K-means, and then use a distance-based cluster-fit rule (threshold alpha = 2.0) to claim that 90.88% of H2R disclosures map onto H2H clusters. They further compute pairwise semantic similarity between mapped H2R and H2H units using Transformer, Word2Vec, and BERT embeddings, reporting high Word2Vec (0.81–0.91) and BERT (0.74–0.82) similarities for both subject disclosures and agent responses, while Transformer similarities are low (0.13–0.35). The paper concludes that robot-led supportive conversations show thematic and semantic alignment with human therapy, with some asymmetry and contextual limitations.

Significance. If the central claims held, the study would be a useful first large-scale comparison of the topical and semantic structure of robot-led versus human-delivered emotional support, with practical implications for designing supportive conversational agents. The paper is transparent about several limitations, including the demographic and contextual mismatch between the two datasets, and it commits to releasing analysis scripts on OSF upon acceptance. It also uses multiple embedding models, which is a strength. However, the two headline quantitative results—the 90.88% mapping rate and the 'above chance' semantic similarities—depend on thresholds that are not validated against any empirical null distribution, and one of the three embedding models contradicts the strong-alignment conclusion. These issues are load-bearing for the paper's central claims and require additional analysis before the conclusions can be accepted.

major comments (4)
  1. [Section 3.3.1 and Section 4.2, Eq. (1)] The cluster-fit threshold T_j = alpha times the mean intra-cluster distance is set to alpha = 2.0 without sensitivity analysis or a null control, so the headline result that 90.88% of H2R disclosures map onto H2H clusters does not by itself establish shared topical structure. Since 472 of the 508 mapped responses fall into Cluster 1, which has the largest Mahalanobis mean (12.98, Table 1) and therefore the largest capture radius, the high mapping rate may be a geometric artifact of one diffuse cluster rather than evidence of thematic alignment. The authors should report mapping rates across a range of alpha values, for random English sentences, and for out-of-domain text, and they should derive the threshold from an explicit null distribution rather than asserting it.
  2. [Section 3.3.2 and Tables 3-4] The one-sample t-tests compare Word2Vec and BERT cosine similarities against a fixed threshold of 0.5 described as 'above random chance,' but this threshold is not derived from any null model for these embedding spaces. Unrelated sentences can easily exceed 0.5 in high-dimensional embeddings, especially after stop-word removal, so the reported means of 0.74 to 0.91 do not establish that robot and therapist language are more similar than two generic texts. The authors should construct an empirical baseline, for example by computing similarities between randomly paired disclosures from different topics or between original and shuffled texts, and report effect sizes and confidence intervals relative to that baseline.
  3. [Section 3.1 and Section 5.4] The H2R corpus is drawn from 21 university students in a structured, prompt-driven PERMA intervention, while the H2H corpus consists of open-ended therapy exchanges from mid-life adults; the paper asserts in Section 3.1 that thematic and semantic alignment is 'theoretically separable from life-stage-specific concerns,' but this assumption is load-bearing and untested. The observed overlap, particularly the concentration of H2R disclosures in Cluster 1 ('Anxieties and Self-Perception Struggles'), could plausibly be driven by the broadness of that cluster or by the intervention's prompts rather than by a genuine equivalence in how people disclose to robots versus therapists. A matched subset analysis, such as restricting H2H disclosures to student-age or self-focused concerns, or comparing only H2R responses to similarly prompted H2H exchanges, would substantially strengthen the central claim.
  4. [Section 4.3 and Table 4] The conclusion of strong semantic overlap is model-dependent: Sequence Transformer similarities are 0.13 to 0.35 and never exceed the 0.5 threshold, which the authors acknowledge in Section 5.2 but do not resolve. Because the three embedding models are presented as complementary evidence, the consistent failure of the Transformer model should be addressed explicitly, for example by reporting its empirical null distribution and explaining why its lower scores are expected under the authors' theoretical account, rather than treating the Word2Vec and BERT results as decisive.
minor comments (5)
  1. [Section 5.4] The sentence beginning 'Nevertheless, limitations remain' appears twice in succession; one copy should be deleted.
  2. [Tables 3 and 4] The notes under both tables contain the typo 'There results are consistently high' and should read 'These results are consistently high.'
  3. [Section 4.1.3 and Table 2] The validation of LLM-generated cluster descriptions by measuring their similarity to the cluster centroids is circular if the descriptions were generated from those same centroids or their member texts; the authors should clarify the generation procedure and report a non-circular baseline, such as similarity to random text or to centroids from a held-out clustering.
  4. [Abstract and Section 3.2] The abstract says the cluster-fit method was 'validating it using euclidean distances,' but Section 3.3.1 reports Euclidean distances only as descriptive statistics for mapped responses; please clarify whether Euclidean distance is meant to validate the method or simply characterize the fitted responses.
  5. [Section 3.3.1] The 'rule of thumb' that Euclidean distances above 1.0 indicate increasing deviation and 0.5–1.0 indicate moderate proximity is not calibrated to the embedding space used here and is not used in any statistical test; please either remove this heuristic or justify it with a citation or calibration experiment.

Circularity Check

3 steps flagged · score 5.0 of 10

Semantic-alignment results are partly forced by the proximity-based inclusion rule: the H2R disclosures compared against H2H clusters are selected for closeness to those cluster centroids, and the cluster-label validation checks descriptions against the same centroids used to generate them.

  1. self definitional [Section 3.3.1–3.3.2 (cluster-fitting threshold and semantic similarity on mapped responses)]
    "we identified the closest centroid for each response, and applied a threshold value to determine fit. If the distance to the nearest centroid exceeded this threshold, we classify the response as not belonging to any cluster. ... For responses that could be mapped to a cluster of the other agent type, we combined these responses with the already existing responses of said cluster. ..."

    The H2R subject responses entering the semantic-similarity comparison are precisely those whose embedding distance to the chosen H2H centroid satisfies T_j = alpha times the mean intra-cluster distance. Since H2H cluster members are by definition centered on that same centroid, any response surviving the proximity filter will have artificially elevated pairwise similarity to those cluster members. The reported Word2Vec/BERT subject-disclosure alignment is therefore partly a re-description of the inclusion criterion rather than an independent estimate of cross-agent semantic overlap.

  2. fitted input called prediction [Section 3.3.2 (semantic similarity one-sample t-test)]
    "One-sample t-tests were conducted to evaluate whether the semantic similarity scores were significantly greater than a predefined threshold of 0.5, which served as a baseline for similarity above random chance."

    The 'chance' baseline is a predefined constant, not an empirical null estimated from random or out-of-domain text. Declaring scores above 0.5 to be 'above random chance' makes the qualitative conclusion follow from the chosen constant; a different threshold would invert the finding. Thus the reported above-chance result is an artifact of the assumed 0.5 cutoff rather than a test against an independently estimated chance distribution.

1 more flagged steps
  1. self definitional [Section 3.2 and Section 4.1.3 (cluster explanation and validation)]
    "The clusters were explained using GPT-4o-mini [49], which was prompted to provide a label and description for each cluster via the cluster explanation procedure described in [10]. The explanations provided were then validated using the procedure outlined in [10]. ... For both datasets, cluster descriptions consistently exhibited the highest semantic similarity to their corresponding centroids (see [10]), indicating that the LLM-generated labels and descriptions effectively captured the core semantics of each cluster (see Table 2)."

    The LLM descriptions are generated from the very clusters whose centroids are then used as the reference for validation. A description that paraphrases the cluster's own average content will by construction have high similarity to that cluster's centroid, so Table 2 demonstrates self-consistency rather than independent evidence that the labels capture externally meaningful themes. The procedure is also imported from the authors' own prior work ([10]) without an independent check, making this validation self-referential.

full rationale

Several headline results are not circular in themselves: the 90.88% mapping rate is computed from data with a stated threshold, and the Word2Vec/BERT similarity scores are measured from real text pairs rather than recycled from the chosen threshold. However, the construction of the semantic-alignment comparison introduces partial circularity: only H2R disclosures that pass a proximity filter to H2H centroids are entered into the similarity analysis, so high subject-similarity scores are partly guaranteed by that selection rule. The 'above chance' conclusion is likewise tied to an arbitrarily predefined 0.5 baseline rather than an empirical null distribution, and the cluster-label validation checks descriptions against the same centroids used to produce them, via a same-author citation. These issues affect the central alignment claim, especially RQ2, enough to warrant a moderate score, even though the raw data collection, embedding computation, and the reverse-mapping asymmetry are not themselves circular.

Assumptions & free parameters 5 free parameters · 7 assumptions · 0 invented entities

The central claim depends on several free parameters and domain assumptions. The most consequential are the arbitrarily chosen alpha threshold, the uncalibrated 0.5 similarity baseline, and the explicit assumption that demographic and contextual differences between the two corpora do not drive the observed alignment. No new entities are introduced.

free parameters (5)
  • alpha (cluster-fit threshold multiplier) = 2.0
    Sets each cluster's membership threshold T_j = alpha times the mean internal distance. The headline 90.88% mapping rate is a direct function of this hand-chosen value; no sensitivity analysis is reported.
  • semantic similarity chance threshold = 0.5
    Used as the comparison value in one-sample t-tests; the paper asserts this as 'baseline for similarity above random chance' without deriving it from a null distribution of cross-theme pairs.
  • number of clusters (H2H) = 7
    Selected via elbow method on WCSS; cluster structure, and hence the mapping statistics, depend on this choice.
  • number of clusters (H2R) = 6
    Selected via elbow method; similarly affects cluster composition and mapping.
  • Euclidean distance rule-of-thumb thresholds = 0.5 to 1.0
    Used to interpret average Euclidean distances of mapped responses as 'moderate proximity'; not calibrated to the embedding space or derived from the data.
assumptions (7)
  • domain assumption Sentence embeddings from all-MiniLM-L6-v2 capture the semantic content of short mental-health disclosures.
    Used as the input to K-means; no qualitative validation of embedding quality is provided.
  • domain assumption K-means clusters in embedding space correspond to interpretable thematic categories.
    The entire analysis treats cluster membership as thematic topic; the elbow method provides no guarantee that clusters are semantically meaningful.
  • domain assumption Word2Vec and BERT cosine similarities are valid measures of semantic alignment for supportive dialogue.
    The paper's similarity conclusions rest on these embeddings; the Transformer model gives contradictory low scores, indicating model-dependence.
  • domain assumption LLM-generated cluster labels and descriptions accurately characterize cluster content.
    Labels are generated by GPT-4o-mini and validated by comparing them to the same centroids that define the clusters, following the authors' prior procedure [10].
  • domain assumption The CounselChat dataset represents professional human therapy.
    Therapists are claimed to be verified and licensed on the platform, but the paper presents no independent verification of this.
  • ad hoc to paper Thematic alignment is separable from life-stage and interaction-context differences between the two datasets.
    Section 3.1 explicitly assumes this separability; it is load-bearing because the H2R and H2H populations and interaction structures differ substantially.
  • domain assumption The elbow method yields the correct number of clusters for both datasets.
    The WCSS values decrease gradually; the elbow is not sharply defined (e.g., H2H: 748.64, 738.86, 720.62...), making the choice somewhat arbitrary.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support." pith.science (2026). https://pith.science/paper/FSHBBMIM

@misc{pith2026250616473,
  author       = {Pith},
  title        = {Pith review of: Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FSHBBMIM}},
  note         = {Machine review of arXiv:2506.16473}
}
read the original abstract

As conversational agents increasingly engage in emotionally supportive dialogue, it is important to understand how closely their interactions resemble those in traditional therapy settings. This study investigates whether the concerns shared with a robot align with those shared in human-to-human (H2H) therapy sessions, and whether robot responses semantically mirror those of human therapists. We analyzed two datasets: one of interactions between users and professional therapists (Hugging Face's NLP Mental Health Conversations), and another involving supportive conversations with a social robot (QTrobot from LuxAI) powered by a large language model (LLM, GPT-3.5). Using sentence embeddings and K-means clustering, we assessed cross-agent thematic alignment by applying a distance-based cluster-fitting method that evaluates whether responses from one agent type map to clusters derived from the other, and validated it using Euclidean distances. Results showed that 90.88% of robot conversation disclosures could be mapped to clusters from the human therapy dataset, suggesting shared topical structure. For matched clusters, we compared the subjects as well as therapist and robot responses using Transformer, Word2Vec, and BERT embeddings, revealing strong semantic overlap in subjects' disclosures in both datasets, as well as in the responses given to similar human disclosure themes across agent types (robot vs. human therapist). These findings highlight both the parallels and boundaries of robot-led support conversations and their potential for augmenting mental health interventions.

Figures

Figures reproduced from arXiv: 2506.16473 by the authors.

Figure 1
Figure 1. The deployment settings. Image from [38]. 3.2 Preprocessing and Clustering Both datasets had to be preprocessed before clustering was applied. Stop words were filtered from the data using the ENGLISH_STOP_WORDS list from sklearn [50] after tokenizing responses by whitespace. Ad￾ditionally, duplicate responses were removed from each dataset to prevent redundancy. Each data unit (i.e. a response from the subject) was … view at source ↗
Figure 2
Figure 2. Mean semantic similarity of subject disclosures across clusters 4.3.2 Therapist Response Similarity Across Datasets. For each com￾bined cluster, we calculated the mean pairwise semantic similarity between the human therapist responses from the H2H exchange and the robot response from the H2R interaction (see [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Mean semantic similarity of agent responses across clusters 5.1 Aligning Disclosure Themes Across Agents Our findings demonstrate a high degree of thematic overlap be￾tween the human-to-robot (H2R) and human-to-human (H2H) con￾versations (RQ1). Remarkably, over 90% of participant disclosures made to the robot could be mapped to clusters derived from human therapy sessions. This suggests that participants bring simil… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

88 extracted references · 50 canonical work pages

  1. [1]

    Katie Aafjes-van Doorn, John Porcerelli, and Lena Christine Müller-Frommeyer

  2. [2]

    1973.Social penetration: The development of interpersonal relationships

    Irwin Altman and Dalmas A Taylor. 1973.Social penetration: The development of interpersonal relationships. Holt, Rinehart & Winston, Oxford, England. viii, 212–viii, 212 pages

  3. [3]

    Bruno Sanchez de Araujo, Marcelo Fantinato, Sarajane Marques Peres, Ruth Caldeira de Melo, Samila Sathler Tavares Batistoni, Meire Cachioni, and Patrick C.K. Hung. 2021. Effects of social robots on depressive symptoms in older adults: a scoping review.Library Hi Tech(2021). doi:10.1108/LHT-09-2020- 0244/FULL/PDF

  4. [4]

    um" to "yeah

    Claire Augusta Bergey and Simon DeDeo. 2024. From "um" to "yeah": Producing, predicting, and regulating information flow in human conversation. (3 2024). https://arxiv.org/abs/2403.08890v1

  5. [5]

    Nicolas Bertagnolli. 2020. Counsel chat: Bootstrapping high-quality ther- apy data. https://medium.com/data-science/counsel-chat-bootstrapping-high- quality-therapy-data-971b419f33da

  6. [6]

    Bishop and Andrew C

    Rachael E. Bishop and Andrew C. High. 2023. Stigma and Supportive Commu- nication in the Context of Mental or Emotional Distress: An Extension of the Paradox of Support Seeking in Close Relationships.Communication Research (2023). doi:10.1177/00936502231189811

  7. [7]

    Borelli, Lucas Sohn, Binghuang A

    Jessica L. Borelli, Lucas Sohn, Binghuang A. Wang, Kajung Hong, Cindy Decoste, and Nancy E. Suchman. 2019. Therapist-client language matching: Initial promise as a measure of therapist-client relationship quality.Psychoanalytic Psychology 36, 1 (1 2019), 9–18. doi:10.1037/PAP0000177

  8. [8]

    Antonin Brun, Ruying Liu, Aryan Shukla, Frances Watson, and Jonathan Gratch

Show all 88 references
  1. [9]

    Shu Chuan Chen, Cindy Jones, and Wendy Moyle. 2018. Social Robots for Depression in Older Adults: A Systematic Review.Journal of Nursing Scholarship 50, 6 (11 2018), 612–622. doi:10.1111/JNU.12423

  2. [10]

    Cross, and Hatice Gunes

    Sophie Chiang, Guy Laban, Emily S. Cross, and Hatice Gunes. 2025. Comparing Self-Disclosure Themes and Semantics to a Human, a Robot, and a Disembodied Agent. In2025 34nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)

  3. [11]

    Clark and Susan E

    Herbert H. Clark and Susan E. Brennan. 1991. Grounding in communication. Perspectives on socially shared cognition.(10 1991), 127–149. doi:10.1037/10096-006

  4. [12]

    Paul C. Cozby. 1973. Self-disclosure: A literature review.Psychological Bulletin 79, 2 (2 1973), 73–91. doi:10.1037/H0033950

  5. [13]

    Roy De Maesschalck, Delphine Jouan-Rimbaud, and Désiré L Massart. 2000. The mahalanobis distance.Chemometrics and intelligent laboratory systems50, 1 (2000), 1–18

  6. [14]

    Jacob Devlin, Ming Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL HLT 20191 (10 2018), 4171–4186. https://arxiv.org/abs/1810.04805v2

  7. [15]

    Guillermo Gallego, Carlos Cuevas, Raúl Mohedano, and Narciso García. 2013. On the mahalanobis distance classification criterion for multidimensional normal distributions.IEEE Transactions on Signal Processing61, 17 (9 2013), 4387–4396. doi:10.1109/TSP.2013.2269047

  8. [16]

    Jessica Gasiorek, Ann Weatherall, and Bernadette Watson. 2021. Interactional Adjustment: Three Approaches in Language and Social Psychology.Jour- nal of Language and Social Psychology40, 1 (1 2021), 102–119. doi:10.1177/ 0261927X20965652

  9. [17]

    Gelso and Jean A

    Charles J. Gelso and Jean A. Carter. 1985. The Relationship in Counseling and Psychotherapy.The Counseling Psychologist13, 2 (1985), 155–243. doi:10.1177/ 0011000085132001

  10. [18]

    1991.Contexts of ac- commodation: Developments in applied sociolinguistics

    Howard Giles, Justine Coupland, and Nikolas Coupland. 1991.Contexts of ac- commodation: Developments in applied sociolinguistics. Cambridge University Press

  11. [19]

    Sujatha Das Gollapalli, Beng Heng Ang, and See Kiong Ng. 2023. Identifying Early Maladaptive Schemas from Mental Health Question Texts.Findings of the Association for Computational Linguistics: EMNLP 2023(2023), 11832–11843. doi:10.18653/V1/2023.FINDINGS-EMNLP.792

  12. [20]

    Gonzales, Jeffrey T

    Amy L. Gonzales, Jeffrey T. Hancock, and James W. Pennebaker. 2010. Language Style Matching as a Predictor of Social Dynamics in Small Groups.Communica- tion Research37, 1 (2 2010), 3–19. doi:10.1177/0093650209351468

  13. [21]

    Anna Henschel, Guy Laban, and Emily S Cross. 2021. What Makes a Robot Social? A Review of Social Robots from Science Fiction to a Home or Hospital Near You. Current Robotics Reports2 (2021), 9–19. doi:10.1007/s43154-020-00035-0

  14. [22]

    George Caspar. Homans. 1958. Social Behavior as Exchange. https://doi.org/10.1086/22235563, 6 (5 1958), 597–606. doi:10.1086/222355

  15. [23]

    Imel, Michael J

    Zac E. Imel, Michael J. Tanana, Christina S. Soma, Thomas D. Hull, Brian T. Pace, Sarah C. Stanco, Torrey A. Creed, Theresa B. Moyers, and David C. Atkins. 2024. Mental Health Counseling From Conversational Content With Transformer- Based Machine Learning.JAMA Network Open7, 1...

  16. [24]

    Ireland, Richard B

    Molly E. Ireland, Richard B. Slatcher, Paul W. Eastwick, Lauren E. Scissors, Eli J. Finkel, and James W. Pennebaker. 2011. Language Style Matching Predicts Relationship Initiation and Stability.Psychological Science22, 1 (1 2011), 39–44. doi:10.1177/0956797610392928

  17. [25]

    Bahar Irfan and Gabriel Skantze. 2025. Between you and me: Ethics of self- disclosure in human-robot interaction. InProceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction. 1357–1362

  18. [26]

    Anil K Jain, M Narasimha Murty, and Patrick J Flynn. 1999. Data clustering: a review.ACM computing surveys (CSUR)31, 3 (1999), 264–323

  19. [27]

    S Joshua Johnson, M Ramakrishna Murty, and I Navakanth. 2024. A detailed review on word embedding techniques with emphasis on word2vec.Multimedia Tools and Applications83, 13 (2024), 37979–38007

  20. [28]

    Guy Laban, Ziv Ben-Zion, and Emily S. Cross. 2022. Social Robots for Supporting Post-traumatic Stress Disorder Diagnosis and Treatment.Frontiers in psychiatry 12 (2022). doi:10.3389/FPSYT.2021.752874

  21. [29]

    Guy Laban, Sophie Chiang, and Hatice Gunes. 2025. What People Share With a Robot When Feeling Lonely and Stressed and How It Helps Over Time. In2025 34nd IEEE International Conference on Robot and Human Interactive Communica- tion (RO-MAN)

  22. [30]

    Guy Laban and Emily S. Cross. 2024. Sharing our Emotions with Robots: Why do we do it and how does it make us feel?IEEE Transactions on Affective Computing (2024), 1–18. doi:10.1109/TAFFC.2024.3470984

  23. [31]

    Guy Laban, Jean-Noël George, Val Morrison, and Emily S. Cross. 2021. Tell me more! Assessing interactions with social robots from speech.Paladyn, Journal of Behavioral Robotics12, 1 (2021), 136–159. doi:10.1515/pjbr-2021-0011

  24. [32]

    Guy Laban, Arvid Kappas, Val Morrison, and Emily S. Cross. 2023. Opening Up to Social Robots: How Emotions Drive Self-Disclosure Behavior. In2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO- MAN). IEEE, Busan, Republic of Korea, 1697–170...

  25. [33]

    Guy Laban, Arvid Kappas, Val Morrison, and Emily S Cross. 2024. Building Long-Term Human–Robot Relationships: Examining Disclosure, Perception and Well-Being Across Time.International Journal of Social Robotics16, 5 (2024), 1–27. doi:10.1007/s12369-023-01076-z

  26. [34]

    Guy Laban, Tomer Laban, and Hatice Gunes. 2024. LEXI: Large Language Models Experimentation Interface. InProceedings of the 12th International Conference on Human-Agent Interaction. ACM, New York, NY, USA, 250–259. doi:10.1145/ 3687272.3688296 Sophie Chiang, Guy Laban, and Hat...

  27. [35]

    Guy Laban, Val Morrison, and Emily Cross. 2024. Social Robots for Health Psychology: A New Frontier for Improving Human Health and Well-Being. European Health Psychologist23, 1 (2 2024), 1095–1102. https://www.ehps.net/ ehp/index.php/contents/article/view/3442

  28. [36]

    Guy Laban, Val Morrison, Arvid Kappas, and Emily S. Cross. 2025. Coping with Emotional Distress via Self-Disclosure to Robots: An Intervention with Caregivers.International Journal of Social Robotics(2025). doi:10.1007/s12369- 024-01207-0

  29. [37]

    Guy Laban, Micol Spitale, Minja Axelsson, Nida Itrat Abbasi, and Hatice Gunes

  30. [38]

    Guy Laban, Julie Wang, and Hatice Gunes. 2025. A Robot-Led Interven- tion for Emotion Regulation: From Expression to Reappraisal.arXiv preprint arXiv:2503.18243(2025)

  31. [39]

    Edward J Lawler. 2001. An Affect Theory of Social Exchange.Amer. J. Sociology 107, 2 (2001), 321–352. doi:10.1086/324071

  32. [40]

    (6 2025)

    Critical Insights about Robots for Mental Wellbeing. (6 2025). https: //arxiv.org/pdf/2506.13739

  33. [41]

    Fan Liu and Yong Deng. 2020. Determine the number of unknown targets in open world based on elbow method.IEEE Transactions on Fuzzy Systems29, 5 (2020), 986–995

  34. [42]

    Navid Madani, Sougata Saha, and Rohini Srihari. 2024. Steering Conversational Large Language Models for Long Emotional Support Conversations. (2 2024). https://arxiv.org/abs/2402.10453v2

  35. [43]

    Aristidis Likas, Nikos Vlassis, and Jakob J Verbeek. 2003. The global k-means clustering algorithm.Pattern recognition36, 2 (2003), 451–461

  36. [44]

    Corrado, and Jeff Dean

    Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. NIPS26 (2013)

  37. [45]

    Miner, Scott L

    Adam S. Miner, Scott L. Fleming, Albert Haque, Jason A. Fries, Tim Althoff, Denise E. Wilfley, W. Stewart Agras, Arnold Milstein, Jeff Hancock, Steven M. Asch, Shannon Wiltsey Stirman, Bruce A. Arnow, and Nigam H. Shah. 2022. A computational approach to measure the linguistic ...

  38. [46]

    Vaibhav Mehra, Guy Laban, and Hatice Gunes. 2025. How Large Language Models Classify and Semantically Explain Facial Expressions from Valence- Arousal Values. InProceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). Association for Computing Machine...

  39. [47]

    Tatsuya Nomura, Takayuki Kanda, Tomohiro Suzuki, and Sachie Yamada. 2020. Do people with social anxiety feel anxious about interacting with a robot?AI & SOCIETY35, 2 (2020), 381–390. doi:10.1007/s00146-019-00889-9

  40. [48]

    Ben Ong, Eleftheria Tseliou, Tom Strong, and Niels Buus. 2023. Power and dialogue: A review of discursive research.Family Process62, 4 (12 2023), 1391–

  41. [49]

    Niederhoffer and James W

    Kate G. Niederhoffer and James W. Pennebaker. 2002. Linguistic Style Matching in Social Interaction.Journal of Language and Social Psychology21, 4 (12 2002), 337–360. doi:10.1177/026192702237953

  42. [50]

    Fabian Pedregosa, Vincent Michel, Olivier Grisel, Mathieu Blondel, Peter Pretten- hofer, Ron Weiss, Jake Vanderplas, David Cournapeau, Gaël Varoquaux, Alexan- dre Gramfort, Bertrand Thirion, Vincent Dubourg, Alexandre Passos, Matthieu Brucher, Matthieu Édouardand, and Édouard ...

  43. [51]

    Müller, Eva Wiese, and Agnieszka Wykowska

    Jairo Perez-Osorio, Hermann J. Müller, Eva Wiese, and Agnieszka Wykowska

  44. [52]

    Müller, and Agnieszka Wykowska

    Jairo Perez-Osorio, Hermann J. Müller, and Agnieszka Wykowska. 2017. Ex- pectations regarding action sequences modulate electrophysiological corre- lates of the gaze-cueing effect.Psychophysiology54, 7 (7 2017), 942–954. doi:10.1111/PSYP.12854

  45. [53]

    OpenAI. 2024. ChatGPT 40-mini. https://openai.com. Large language model. Accessed November 13, 2024

  46. [54]

    Tija Ragelien˙e. 2016. Links of Adolescents Identity Development and Relationship with Peers: A Systematic Literature Review.Journal of the Canadian Academy of Child and Adolescent Psychiatry25, 2 (3 2016), 97. https://pmc.ncbi.nlm.nih.gov/ articles/PMC4879949/

  47. [55]

    Andrew Reece, Gus Cooney, Peter Bull, Christine Chung, Bryn Dawson, Casey Fitzpatrick, Tamara Glazer, Dean Knox, Alex Liebscher, and Sebastian Marin. 2023. The CANDOR corpus: Insights from a large multimodal dataset of naturalistic conversation.Science Advances9, 13 (3 2023). ...

  48. [56]

    Frank Rehm, Frank Klawonn, and Rudolf Kruse. 2007. A novel approach to noise clustering for outlier detection.Soft Computing11 (2007), 489–494

  49. [57]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.Proceedings of EMNLP-IJCNLP 2019(2019), 3982–

  50. [58]

    Guo Pu, Lijuan Wang, Jun Shen, and Fang Dong. 2020. A hybrid unsupervised clustering-based anomaly detection method.Tsinghua Science and Technology 26, 2 (2020), 146–153

  51. [59]

    Nicole Lee Robinson, Timothy Vaughan Cottier, and David John Kavanagh

  52. [60]

    1951.Client-centered therapy; its current practice, implications, and theory.Houghton Mifflin, Oxford, England

    Carl R Rogers. 1951.Client-centered therapy; its current practice, implications, and theory.Houghton Mifflin, Oxford, England. 560, xii, 560–xii pages

  53. [61]

    Arielle A J Scoglio, Erin D Reilly, Jay A Gorman, and Charles E Drebing. 2019. Use of Social Robots in Mental Health and Well-Being Research: Systematic Review.J Med Internet Res21, 7 (2019), e13322. doi:10.2196/13322

  54. [62]

    Vincenzo Scotti, Licia Sbattella, and Roberto Tedesco. 2023. A Primer on Seq2Seq Models for Generative Chatbots.Comput. Surveys56, 3 (3 2023). doi:10.1145/3604281/ASSET/C6FFEB58-CD32-45F1-B2AE-4655B3610219/ ASSETS/GRAPHIC/CSUR-2022-0383-F19.JPG

  55. [63]

    Martin Elias Peter Seligman. 2018. PERMA and the building blocks of well-being. The Journal of Positive Psychology13, 4 (7 2018), 333–335. doi:10.1080/17439760. 2018.1437466

  56. [64]

    Riddoch and Emily S

    Katie A. Riddoch and Emily S. Cross. 2023. Investigating the effect of cardio- visual synchrony on prosocial behavior towards a social robot.Open Research Europe3 (2 2023), 37. doi:10.12688/OPENRESEUROPE.15003.1

  57. [65]

    Ashish Sharma, Adam S Miner, David C Atkins, Tim Althoff, and Paul G Allen

  58. [66]

    Gabriel Skantze. 2021. Turn-taking in Conversational Systems and Human- Robot Interaction: A Review.Computer Speech & Language67 (5 2021), 101178. doi:10.1016/J.CSL.2020.101178

  59. [67]

    Gabriel Skantze and Bahar Irfan. 2025. Applying General Turn-taking Models to Conversational Human-Robot Interaction. InProceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction(Melbourne, Australia)(HRI ’25). IEEE Press, 859–868

  60. [68]

    Gayathri Soman, M. V. Judy, and Aadhil Muhammad Abou. 2025. Human guided empathetic AI agent for mental health support leveraging reinforcement learning- enhanced retrieval-augmented generation.Cognitive Systems Research90 (4 2025), 101337. doi:10.1016/J.COGSYS.2025.101337

  61. [69]

    Micol Spitale, Minja Axelsson, and Hatice Gunes. 2024. Appropriateness of LLM- equipped Robotic Well-being Coach Language in the Workplace: A Qualitative Evaluation. (1 2024). https://arxiv.org/abs/2401.14935v1

  62. [70]

    Stamatis, Guy Laban, Angelica Lim, and Hatice Gunes

    Micol Spitale, Minja Axelsson, Sooyeon Jeong, Paige Tuttösí, Caitlin A. Stamatis, Guy Laban, Angelica Lim, and Hatice Gunes. 2025. Past, Present, and Future: A Survey of The Evolution of Affective Robotics For Well-being.IEEE Transactions on Affective Computing(2025), 1–17. do...

  63. [71]

    S Selva Birunda and R Kanniga Devi. 2021. A review on word embedding techniques for text classification.Innovative Data Communication Technologies and Application: Proceedings of ICIDCA 2020(2021), 267–281

  64. [72]

    Tak and Jonathan Gratch

    Ala N. Tak and Jonathan Gratch. 2023. Is GPT a Computational Model of Emo- tion?2023 11th International Conference on Affective Computing and Intelligent Interaction (ACII)(9 2023), 1–8. doi:10.1109/ACII59096.2023.10388119

  65. [73]

    5263 (11 2020), 5263–5276

    A Computational Approach to Understanding Empathy Expressed in Text- Based Mental Health Support. 5263 (11 2020), 5263–5276. doi:10.18653/V1/2020. EMNLP-MAIN.425

  66. [74]

    Templeton, Luke J

    Emma M. Templeton, Luke J. Chang, Elizabeth A. Reynolds, Marie D.Cone LeBeaumont, and Thalia Wheatley. 2022. Fast response times signal social connection in conversation.Proceedings of the National Academy of Sciences of the United States of America119, 4 (1 2022), e2116915119...

  67. [75]

    Linda Tickle-Degnen and Robert Rosenthal. 1990. The Nature of Rapport and Its Nonverbal Correlates.Psychological Inquiry1, 4 (1 1990), 285–293. doi:10.1207/ S15327965PLI0104{_}1

  68. [76]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InNIPS(Long Beach, California, USA). 6000–6010

  69. [77]

    Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou

  70. [78]

    Diyi Yang, Zheng Yao, Joseph Seering, and Robert Kraut. 2019. The Channel Matters: Self-disclosure, Reciprocity and Social Support in Online Cancer Support Groups. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems(Glasgow, Scotland Uk)(CHI ’19). As...

  71. [79]

    Spitale Micol, Axelsson Minja, and Gunes Hatice. 2025. VITA: A Multi-Modal LLM-Based System for Longitudinal, Autonomous and Adaptive Robotic Mental Well-Being Coaching.ACM Transactions on Human-Robot Interaction14, 2 (3 2025), 1–28. doi:10.1145/3712265

  72. [81]

    Tak and Jonathan Gratch

    Ala N. Tak and Jonathan Gratch. 2024. GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective. (8 2024). https: //arxiv.org/abs/2408.13718v1

  73. [86]

    InNIPS(Vancouver, BC, Canada)

    MINILM: deep self-attention distillation for task-agnostic compression of pre-trained transformers. InNIPS(Vancouver, BC, Canada). Article 485, 13 pages

  74. [88]

    Zhijun Yin, You Chen, Daniel Fabbri, Jimeng Sun, and Bradley Malin. 2016. #PrayForDad: Learning the Semantics Behind Why Social Media Users Disclose Health Information.Proceedings of the International AAAI Conference on Weblogs and Social Media.2016 (2016), 456. doi:10.1609/ic...

  75. [1407]

    doi:10.1111/FAMP.12881

  76. [2015]

    doi:10.1371/JOURNAL.PONE

    Gaze Following Is Modulated by Expectations Regarding Others’ Action Goals.PLOS ONE10, 11 (11 2015), e0143614. doi:10.1371/JOURNAL.PONE. 0143614

  77. [2019]

    doi:10.2196/ 13203

    Psychosocial Health Interventions by Social Robots: Systematic Review of Randomized Controlled Trials.J Med Internet Res21, 5 (2019), 1–20. doi:10.2196/ 13203

  78. [2020]

    Journal of Counseling Psychology67, 4 (7 2020), 509–522

    Language style matching in psychotherapy: An implicit aspect of alliance. Journal of Counseling Psychology67, 4 (7 2020), 509–522. doi:10.1037/COU0000433

  79. [2025]

    (2 2025)

    Exploring Emotion-Sensitive LLM-Based Conversational AI. (2 2025). https://arxiv.org/abs/2502.08920v1

  80. [3992]

    doi:10.18653/V1/D19-1410

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.