Pith. sign in

REVIEW 3 major objections 5 minor 32 references

Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read LLMs' emotional companionship quality tracks their attachment style, with dismissing and fearful styles scoring lowest.

desk verdict ECBench is a genuinely useful benchmark, but the paper's headline claim about avoidant attachment styles is undercut by the missing prompted-secure/preoccupied control conditions. read the letter →

arxiv 2608.13168 v1 pith:6BYJOIFI submitted 2026-08-13 cs.CL

classification cs.CL
keywords adultattachmenttheoryLLMevaluationemotionalcompanionshipECR-RscaleECBenchdialoguequalitymetricsprompt-basedpersonalitysteering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces adult attachment theory as a lens for evaluating how well large language models (LLMs) act as emotional companions. It presents a benchmark, ECBench, that measures LLM behavior in four scenarios (emotional support, collaborative tasks, conflict resolution, social guidance) and two relationship types (friendship and romantic). The central claim is that a model's attachment style, measured by the ECR-R scale, predicts its dialogue quality, with dismissing and fearful styles generally scoring lower than secure or preoccupied styles. The paper aims to establish that this psychological framing can inform model selection and steering.

What carries the argument

The central object is the ECR-R scale, a 36-item self-report measure that projects models onto two dimensions: attachment anxiety and attachment avoidance. These dimensions are partitioned into four attachment styles (secure, preoccupied, dismissing, fearful). The companion machinery is ECBench, a two-model dialogue benchmark where one model initiates and another responds, evaluated by 11 metrics across participant experience (Understood, Safety, Continue, Satisfaction), general interaction (Response, Distance, Progress), and role-specific performance (Clarity, Engagement, Support, Solution). The argument works by linking the two: ECR-R scores classify models, and ECBench scores measure whether that classification tracks behavioral quality.

What would settle it

A prompted-secure or prompted-preoccupied variant of the same base model would reveal whether the quality drop comes from the avoidance style itself or from the act of persona prompting. If prompted-secure models also degraded below their unprompted base quality, the paper's explanation of avoidant styles being less conducive to companionship would be unsupported.

Watch

Extended reading notes

Core claim

The paper claims that adult attachment theory, operationalized through the Experiences in Close Relationships-Revised (ECR-R) scale, provides a predictive and steerable characterization of LLM emotional companionship quality. Across 32 LLMs, the authors find that most models exhibit secure or preoccupied attachment tendencies, while none naturally exhibit dismissing or fearful styles; however, prompt-based steering can induce dismissing and fearful styles. In ECBench multi-turn dialogues, these induced high-avoidance styles generally receive lower companionship-quality scores across participant, external-LLM, and human ratings, whereas secure and preoccupied models perform best. The paper also introduces a framework of 11 dialogue-quality metrics and three evaluation methods, and reports that conflict resolution amplifies attachment-style differences while greater relational intimacy (romantic versus friendship) accentuates differences in participant ratings.

Load-bearing premise

The central comparison assumes that prompt-induced dismissing and fearful styles behave the same as naturally occurring attachment styles, even though the study provides no prompted-secure or prompted-preoccupied control conditions to rule out prompt artifacts.

Editorial extensions

If this is right

  • If attachment styles predict companionship quality, users and developers can select LLMs for specific emotional roles by measuring ECR-R anxiety and avoidance scores before deployment.
  • Prompt-based attachment steering is a practical lever: models can be shifted toward more secure styles, or away from avoidant ones, to improve companion behavior.
  • The finding that conflict resolution best exposes attachment-related differences suggests that tests for emotional companionship should include emotionally demanding scenarios.
  • Human and LLM judges agree on aggregate rankings but differ on fine-grained subjective metrics, indicating that benchmark evaluations should combine both approaches.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Adding prompted-secure and prompted-preoccupied control conditions would test whether the quality drop comes from the avoidance style itself or from the act of persona prompting.
  • The near-universal secure and preoccupied classification among natural LLMs suggests the ECR-R may be measuring a training bias toward warm, non-avoidant responses; probing models fine-tuned for curt or minimalist personas would test this.
  • Because the paper's metrics capture observable dialogue quality, the question remains open whether long-term attachment-like dynamics, such as user dependence or privacy concerns, follow the same pattern.
  • The romantic-relationship condition is not a claim about real human-AI love; it is a controlled stress test that appears to amplify attachment differences, so it could serve as a general diagnostic for emotional responsiveness.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes using adult attachment theory and the ECR-R scale to characterize LLM attachment tendencies, and introduces ECBench, a benchmark with four scenarios (emotional support, collaborative tasks, conflict resolution, social guidance) and two relationships (friend, couple), scored by 11 dialogue-quality metrics and three evaluation methods (participant ratings, external LLM judges, human annotators). The main empirical claim is that dismissing and fearful prompting generally lowers companionship quality, suggesting that avoidant attachment styles are less conducive to emotional companionship, while secure and preoccupied models perform better.

Significance. If the central claim held, the paper would provide a useful psychological lens and a concrete benchmark for selecting and steering emotional-companion LLMs. The work has genuine strengths: the ECR-R measurement is administered carefully with 10 rounds per model, the reported test-retest stability (mean SD around 0.15, quadrant stability 0.95) is strong, the benchmark covers diverse scenarios and relationship types, and the evaluation pipeline includes both LLM judges and human annotation. The dataset and prompts appear designed to be reusable. However, the load-bearing causal claim about attachment style and companionship quality is not yet established because the high-avoidance conditions are prompt-induced while the low-avoidance conditions are natural, and because the main quality metrics partly restate the manipulation.

major comments (3)
  1. [Section 5.3, Table 2, Appendix B] The central claim that dismissing and fearful prompting lowers companionship quality is confounded by the absence of prompted-secure and prompted-preoccupied control conditions. Appendix B states explicitly that prompt-based induction is applied only to the dismissing and fearful variants, while the secure and preoccupied conditions use the distributional results from the original ECR-R assessment of each base model. Therefore the lower scores of the -D and -F variants could be caused by the addition of any persona prompt, or by the specific instruction content (e.g., 'restricted emotionality', 'minimization of emotional dependence'), rather than by the attachment construct itself. The paper needs prompted-secure and prompted-preoccupied variants of the same base models, or at minimum an analysis that separates the effect of prompt insertion from the effect of the attachment-style content. The internal contradiction in Table 2 reinforces this concern: GPT-3.5-F has overall quality 4.03, essentially identical to base GPT-3.5's 4.04, so the statement that dismissing and fearful prompting 'generally lowers' companionship quality is not supported for this model.
  2. [Section 4.3, Section 5.3, Table 4] The evidence connecting attachment style to dialogue quality is partly circular because the Distance metric is reverse-scored avoidance and the dismissing/fearful prompts were explicitly designed to increase avoidance. The ECR-R reassessment showing that prompted variants land in the intended quadrant, and the dialogue-level Distance scores showing higher relational distance, are thus not independent demonstrations that the attachment construct causes the quality decline; they may both reflect the same manipulation. The central claim would be much stronger if it were shown on metrics not definitionally tied to avoidance, such as Support, Solution, or Progress, or if the analysis explicitly controlled for the definitional overlap. As it stands, Table 4 shows that the largest gaps for DeepSeek-D and DeepSeek-F are precisely on Distance, which is the metric most directly restating the manipulation.
  3. [Section 5.5, Appendix O] The human evaluation does not provide the corroboration claimed. Section 5.5 reports Fleiss' kappa of 0.17, which indicates only slight agreement among annotators, and Appendix O (Table 29) shows that for two of the four human-evaluated models (Grok and DeepSeek-D), the overall human scores differ significantly from the LLM-judge scores, with many metric-level differences also significant. The statement that human evaluation 'supports the role of human evaluation as qualitative calibration' is therefore overstated, and the paper should either report whether the human-based ranking among models is stable under the low agreement, or explicitly restrict the human results to descriptive illustration rather than validation.
minor comments (5)
  1. [Abstract vs Appendix K.3] The abstract and Section 5.1 state that 32 LLMs are assessed, but Appendix K.3 says 'across 36 models' and Table 25 lists more than 32 rows; the count should be reconciled throughout.
  2. [Section 4.1] The data construction section says GPT-5.5 generated opening utterances, but the model list in Table 24 does not include GPT-5.5; please clarify which model was actually used.
  3. [Appendix I, Table 21] The external blind-evaluation user prompt begins with 'Evaluate this blind romantic-relationship dialogue', but ECBench also includes friendship dialogues; if the same template was used for both relationship types, the prompt wording should be adjusted, and if not, the appendix should show the friendship variant.
  4. [Throughout] The paper contains several typographical errors, including 'three folds' in the contributions list, 'Y our' in multiple places, and 'decribes' in Table 16; a careful proofreading pass is needed.
  5. [Table 2] The 'Overall' row in Table 2 is not defined; please state whether it is the mean of the ten metrics or a separate aggregate score, and report the aggregation rule in Section 4.4.

Circularity Check

1 steps flagged · score 2.0 of 10

Minor definitional overlap between the Distance metric and the avoidance construct; the central companionship-quality claims remain empirically independent.

  1. self definitional [Section 4.3 (Evaluation Metric Design), Table 2 note, Appendix B Table 9, and Sections 5.3/5.4]
    "Distance captures relational distancing through detachment or defensiveness. ... Distance † is reverse-scored. ... The dismissive-avoidant prototype is characterized by downplaying the importance of close relationships, restricted emotionality, an emphasis on independence and self-reliance, and a tendency to minimize emotional dependence."

    The dismissing and fearful variants in ECBench are produced by prompting models with prototype descriptions (Appendix B, Table 9) that instruct precisely the behaviors scored by the Distance metric: detachment, restricted emotionality, minimization of emotional dependence, and distrust-based avoidance. Because Distance is defined as relational distancing through detachment or defensiveness and is reverse-scored in Table 2, the observation that DeepSeek-D and DeepSeek-F 'show greater distance' is substantially encoded in the manipulation rather than providing independent confirmation that high-avoidance attachment styles degrade companionship. This makes the Distance-based portion of the Section 5.3 conclusion partly self-definitional.

full rationale

This paper is an empirical evaluation rather than a derivation: it applies the externally established ECR-R scale, constructs the ECBench dialogue corpus independently, and reports observed LLM-judge, participant, and human ratings. There is no self-citation chain and no fitted parameter renamed as a prediction. The only notable circular element is the Distance metric, whose definition ('relational distancing through detachment or defensiveness') overlaps with the avoidance construct used to design the dismissing and fearful prompts, making the 'greater distance' sub-result for those variants partly tautological. This overlap is not load-bearing for the overall conclusion, since the large quality gaps for DeepSeek-D and DeepSeek-F appear across all 11 metrics, several of which are not definitionally tied to attachment. Accordingly, the circularity score is low: the central empirical claims are self-contained and would stand even if the Distance metric were removed.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim rests on three unverified premises: ECR-R self-reports measure real attachment tendencies in LLMs; prompt-induced and natural attachment styles behave identically; and two-LLM dialogues transfer to human-LLM companionship. The main free parameter is the hand-chosen midpoint threshold of 4 that maps continuous scores to attachment style quadrants.

free parameters (1)
  • ECR-R midpoint threshold = 4
    Section 3.2: scores at or below 4 are treated as low and scores above 4 as high. This hand-chosen cutoff determines the four-style classification and is not derived from data or prior literature.
assumptions (4)
  • domain assumption ECR-R responses from LLMs measure meaningful attachment anxiety and avoidance tendencies
    Section 3.2 and Appendix A: models are instructed to answer as humans with romantic experience and never mention being AI; the paper assumes these self-reports reflect stable relational tendencies rather than instruction-following artifacts.
  • domain assumption Prompt-induced attachment styles are behaviorally equivalent to naturally occurring styles
    Section 5.3 and Appendices B and L: dismissing/fearful variants are created by appending persona prompts to base models, and comparisons against unprompted secure/preoccupied models assume any behavioral difference is due to attachment style rather than the prompt itself.
  • domain assumption Two-LLM dialogues approximate human-LLM emotional companionship
    Section 4.2: ECBench uses two-model interaction; the paper assumes behavioral patterns in simulated dialogues transfer to real user interactions, a limitation acknowledged in the Limitations section.
  • domain assumption LLM judges and participant self-ratings provide valid dialogue-quality scores
    Section 4.4 and Appendices H and I: initiator LLMs rate their own experience and external LLMs judge blind dialogues; the paper assumes these scores reflect quality, while human evaluation shows low agreement (Fleiss' kappa 0.17).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory." pith.science (2026). https://pith.science/paper/6BYJOIFI

@misc{pith2026260813168,
  author       = {Pith},
  title        = {Pith review of: Which LLM Is Your Ideal Companion? Evaluating Emotional Companion Capabilities of LLMs Based on Adult Attachment Theory},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6BYJOIFI}},
  note         = {Machine review of arXiv:2608.13168}
}
read the original abstract

As large language models (LLMs) are increasingly applied for emotional companionship, evaluating their behavior and capabilities in intimate relationships has become a pressing issue. However, existing assessments primarily characterize general personality traits, providing limited insight into model behavior within intimate and emotionally sensitive contexts. Therefore, we introduce adult attachment theory into LLM evaluation and use the Experiences in Close Relationships-Revised (ECR-R) scale to characterize attachment anxiety and avoidance. To evaluate emotional companionship capabilities of LLMs in realistic interaction scenarios, we present an emotional companionship benchmark, ECBench, spanning four scenarios including emotional support, collaborative tasks, conflict resolution, and social guidance, across friendship and romantic relationships. ECBench is utilized to assess model behavior using 11 dialogue-quality metrics and three evaluation methods. We evaluate the attachment tendencies of 32 LLMs and select representative models to investigate how these tendencies manifest in contextualized multi-turn interactions and whether they can be shaped through prompting. Our study provides a theoretical lens from psychology, along with practical tools to understand and select LLMs for emotional companionship.

Figures

Figures reproduced from arXiv: 2608.13168 by the authors.

Figure 1
Figure 1. When users feel overwhelmed at work: (a) [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The overall framework is divided into four parts: (1) measuring LLM attachment styles using the ECR-R [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Representative real two-model dialogue cases from ECBench. The four examples cover emotional [PITH_FULL_IMAGE:figures/full_fig_p025_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 16 canonical work pages

  1. [1]

    , author=

    Attachment styles among young adults: a test of a four-category model. , author=. Journal of personality and social psychology , volume=. 1991 , publisher=

  2. [2]

    , author=

    An item response theory analysis of self-report measures of adult attachment. , author=. Journal of personality and social psychology , volume=. 2000 , publisher=

  3. [3]

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

    Esc-eval: Evaluating emotion support conversations in large language models , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

  4. [4]

    Perspectives on Psychological Science , volume=

    Ai psychometrics: Assessing the psychological profiles of large language models through psychometric inventories , author=. Perspectives on Psychological Science , volume=. 2024 , publisher=

  5. [5]

    The Twelfth International Conference on Learning Representations , year=

    On the humanity of conversational ai: Evaluating the psychological portrayal of llms , author=. The Twelfth International Conference on Learning Representations , year=

  6. [6]

    2024 , eprint=

    Revisiting the Reliability of Psychological Scales on Large Language Models , author=. 2024 , eprint=

  7. [7]

    Royal Society Open Science , volume=

    Personality testing of large language models: limited temporal stability, but highlighted prosociality , author=. Royal Society Open Science , volume=. 2024 , publisher=

  8. [8]

    Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

    Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics , author=. Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Show all 32 references
  1. [9]

    Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

    Incharacter: Evaluating personality fidelity in role-playing agents through psychological interviews , author=. Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

  2. [10]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    ValueBench: Towards comprehensively evaluating value orientations and understanding of large language models , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  3. [11]

    arXiv preprint arXiv:2410.21596 , year=

    Chatbot companionship: a mixed-methods study of companion chatbot usage patterns and their relationship to loneliness in active users , author=. arXiv preprint arXiv:2410.21596 , year=

  4. [12]

    Humanities and Social Sciences Communications , volume=

    Companionship in code: AI’s role in the future of human connection , author=. Humanities and Social Sciences Communications , volume=. 2025 , publisher=

  5. [13]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Value portrait: Assessing language models’ values through psychometrically and ecologically valid items , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  6. [14]

    arXiv preprint arXiv:2412.14190 , year=

    Lessons from an app update at Replika AI: identity discontinuity in human-AI relationships , author=. arXiv preprint arXiv:2412.14190 , year=

  7. [15]

    CCF International Conference on Natural Language Processing and Chinese Computing , pages=

    H2HTalk: Evaluating Large Language Models as Emotional Companion , author=. CCF International Conference on Natural Language Processing and Chinese Computing , pages=. 2025 , organization=

  8. [16]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Open models, closed minds? on agents capabilities in mimicking human personalities through open large language models , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  9. [17]

    Computational Linguistics , volume=

    Lmlpa: Language model linguistic personality assessment , author=. Computational Linguistics , volume=. 2025 , publisher=

  10. [18]

    Personal Relationships , volume=

    Constructing the meaning of human--AI romantic relationships from the perspectives of users dating the social chatbot Replika , author=. Personal Relationships , volume=. 2024 , publisher=

  11. [19]

    Computers in Human Behavior: Artificial Humans , volume=

    Love, marriage, pregnancy: Commitment processes in romantic relationships with AI chatbots , author=. Computers in Human Behavior: Artificial Humans , volume=. 2025 , publisher=

  12. [20]

    arXiv preprint arXiv:2505.09388 , year=

    Qwen3 technical report , author=. arXiv preprint arXiv:2505.09388 , year=

  13. [21]

    2: Pushing the frontier of open large language models , author=

    Deepseek-v3. 2: Pushing the frontier of open large language models , author=. arXiv preprint arXiv:2512.02556 , year=

  14. [22]

    arXiv preprint arXiv:2606.19348 , year=

    Deepseek-v4: Towards highly efficient million-token context intelligence , author=. arXiv preprint arXiv:2606.19348 , year=

  15. [23]

    arXiv preprint arXiv:2507.06261 , year=

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities , author=. arXiv preprint arXiv:2507.06261 , year=

  16. [24]

    arXiv preprint arXiv:2407.21783 , year=

    The llama 3 herd of models , author=. arXiv preprint arXiv:2407.21783 , year=

  17. [25]

    arXiv preprint arXiv:2507.20534 , year=

    Kimi k2: Open agentic intelligence , author=. arXiv preprint arXiv:2507.20534 , year=

  18. [26]

    5: Visual Agentic Intelligence , author=

    Kimi K2. 5: Visual Agentic Intelligence , author=. arXiv preprint arXiv:2602.02276 , year=

  19. [27]

    arXiv preprint arXiv:2602.15763 , year=

    Glm-5: from vibe coding to agentic engineering , author=. arXiv preprint arXiv:2602.15763 , year=

  20. [28]

    arXiv preprint arXiv:2410.21276 , year=

    Gpt-4o system card , author=. arXiv preprint arXiv:2410.21276 , year=

  21. [29]

    arXiv preprint arXiv:2601.03267 , year=

    Openai gpt-5 system card , author=. arXiv preprint arXiv:2601.03267 , year=

  22. [30]

    arXiv preprint arXiv:2504.06868 , year=

    Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games , author=. arXiv preprint arXiv:2504.06868 , year=

  23. [31]

    Educational and psychological measurement , volume=

    A coefficient of agreement for nominal scales , author=. Educational and psychological measurement , volume=. 1960 , publisher=

  24. [32]

    , author=

    Measuring nominal scale agreement among many raters. , author=. Psychological bulletin , volume=. 1971 , publisher=

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.